← All writing

I tried Jev. The interesting part was where I disagreed with it.

A personal note on using Jev to sort public problem signals, overruling a misleading fit score, and leaving the real decisions with people.

I wanted to know which problems were worth looking into for Kaliits. It is easy to find posts about automation. It is harder to tell whether the person has a real problem, is describing an ordinary responsibility, or is selling their own software.

That was the question behind my Jev experiment. Could a model help me sort those signals without turning every mention of a problem into a supposed lead?

A smaller job for AI

TypeSafe describes Jev as a model that takes a state and predefined questions, then returns typed decisions with probabilities. That caught my attention because I did not need another paragraph telling me that logistics has huge potential. I needed a way to ask the same focused questions across a batch of posts.

In my experiment, the state was public post context. The questions covered the main problem category, whether the account was firsthand, whether it was vendor promotion, and how well the workflow might fit Kaliits.

The model helped organize the material. It did not decide whom to contact or what to build.

What I actually tested

On 25 September 2026, I reviewed a batch of 16 public LinkedIn and Reddit items, used Jev to classify them, and checked the inclusions manually. This was a batch experiment, not a live feed. No outreach was sent as part of that research.

The material included hiring posts, practitioner commentary, vendor promotions and operator questions. Those differences mattered more than a single fit score.

A hiring post can tell me that people check documents, update systems and coordinate shipments. It does not tell me that those people need my software. A vendor describing a painful process might have useful language, but their sales post is not independent evidence that someone else wants to buy.

The score I needed to overrule

One post described delays associated with inspections. Jev gave it a high fit score.

The pain could be real while the proposed software fit was weak. An internal application cannot make a required physical inspection disappear. It might help with preparation, visibility or updates, but that is a different claim.

I excluded that signal from the solvable-problem count. That disagreement was useful: it showed me a distinction my questions and review process needed to preserve.

Does this sound painful? Can we influence the cause? Does the person actually want help? Those are separate questions. Combining them into one score makes the research look cleaner while hiding the decision.

What changed in my thinking

Document readiness and ownership of the next action emerged as themes worth investigating. The idea became narrower: receive one type of dossier, find missing information, make the responsible person visible, and prepare the next step for review.

That was a hypothesis. The batch did not establish demand, frequency, losses or willingness to pay. It also did not justify abandoning every other idea.

What I gained was a better set of questions for conversations. Show me the last case that needed chasing. What was missing? Who noticed? Who owned the next action? What does the current software already do?

Where I would use Jev again

I would use it where the possible judgments are clear but the language is messy. I would keep the original source beside the result and review disagreements before trusting a larger batch.

For exact quantities, permissions and calculations, I would use code. For writing, I would use a generative model. For deciding whether a problem deserves weeks of work, I would still need people and actual cases.

A typed answer is convenient for software. It can still be the wrong answer.

The interesting outcome of this experiment was not a list of customers. It was a clearer view of what I knew, what I was guessing, and which questions deserved a conversation.

TypeSafe's introduction to Jev

The experiment described here is the dated 25 September research pass. It is not a fresh market survey or an independent benchmark of Jev's speed, price or accuracy.

If this connects to an actual workflow you want to discuss, Kaliits is where I handle that conversation.

All writing ↗