Giving AI smaller questions
A project produces a lot of information that should not become permanent documentation. A passing suggestion. A build that is still running. A decision nobody has actually accepted yet.
It also produces things worth keeping: why a tool behaves the way it does, which approach won, a preference that would save another round of explaining.
I keep reusable notes in a hub repository that coordinates work across my projects. Deciding what goes in there is a small judgment every time. Does this observation deserve a record, and does the evidence support what the record would say?
That is one of the places where I have started using Jev, TypeSafe’s model for focused judgments. The other experiments in this post are project names, finding shell commands and choosing blog stories.
All of them are early, and they all split the work the same way. Code gathers the evidence. Jev answers one bounded question. The caller decides what happens next. Generating names or writing prose is a job for a larger model.
Give the judgment a place in the workflow.
Tools gather evidence and supply it to Jev. Jev returns a focused judgment. The caller then reviews the answer and either takes the next step or asks for more evidence.
View Mermaid source
flowchart TD
A[Tools gather evidence] --> B[Jev answers a focused question]
B --> C[Code and reviewer inspect the answer]
C --> D[Take the next step]
C --> E[Review or gather more evidence]
E --> A Give the model something it can answer
TypeSafe describes Jev as a “System One” model: it makes focused judgments on the context you give it and returns typed answers and probabilities rather than explanations or prose.
“Typed” means the answer has a defined shape. If I supply four categories, the result is one of those four. Code can use that directly without fishing a decision out of a paragraph.
The supplied context is called state. In my integrations, Jev cannot follow a repository link or fetch anything it is missing. The calling tool has to bring the evidence with it.
I use three question types. Choice picks from known options. Score rates something on explicitly ordered levels. Noul, despite the odd name, does something simple: it gives a probability for a yes/no question.
The answer starts with the question.
Open a question type to see how it fits these tools.
ChoiceWhich option fits?
“Which story is worth developing next?”
- Supply
- Candidate briefs, their source evidence, and the articles already on the blog.
- Receive
- One of the supplied story options, or “none”, with probabilities for the options.
- Use it
- Use the recommendation to choose what to investigate or draft. The editor still decides the angle.
ScoreHow far along a scale?
“How ready is this story to draft?”
- Supply
- The same candidate evidence, plus ordered descriptions: idea only → evidence missing → draftable → concrete example.
- Receive
- A position along those defined levels, with a distribution over them.
- Use it
- Compare readiness and identify evidence gaps. A relative favorite can still be a weak candidate.
NoulIs this supported?
“Does the evidence support this proposed note?”
- Supply
- The candidate note and the observations or user statements offered as evidence.
- Receive
- The probability of “yes”, from 0 to 1. Near 0.5 means uncertainty, not a medium-strength fact.
- Use it
- Combine it with the other capture checks. An agent verifies the evidence before writing a record.
The distinctions matter. “Which candidate is best?” and “Is this candidate good enough?” are different questions. So are “How serious is this?” and “How likely is this to be serious?” A yes probability of 0.5 means uncertainty, not medium severity. The Noul documentation explains how to read it.
Independent questions can share the same state in one batch. They cannot see each other’s answers; my code combines them afterwards.
A tidy answer shape does not make the facts I supplied true. That part stays with me.
Is this note worth keeping?
My knowledge-capture tool is a Rust command-line program. It takes a candidate note, its supporting evidence and the existing records, then asks Jev three questions: what kind of record is this, is it worth keeping, and does the evidence support it?
The categories are fact, decision, preference and skip. Code combines the three answers into capture, review or skip.
Three answers, one recommendation, nothing written yet.
- Gather
Collect the candidate
The command receives a proposed note, its supporting evidence, and the records the hub already keeps.
- Ask
Three bounded questions
What kind of record is this, is it worth keeping, and does the evidence support it?
- Combine
Capture, review, or skip
Code turns the three typed answers into a single recommendation.
- Verify
Check before writing
The assistant re-reads the actual evidence, and whether I accepted the decision, before a record exists.
Even “capture” writes nothing. The command returns a recommendation. The coding assistant then re-reads the actual evidence and, where it matters, checks whether I accepted the decision before a record exists. The public capture workflow describes that handoff.
The integration is live, but I have not tuned the thresholds yet. In the first sanitized examples even the good candidates came back as “review”. That says the settings are cautious, not how accurate the model is.
I like that boundary. The model can flag a note worth looking at. Checking the evidence stays a separate step.
Picking a project name
Naming a project splits the same way.
My naming workflow uses Codex to generate candidates. A Rust tool checks the relevant registries and APIs. Jev then rates naming fit, whether the supplied evidence points at a confusing product in the same category, and which candidate to recommend. “None” is always one of the options.
Jev does not generate the names here. It also cannot turn a failed registry request into an availability result. Unknown stays unknown.
In a live comparison of three candidates it preferred Taplume. Asking about absolute fit made that a lot less flattering:
First place is not an approval stamp.
Which of these would you pick?
Preferred
Jev’s favorite among the three supplied candidates.
How good is the fit?
Weak to acceptable
The fit assessments left room for a better shortlist.
Decision left open. No name was accepted.
I kept that result as it is. Winning a comparison is not the same as clearing the bar. A shortlist can have a first place and still not contain a name I should use.
Later web searches turned up historical listings for a similarly named Android app in third-party catalogs. Those were leads, not verified first-party evidence of a conflict. No name was accepted.
The tool gives me a shortlist. It does not give me legal clearance.
Finding a shell command
Quirl is my shell project. A shell interprets what you type into a terminal, including the flags that only become memorable after you have looked them up for the third time.
For a command-discovery experiment, I gave Jev descriptions of 62 known commands. It could choose one, or return NONE, CLARIFY or COMPOSE when nothing fit, when it needed more information, or when the request called for several commands.
It only selected options. It did not assemble command lines or run anything, and it is not wired into the shell.
One case asked for recursive directory sizes. The expected answer and Jev disagreed. Open the catalog evidence before deciding which one was wrong:
The answer key needed another look.
“Show recursively how much disk space each subdirectory consumes, not the directory entry size.”
NONE
The test said no catalog option fit.
tree
The model picked a known command.
What was already in the catalog?
tree --duDirectory totals in a tree vieweza --total-sizeDirectory totals in a listingThe directory-total intent had support. The correction belonged in the expected answer. Exact filesystem accounting was outside this test.
The catalog did contain options that supported directory totals. The expected answer was wrong, so I corrected it and kept the original next to the correction. The disagreement had found a bug in the test.
The benchmark report keeps the one unresolved case visible next to the validated results:
What the test actually accounted for.
- Unique cases
- 112
- The denominator for this experiment.
- Acceptable results
- 111
- Validated after the label correction.
- Unassessed response
- 1
- The test runner rejected it; it is not counted as a success.
That is why I keep the evidence next to the recommendation. “The model disagreed” is the start of an investigation, and the command descriptions are what let me finish it.
Choosing a blog story
The draft behind this article went through one more experiment.
A new storytelling workflow gathers repository evidence and checks the existing posts and drafts. It hands Jev summarized story angles and coverage, then asks for a recommendation and a readiness rating per angle.
A pilot compared four possible stories. Jev recommended the Quirl benchmark story, and Codex wrote the draft. Then the question changed from “which event is interesting?” to “what does a reader need in order to care?”
These four steps are how the article you are reading came about, including the point where my own feedback changed its scope.
A promising story still needs a reader.
- JEV · SELECT
A benchmark answer was wrong
A concrete disagreement made the Quirl experiment the recommended story.
- CODEX · DRAFT
Explain the interesting case
The first draft followed the benchmark correction.
- MY FEEDBACK · REFRAME
Why was Jev there at all?
Introduce Jev, show how I use it across projects, and give a new reader enough context.
- REVISED DRAFT · LOCAL REVIEW
Giving AI smaller questions
The benchmark becomes one example alongside notes, names, and blog stories.
The first recommendation was useful. The benchmark disagreement made it into this article. It just needed a wider frame, because before a reader can care about Jev choosing a shell command, I have to explain what Jev is doing in my projects at all.
Keep the question small
The use I keep coming back to is a small decision between gathering information and acting on it. Is this note worth keeping? Does this name fit? Which command deserves a closer look? Which piece of work could become a story?
What these have in common is that I can put the evidence next to the answer and look at the disagreement. Where to set the cutoff for trusting an answer still needs testing. A confidence value says how concentrated the model’s answer is. It cannot tell me the question was a good one.
For the next integration I would start the same way: one recurring judgment, the context it needs, and a way to leave it unresolved. A name can stay unchosen. A note can wait for review. And a perfectly reasonable story recommendation can turn into a different article.