Skip to main content
Back to posts

💡 Giving AI smaller questions

A project produces a lot of information that should not become permanent documentation. A passing suggestion. A build that is still running. A decision nobody has actually accepted yet.

It also produces things worth keeping: why a tool behaves the way it does, which approach won, a preference that would save another round of explaining.

I keep reusable notes in a hub repository that coordinates work across my projects. Deciding what goes in there is a small judgment every time. Does this observation deserve a record, and does the evidence support what the record would say?

That is one of the places where I have started using Jev, TypeSafe’s model for focused judgments. The other experiments in this post are project names, finding shell commands and choosing blog stories.

All of them are early, and they all split the work the same way. Code gathers the evidence. Jev answers one bounded question. The caller decides what happens next. Generating names or writing prose is a job for a larger model.

One pattern / Different jobs

Give the judgment a place in the workflow.

Tools gather evidence and supply it to Jev. Jev returns a focused judgment. The caller then reviews the answer and either takes the next step or asks for more evidence.

Diagram description and source remain available.
View Mermaid source
flowchart TD
  A[Tools gather evidence] --> B[Jev answers a focused question]
  B --> C[Code and reviewer inspect the answer]
  C --> D[Take the next step]
  C --> E[Review or gather more evidence]
  E --> A
The division of work in these integrations. Jev does not fetch the evidence or carry out the next action.

Give the model something it can answer

TypeSafe describes Jev as a “System One” model: it makes focused judgments on the context you give it and returns typed answers and probabilities rather than explanations or prose.

“Typed” means the answer has a defined shape. If I supply four categories, the result is one of those four. Code can use that directly without fishing a decision out of a paragraph.

The supplied context is called state. In my integrations, Jev cannot follow a repository link or fetch anything it is missing. The calling tool has to bring the evidence with it.

I use three question types. Choice picks from known options. Score rates something on explicitly ordered levels. Noul, despite the odd name, does something simple: it gives a probability for a yes/no question.

Three question shapes

The answer starts with the question.

Open a question type to see how it fits these tools.

ChoiceWhich option fits?

“Which story is worth developing next?”

Supply
Candidate briefs, their source evidence, and the articles already on the blog.
Receive
One of the supplied story options, or “none”, with probabilities for the options.
Use it
Use the recommendation to choose what to investigate or draft. The editor still decides the angle.
How Choice works
ScoreHow far along a scale?

“How ready is this story to draft?”

Supply
The same candidate evidence, plus ordered descriptions: idea only → evidence missing → draftable → concrete example.
Receive
A position along those defined levels, with a distribution over them.
Use it
Compare readiness and identify evidence gaps. A relative favorite can still be a weak candidate.
How Score works
NoulIs this supported?

“Does the evidence support this proposed note?”

Supply
The candidate note and the observations or user statements offered as evidence.
Receive
The probability of “yes”, from 0 to 1. Near 0.5 means uncertainty, not a medium-strength fact.
Use it
Combine it with the other capture checks. An agent verifies the evidence before writing a record.
How Noul works
Examples of the questions used by the editorial and knowledge tools, simplified for explanation. This figure makes no API calls and displays no invented model scores.

The distinctions matter. “Which candidate is best?” and “Is this candidate good enough?” are different questions. So are “How serious is this?” and “How likely is this to be serious?” A yes probability of 0.5 means uncertainty, not medium severity. The Noul documentation explains how to read it.

Independent questions can share the same state in one batch. They cannot see each other’s answers; my code combines them afterwards.

A tidy answer shape does not make the facts I supplied true. That part stays with me.

Is this note worth keeping?

My knowledge-capture tool is a Rust command-line program. It takes a candidate note, its supporting evidence and the existing records, then asks Jev three questions: what kind of record is this, is it worth keeping, and does the evidence support it?

The categories are fact, decision, preference and skip. Code combines the three answers into capture, review or skip.

Knowledge capture / Hub

Three answers, one recommendation, nothing written yet.

  1. Gather

    Collect the candidate

    The command receives a proposed note, its supporting evidence, and the records the hub already keeps.

  2. Ask

    Three bounded questions

    What kind of record is this, is it worth keeping, and does the evidence support it?

  3. Combine

    Capture, review, or skip

    Code turns the three typed answers into a single recommendation.

  4. Verify

    Check before writing

    The assistant re-reads the actual evidence, and whether I accepted the decision, before a record exists.

The recorded shape of the capture command. It describes the handoff between the tool and the assistant, not a measured success rate.

Even “capture” writes nothing. The command returns a recommendation. The coding assistant then re-reads the actual evidence and, where it matters, checks whether I accepted the decision before a record exists. The public capture workflow describes that handoff.

The integration is live, but I have not tuned the thresholds yet. In the first sanitized examples even the good candidates came back as “review”. That says the settings are cautious, not how accurate the model is.

I like that boundary. The model can flag a note worth looking at. Checking the evidence stays a separate step.

Picking a project name

Naming a project splits the same way.

My naming workflow uses Codex to generate candidates. A Rust tool checks the relevant registries and APIs. Jev then rates naming fit, whether the supplied evidence points at a confusing product in the same category, and which candidate to recommend. “None” is always one of the options.

Jev does not generate the names here. It also cannot turn a failed registry request into an availability result. Unknown stays unknown.

In a live comparison of three candidates it preferred Taplume. Asking about absolute fit made that a lot less flattering:

Project naming / Two different questions

First place is not an approval stamp.

ONE CANDIDATETaplumeSame name. Two useful answers.
RELATIVE CHOICE

Which of these would you pick?

Preferred

Jev’s favorite among the three supplied candidates.

ABSOLUTE FIT

How good is the fit?

Weak to acceptable

The fit assessments left room for a better shortlist.

Decision left open. No name was accepted.

From the live three-candidate comparison. These are qualitative findings, not a measured scale or a claim that the name is available.

I kept that result as it is. Winning a comparison is not the same as clearing the bar. A shortlist can have a first place and still not contain a name I should use.

Later web searches turned up historical listings for a similarly named Android app in third-party catalogs. Those were leads, not verified first-party evidence of a conflict. No name was accepted.

The tool gives me a shortlist. It does not give me legal clearance.

Finding a shell command

Quirl is my shell project. A shell interprets what you type into a terminal, including the flags that only become memorable after you have looked them up for the third time.

For a command-discovery experiment, I gave Jev descriptions of 62 known commands. It could choose one, or return NONE, CLARIFY or COMPOSE when nothing fit, when it needed more information, or when the request called for several commands.

It only selected options. It did not assemble command lines or run anything, and it is not wired into the shell.

One case asked for recursive directory sizes. The expected answer and Jev disagreed. Open the catalog evidence before deciding which one was wrong:

Quirl / Inspect the disagreement

The answer key needed another look.

“Show recursively how much disk space each subdirectory consumes, not the directory entry size.”
EXPECTED ANSWER

NONE

The test said no catalog option fit.

JEV’S SELECTION

tree

The model picked a known command.

What was already in the catalog?
tree --duDirectory totals in a tree view
eza --total-sizeDirectory totals in a listing

The directory-total intent had support. The correction belonged in the expected answer. Exact filesystem accounting was outside this test.

One recorded command-selection case. The original mismatch was preserved when the expected answer was corrected. No command was executed by the experiment.

The catalog did contain options that supported directory totals. The expected answer was wrong, so I corrected it and kept the original next to the correction. The disagreement had found a bug in the test.

The benchmark report keeps the one unresolved case visible next to the validated results:

Quirl / Exploratory command selection

What the test actually accounted for.

Unique cases
112
The denominator for this experiment.
Acceptable results
111
Validated after the label correction.
Unassessed response
1
The test runner rejected it; it is not counted as a success.
Handwritten exploratory cases, after correcting the expected answer above. This is not an accuracy estimate for everyday shell use.

That is why I keep the evidence next to the recommendation. “The model disagreed” is the start of an investigation, and the command descriptions are what let me finish it.

Choosing a blog story

The draft behind this article went through one more experiment.

A new storytelling workflow gathers repository evidence and checks the existing posts and drafts. It hands Jev summarized story angles and coverage, then asks for a recommendation and a readiness rating per angle.

A pilot compared four possible stories. Jev recommended the Quirl benchmark story, and Codex wrote the draft. Then the question changed from “which event is interesting?” to “what does a reader need in order to care?”

These four steps are how the article you are reading came about, including the point where my own feedback changed its scope.

This article / From selection to framing

A promising story still needs a reader.

  1. JEV · SELECT

    A benchmark answer was wrong

    A concrete disagreement made the Quirl experiment the recommended story.

  2. CODEX · DRAFT

    Explain the interesting case

    The first draft followed the benchmark correction.

  3. MY FEEDBACK · REFRAME

    Why was Jev there at all?

    Introduce Jev, show how I use it across projects, and give a new reader enough context.

  4. REVISED DRAFT · LOCAL REVIEW

    Giving AI smaller questions

    The benchmark becomes one example alongside notes, names, and blog stories.

The actual sequence behind this draft. Jev selected material; Codex wrote prose; my feedback changed what the article needed to explain.

The first recommendation was useful. The benchmark disagreement made it into this article. It just needed a wider frame, because before a reader can care about Jev choosing a shell command, I have to explain what Jev is doing in my projects at all.

Keep the question small

The use I keep coming back to is a small decision between gathering information and acting on it. Is this note worth keeping? Does this name fit? Which command deserves a closer look? Which piece of work could become a story?

What these have in common is that I can put the evidence next to the answer and look at the disagreement. Where to set the cutoff for trusting an answer still needs testing. A confidence value says how concentrated the model’s answer is. It cannot tell me the question was a good one.

For the next integration I would start the same way: one recurring judgment, the context it needs, and a way to leave it unresolved. A name can stay unchosen. A note can wait for review. And a perfectly reasonable story recommendation can turn into a different article.