THE DISTINCTION / STATE, EXECUTION, AND EVIDENCE

Tools are part
of the picture.

Giving an agent custom functions does not make it a different kind of system. The meaningful questions are where exact state lives, how results reach the answer, and who checks the computation.

01 / TWO DIFFERENT FAILURE MODES

Right result.
Wrong rendering.

A tool computes an invoice total of $1,283.47. If the model rewrites it as $1,238.47, the calculation was correct but the answer was corrupted. A registry reference rendered directly closes that particular path.

Right execution.
Wrong meaning.

A model selects last month’s invoice instead of this month’s. The tool can execute perfectly and still answer the wrong question. Exact rendering cannot fix that. Binding checks, behavioral evidence, and abstention address a separate problem.

02 / COMPARE THE CONTRACTS

What changes
around the call?

Illustrative tool loop versus our integration milestone and research target. Tool calling can also be engineered with these protections.
ConcernA basic tool loopFirst ExactAI milestoneResearch target
Custom functionsModel selects a callable.Qwen selects typed application functions.ExactNet composes dynamically supplied typed operators.
Intermediate stateTool results return to model context.Typed registry objects persist across turns; context uses descriptors and references.Intermediate program values remain in runtime state.
Multi-step workModel chooses successive tool calls.One session coordinates Qwen, functions, and built-in ExactNet operations.One semantic delegation can produce an entire exact program.
Final answerModel may write values into prose.ExactIR references and formatting materialize exact spans.The same rendering boundary is retained.
ChecksDepend on the application and tools.References, signatures, result types, execution errors, and provenance.Add semantic types, behavioral checks, and calibrated execute/defer decisions.
New functionsExpose descriptions to the model.No ExactNet retraining; Qwen owns selection.Test unseen-function composition by ExactNet without new output classes.
Private valuesExposure depends on tool arguments, returned data, and deployment.Raw PII can stay in a local registry; use opaque IDs and minimal descriptors.Preserve that boundary through program synthesis and execution.

These are architectural choices, not capabilities forbidden to tool-calling agents. A well-designed tool system can preserve typed state, validate arguments, and render results directly. Our hypothesis is that packaging these boundaries with a specialist for exact programs can make them easier to adopt and more effective. That needs comparative evidence.

Privacy is also a deployment choice, not an exclusive advantage over tools. Local models and local functions can keep private information on your system. References help minimize disclosure, but identifying prompts, descriptors, logs, or outputs can still leak it. See the local-data design →

03 / FROM CALL SELECTION TO PROGRAM SELECTION

Delegate the subproblem.
Keep its state intact.

Illustrative iterative tool loop

LLM → invoice_total(subtotal, tax)
    ← result value
LLM → subtract(current, previous)
    ← result value
LLM → compose answer

Each return is another opportunity for the application to preserve or lose the exact-value contract.

Future ExactAI target

LLM → “Calculate and compare invoices”
ExactNet → candidate typed program
Verifier → accept / clarify / defer
Runtime → execute accepted program
LLM ← result references
Renderer → materialize exact spans

Moving an exact subproblem into one program could reduce repeated large-model orchestration. Latency and cost improvements remain hypotheses to measure.

04 / WHAT WE ARE NOT CLAIMING

Useful boundaries.
Honest limits.

We did not invent function calling

Calculators, typed representations, references, program synthesis, and external execution are building blocks. Custom functions alone are not a novelty claim.

Verification is not omniscience

A type-correct program can be semantically wrong. Behavioral contracts cover specified properties and cases; they do not automatically recover an ambiguous user’s intention.

Exact values do not make all prose factual

A trace explains the source and computation of an exact span. It does not certify every sentence around it. Literal detection is an incomplete guardrail.

The target is not the current prototype

Verified synthesis, RLBC, calibrated abstention, and unseen-function composition belong to the research roadmap. The first milestone is the integrated conversational wrapper.

Read how we plan to test the distinction →