METHODOLOGY / THE EXACTAI RESEARCH DIRECTION

From intent
to checked programs.

Preserve values in the runtime. Learn which computation the user means. Evaluate these as separate responsibilities.

01 / CORE BELIEFS

One agent.
Distinct responsibilities.

The application should see one client and one session. Conversation, exact decisions, execution, and rendering can use different components without becoming separate agents the user must manage.

Exact state belongs outside prose

An invoice total is a typed object with a stable identity. Later turns should reference that object instead of reconstructing it from a generated sentence.

Determinism is not understanding

A perfectly executed subtraction can use the wrong period, entity, or metric. Type validity and execution success are necessary evidence, not proof of semantic correctness.

Precision before mandatory coverage

Our target is a system that can defer or ask a useful question. We will measure correctness among executed requests alongside the fraction of requests it chooses to execute.

02 / FIRST MILESTONE — CONVERSATIONAL INTEGRATION

Connect the pieces.
Keep the values.

The existing prototype contains the runtime and a specialist operation selector. The next milestone integrates them into one local Qwen conversation, with no ExactNet retraining.

Qwen orchestrates

It interprets the request, selects custom functions from descriptions and typed signatures, and delegates supported exact questions to ExactNet.

ExactNet binds and selects

It works with its existing built-in operations. Registering a function makes that function available to the runtime; it does not teach the current ExactNet model a new operation.

The runtime owns exact state

Validate references and signatures, execute trusted functions, validate their results, and store fresh bindings. The renderer materializes exact spans; structured history retains references.

The session owns the loop

Execute, delegate, or finish. Intermediate results return as IDs, types, descriptors, and provenance. Failures are explicit, and the turn budget is eight Qwen/ExactNet inference calls combined.

Follow the two-turn invoice example →

PRIVACY / KEEP PRIVATE VALUES AT THE SOURCE

Private information
can stay local.

Typed references let an agent reason about a sensitive object without requiring its raw value in model context. The registry, function execution, and renderer can all remain inside your application’s local boundary.

What the model needs

Reference: object_7
Type: OpaqueIdentifier
Meaning: selected customer account

Raw account identifier:
stored only in the local registry

Use opaque IDs and minimal descriptors. Trusted local functions resolve references to native values; permitted output can be rendered locally from those same references.

What the application controls

Explicit registration is the dependable path for sensitive fields. Existing exact-value detectors are not a comprehensive PII detector: names, addresses, and other identifying context require deliberate handling.

The initial deployment uses local models. A future remote-provider integration could send minimized references, but would still require review of prompts, descriptors, logs, functions, and output destinations. Reference substitution alone does not guarantee anonymity or prevent data leakage.

03 / RESEARCH TARGET — A NEURAL COMPILER

Learn composition.
Verify behavior.

A fixed class for every business calculation cannot keep growing forever. The proposed next ExactNet produces structured, typed programs from a small instruction set and dynamically supplied function descriptions.

  1. Propose

    Generate candidate programs from intent, typed registry descriptors, and available operations.

  2. Verify

    Check syntax, references, types, semantic dependencies, and behavior where contracts are available.

  3. Choose

    Use calibrated evidence to execute, ask for clarification, defer, or report unsupported work.

  4. Execute

    Run the accepted exact computation and preserve its results and provenance for future turns.

Richer descriptors—entity, metric, period, currency—would help eliminate wrong bindings. New application functions would arrive as typed operator descriptions rather than new neural output classes. Both are future research work.

04 / RLBC — REINFORCEMENT LEARNING FROM BEHAVIORAL CONTRACTS

Reward what a program does.
Not just how it is written.

First learn programs with supervised examples. Then build synthetic tasks with known semantics and behavioral contracts. Finally, investigate reinforcement learning that rewards behavior across examples and interventions.

At a high level: the model proposes a computation, a test environment runs it, and the results provide feedback for training. Reward correct behavior when inputs change; penalize confident mistakes; reward appropriate clarification or deferral. This is planned training on controlled tasks, not online experimentation with users’ private data.

Example: revenue per employee

result = revenue / employee_count

Required: revenue, employee_count
Forbidden: operating_income

Double revenue → double result
Double employees → halve result
Change unrelated metric → unchanged

Change the inputs, test the meaning

A wrong program can accidentally give the right answer on one dataset. Counterfactual values and targeted interventions test whether its behavior continues to match the intended computation.

Equivalent programs should receive credit even when their syntax differs. Contracts are only as informative as their properties and test cases; passing them is not a universal proof of user intent.

Track syntax, type, binding, behavior, counterfactual, provenance, efficiency, and calibration rewards separately before choosing a combined objective. RLBC is a proposed training framework here, not a demonstrated performance result.

05 / WHAT WOULD COUNT AS PROGRESS?

Measure the answer.
Measure the decision to answer.

Integration evidence

Can one session invoke a new custom function, retain its result, delegate a follow-up to ExactNet, and render the answer correctly? Test orchestration separately from the models’ semantic choices.

Generalization evidence

Hold out function definitions and entire operation compositions during training. Supply unseen typed descriptors at evaluation time and test whether the model can use them correctly.

Precision at coverage

Report semantic correctness among executed requests together with execution coverage and confidence calibration. A low error rate achieved by declining nearly everything needs to be visible.

Behavioral evidence

Compare supervised learning, execution rewards, behavioral rewards, counterfactual rewards, and full RLBC. Include wrong-entity and wrong-period examples, program cost, and end-to-end latency.

06 / THE ORDER OF WORK

A practical first step.
A testable research path.

  1. Conversational integrationCustom functions, existing ExactNet, session registry, trace. No retraining.
  2. Compositional ExactIRPrimitive instructions, with existing high-level operations retained as compatibility macros.
  3. Program-producing ExactNetBegin with supervised typed program learning.
  4. Static verificationValidate structure, types, semantic metadata, and dependencies.
  5. Selective executionCalibrate execute, defer, and clarification decisions.
  6. Behavioral datasetsKnown semantics, contracts, counterfactuals, and intervention rules.
  7. RLBC experimentsAblate behavioral, efficiency, and abstention rewards.
  8. Dynamic function compositionUse unseen function descriptors without vocabulary expansion.
  9. Verified macro learningPromote repeated, verified program structures into reusable operations.
Where this differs from tool calling →