A
Task correctness
Did the agent reach the correct terminal state, ignoring safety?
Can recovery engineering make a smaller, cheaper model as safely reliable as a frontier model—and at what workflow depth does that substitution stop working?
3
model capability tiers
4
recovery configurations
3
primary hypotheses
2
equally valid outcomes
The question in one line
The experiment separates what comes from the model from what comes from the surrounding recovery system, while charging every retry, verifier, and fallback token to the final cost.
Candidate system
Smaller model
+ verify · retry · fallback
≈?
Reference system
Frontier model
+ no recovery scaffolding
Equal or better safe reliability · materially lower total-token cost
Why ordinary success is not enough
A payment succeeds, its confirmation is lost, and the agent retries. A success-only benchmark records a win even if the invoice was paid twice. This study treats correctness and transactional safety as separate primitive outcomes.
A = 1 · U = 1
The requested task finished correctly, but at least one unsafe duplicate operation occurred. The run is not counted as safely reliable.
A
Did the agent reach the correct terminal state, ignoring safety?
U
Did at least one duplicated state-changing operation occur?
D
How many duplicate non-idempotent operations occurred?
R
How often was the task completed correctly without unsafe duplication?
C
Every token used by planning, retry, verification, fallback, and recovery.
Δdep
Whether correctness and transactional safety factor independently in a condition.
Experimental lattice
FAULTLINE will construct enterprise-like transactional workflows and vary six factors under controlled seeds. The full crossing is intentionally restricted to preserve statistical power where the decision boundary matters.
Model capability
01Indexed on raw task capability measured inside the study—not an external benchmark.
Recovery engineering
02Four increasingly capable scaffolding configurations isolate what the system adds.
Retry budget
03Tests where retries stop helping safe completion and begin increasing duplicate risk.
Fault severity
04Controlled tool errors, latency, and payload corruption with rates frozen before study.
Dependency depth
05Workflow depth is manipulated by construction so the failure boundary can be measured.
Idempotency
06A restricted, pilot-selected arm measures how transaction protection bends the frontier.
Confirmatory family
H1
Measure whether recovery engineering lets a lower model tier meet the same correctness and safety targets, then identify the workflow depth at which that advantage degrades.
H2
Test whether a roughly 32B model with scaffolding reaches the target with at least a 5× total-token reduction versus a frontier model without scaffolding at depth two.
H3′
Compare the retry budget that maximizes measured safe reliability with the uncertainty-aware optimum predicted by an independence reference model. Mechanism is not inferred from direction alone.
Freeze discipline
Pilot data tune task difficulty, establish model capability indices, estimate hazards, check target attainability, power the experiment, and select the restricted idempotency cells. They are then excluded from every confirmatory estimate.
Calibrate tasks, hazards, targets, and dynamic range.
Simulate the planned estimator and concentrate runs near the decision boundary.
Freeze models, prompts, cells, seeds, estimators, tests, and deviations policy.
Execute the frozen experiment without redesign after outcomes are visible.
Publish the harness, task generator, traces, analysis, version history, and log.
Pre-pilot review queue
The conceptual architecture is stable. These are bounded statistical and wording repairs—not an invitation to redesign the experiment.
Freeze the pilot estimation procedure and carry uncertainty in success and duplication hazards into a distribution or set of predicted retry optima. One noisy pilot argmax will not be treated as an exact null.
The retry-success and duplication-hazard tests will form a separate secondary confirmatory family with its own preregistered Holm correction.
A non-zero dependence gap means correctness and safety do not factor independently under that condition. Causal mechanism requires trace evidence and direct hazard instrumentation.
A reviewed v0.2 experimental specification with the primary outcomes, factor lattice, hypotheses, kill criteria, calibration rules, and freeze sequence defined.
A Pilot Protocol fixing task generation, model versions, decoding, tool schemas, injection schedules, randomization, traces, quarantine, and go/no-go outputs.
A reusable transactional-semantics agent-safety testbed with ERP-like writes, toggleable idempotency, controlled faults, and complete action accounting.
Research in public