Skip to content
LogoSamir Sawarkar
Current researchv0.2 · Pre-pilot · Not frozen

Capability–Reliability Substitution in Enterprise Agent Systems

Can recovery engineering make a smaller, cheaper model as safely reliable as a frontier model—and at what workflow depth does that substitution stop working?

3

model capability tiers

4

recovery configurations

3

primary hypotheses

2

equally valid outcomes

The question in one line

Buy a bigger brain—or engineer a safer system?

The experiment separates what comes from the model from what comes from the surrounding recovery system, while charging every retry, verifier, and fallback token to the final cost.

Candidate system

Smaller model

+ verify · retry · fallback

≈?

Reference system

Frontier model

+ no recovery scaffolding

Equal or better safe reliability · materially lower total-token cost

Why ordinary success is not enough

An agent can complete the task and still do damage.

A payment succeeds, its confirmation is lost, and the agent retries. A success-only benchmark records a win even if the invoice was paid twice. This study treats correctness and transactional safety as separate primitive outcomes.

A = 1 · U = 1

The requested task finished correctly, but at least one unsafe duplicate operation occurred. The run is not counted as safely reliable.

A

Task correctness

Did the agent reach the correct terminal state, ignoring safety?

U

Unsafe duplication

Did at least one duplicated state-changing operation occur?

D

Duplicate severity

How many duplicate non-idempotent operations occurred?

R

Safe reliability

How often was the task completed correctly without unsafe duplication?

C

Complete token cost

Every token used by planning, retry, verification, fallback, and recovery.

Δdep

Outcome dependence

Whether correctness and transactional safety factor independently in a condition.

Experimental lattice

Manipulate the system; do not merely observe it.

FAULTLINE will construct enterprise-like transactional workflows and vary six factors under controlled seeds. The full crossing is intentionally restricted to preserve statistical power where the decision boundary matters.

Model capability

01

Small · Mid · Frontier

Indexed on raw task capability measured inside the study—not an external benchmark.

Recovery engineering

02

None · Retry · Verify + retry · Verify + retry + fallback

Four increasingly capable scaffolding configurations isolate what the system adds.

Retry budget

03

0 · 1 · 2 · 4 · 8

Tests where retries stop helping safe completion and begin increasing duplicate risk.

Fault severity

04

None · Low · High

Controlled tool errors, latency, and payload corruption with rates frozen before study.

Dependency depth

05

2 · 4 · 8 · conditionally 6

Workflow depth is manipulated by construction so the failure boundary can be measured.

Idempotency

06

Off · On

A restricted, pilot-selected arm measures how transaction protection bends the frontier.

Confirmatory family

Three predictions, committed before the locked run.

H1

Substitution exists—and weakens with depth.

Measure whether recovery engineering lets a lower model tier meet the same correctness and safety targets, then identify the workflow depth at which that advantage degrades.

H2

The mid tier can reach the target frontier.

Test whether a roughly 32B model with scaffolding reaches the target with at least a 5× total-token reduction versus a frontier model without scaffolding at depth two.

H3′

Observed retry optimum departs from the null model.

Compare the retry budget that maximizes measured safe reliability with the uncertainty-aware optimum predicted by an independence reference model. Mechanism is not inferred from direction alone.

Primary family-wise error is fixed at α = 0.05 and allocated across H1, H2, and H3′. Capability displacement is reported ordinally across the three tiers—never converted into a falsely precise continuous number.

Freeze discipline

The pilot calibrates the ruler. It does not answer the paper.

Pilot data tune task difficulty, establish model capability indices, estimate hazards, check target attainability, power the experiment, and select the restricted idempotency cells. They are then excluded from every confirmatory estimate.

  1. 01

    Quarantined pilot

    Calibrate tasks, hazards, targets, and dynamic range.

  2. 02

    Power + allocation

    Simulate the planned estimator and concentrate runs near the decision boundary.

  3. 03

    Preregister v1.0

    Freeze models, prompts, cells, seeds, estimators, tests, and deviations policy.

  4. 04

    Locked confirmatory run

    Execute the frozen experiment without redesign after outcomes are visible.

  5. 05

    Release

    Publish the harness, task generator, traces, analysis, version history, and log.

Pre-pilot review queue

Three corrections remain before the first pilot run.

The conceptual architecture is stable. These are bounded statistical and wording repairs—not an invitation to redesign the experiment.

Required before pilot
01

Propagate uncertainty into H3′

Freeze the pilot estimation procedure and carry uncertainty in success and duplication hazards into a distribution or set of predicted retry optima. One noisy pilot argmax will not be treated as an exact null.

02

Give secondary tests their own error control

The retry-success and duplication-hazard tests will form a separate secondary confirmatory family with its own preregistered Holm correction.

03

Keep dependence non-causal

A non-zero dependence gap means correctness and safety do not factor independently under that condition. Causal mechanism requires trace evidence and direct hazard instrumentation.

What exists now

A reviewed v0.2 experimental specification with the primary outcomes, factor lattice, hypotheses, kill criteria, calibration rules, and freeze sequence defined.

What comes next

A Pilot Protocol fixing task generation, model versions, decoding, tool schemas, injection schedules, randomization, traces, quarantine, and go/no-go outputs.

What the instrument contributes

A reusable transactional-semantics agent-safety testbed with ERP-like writes, toggleable idempotency, controlled faults, and complete action accounting.

Research in public

The goal is not an impressive claim. It is a result that survives skeptical review.