Skip to content

The OxiDex AI Harness

OxiDex is a Rust reimplementation of ExifTool. The AI harness — referred to internally as "the fleet" — is an autonomous system that runs continuously against the repository to find metadata tags that real ExifTool reports but OxiDex does not, and to write the parser code that closes those gaps.

This document describes what the harness is, how it works end to end, and what it has actually produced. Every number below is followed by the file it was counted from and the window it covers, so it can be reproduced. Where a figure cannot be derived from available records, it is labelled not measured rather than estimated.

Reach for transcription first

The harness is not the default route to coverage. Most of the remaining gap is data that already exists in machine-readable form in ExifTool's own tables, and generating from it is deterministic, verifiable, and roughly five orders of magnitude cheaper — 27,747 tag entries extracted in 1.3 seconds, against this harness's ~112 model calls per delivered tag.

See Transcription. Before starting harness work on a gap, confirm the knowledge is not already written down somewhere; if it is, that is transcription work and belongs there instead.

A note on numbers

This project has a documented history of inflated claims. The README once advertised 32,677 tags and "complete parity" when the real figures were 16,677 and ~58%; that had to be corrected in PR #199. The measurement traps that produced those errors are described in Counting rules, and every headline figure here is derived under them.


What it is

The harness is a dispatcher process that runs a pool of parallel worker processes. Each worker takes one format (JPEG, NEF, CR2, …), asks a large language model to write a patch that makes OxiDex extract a tag it currently misses, and then tries very hard to disprove that the patch worked. Only a patch that survives every check gets committed.

The two programs are:

ComponentFileSize
Dispatcherscripts/parallel_model_fix_loop.py2,960 lines
Workerscripts/model_fix_loop.py~8,200 lines

The dispatcher holds a singleton flock at ~/.oxidex/logs/dispatcher.lock, allocates worker slots across formats in proportion to each format's open-gap count, and spawns each worker in its own process group and its own git worktree. Per-tag assignment does not happen in the dispatcher — each worker claims individual tags under a separate lock, so two workers never work the same tag.

What it is for

Finding and filling missing tags. The ground truth is real ExifTool: a comparison harness runs both binaries over the same sample file and diffs the tag sets. A tag ExifTool emits that OxiDex does not is a gap. A tag both emit with different values is a value difference. Both are work items.

The harness exists because this work is large, repetitive, and individually small — most fixes are a handful of lines wiring one tag id to one field in one table — but each one still requires reading ExifTool's Perl source to get the tag id, type, and conversion right.


How it works, end to end

One attempt on one tag passes through the following stages. The order matters, and it is not the obvious one: the cheap checks run first, the two expensive ones (the reviewer model call and the full workspace test suite) run last, and the commit is the very last step of all.

Line references are to scripts/model_fix_loop.py.

Before any model call is spent

StageWhereNotes
Claim the tag under flockrun_tag_loop.claim, 7116–7210Also filters blacklisted, already-landed, and cross-tier tags
Reachability check7248–7278, default_format_reachable 4863–4935Runs OxiDex on the sample; if the parser produces no format-specific output at all, the tag is skipped with zero model calls
Snapshot the ExifTool baseline8054The "before" side of the later gap comparison

The reachability check exists because of a real failure class: a parser that was written but is never actually invoked. Every model call spent on such a format is wasted, because no patch to that parser can change the output.

Per repair round

The round budget is not fixed. It starts at max_repair_rounds (default 5) and each substantive reviewer rejection extends it by one, up to max_review_rounds (default 5) extra — so a candidate the reviewer keeps arguing with gets up to 10 rounds of back-and-forth, while one that merely fails to compile still gets 5. A reviewer that could not be reached (review call failed: …) or whose reply carried no verdict (unparseable review verdict: …) has not judged anything: those rounds extend nothing and leave no trace in the rejection history (REVIEW_INFRA_PREFIXES), so a reviewer-side outage cannot double the fleet's per-candidate spend. See "Why the reviewer gets a bigger budget" below.

Each round:

#GateWhereRejects with
1Model call, extract a unified diffattempt_build 4447–4628no diff in model response
2git apply tolerance ladder4631, git_apply_with_rung 1103resend; on exhaustion 4656
3cargo build4643resend; on exhaustion 4656
4Compare against ExifTool (fresh comparison run)recheck_fn, 5439produces data, not a verdict
5Structural gate — no new OxiDex-only tags introduced5459–5469introduced new oxidex-only tag(s): …
6Gap-count gate5471–5526gap count did not decrease
7Targeted tests (cargo test --lib for the format)5528–5537targeted tests (…) regressed
8Duplicate-insertion check5539–5545, detect_duplicate_tag_insertion 1238duplicate: a handler for … already exists elsewhere
9Evidence gathering (live re-extraction + emission scan)5550–5557never rejects — wrapped in try/except, degrades to empty
10Reviewer model call (APPROVE / REJECT / UNVERIFIABLE)5559–5562rejected by review: …and a substantive rejection buys one extra round (an unreachable reviewer or unparseable verdict buys none)
11Full cargo test --workspace5564–5578cargo test --workspace regressed
12Build evidence trailers_build_fix_gap_trailers 5130
13git commit5610–5614

Two ordering decisions are deliberate and worth naming:

  • Gate 8 precedes gate 10. The duplicate check is cheap and local; the reviewer call costs a model request. The code comment at 5259–5269 states the intent directly: run the duplicate check "BEFORE spending a reviewer call on it."
  • Gate 11 follows gate 10. The full workspace suite is the single most expensive check, so it runs only on a patch that has already convinced the reviewer.

Gate 9 is intentionally not a gate. Evidence gathering feeds the reviewer better context, but a failure to collect it must not sink an otherwise good patch, so both collectors degrade to an empty string.

Why the reviewer gets a bigger budget

max_review_rounds is a separate budget rather than a larger max_repair_rounds, because the two failure populations are not alike:

  • A build failure, gap-count miss or test regression is the model failing against a machine that already told it exactly what was wrong. Five tries at that is generous; a sixth is usually the same wrong idea again, and raising the shared budget would double the cost of every doomed target in the fleet.
  • A reviewer rejection is different in kind. The patch compiled, the gap count moved, and the targeted tests passed — every mechanical gate agreed. What is left is a judgment call about genuineness (hardcoded sample values, double emission, invented fixtures — the REVIEW_CHECKLIST C1–C5 items), and that is exactly the argument worth having more than once: the fixer is close, and the objection is specific enough to act on.

So the extra rounds are spent where the conversation is productive rather than spread evenly over failures that are not.

Each retry hands back every rejection so far, not just the newest, explicitly marked as still binding. Without that, a fixer shown only the latest objection satisfies it by reintroducing whatever it was rejected for three rounds ago, and then oscillates between the two until the budget is gone — a failure mode that a longer loop makes worse, not better. The retry message also states that the working tree was reverted (so the next diff must apply to the original files, not to the rejected patch) and reminds the fixer that re-reading files is free. Reviewer infrastructure failures are excluded from that history: relaying review call failed: [Errno 54] … as a "still binding" objection would demand the fixer satisfy a connection error with a diff, so those rounds get a plain "the reviewer could not be reached, resend" handoff instead.

Reading files is free

The fixer may send as many REQUEST: <path> turns as it likes; productive ones cost it nothing and it is told so, in words, on every answer. This matters in both directions — a model that cannot see a budget cannot ration it, but a model that can see one will ration against it, so staying silent about the change would leave the fixer under-investigating against a cap that no longer exists.

What is still bounded is unproductive investigation:

KnobDefaultCounts
max_request_turns20REQUESTs that bought nothing: an unresolvable path, a range starting past the end, or a re-ask of content still visible in the conversation
max_request_repeats3Identical REQUESTs before the answer is replaced by a pivot nudge (unless compaction elided the content — then it is simply re-served, free)
max_request_turns_ceiling250Total REQUESTs per attempt, free ones included — a runaway backstop, surfaced to the model only at the moment it fires

max_request_turns kept its name so existing config.toml files keep working; what changed is what it counts. The split is decided in two steps: request_answer_served, a whitelist of the three answer shapes resolve_request uses for real content (an unrecognized shape is charged rather than waved through, which is the side that stays bounded), and then a visibility check on the payload itself — a served answer is charged as a re-ask when those exact bytes are already in the conversation, whatever the request string looked like. That closes the string-variant hole (clamped ranges, ./-prefixed paths and re-ranged samples all reach byte-identical payloads through distinct strings) and it makes compaction self-consistent: the elision stub says "Re-REQUEST it if still needed", and once the payload is stubbed out it is no longer visible, so obeying that instruction is free. Distinct line ranges of one file stay free when they show new lines; a range that resolves to exactly what was already shown does not.

The ceiling exists because "free" is not "infinite": a model walking a generated directory file by file would otherwise spend an unbounded number of paid calls, and no amount of raising max_request_turns would stop it, since those reads are not charged against it at all. Hitting it produces its own reason string (… (hit the 250-REQUEST safety ceiling)) so it is distinguishable in the lessons ledger from an exhausted wasted-request budget.

The git apply tolerance ladder

Model-authored diffs frequently have slightly wrong line numbers or whitespace. Rather than discard them, gate 2 retries progressively looser, stopping at the first rung that applies:

text
1. exact                       git apply --recount -
2. ignore-whitespace           git apply --recount --ignore-whitespace -
3. context1                    git apply --recount -C1 -
4. context1-ignore-whitespace  git apply --recount -C1 --ignore-whitespace -
5. 3way                        git apply --3way --recount -

rung=None in the log means every rung was tried and none applied. The recorded error is the strict rung's stderr, because that one names the context git actually searched for.

Rung 5 is special-cased: --3way implies --index, so _restore_after_three_way (1062–1100) un-stages afterwards on both success and failure. Without it, a staged change would survive the next round's git checkout -- . and silently leak into an unrelated attempt.

Measured, all diffs carrying a rung field (5,247 of 7,886 rows in ~/.oxidex/logs/model-fix-diffs/manifest.log; the remaining 2,639 predate the field):

RungDiffsShare of applied
exact2,01874.4%
context154720.2%
ignore-whitespace782.9%
context1-ignore-whitespace712.6%
applied, total2,714
None (all rungs failed)2,533

The looser rungs rescue 696 diffs, 25.6% of everything that applied. The ladder earns its keep — a strict-only gate would have discarded a quarter of all working patches.

Model selection

There is no fixed model. pick_model_fn defaults to random.choice over the configured pool and is re-drawn before every individual call (attempt_build 4445), so a single repair conversation can span several models and even several providers. There is no within-call failover: call_model retries the one model it was handed, and only when the whole ladder is spent does the attempt fail — at which point the next draw is a fresh uniform pick. As config.toml puts it, the pool is the retry.

A fleet-wide rate governor sits inside the retry loop, but it is now only a steady-state rpm budget: a token bucket in one flock-guarded file, refilling at governor_calls_per_minute and capped at governor_burst. It shapes how fast the fleet is collectively allowed to go. It does not react to failure.

Superseded (2026-08-02). It used to. A rate-limit response set a global cooldown that paused every worker, growing exponentially per consecutive limited outcome (30s, 60s, 120s, 240s, capped at 300s). That was wrong in both directions:

  • Against an rpm limit it phase-locked the fleet. Every worker was released at the same instant, emitted a synchronised burst, tripped the limit together and parked together — exactly the shape a rate limiter is built to reject.
  • Against a cost-window cap it was futile. A 300s ceiling against a weekly cap meant the fleet woke, was rejected and parked again, every five minutes, for days.

What replaced it, all inside call_model:

  • Classification first. classify_429 reads the gateway's theclawbayError discriminator before deciding anything, yielding rpm, window_cap (5h_cost_limit_reached) or terminal_cap (weekly_cost_limit_reached, invalid_api_key). Anything unparseable is rpm, because wrongly retrying a permanent condition costs one park interval while wrongly giving up on a transient one costs the work.
  • A cost cap is never retried. Retrying cannot make budget appear. The call raises ModelQuotaExhausted immediately and parks that endpoint in that process (endpoint_park), so the worker rides the window out and resumes unattended when it rolls over, without touching any other worker.
  • rpm rejections retry per-worker, with full jitter (delay + uniform(0, delay)) and honour Retry-After when present. The jitter is not rate shaping — it is what stops N workers rejected in the same instant from retrying in the same instant.

No code path pauses more than one worker. max_retries is consequently free: it stays at its configured 1000, riding out a genuine outage, because a long ladder can no longer park the fleet.

Every attempt is also appended to ~/.oxidex/logs/model-calls.jsonl with its role, model, endpoint, error class and latency — a structured superset of the manifest.log line, so the error-class breakdown below can be answered without parsing prose.

What the gateway actually returns

Probed live on 2026-08-02 against api.theclawbay.com. Re-derive rather than trust; the catalogue and the routes both move.

  • GET /v1/models — 42 models. The reasoning-capable OpenAI tier is gpt-5.6, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-5.5, gpt-5.4, gpt-5.4-mini. Cross-vendor models reachable on the same key and the same OpenAI-compatible route include claude-opus-4-8, claude-sonnet-4-6, claude-haiku-4-5, gemini-3.1-pro-preview, glm-5.2, kimi-k2.7-code, deepseek-v4-pro. Reviewer independence therefore costs no integration work at all — it is a model name.

  • There is no /quota endpoint. /quota, /v1/quota and /api/quota all 404. GET /v1/usage exists and returns {aggregation_timestamp, n_context_tokens_total, n_generated_tokens_total} — token counts, not cost windows and not remaining headroom. Any design that assumes it can poll remaining 5-hour or weekly budget before dispatching needs rethinking: today the only signal that a window is spent is the 429 itself, after the fact.

  • The error envelope, captured verbatim:

    json
    {"error": "invalid request", "code": "upstream_rejected",
     "theclawbayError": {"requestId": "…", "category": "internal",
                         "code": "upstream_rejected",
                         "userMessage": "The upstream model provider rejected this request.",
                         "retryable": false, "retryAfterSeconds": null,
                         "nextAction": "Review the request and retry…"}}

    Two things in it contradict a reasonable first guess, and both were bugs in the classifier until this response was actually looked at:

    1. retryAfterSeconds is in the body, not a Retry-After header. A classifier that parses only headers silently ignores the server's own instruction — the one case that beats every backoff curve. call_model now takes the larger of the two when both are present, because undershooting just earns another 429.
    2. retryable: false does not mean "budget spent." The gateway uses it for ordinary upstream failures, as above. Treating every non-retryable 429 as a cost cap would park a perfectly healthy endpoint for an hour over one upstream hiccup. A cap must be quota-shaped; a rejected key is terminal on its own evidence; everything else that says "do not retry" fails the call and leaves the endpoint alone.

Why the commit being last matters

git_commit_fn is called exactly once in fix_gap, at line 5610, after every gate above. The commit is also what stages anything at all:

python
subprocess.run(["git", "add", "-A"], cwd=repo_root, check=True)
argv = ["git", "commit", "-m", message]

git_apply_with_rung (1117–1119) applies to the working tree only — no --index, no --cached. Until line 5610 executes, a candidate fix exists solely as unstaged, uncommitted working-tree modifications inside that worker's worktree.

So an interrupted run loses exactly the work that already passed its gates. Concretely, a killed worker loses:

  • the applied edit itself (unstaged working-tree content);
  • every gate result — built, diff, remaining, post_match, live_evidence, approved, and the entire model conversation — all process-local Python locals, never persisted;
  • the cargo build and the full cargo test --workspace runs, which are the two most expensive things the harness does.

Three mechanisms then destroy the working tree: the dispatcher's shutdown handler SIGKILLs every worker process group (868–880), so there is no chance to commit; the next round's clean_worktree runs git checkout -- . and git clean -fd; and reap_orphan_worker_pgids kills any survivor at the next dispatcher startup.

The only durable residue is the raw diff text. logging_git_apply (7811–7831) writes every diff — applied or rejected — to ~/.oxidex/logs/model-fix-diffs/ before returning, and its own comment calls this "the only durable record of what was tried." Re-landing such a diff means re-running apply, build, comparison, both test suites, and the reviewer call from scratch.

There is no checkpointing of a passed-gates-but-uncommitted candidate anywhere. The tag-state file is only written after fix_gap returns, so a killed worker records nothing; its claim simply goes stale after two hours and becomes re-claimable. The claim heartbeat thread is a daemon thread specifically so that a dying worker cannot keep its own claim alive.

This is the harness's sharpest design cost. The cheapest possible improvement is not a faster model — it is a checkpoint between gate 11 and gate 13.


Results

Counting rules, and the traps they avoid

Four distinct numbers in this repository look like "tags closed" and only one of them is.

1. Tag: trailers are tags assigned, not tags delivered. Each worker is handed a cluster of up to max_cluster_tags = 6 tags and emits a Tag: trailer for each. It commits if it closes any of them. Commit 2efe807e (PR #185) carries 12 Tag: trailers and delivered 2 tags. Counting trailers overstates by up to 6x on individual commits and by 1.8x overall.

2. wire N missing tags in the commit subject is the delivered count. It is computed at line 5579 as closed = gap["gap_count"] - remaining, where remaining comes from re-running the full ExifTool comparison after the patch. It is a measurement, not a claim.

3. Verified: recheck-pass gaps=A->B is the same measurement, recorded independently. Across all 23 measured fix blocks, A - B equals wire N exactly, with zero mismatches — the two records corroborate each other.

4. N tag fix(es) in a sweep commit subject counts worker commits, not tags. Sweep #180's subject says "3 tag fix(es)" and it delivered 14 tags.

There is also a double-counting trap. Summing wire N naively across all commit bodies gives 228 — and that figure is wrong. GitHub squash-merges write one * <subject> bullet per constituent commit, and stacked PRs re-list their parent's commits. Commit b74ec52c (PR #41) lists 11 wire 1 missing tags bullets while changing zero files under src/ — proof those bullets belong to a parent branch, not to that commit. After de-duplicating identical fix blocks, 132 raw blocks collapse to 40.

Gaps closed

67 tag gaps closed with measured evidence, across 23 fix blocks, 13 formats, over 2026-07-25 → 2026-07-30.

Count
Tags assigned to those fixes (Tag: trailers)119
Distinct tag keys named103
Tags delivered (wire N, corroborated by gaps=A->B)67
Delivery rate56%

Formats touched: CR2, ELF, HEIC, JPEG, NEF, PDF, PE, PSD, RAR, RW2, TTF, X3F, XMP.

Method: parsed every fix(<fmt>): wire N missing tags block out of git log origin/main commit bodies, keeping only blocks carrying a Verified: recheck-pass gaps=A->B trailer, then de-duplicated on (format, wire count, pool, tag list, closed count, worker) to remove stacked-PR re-listing.

Not measured: the pre-2026-07-25 era. Fix commits before that date predate the evidence trailers, so there is no per-commit measurement to read, and the stacked-PR re-listing above makes their bullets impossible to de-duplicate reliably. Aggressive de-duplication yields 36 tags; the raw bullet count yields 93. The true figure is somewhere in that range and cannot be narrowed from the available records. It is excluded from the 67 above rather than estimated.

Current coverage

Three coverage figures are in circulation and they measure different things. All three are correct; none of them substitutes for another.

FigureValueWhat it measures
Repo-wide definitions count16,684 tag definitions (no parity ratio — see AGENTS.md, "Closing an ExifTool coverage gap")Tags known to the OxiDex tag database, across 140+ formats
Live comparison snapshot738 of 1,511 tags = 48.8%13 formats, one sample file each, /tmp/fin_*.json generated 2026-07-31T01:19Z
Full-corpus audit (JPEG)60.8% of tag instances per file, median file 84.3%4,085 JPEGs, PR #203

Per-format, from the 2026-07-31 snapshot:

FormatMatchedMissingValue diffsExifTool totalCoverage
MachO6006100.0%
ISO13101492.9%
TTF22402684.6%
NEF141382520469.1%
X3F291304269.0%
PSD622819168.1%
CR2102781119153.4%
PDF220450.0%
XMP156103148.4%
RW26885015344.4%
DNG1071411726540.4%
JPEG1332241337035.9%
MRW3872411433.3%
Total738692811,51148.8%

Coverage percentages are key-set statements, not per-file statements

The comparison harness collapses every tag to one entry per family:name across the whole corpus, so presence is unioned and values are compared across different files. Adding sample files can lower a percentage with no code change at all. Measured cost on CR2 (PR #203): 118 tags match when two samples are measured separately, 94 when measured together.

JPEG is the clearest illustration — the same parser scores 21.9% (corpus-collapsed), 35.9% (this single sample), 60.8% (real per-file mean), or 84.3% (median file) depending only on how you count. Always read the absolute counts beside the percentage.

Models tried

Eight distinct model identifiers appear in ~/.oxidex/logs/model-fix-requests/manifest.log over 2026-07-22 → 2026-07-30. Note that DeepSeek-V4-Pro and deepseek/deepseek-v4-pro are the same model reached through two different providers, which use different id casing.

Model idCallsNotes
gpt-5.6-sol18,607Current sole pool member
deepseek/deepseek-v4-pro19,555OpenRouter id form
gpt-5.6-terra16,271Prior incumbent
Kimi-K2.65,213
gpt-5.52,792
DeepSeek-V4-Pro2,675Same model, different provider/casing
GLM-5.22,169
openrouter/free13Meta-endpoint, not a model — routes at random per request across ~15 free backends

openrouter/free deserves the caveat. It is OpenRouter's Free Models Router: each request lands on a randomly chosen backend from roughly fifteen, filtered to those supporting the features the request uses. You do not get to choose, and you get a different model every call. It was adopted because per-request availability was excellent (30 concurrent requests in 12.1s, zero 429s) but it was barely exercised — 13 calls — before the pool moved on.

Per-model patch apply rate

How often a model's diff applied cleanly to the working tree. This is the purest available measure of patch quality: it only counts calls that actually produced a diff, so it is not contaminated by provider outages.

ModelDiffs producedAppliedRejectedApply rate
gpt-5.6-sol64951913080.0%
gpt-5.6-terra71541776.1%
gpt-5.551833218664.1%
Kimi-K2.659832727154.7%
deepseek/deepseek-v4-pro4,8802,3082,57247.3%
DeepSeek-V4-Pro66830036844.9%
GLM-5.2134607444.8%
openrouter/free41325.0%
All7,5223,9013,62151.9%

Method: ~/.oxidex/logs/model-fix-diffs/manifest.log records applied=True|False per diff but carries no model= field, so each diff row was joined to the nearest preceding successful phase=fixer call for the same worker in ~/.oxidex/logs/model-fix-requests/manifest.log, within a 30-minute window. 7,522 of 7,886 diff rows resolved; 358 predate the worker= field and 6 had no matching call.

The small-sample rows are weak evidence. gpt-5.6-terra (71 diffs) and openrouter/free (4 diffs) should not be ranked against deepseek/deepseek-v4-pro (4,880 diffs) as though the numbers carry equal weight.

Per-model tag attribution

The (via …) string is the pool roster, not the producing model

Line 5611–5612 builds the commit subject as:

python
f"fix({fmt.lower()}): wire {closed} missing tags "
f"(via {'/'.join(m['name'] for m in config['models'])})",

That joins every member of the configured pool, not the model that produced the diff. Since pick_model_fn re-draws per call, a fix committed under a six-member pool could have come from any of the six — and _build_fix_gap_trailers takes no model argument, so there is no per-commit record of the actual producer either. Attribution is therefore unambiguous only for single-member pools.

Both columns count as valid attribution: a tag the model produced directly, and a tag a human or agent then corrected before landing, are both attributable to the model that produced the original patch. They are shown separately so the distinction stays visible.

Pool roster in commitAttributionTags directTags hand-modifiedTotalAssigned
deepseek/deepseek-v4-prounambiguous30255588
Kimi-K2.6 / DeepSeek-V4-Proambiguous, 2 models80822
gpt-5.5 / GLM-5.2 / Kimi-K2.6 / DeepSeek-V4-Proambiguous, 4 distinct models3033
gpt-5.6-solunambiguous1016
Total422567119

"Hand-modified" is the set of tags landed through a commit that explicitly records human or agent correction of the model's patch:

  • 5b5f2dfb — "land 4 hand-verified fixes from the patch archive (13 real gaps closed)"
  • 4e3b9515 — "correct fabricated enum constants that passed the recheck"
  • 86518989 — validator-flagged APP12 tags, re-verified by hand
  • 872b8d89 — validator flag was a false positive, corrected by hand

That last pair matters: the harness's own validator produced both false positives (872b8d89) and false negatives — 4e3b9515 caught fabricated enum constants that had passed the automated recheck. The gate stack is strong but not sound, which is why the hand-modified column is not zero.

Not measured: per-model attribution for the 11 tags landed under multi-model pools. The commit record does not identify which pool member produced each diff, and the request manifest cannot be joined to a commit (it records calls, not commits). This would require adding the producing model to _build_fix_gap_trailers going forward.

Failure taxonomy

Failures split into two layers that must not be added together: transport failures, where the API call itself did not return, and semantic failures, where a reply arrived but the proposed fix did not survive the gates.

Transport layer

From ~/.oxidex/logs/model-fix-requests/manifest.log, 2026-07-22 → 2026-07-30. This file is the authoritative record of call outcomes: each line ends in OK, contains ERROR=, or is a RETRY line. RETRY lines are counted separately — they are neither a success nor a terminal failure, and folding them into either skews the rate.

Totals: 67,296 terminal outcomes — 34,643 OK, 32,653 ERROR — plus 9,613 RETRY lines.

FailureCount
429 rate limit27,662
403 Forbidden4,504
Model returned an empty reply144
DNS resolution failure130
Read operation timed out77
Connection reset by peer57
Malformed / non-JSON response body24
Missing choices in response16
503 Service Unavailable15
Remote closed connection without response9
502 Bad Gateway4
deadline_seconds exceeded while streaming3
409 Conflict2
Other socket errors5

* This number cannot be broken down further. These 27,662 are an undifferentiated pile: the harness that produced them did not read the theclawbayError discriminator, so nothing recorded whether a given 429 was an rpm rejection or a spent cost budget — and those demand opposite responses. That is the baseline the 2026-08-02 model-layer change is measured against, and model-calls.jsonl is what makes the successor number decomposable.

Do not read all-time per-model failure rates as model quality

They are artifacts of two provider outage days, not model behaviour. Per day:

ModelDateCallsFailure rate
gpt-5.6-sol2026-07-2315,27190.6%
gpt-5.6-terra2026-07-2314,90293.3%
deepseek/deepseek-v4-pro2026-07-288,41453.9%
gpt-5.6-sol2026-07-302,8770.2%
deepseek/deepseek-v4-pro2026-07-279,7160.5%
GLM-5.2all2,1690.1%
Kimi-K2.6all5,2130.3%
gpt-5.5all2,7920.5%

On 2026-07-23 both gpt-5.6 variants failed at 90–93% simultaneously — that is a provider-wide outage, and it alone accounts for essentially all 27,662 rate-limit errors. The 2026-07-28 deepseek figure is the 4,504 403s, also provider-side. The same gpt-5.6-sol that looks like a 75% failure model all-time ran 2,877 calls at 0.2% on 2026-07-30.

A flat all-time table would rank models by which days they happened to be on shift. Report per day, or separate transport failures from patch quality, or do not rank at all.

Semantic layer

From ~/.oxidex/logs/lessons.jsonl, 14,546 records. Each row is one recorded outcome with an event classification:

EventCountMeaning
infra4,814Model call failed — the transport failures above, charged to no tag
critique4,417A failed attempt was analysed to inform the next round
build_failed3,196Patch applied but cargo build failed
gap_not_closed826Built and ran, but the gap count did not decrease
structural386Structural lesson recorded about a format
review_rejected360Reviewer model rejected the patch
wrong_value226Tag emitted, but the value disagrees with ExifTool
fixed171Gate stack passed, commit created
test_regressed139An existing test broke
duplicate11A handler for that tag already existed elsewhere

By reason string:

ReasonCount
No diff in model response2,172
Gap count did not decrease833
No working fix after repair attempt (apply/build exhaustion)800
Patch did not apply (rejected diffs, from the diffs manifest)3,687

"No diff in model response" at 2,172 is the largest single controllable failure — a model replying in prose, or with a mangled fence, when a unified diff was required. It has four sub-forms in the source: a plain failure to emit one, and exhaustion of the request budget, the verify budget, or the patch-chunking safety limit.

Note that fixed = 171 here but only 67 tags landed on main with evidence. The two count different things: lessons.jsonl records every worker-local commit including those on branches that were never swept into main, superseded, or discarded. It is not a substitute for the git record.


Reproducing these numbers

All figures come from four sources. The two manifests are authoritative for call and patch outcomes; the git log is authoritative for what actually landed.

text
~/.oxidex/logs/model-fix-requests/manifest.log   call outcomes (OK / ERROR= / RETRY)
~/.oxidex/logs/model-fix-diffs/manifest.log      patch apply outcomes (applied=True|False)
~/.oxidex/logs/lessons.jsonl                     semantic outcome per attempt
git log origin/main                              what landed, with evidence trailers
/tmp/fin_*.json                                  per-format coverage snapshots

Manifest line shapes:

text
<ts> phase=fixer worker=<FMT> tier=T1 model=<model> prompt_chars=N elapsed=Xs reply_chars=N OK
<ts> phase=fixer worker=<FMT> tier=T1 model=<model> prompt_chars=N elapsed=Xs ERROR=<reason>
<ts> phase=fixer worker=<FMT> tier=T1 model=<model> RETRY <message>

<ts> worker=<FMT> applied=True|False rung=<rung> file=<name> apply_msg='<msg>'

Counting call outcomes correctly — note the three-way split:

bash
awk '$0 !~ /RETRY/ { total++
       if ($0 ~ / OK$/) ok++
       else if ($0 ~ /ERROR=/) err++ }
     END { printf "calls=%d OK=%d ERROR=%d OK%%=%.1f\n", total, ok, err, 100*ok/total }' \
  ~/.oxidex/logs/model-fix-requests/manifest.log

Patch apply rate:

bash
grep -oE 'applied=(True|False)' ~/.oxidex/logs/model-fix-diffs/manifest.log | sort | uniq -c

Known-bad measurement approaches

Do not use any of these. Each has produced a wrong answer in this project before:

  • pgrep -fc on worker processes, or pairing request files against response files. Both give wrong failure rates. The manifest is the only reliable source.
  • failrate.py in the session scratchpad. PR #204 changed the log format and its parser stopped matching OK lines; it now reports 100% failure in every window.
  • Commit trailer counts. They report tags assigned, not tags delivered.
  • Naively summing wire N across all commit bodies. Squash-merge bullets re-list parent-branch commits; this yields 228 instead of the de-duplicated figure.
  • A coverage percentage without its absolute counts. The denominator moves when samples are added.

Summary of what is not measured

Stated explicitly, because an honest gap is more useful than a plausible invention:

FigureWhy it cannot be derived
Tags closed before 2026-07-25Predates evidence trailers; stacked-PR bullets cannot be de-duplicated. Bounded to 36–93, not narrowable
Per-model attribution for 11 tagsLanded under multi-model pools; the commit records the pool roster, not the producer
Which of the ~15 openrouter/free backends served any given callThe router does not report the chosen backend, and it re-routes per request
Cost per landed tagToken usage is logged to cache-stats.log only when the provider reports it; several providers do not
Wall-clock time per landed tagPer-call elapsed is recorded, but build and test time is not attributed to a tag
Whether the 1,000/day free-tier cap binds in practiceNot measurable without spending the budget

See also

  • CORPUS_AUDIT.md — the full-corpus coverage audit that established the per-file JPEG figures used above.

Released under the GPL-3.0 License.