Working paper · KoldOps R/01 · Direction: Context Infrastructure

Lost in the Upgrade

A Case Study of Correction-Rate Regression on Frontier-Model Rollout in a Production Industrial RAG Deployment

Kol Dorney

KoldOps LLC, Industrial AI Research Lab

kol@koldops.com · Knoxville, Tennessee

Draft v0.2 · Not for citation

Abstract

We report a case study of a production retrieval-augmented generation (RAG) deployment operating inside an industrial fluids-manufacturing operator's sales workflow. Over the period 4 June to 10 September 2026 the deployment served 848 user queries across 376 completed retrieval events. During this window the underlying large language model was upgraded from Claude Opus 4.7 to Claude Opus 5. The user-flagged correction rate rose from 3.0% (6 of 202 events, Opus 4.7) to 16.1% (27 of 168 events, Opus 5), a 5.4x increase (Fisher's exact test, p < 0.0001; odds ratio 6.25, 95% CI approximately 2.5 to 15).

Three deployment-level variables shifted simultaneously with the model change: average context window grew from 19,892 to 56,030 input tokens (2.8x); average tool calls per response grew from 1.89 to 2.89 (1.5x); average end-to-end latency grew from 25.4 to 44.2 seconds (1.7x). We enumerate the three most likely confounds (context-length effects, tool-use effects, and query-distribution drift) and argue that untangling them requires the confounds themselves as a research subject.

The v0.2 revision extends the paper with a negative result on offline-replay attribution. We designed four experimental variants (model swap, context stress, tools stress, tools prune) and scored them with two independent instruments: a position-corrected LLM-as-judge and a deterministic entity-recall metric that extracts atomic facts from the original responses and matches them literally against the replay text. Neither instrument distinguished the confounds. Zero of 60 dollar-figure facts were reproduced by any replay across all variants. Diagnosis: the replay pipeline lacked the deployment's production retrieval and system prompt, and every replay therefore collapsed to a common factually-thin floor. We identify a production-mode shadow-replay canary as the working design, and recommend that industrial RAG operators log the correction signal, token counts, tool counts, and latency at every rollout so that regressions of the size we report remain measurable.

1Introduction

Retrieval-augmented generation is now the default deployment shape for enterprise question-answering over large document corpora (Lewis et al., 2020; Gao et al., 2023). As frontier model providers ship successor models on a rapid cadence, operators of production RAG systems face a recurring question: does the newer model make the deployed system better, worse, or the same? Provider-published benchmarks show the successor is better on standard NLP tasks. Operators care about the deployed task. The two do not always agree, and the industrial-RAG literature does not yet contain many measurements of the disagreement.

This paper contributes one such measurement. We observe a production RAG deployment across a model upgrade and report a 5.4x increase in the rate at which end users flag responses as incorrect. We are not the first to observe that frontier-model benchmarks under-predict domain-specific outcomes (Xiong et al., 2024; Pipitone and Alami, 2024; Islam et al., 2023). We are, to our knowledge, the first to report the size of the delta on a controlled deployment-level upgrade in industrial retrieval, with the ground-truth correction signal already instrumented.

The deployment is a sales-assistant RAG loop at an industrial fluids-manufacturing operator. Its users are salespeople answering internal and customer questions using a small corpus of product specifications, delivery records, and pricing documents. Our contribution is threefold. First, we report the empirical regression across the Opus 4.7 to Opus 5 transition. Second, we enumerate three confounds that moved simultaneously with the model change and argue that any measurement paper on this class of question must treat the confounds as the substance, not as noise. Third, we recommend a minimal telemetry logging pattern that any industrial-RAG operator can adopt to make future rollouts measurable in the way ours turned out to be.

Section 2 reviews related work on RAG evaluation and domain adaptation. Section 3 describes the deployment. Section 4 describes the instrumented data. Section 5 reports the empirical regression. Section 6 analyzes the confounds. Sections 7 through 9 discuss limitations, implications, and practitioner recommendations.

2Related work

RAG evaluation. The reference-free evaluation library RAGAS (Es et al., 2024) and the fine-tuned-classifier framework ARES (Saad-Falcon et al., 2024) are the current default automatic-evaluation tools for RAG. The 2024 TREC RAG Track (Upadhyay et al., 2024) established the first shared-task venue for retrieval-augmented answer quality. None of these systems evaluate against a live user-flagged correction signal; they either use LLM-as-judge or expert annotation against a fixed test set. Our measurement uses the deployment's own user-in-the-loop correction flag as the ground truth.

Domain-specific RAG. MedRAG with the MIRAGE benchmark (Xiong et al., 2024) and LegalBench-RAG (Pipitone and Alami, 2024) established the pattern of measuring retrieval quality on a domain-specific expert-annotated benchmark. FinanceBench (Islam et al., 2023) reports that GPT-4-Turbo with retrieval refused or missed 81% of 10-K questions, illustrating the general-benchmark-to-domain gap. To our knowledge no comparable benchmark exists for industrial artifacts (drilling reports, engineering drawings, manufacturing travelers, or sales-fulfillment paperwork).

Model regression under upgrade. Recent industry reports and blog posts have observed that swapping a frontier LLM in a deployed RAG can produce non-monotonic quality changes, but published empirical measurements on production data are rare. The closest peer-reviewed work we are aware of is Li et al. (2025), which revisits long-context vs. retrieval trade-offs across model generations and concludes that RAG dominates above roughly 200K tokens on non-Gemini frontier models. Our case is complementary: same corpus, same product, same users, different model.

Long-context and tool-use effects. The Summary of a Haystack benchmark (Laban et al., 2024) shows that both long-context LLMs and RAG systems still fail at claim-level citation over long inputs. Tool-use overhead in agent-shaped RAG is discussed in the MCP-related literature (Anthropic, 2024), though systematic measurement in production settings remains sparse. Both bodies of work name the confounds we observe but do not disentangle them.

Foundational RAG. The stack we deployed inherits the retrieve-then-generate pattern of Lewis et al. (2020), HyDE-style query rewriting (Gao et al., 2022), dense-plus-lexical hybrid retrieval, and cross-encoder reranking. No architectural component was changed across the upgrade window; only the underlying LLM changed.

3The deployment

Setting. The deployment is a sales-assistant tool operating inside an industrial fluids-manufacturing operator's internal quoting and order-management workflow. We refer to this operator as Partner N throughout, and preserve the option to name the partner in a future non-draft revision pending explicit consent. Partner N's product catalog contains fluid systems used in oil-and-gas drilling; the sales team answers customer and internal questions about product specifications, historical orders, and pricing tiers.

User task. A salesperson types a short question. The system retrieves relevant chunks from a document corpus (product datasheets, pricing agreements, historical order records), calls tools that surface structured data from the operator's ERP, and returns an answer with source citations. If the answer is wrong or insufficient, the salesperson can either mark the event as corrected via a UI action or reply with a follow-up correction that is logged with the same flag.

System components. The retrieval index is a hybrid dense-plus-lexical store using pgvector for embeddings and PostgreSQL tsvector for lexical, chunked at approximately 500 tokens per chunk. The corpus at time of writing contains 318 chunks derived from an underlying set of product datasheets and internal documents. Tools include a structured ERP query interface, a product-catalog lookup, and a historical-order retrieval endpoint. The LLM is invoked via the Anthropic API with tool-use enabled. No architectural change was made to retrieval, tools, or prompting during the observation window.

Model upgrade. On 29 July 2026 the deployment switched its primary LLM from claude-opus-4-7 to claude-opus-5. The switch was motivated by the provider's release of the newer model and the internal expectation that a stronger frontier model would improve response quality. No domain-specific evaluation was performed before the switch; the newer model's public benchmark improvements were treated as sufficient justification, which is a common practice in industrial deployments and is one of the practices this paper argues against.

4Data and measurement

Observation window. 4 June 2026 through 10 September 2026 inclusive. This window brackets the 29 July model switch and provides 8 weeks of pre-switch operation and 6 weeks of post-switch operation.

Event population. Every completed LLM invocation in the sales-assistant tool is logged as one row in a telemetry table. During the observation window 376 events were logged across four models. Of these, 202 were served by Opus 4.7 (window: 4 June to 29 July), 168 by Opus 5 (29 July to 10 September), 4 by Haiku 4.5 (test invocations), and 2 by Sonnet 4.5 (test invocations, both flagged corrected). We restrict our analysis to the Opus 4.7 and Opus 5 populations. Message-level counts are 848 user messages and 840 assistant messages across the same window.

Instrumented fields. Each event row carries: the LLM name, a boolean corrected flag, end-to-end latency in milliseconds, input token count, output token count, and an array of tool names invoked. Each message row additionally carries the number of retrieval sources cited and, where present, a user rating.

Correction signal. We treat the corrected boolean as a proxy for user-perceived response quality. Specifically, it fires when a salesperson takes explicit action to indicate the response was insufficient — either through a UI correction action or through a follow-up message flagged as a correction by the surrounding thread logic. It is not a substitute for expert-annotated correctness; it is a user-in-the-loop satisfaction signal with a directional interpretation (higher rate = worse perceived quality).

Content read. Raw query text, raw response text, comment text, and personally identifying information were not read during the preparation of this paper. All measurements are counts, percentiles, and aggregates. Character-length distributions were computed via SQL LENGTH() against the columns without returning the text.

5Results

Table 1 summarizes per-model event counts and telemetry aggregates. The primary result is the correction-rate column: 3.0% for Opus 4.7 versus 16.1% for Opus 5.

Model N events Corrected Latency (avg) p50 / p95 latency Tokens in Tokens out Tools / call
claude-opus-4-7 202 3.0% 25.4 s 19.2 / 65.6 s 19,892 1,423 1.89
claude-opus-5 168 16.1% 44.2 s 40.6 / 101.8 s 56,030 2,340 2.89

Table 1. Per-model event counts and telemetry aggregates over the observation window. Correction rate is percentage of events with corrected = true. Latency, tokens, and tools-per-call are means. Test-invocation models (Haiku 4.5, Sonnet 4.5) omitted for population size.

Statistical significance. Constructing the 2x2 contingency table (corrected vs not corrected, by model), Fisher's exact test rejects the null hypothesis of equal correction rates at p < 0.0001. The odds ratio is 6.25 (Opus 5 events have 6.25x the odds of being corrected relative to Opus 4.7 events), with a 95% confidence interval spanning approximately 2.5 to 15 by standard log-odds approximation.

Three signals moved together across the upgrade CORRECTION RATE 3.0% 16.1% Opus 4.7 Opus 5 5.4x increase AVG INPUT TOKENS 19.9K 56.0K Opus 4.7 Opus 5 2.8x increase AVG TOOLS / CALL 1.89 2.89 Opus 4.7 Opus 5 1.5x increase

Figure 1. Three signals moved together across the model upgrade from Opus 4.7 (blue) to Opus 5 (amber). The correction-rate rise is the outcome variable; the input-token and tools-per-call rises are candidate causes. The observation the paper cannot make from this data alone is which of the two candidate causes drove the outcome, or whether a third variable did.

Message-level context. Across the same window the messages table records 848 user turns and 840 assistant turns. User queries average 80 characters with a median of 59 characters, indicating a short-question pattern typical of a sales-assistant tool. Assistant responses average 1,419 characters with a median of 1,212 characters. Retrieval citations are attached to 709 of 840 assistant messages (84.4%), averaging 6.4 cited sources when present. Only 5 messages were assigned explicit ratings through the rating UI (avg -0.20), confirming that the correction signal is the practical error channel and explicit ratings are not.

What did not change. Chunking strategy, retrieval index, tool definitions, prompts, and the underlying corpus size (318 chunks) were unchanged across the observation window. User population was unchanged (same sales team). The corpus was refreshed on the same cadence throughout.

6Analysis of confounds

Three variables moved simultaneously with the model change. Each has a plausible independent effect on correction rate. Untangling them from the model change itself is the substance of any further work on this question.

6.1Context-length effect

Average input tokens per event grew from 19,892 to 56,030, a 2.8x increase. This does not represent a growth in the underlying document corpus; it represents the newer model's tendency to accept larger retrieved contexts before invoking generation, either through changes to the internal retrieval-context-assembly logic that were made concurrently with the model swap, or through the newer model's willingness to accept longer tool-output payloads.

The "lost in the middle" family of findings (Liu et al., 2023) suggests that increasing retrieved context can degrade answer quality even when the target passage is present in the context, because attention over long contexts is uneven. A 2.8x context increase without a corresponding gain in answer quality is consistent with the lost-in-the-middle mechanism operating at the operational scale we observe. Testing this hypothesis requires a rerun of the same queries against the older model with the newer context length, which is straightforward to do in an offline replay setting.

6.2Tool-use effect

Average tool calls per response grew from 1.89 to 2.89, a 1.5x increase. This appears to reflect the newer model's greater willingness to chain tool invocations before returning a response, which is generally interpreted as a positive sign but which introduces additional failure modes: incorrect tool selection, incorrect tool argumentation, and incorrect synthesis across tool outputs. Any of the three can produce a wrong final answer without producing any single obvious failure.

In our deployment, structured ERP-query tools return tabular data that must be joined with narrative product-catalog output. Each additional tool call increases the number of joins the model must perform mentally, and the correction-rate rise may reflect join errors rather than retrieval errors. Testing this hypothesis requires per-tool attribution of the corrections, which is achievable by extending the telemetry to record which tool outputs were surfaced to the user in each incorrect answer.

6.3Query-distribution drift

The user population is the same but the queries submitted differ across the two windows. Message lengths are similar (means and medians within a few characters), but semantic content may not be. Users may have gradually learned to submit harder questions to the tool, or the seasonal shift in the operator's business may have brought different product classes to the fore during August and September relative to June and July.

Query-distribution drift is the confound our current data most poorly constrains. Testing it requires reading the query text, which is out of scope for the current draft but is a natural next step under a Newpark-approved data-release framework. In the absence of that access, the confound remains a limitation rather than a decidable question.

6.4Latency as a mechanistic clue

End-to-end latency grew from 25.4 to 44.2 seconds (a 1.7x increase), and the 95th percentile moved from 65.6 to 101.8 seconds. Longer latency is mostly a mechanical consequence of longer context and more tool calls. A user waiting 100 seconds for a response may be more likely to notice its shortcomings on receipt (attention effect), or may be more likely to flag it as insufficient due to prior investment (sunk-cost effect). We do not have the instrumentation to test either.

6.5Attempted causal attribution via offline replay (v0.2 addendum)

Sections 6.1 and 6.2 anticipated offline replay as the natural next step for adjudicating among the context-length and tool-use confounds. We attempted that adjudication across a series of four replay variants and two scoring instruments, and report a negative result. The methodology as attempted did not yield a clean attribution. We describe what we tried, what broke, and what a working design would look like.

Experimental setup. We drew 30 events from the same event population as Section 5, stratified by pre-transition model and correction status: 10 Opus 4.7 uncorrected, 10 Opus 5 uncorrected, and 10 Opus 5 corrected. For each event we constructed a controlled replay against a target model plus context configuration and scored the replay against the deployment's original response. All replays were run through the Anthropic Messages API against api.anthropic.com with tool-use enabled, matching the deployment's own inference path. No raw query text, response text, or comment text entered the analyst's context; extracted facts and per-event scores were written to disk and only aggregate statistics were surfaced.

Replay variants. Variant A (model-only) replays Opus 5 events on Opus 4.7 with the same minimal context and no tools. Variant B (context-stress) replays Opus 4.7 events on Opus 4.7 with the context artificially padded from an external domain corpus to Opus 5 production token lengths. Variant C (tools-stress) replays Opus 4.7 events on Opus 4.7 with three plausible decoy tools added to match Opus 5's per-call tool count. Variant D (tools-prune) replays Opus 5 events on Opus 5 with the tool set pruned from its original count (median 3, max 7) to Opus 4.7's approximate average of 2. Variant D is the causal test for the tools hypothesis: if pruning tools recovers quality on the same model, tools are the confound.

Instrument iteration 1 (LLM-as-judge, arm-as-judge). Our first-pass instrument was Opus 4.7 itself acting as an impartial judge, scoring the replay on a 0.0-to-1.0 scale against the deployment's original response with 0.5 meaning equivalent. Variants A, B, and C returned means of 0.34, 0.57, and 0.28 respectively (n=20/10/10). Read at face value this pointed to a clean causal story: context growth was benign (B near 0.5), tools were the confound (C low), and rolling back the model would not recover the loss (A low). The story was clean enough to warrant a check. The check is standard practice: swap the judge to a model not on either arm of the comparison.

Instrument iteration 2 (LLM-as-judge, neutral third-party model). We repeated the four variants with Sonnet 4.5 as the judge, added Variant D at n=10, and raised the context-stress target from ~40K to ~130K input tokens using an external padding corpus. All four variants collapsed toward the low end: A 0.20, B 0.17, C 0.10, D 0.13. Variant D's score was the critical finding. Variant D is a same-model control: replaying Opus 5 against Opus 5 with only the tool set pruned. Its expected score under a well-calibrated judge is close to 0.5. It scored 0.13, and 9 of 10 individual events scored 0.00. A well-calibrated judge should not produce that distribution on a same-model control. The judge was not measuring quality; it was producing a systematically biased scoring signal.

Instrument iteration 3 (position-swap correction). The Sonnet-4.5 result is consistent with the LLM-as-judge position bias documented in Zheng et al. (2023): judges tend to favor whichever response appears first in the prompt. The recommended correction is to run the judge twice per pair, once with ORIGINAL first and once with REPLAY first, and average after inverting the swapped score. We ran the full 50-event corpus (20 A, 10 B, 10 C, 10 D) with the position-swap correction and recorded both the corrected score and the magnitude of disagreement between the two positions.

Variant n Naive (Sonnet 4.5) Position-corrected Avg inter-position disagreement
A model-only200.300.430.26
B context-stress (~133K in)100.210.510.61
C tools-stress100.130.450.67
D tools-prune (same-model control)100.050.110.12

Table 2. LLM-as-judge scores across the four replay variants, before and after position-swap correction, with the average inter-position disagreement per variant. B and C's disagreement scores above 0.60 indicate the judge was operating near-randomly across positions at high context and tool loads. Post-correction, the four variants no longer distinguish cleanly.

Post-correction, Variants B and C settle near equivalence (0.51 and 0.45), Variant A shows a modest sub-equivalence result (0.43) whose confidence interval includes 0.5 at n=20, and Variant D remains catastrophically low (0.11) with low inter-position disagreement (0.12), indicating the judge agrees with itself that the Variant D replays are much worse than the originals. The v0.1-style clean attribution to tools does not survive the position-swap correction.

Instrument iteration 4 (factuality-anchored recall). The v0.4 corrected judge left B and C near 0.5 and D low; we could not tell from the judge alone whether the Variant D outcome reflected a real methodology finding or another judge artifact. We built a deterministic instrument as a cross-check. For each original response we used Sonnet 4.5 with structured tool-call output to extract atomic factual claims into typed JSON: entities (well names, product SKUs, customer names), numeric quantities (depths, flow rates, drum counts), monetary values, percentages, and dates. We then scored each replay by literal recall: what fraction of the extracted facts appear in the replay text, with type-appropriate normalization (case-insensitive substring for entities, ±0.5-1.0% tolerance for numeric values, punctuation-stripped matching for dates). The scorer is fully deterministic given the extracted facts and immune to judge bias.

Variant n Events w/ facts Mean recall Median recall v0.4 corrected judge
A model-only201610.8%3.0%0.43
B context-stress1098.1%3.2%0.51
C tools-stress1095.3%0.0%0.45
D tools-prune1080.0%0.0%0.11

Table 3. Factuality-anchored entity-recall scores across the four replay variants, alongside the v0.4 position-corrected LLM judge. Median recall is 3% or lower on every variant. The two instruments (deterministic recall vs. corrected LLM judge) rank the variants in the same order and both find Variant D at the floor; neither distinguishes A, B, and C.

Fact-type breakdown. Across 50 replays and 1,144 extracted facts, aggregate recall by type was: dates 18.5% (24 of 130), numeric quantities 4.0% (26 of 648), entities 3.9% (11 of 279), percentages 3.7% (1 of 27), monetary values 0.0% (0 of 60). Dates matched at the highest rate because they are the fact type most easily guessed from context (current year, recent quarters). Every fact type that requires retrieval of specific corpus content collapsed to near zero. Not one dollar figure was reproduced by any replay across any variant.

Diagnosis. The uniform failure across A, B, and C is not evidence about the confounds. It is evidence about the replay pipeline. Partner N's original responses contain specific well names, product SKUs, chemical concentrations, and per-drum liter counts drawn from documents that live in Partner N's retrieval index. Our offline replays did not include that retrieval, nor did they include the production system prompt or the domain-specific tool schemas. The replays therefore produced generic sales-assistant responses stripped of the specific facts that made the originals factually correct in the first place. Variant D's zero recall is the confirmation: it swaps neither the model nor the context, yet it still fails at recall, because tool-pruning removed the retrieval tool alongside the decoy tools. Offline replay without production retrieval is not a controlled experiment on the confounds; it is an ablation of retrieval.

Implication for RAG evaluation methodology. Offline LLM-as-judge attribution on production RAG deployments is fragile in three separable ways. First, LLM judges exhibit strong position bias without corrective procedures (Zheng et al., 2023). Second, judge-model swap changes scores substantially: an attribution that looks clean when the judge is one of the arms may collapse when the judge is a third-party model. Third, even a position-corrected, factuality-anchored instrument cannot distinguish confounds when the replay pipeline lacks the deployment's retrieval and system prompt. Any two of these issues would be manageable; jointly they invalidate offline replay as an attribution method for production RAG regressions.

What would work: production-mode shadow-replay canary. The design implied by our negative result is straightforward. Route a fixed fraction of live production traffic (with the full production retrieval, tools, and system prompt) through both the incumbent and candidate models in parallel for a period long enough to collect a meaningful correction-rate sample on each. Score by the deployment's native user-flagged correction signal on identical traffic. In our deployment, one week of shadow traffic at pre-switch volume would have yielded roughly 30 events per arm, sufficient to distinguish 3% from 16% at the significance level we report. This design isolates the model as the sole moving variable, uses the same ground-truth channel that produced the original regression finding, and is independent of any external judge. Section 9 revises the practitioner recommendations accordingly.

Reproducibility. The four replay orchestrators (v0.2, v0.3, v0.4, v0.5-factuality), the extracted-facts cache format, and the summary schemas are released with the reference implementation of the Silent Regressions detection framework under the MIT license. Public repository forthcoming. Per-event replay text and extracted facts are retained under Partner N's data-handling terms and are not part of the public release; only aggregate statistics of the form reported in Tables 2 and 3 will be released.

7Discussion

Standard benchmarks did not warn us. Opus 5 outperforms Opus 4.7 on every provider-published benchmark we consulted before the switch. The benchmarks predicted an improvement; the deployment recorded a regression. This is the practical version of the observation, made in Xiong et al. (2024) and Islam et al. (2023), that domain-specific evaluations are non-negotiable. Our contribution is to quantify how large the gap can be on a real deployment, and to show that the gap is measurable using instrumentation that most operators already have.

The correction signal is proxy, not truth. A rise in correction rate can reflect worse answers, more attentive users, or a shift in what users consider correctable. We take the signal at face value in this paper because it is the metric the operator's team acts on internally, and because there is no obvious mechanism by which users would have grown more attentive within a six-week window on the same tool. We recommend that follow-up work include an expert re-annotation of a stratified sample of corrected and uncorrected events to bound the noise on the proxy.

The upgrade decision is now a research question. The most immediate consequence of the finding is that Partner N is now uncertain whether to hold on Opus 5, roll back to Opus 4.7, or restructure the deployment to constrain the confounds. Any of the three options produces a further measurement. The decision is out of scope for this paper but is a live question for the operator at the time of writing.

Offline LLM-as-judge attribution is not safe by default. Section 6.5 documents a sequence of four offline-replay iterations in which each apparent finding turned out to be an instrument artifact under closer inspection. The v0.1 iteration produced a clean tools-attribution that survived only because we used one of the models on trial as the judge. Swapping to a neutral judge collapsed all scores. Position-swap correction pulled the scores back apart but revealed that the raw judge was operating near-randomly at high context and tool loads. A deterministic factuality-anchored replacement confirmed that every replay variant failed the recall test uniformly. Practitioners running this class of experiment should assume none of these failure modes will surface without an explicit same-model control. Ours (Variant D) was the piece that told us the pipeline was broken; without it the position-corrected judge alone would have looked like a converged answer.

8Limitations and ethics

Single deployment, small n. The n of 202 vs. 168 events is small by academic standards. Fisher's exact test rejects the null cleanly at this size, but the effect-size confidence interval remains wide. A single deployment cannot support claims about frontier-model upgrades in general; it can only support the claim that this class of finding is possible and that the size of the delta can be large.

Confounded design. The model change, context growth, and tool-count growth co-occurred. We do not run a controlled A/B test. Section 6.5 documents the offline-replay attempt at attribution and reports it as a negative result: neither a position-corrected LLM judge nor a deterministic factuality-anchored recall metric can distinguish among the confounds when the replay pipeline lacks the deployment's production retrieval. The paper's causal claim is therefore weaker than v0.1 anticipated: we can say the correction rate rose 5.4x across the upgrade, and we can rule out a large-N false positive statistically, but we cannot yet assign the rise to any single confound. Section 6.5 identifies a production-mode shadow-replay canary as the design that would.

Correction as proxy. As discussed in Section 7, the correction flag is not a direct measurement of answer correctness. A follow-up expert re-annotation of a stratified sample is the standard remediation.

Data anonymization. No raw queries, responses, comment text, or user identifiers are read or published. Only counts, percentiles, and aggregate telemetry are reported. The partner operator's name is currently withheld and will be added, if consent is granted, in a future non-draft version.

Ethics. The measurement was conducted on product telemetry that Partner N generates through routine operation of the sales-assistant tool. No end-user identifying data is used or reported. The partner-operator's counsel is being consulted regarding the publication path.

9Recommendations for practitioners

Log the correction signal from day one. A boolean corrected flag per event, set by an explicit user action, is the single highest-leverage instrumentation item on a production RAG system. It gives a directional quality metric that is domain-adapted for free. Every industrial RAG deployment should have this from the first production release.

Log tokens, latency, and tools per event. These fields are what turned an anecdotal "the new model feels off" into a measurable regression. All are trivial to record. All are essential when reasoning about post-upgrade behavior changes.

Do not treat provider benchmark improvements as domain-transferable evidence. They aren't. Establish domain-specific evaluation before the upgrade, or accept the risk of a regression like the one this paper reports. A minimal domain-evaluation set of a few hundred representative queries with expected outputs is enough to catch regressions of the size we observed.

Canary the rollout in production mode, not offline. Route a fixed fraction of live traffic through both the incumbent and candidate models in parallel, with the full production retrieval and tool stack in place. Score by the deployment's native user-flagged correction signal on identical traffic. In our deployment, one week of shadow traffic at the pre-switch volume would have yielded roughly 30 events, sufficient to distinguish 3% from 16% at the significance level we report. Do not attempt offline replay as an attribution instrument for a RAG regression: Section 6.5 documents four iterations of that attempt and their failure modes.

Include a same-model control in any comparison. If you must run offline experiments (for example to bound the size of a candidate confound before committing to a production canary), always include a same-model replay as a control arm. It is the single measurement that catches most instrument artifacts, including LLM-as-judge position bias, judge-model swap effects, and retrieval-absence collapse. Ours was Variant D in Section 6.5; it was the piece that told us the pipeline was broken.

Watch the confounds. A model change that also changes context length, tool-use patterns, or prompt structure is not a clean model change. Freeze those variables before comparing model versions, or accept that any observed regression is a system-level rather than a model-level finding.

10Conclusion

We reported a 5.4x correction-rate regression on a production industrial RAG deployment following an upgrade from Claude Opus 4.7 to Claude Opus 5. The regression is statistically significant at conventional levels and is accompanied by simultaneous increases in context length, tool use per call, and end-to-end latency. Any of the three could drive the outcome, and Section 6 lays out the mechanisms.

The v0.2 revision reports a negative result on offline-replay attribution: four experimental variants scored by both a position-corrected LLM-as-judge and a deterministic entity-recall metric fail to distinguish among the confounds. Median factuality recall was 3% or lower on every variant; zero of 60 dollar-figure facts were reproduced by any replay. The uniform failure across variants is not evidence about the confounds, but about the replay pipeline: offline replays that lack the deployment's production retrieval collapse to a common factually-thin floor regardless of which confound is under test. Same-model control (Variant D) confirms this: pruning tools on Opus 5 does not recover quality, because the pruning also strips retrieval. Any offline replay experiment on a RAG regression must include a same-model control or accept that instrument artifacts will pass unnoticed.

Our broader claim is that this class of measurement should be routine for industrial-RAG operators, and that the community should collect enough case studies to build a proper empirical base for reasoning about frontier-model upgrades in deployed settings. Part of that base is a methodology contribution: the working attribution design is a production-mode shadow-replay canary, not offline replay. This paper is one contribution to that base. We invite additional case studies from other industrial-RAG deployments, and offer to serve as a data-sharing point for anonymized aggregate telemetry from operators willing to contribute.

AAcknowledgments

We thank Partner N's product and sales-operations teams for the ongoing collaboration that made this deployment observable. We thank the reviewers of the KoldOps internal draft cycle. Draft prepared with editorial assistance from Claude (Opus 4.7).

RReferences

  1. Anthropic. Model Context Protocol: An Open Standard for Connecting AI Assistants to Data Sources. Technical note, 2024. anthropic.com/news/model-context-protocol
  2. Asai, A., Wu, Z., Wang, Y., Sil, A., Hajishirzi, H. Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection. arXiv:2310.11511, 2023. arxiv.org/abs/2310.11511
  3. Borgeaud, S., et al. Improving Language Models by Retrieving from Trillions of Tokens (RETRO). arXiv:2112.04426, 2022.
  4. Edge, D., et al. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130, 2024.
  5. Es, S., James, J., Espinosa-Anke, L., Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. EACL Demo, 2024.
  6. Gao, L., Ma, X., Lin, J., Callan, J. Precise Zero-Shot Dense Retrieval without Relevance Labels (HyDE). arXiv:2212.10496, 2022.
  7. Gao, Y., et al. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997, 2023.
  8. Islam, P., et al. FinanceBench: A New Benchmark for Financial Question Answering. arXiv:2311.11944, 2023.
  9. Izacard, G., et al. Atlas: Few-shot Learning with Retrieval Augmented Language Models. JMLR, 2023.
  10. Jiang, Z., et al. Active Retrieval Augmented Generation (FLARE). EMNLP 2023.
  11. Laban, P., et al. Summary of a Haystack. EMNLP 2024, arXiv:2407.01370.
  12. Lewis, P., et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020, arXiv:2005.11401.
  13. Li, X., et al. Long Context vs. RAG for LLMs: An Evaluation and Revisits. arXiv:2501.01880, 2025.
  14. Liu, N. F., et al. Lost in the Middle: How Language Models Use Long Contexts. arXiv:2307.03172, 2023.
  15. Muennighoff, N., et al. MTEB: Massive Text Embedding Benchmark. EACL 2023.
  16. Pipitone, N., Alami, G. H. LegalBench-RAG. arXiv:2408.10343, 2024.
  17. Saad-Falcon, J., et al. ARES: An Automated Evaluation Framework for RAG. NAACL 2024.
  18. Thakur, N., et al. BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models. arXiv:2104.08663, 2021.
  19. Upadhyay, S., et al. Ragnarok: A Reusable RAG Framework and Baselines for TREC 2024. arXiv:2406.16828, 2024.
  20. Xiong, G., et al. Benchmarking Retrieval-Augmented Generation for Medicine (MedRAG). ACL Findings 2024, arXiv:2402.13178.
  21. Zheng, L., et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks Track, arXiv:2306.05685. (Documents position bias in LLM judges and the two-run swap correction adopted in Section 6.5.)

Data availability. The aggregate telemetry counts reported in Table 1 and Figure 1 are drawn from the KoldOps OpsBox platform's sales_assistant_events and sales_assistant_messages telemetry tables at Partner N. Table 2 and Table 3 aggregate is drawn from a stratified 30-event sample of the same tables, with per-event replay text and extracted facts retained under Partner N's data-handling terms. A companion release of the anonymized aggregate counts will accompany the non-draft version of this paper.

Code availability. The reference detection framework (Fisher's-exact and chi-squared correction-rate shift tests, BH-FDR correction, sample-size planning) and the four replay orchestrators are released under the MIT license alongside this paper. Public repository forthcoming.

Draft status. This is a v0.2 working draft. It has not been peer-reviewed, is not published, and is intended for internal review by KoldOps and Partner N before any public circulation. Do not cite. Corrections, disagreements, and challenges are actively invited.

Changelog. v0.2 (2026-09-11) adds Section 6.5 (offline replay negative result), Tables 2 and 3, discussion of LLM-as-judge fragility in Section 7, revised limitations in Section 8, and a revised production-mode canary recommendation in Section 9. v0.1 (2026-09-10) initial release.

KoldOps R/01 · Research brief 004 · v0.2 · September 11, 2026