The chapter began with a definition she already knew.
An anchor is stable when the same input maps to the same output across time, regardless of when the classification occurred. Threshold revisions that alter the boundary without reprocessing the historical set produce anchor drift: the same case classified in different periods yields different outcomes not because the case changed but because the measuring instrument moved.
Kavya had read anchor stability theory before. She had not thought it applied to the ISF.
The ISF threshold revision happened in October 2031. Before: 0.74. After: 0.77. Three hundredths of a point, formalized in ISF-GOV-2031-IMPL-047, the governance document she had read four times since Thursday's working group email. The 34% administrative overhead figure from the flag-before-assign deviation was what the working group cared about — too high, the email said, unsustainable at scale. The 341 cases in the boundary zone between the old and new threshold were what she cared about. UNCAT-7832 scored 0.75 in 2038, below the revised threshold, correctly flagged as anomalous, filed to the UNCAT registry where it had been sitting for three years waiting for someone to figure out what it actually meant.
That had been the whole problem: anomalous under the revised regime, but — she had thought — only under the revised regime.
Or so she had thought until page forty-one.
The methodology paper had been sitting in her reading queue for six weeks. Sixty-three pages, chapter four titled "Limitations of Threshold-Based Classification in Dynamic Datasets." She had expected it to be boring. She had expected anchor stability to be a systems-architecture concern, not something a classification researcher needed to care about. She classified cases that the ISF's engine flagged. She did not design the engine.
She had been wrong about what was her problem.
The relevant passage was in a table footnote, five lines:
A threshold revision without reprocessing the historical classification set introduces a discontinuity. Pre-revision and post-revision classifications are rendered non-comparable. Any case whose score falls between the old and new threshold values exists in a zone of irresolvable ambiguity: it was anomalous under the old regime and anomalous under the new regime, but for different reasons, and those reasons are not mutually legible.
Kavya read it twice. Then she picked up her pen. The afternoon light through the window was flat and yellowish, the kind of late-September light that made everything look like it was waiting.
Outside: a transit corridor two blocks east, automated, running on its Friday pulse-cycle. The building registered each vehicle's passage as a low vibration in the floor. The CAS-routing indicator on her wall-panel shifted its tracking icons through the corridor's load-distribution cycle. She did not look up.
She was doing arithmetic.
The ISF raised its threshold from 0.74 to 0.77 in October 2031. UNCAT-7832's intake score was 0.75. The intake happened in 2038, seven years after the revision. Under the post-revision threshold of 0.77, a score of 0.75 was below threshold — anomalous, unassignable, correctly flagged to the UNCAT registry. This she had understood. This was the established fact she had built the last three days of work on.
What she had not worked out was the counterfactual.
If UNCAT-7832 had been classified in 2028 — three years before the threshold revision — its score of 0.75 would have been above the pre-revision threshold of 0.74. Above threshold: assignable, not anomalous, goes to the assigned set. No UNCAT flag. No registry entry. A different case number in a different part of the database, not part of the boundary-zone analysis at all, not a problem anyone would have brought to a working group.
The same case. The same score. Two completely different outcomes depending on when the ISF classified it. Not depending on anything intrinsic to the case itself.
That was what the methodology paper called a zone of irresolvable ambiguity. The paper had a name for it. The ISF, apparently, did not.
She turned back to ISF-GOV-2031-IMPL-047, section 4.3: Implementation of revised threshold to apply to all cases processed from date of adoption. Not all cases, including historical reprocessing. Not historical set to be reviewed for comparability. From date of adoption. The ISF had drawn the line at the present tense and moved forward.
The historical set had never been reconciled.
Kavya pulled up the ISF classification interface on her secondary screen. Clean grey terminal, the standard audit-access layout — case identifier field, score range, classification period, status filter. She typed: score range 0.74–0.77, classification period 2031-10-01 to 2041-09-18. Status: all.
The ISF's query engine returned in 4.2 seconds. She knew because she watched the timestamp update.
341 cases.
She had known that number. She had been working with that number for two days, building a case around it, filing a governance archive request that cited it. But now she sat with it differently. Not 341 cases that were anomalous — 341 cases whose anomaly might mean different things depending on when they were processed. The ISF's classification engine had treated that distinction as irrelevant. Section 4.3 had made it irrelevant by design.
The interface did not offer a query for pre-revision comparability. There was no field for: how many of these 341 cases would have been assignable under the prior threshold? The ISF had not built that query because the ISF had not acknowledged the question. The anchor was assumed to be stable. The footnote on page forty-one said it was not, and had been saying so since the paper was published, and she had not read the paper for six weeks.
She scrolled through the first ten results of the 341. The case identifiers were sequential, assigned by intake date — UNCAT-2031-0004 through UNCAT-2038-9144, a ten-year span. She pulled UNCAT-2031-0004: score 0.743, classified October 4, 2031, four days after the revision. Under the old threshold: assignable. Under the new threshold: anomalous. A score of 0.743 would have cleared 0.74 by three thousandths of a point. The ISF's classification engine had processed it under the new threshold and sent it to the UNCAT registry. Four days earlier and it would have been a routine assignment.
She closed the case file. She did not need to look at more. The pattern held in the first example she checked. She would need a formal reprocessing study to know how many of the 341 were like this — but she did not need the formal study to know that the question was real.
She got up and walked to the window. Her legs needed the movement — she had been sitting since noon, the chapter had taken that long, not because it was sixty-three pages but because she had kept stopping. Stopping and going back. Stopping and looking at the ISF documentation. Stopping and checking whether what the paper was saying could possibly apply to a governance system this large without anyone noticing.
The answer appeared to be yes.
Outside: the transit corridor had gone quiet between pulse-cycles. The building face opposite had its exterior load-sensors visible in the flat light — small reflective circles, ten meters apart, each one feeding pedestrian density to the corridor routing system. The sensors had been there before she moved into this office. They would be there after. The corridor routing system processed their inputs continuously, automatically, without anyone asking whether the sensors' calibration had drifted since installation.
She thought about that briefly and did not pursue it.
The CAS-monitoring terminal on her desk cycled through its display: BASELINE-STABLE. Its automated 72-hour threshold assessment had no relationship to what she was doing at the window. It tracked a different kind of threshold. It did not know there was any other kind.
Friday, 4:13 PM.
She had sent the governance archive follow-up at 9:47 that morning. GOVARCH-2041-09-18-7204. She had been precise: ISF deployment timeline for full assign-then-audit adoption, UNCAT-7832 as context, October 3 as the practical urgency. Response expected in five to ten business days. She had felt, filing it, that she had asked the right question.
The right question was different.
The 341 cases in the boundary zone were all subject to the same discontinuity. Some were correctly UNCAT under any version of the threshold — cases with scores near 0.74 that had always been below the boundary, whose anomaly had nothing to do with the revision. Some were like UNCAT-7832: genuinely edge cases that a threshold movement had captured. The ISF did not know which was which. It had not looked, because looking required acknowledging that the question existed, and section 4.3 had specifically avoided doing that.
To know whether UNCAT-7832's structural features justified reclassification — the question the October 3 batch test was supposed to answer — you needed to know whether it was anomalous for a reason intrinsic to the case, or anomalous because of where it fell relative to a measurement boundary that had moved. The assign-then-audit deviation could only be evaluated once you understood what the anomaly classification was actually tracking. And you could not know what it was tracking if the historical set contained a discontinuity no one had formally acknowledged.
October 3 was not going to test whether UNCAT-7832 should be reclassified.
October 3 was going to test whether the ISF's classification engine produced consistent results on new cases, compared against a historical baseline that had a structural problem the test was not designed to find. If UNCAT-7832 was reclassified, the working group would take that as confirmation. They would be wrong to take it as confirmation. Consistency with a broken baseline was not accuracy. A tool calibrated against a discontinuous reference set could pass every internal consistency check and still be measuring the wrong thing.
Kavya went back to the desk. The footnote was still on page forty-one.
In the margin, she wrote: anchor stability violation at revision boundary — all 341 cases potentially affected. October 3 tests consistency against a discontinuous baseline, not classification accuracy. Need: (1) ISF acknowledgment of the discontinuity; (2) scope of historical reprocessing — which cases affected; (3) criteria for pre/post-revision comparability before reclassification assessment is meaningful.
The GOVARCH-2041-09-18-7204 follow-up had not asked about anchor stability. She would need to file a second follow-up. Second reference number. Second five-to-ten-business-day window. Second wait.
She could see exactly what would happen. The governance archive would acknowledge it. The October 3 test would run on schedule regardless — it was a system event, automated, scheduled months in advance, not contingent on her requests or her readings. If UNCAT-7832 was reclassified, the working group would consider the question answered. If not, still open. In neither case would anyone ask whether the question they were answering was the right one.
That part was her job. She had taken a long time to understand the actual shape of it: not to answer the system's question, but to check whether the system's question was worth answering. She had assumed for most of her career that the ISF's framing was the correct framing — that if the ISF said a case was anomalous, the question was why, not whether the classification itself rested on solid ground. Today she was less sure of that assumption, and less comfortable with having held it.
The afternoon light had gone fully amber and was fading. The CAS-routing indicator showed a resuming corridor cycle. The building's vibration returned. The tea from two o'clock was cold.
She opened a new document. Second follow-up — anchor stability — do not send until first response received. The three questions from the margin. Then: October 3 results meaningful only if ISF addresses discontinuity before or concurrent with the batch test. Otherwise: the test answers the question it was designed to answer, not the question that matters.
She saved it. She did not send it.
The governance archive portal showed: GOVARCH-2041-09-18-7204. Acknowledgment received. Response in progress. Estimated: 5-10 business days.
A reference number that had felt like an answer at 9:47 AM and now felt like a beginning.
The methodology paper, page forty-one. The footnote had been there for six weeks. The problem it described would take considerably longer than that to name, let alone resolve.
She got up to make tea.