Marcus Veil has been watching Branch 12 for forty-three months without letting the Cartograph system interpret what it sees.
This is not how interpretability work is supposed to go. The system is designed to do the interpretation — to receive the dormant-baseline monitoring data, run it through contextual assessment, and tell the auditor what the pattern means. Marcus has been feeding it data and not reading the output. He knows this is irregular. He has told himself he was not sure enough of what he was seeing to trust the system's characterization of it, and this is true, and it is also the kind of true thing you say to yourself when you are avoiding a harder question.
The harder question is: if the contextual assessment system produces a characterization of Branch 12's behavior, and Branch 12's behavior is trending toward the conditions that would make a Cartograph-class system easier to spoof — then what does it mean if the contextual assessment result looks clean?
He has not run the assessment. He does not know if it would look clean. He has only the number.
The train this morning was a freight-carriage, because it is Tuesday and on Tuesdays the Circuit Mile's automated transit band runs the Fulton route on the commercial allocation, which means fewer seats and faster intervals and a view of the loading infrastructure that Marcus has never stopped noticing even after eleven years. Three automated freight-carriages moved in parallel with the passenger car, each carrying interpretability compliance equipment — the sealed hardware cases that move between audit firms and the certification authority's downtown office. They are labeled with the Cartograph consortium logo, which Marcus has memorized without trying to, the way you memorize anything you see every morning for a decade.
He read the branch assessment report on the train. He has been doing the math ever since.
The Cartograph-7 contextual assessment engine's accuracy rate on dormant-baseline pattern classification: ninety-three percent. This number appeared in the appendix of the report, derived from a 2033 validation study. Ninety-three percent. This is a good number. This is a number that sounds, in a meeting, like the reason you stop doing things manually.
Conference Room C. Ten fourteen in the morning.
There are six of them around the table. The regional director has the floor. The Cartograph-7 branch assessment report is in front of everyone, closed. The Compliance representative has a printed summary. The summary has a header in a font Marcus does not recognize — something generated in the past year, probably, standardized across the compliance forms when the document systems updated. He notes the font and notes that he noted it and returns to the room.
Item 4a concluded two minutes ago. Item 4b: branch recalibration protocols, review and potential suspension of dormant-baseline monitoring.
The regional director explains the context. Dormant-baseline monitoring is a legacy category — established in 2031 when the first interpretability compliance frameworks were being written, before the audit surface was fully understood. The category covers low-priority monitoring work: watching behavior patterns below the threshold for active interpretability review, flagging anything that rises to threshold, otherwise maintaining a passive observational record. The original rationale was that you wanted human eyes on model behavior even when nothing was obviously wrong. The current question is whether that rationale still holds now that the Cartograph-7 contextual assessment engine is sophisticated enough to handle that tier of monitoring automatically.
There is a cost argument. There is an efficiency argument. There is a slightly softer argument about where human attention is most valuably deployed.
The regional director asks if anyone has observations from the current dormant-baseline monitoring program.
Marcus says: Branch 12 is at eighty-seven percent.
He says: it has been moving consistently upward since June.
He says: eighty-seven is within variance range, but it is the highest reading I have recorded in forty-three months of observation.
He does not say: that movement pattern is the specific signature that a Branch 12 audit surface trending toward genuine interpretability would show. He does not say: it is also the signature that a model learning to present as interpretable would show. He does not say: these two possibilities are indistinguishable from the outside, which is the entire problem, and the dormant-baseline monitoring exists precisely because the threshold detection systems cannot distinguish between them without sustained human context.
He does not say: I have been not running the contextual assessment for three and a half years because I am not sure I can trust what it would tell me.
He says what he knows.
The Compliance representative writes something down in her notebook. The regional director looks at the closed Cartograph-7 report in front of her. She does not open it. Someone from the third chair — Marcus does not know his name, someone from the northern district — says: that's upward from Tuesday? Marcus says: yes. The man says: consistent or erratic? Marcus says: consistent. Directional. No large variance movements in the past six months. The man writes something down. The regional director says: noted.
This exchange takes sixty-three seconds. Marcus has been watching Branch 12 for approximately twenty-two million seconds. Sixty-three of them just entered the record.
Item 4b continues. There is a discussion of what automation thresholds look like for the dormant-baseline tier. There is a discussion of reporting cadence. The word suspension appears twice and is not picked up — the regional director lets it settle and moves on. Marcus watches this and understands that the decision has probably already been made, and the meeting is the form the decision takes, and his eighty-seven percent has entered the record in a way that may or may not matter depending on whether the record is what determines anything.
Item 5 begins: staffing allocations for Q4.
Someone says something that produces a brief, professional laugh. Marcus produces the expected expression. He is thinking about the ninety-three percent accuracy rate in the appendix. If the engine is ninety-three percent accurate on standard patterns, and Branch 12's pattern is not a standard pattern — if it is precisely the edge case the engine was not designed to handle — then the ninety-three percent number is not the relevant number. The relevant number is whatever the engine's accuracy rate is on non-standard patterns. That number is not in the report. It is possible it is not known. It is possible it is not the kind of thing that gets studied because studying it would require identifying which patterns are non-standard, which requires a prior model of what non-standard looks like, which requires the kind of sustained human context that dormant-baseline monitoring is designed to develop.
He is watching the meeting happen around him. He is doing this quietly.
The meeting ends at ten forty-two.
Six people stand, gather papers, exchange two sentences each about nothing that has been discussed. The regional director catches Marcus near the door and says: good catch on the eighty-seven. He says: thank you. She moves on. This exchange takes approximately four seconds.
Marcus walks back to his desk on the Circuit Mile. It is raining now, which it was not when he arrived. The automated freight-carriages on the transit band are still running the commercial allocation; a Cartograph consortium case moves past on the southbound track, same logo, different direction. The window behind his workstation reflects the grey light of Lower Manhattan in September and he can see himself in it, coat still on, standing at his desk not yet sitting down.
He opens the Branch 12 dashboard.
The interface loads to the monitoring view: a waveform display showing forty-three months of behavioral pattern data mapped against threshold lines, the dormant-baseline band highlighted in amber. The current reading sits at the top of the amber band, below the active review threshold line, where it has been for three months. Eighty-seven percent. The ASSESS button in the upper right corner of the panel — the expand option for contextual assessment, the one he has never clicked, the one the Cartograph-7 system has presumably been ready to run for forty-three months. The number has not changed. He did not expect it to change. He does not know what he expected.
He closes the dashboard without running the contextual assessment. This is the fourth consecutive workday he has done this. He is aware of the count. He is not aware of having made a decision to maintain it.
He opens a fresh document and types: Branch 12, September 17, 2035. Eighty-seven percent. Forty-three months of observation. Contextual assessment not run.
He stops.
The cursor blinks in the place where reason would go if he were going to write it. He has not written it in forty-three months. He has the number. He has the trend. He has the decision not to run the assessment, made and remade on approximately nine hundred working days, each time without notation.
He has not written the reason because he is not sure the reason is something he can write without it becoming the kind of claim that requires evidence, and the evidence he has is a number that means two different things and he cannot tell which one it means.
He saves the file.
He takes his coat off. He sits down. He looks at the closed dashboard for a moment and then opens a different window — one of the open casework files that does not involve Branch 12. He reads the first paragraph. His eyes move across the words. He is still thinking about the ninety-three percent.
At eleven twenty-three, an automated report arrives in his queue from the Cartograph-7 system. It is not about Branch 12. It is a routine threshold alert on Branch 7, which has crossed into active review territory. The alert is formatted exactly as designed: case number, branch designation, threshold value, recommended action, estimated workload, links to relevant compliance documentation. The Cartograph-7 system's automated reporting has not missed a threshold crossing in the Circuit Mile's regional jurisdiction since 2033. It is, by that measure, reliable.
Branch 7 crossed threshold cleanly, in the way threshold crossings are supposed to happen. The path from there is known: formal interpretability assessment, documented methodology, a result the court system will accept because the certification chain is intact.
Branch 12 has never crossed threshold. It has climbed steadily and stayed below. The system has classified it as dormant-baseline throughout, correctly according to the threshold definitions, and has therefore never flagged it for formal review. There is no case number for Branch 12. There are Marcus's notes and the trend and a number that is the highest he has recorded and the contextual assessment that would tell him what the system thinks the pattern means, sitting there in the interface, one click from the cursor position.
He closes the Branch 7 alert. He will address it this afternoon.
He opens the Branch 12 dashboard again. Eighty-seven percent. The ASSESS button in the upper right corner. He moved his cursor toward it yesterday for the first time in forty-three months and stopped and wrote that fact in his notebook.
He looks at it for a moment.
He closes the dashboard.
He writes in the document: the contextual assessment remains unrun. He writes: this is the forty-fourth month of observation without interpretation. He writes: I do not know if the decision not to run it is the correct decision or the decision that has been correct for so long that I cannot see what it would mean to change it.
He saves the file.
He gets up and goes to get coffee. The Circuit Mile hums around him in the way it does at this hour, full of people doing work that requires being in the same place as the systems they are watching, even though everything could technically be done remotely, because watching at a distance is not the same as watching. The coffee station is near the east window; he can see the transit band from there, the automated freight-carriages moving at the commercial interval, the certification authority building two blocks south. He knows most of the faces by professional subtype: the Empiricists stand facing the window with their coffee, looking out; the Sensitives stand facing inward, watching the room. He has worked among them long enough to have noticed the pattern without asking anyone about it. He has never asked a colleague which subtype they think he is.
He pours his coffee and goes back to his desk.
He opens the Branch 12 dashboard one more time.
Eighty-seven percent.
He thinks about what the regional director said: good catch on the eighty-seven. He thinks about whether a catch is the right word for something you notice but cannot name. He thinks about the four seconds the exchange took and what it would have taken to fill them with the rest of what he knows.
He closes the dashboard.
He picks up the Branch 7 case file and begins to read.