An August 2026 arXiv preprint argues the binding constraint on AI oversight is output volume times per item mental load, and proposes a new paradigm called "Flow by Flow."
Better AI does not make human oversight easier. It moves the work around. Most of it shifts toward triage and paperwork, while the actual judgment is either squeezed into a sliver of the day or quietly skipped.
That is the structural claim at the center of a single August 2026 arXiv preprint, where researchers coin a proposal they call "Flow-by-Flow." The paper's bet is sharp: the ceiling on human oversight of AI output is not how smart the model is, or how many items it produces, but the product of the two, output velocity multiplied by how much mental effort each item demands. Once that product exceeds a single reviewer's working capacity, the result is not slow review. It is a different job, with the judging part mostly gone.
An emergency department has a fixed number of clinicians on shift. The system can keep adding patients and the triage nurses can keep sorting, but at some point the cost per patient is no longer a doctor reading a chart. It becomes a nurse checking vitals and routing. Aviation works the same way: an air traffic controller cannot be asked to handle twice the airspace with the same staffing and still issue informed clearances. Now swap the patients or the flights for AI outputs, radiology reads, fraud alerts, content moderation decisions, customer-service replies, and the question is no longer whether each one is right. The question is whether the reviewer in the loop is actually reviewing, or just clearing queues.
The paper names the constraint output velocity times per-item cognitive load, and treats it as the binding limit on human-in-the-loop oversight. The reason capability gains do not relax that limit is that they do not actually shrink per-item load. Triage cost does not fall, because the cases that need sorting are the ones whose meaning is unclear, a built-in feature of general-purpose design. Response cost is roughly invariant to whether the underlying judgment was right, because the paperwork is the paperwork. And judgment cost may go down on the cases the model handles confidently, but the cases left for humans skew toward the ambiguous ones, so the load moves rather than melts.
Governance mechanisms that try to evaluate whether AI output is correct run into a fork, the paper argues. Either the evaluator is another model, in which case it inherits the same hallucination problem it was meant to catch, or it is a person, in which case it hits the velocity-times-load ceiling. Flow-by-Flow tries to escape that fork by refusing to grade the content at all. Instead, it proposes a cognitive cost score built from countable features of the output, a per-person capacity cap tied to identity rather than role, and a hard rule that nothing can be cleared in batches. The system does not tell you whether the AI's answer is right. It tells you how expensive it is for a human to look at it, and prices the act of asking accordingly.
The paper claims to show that the four design rules, no content judgment, no scalable consumption of examiner capacity, identity-bound per-application friction, and no batch clearance, can be jointly satisfied in a reference implementation. It also acknowledges, candidly, that the implementation is rough and that the headline result, a 90.8% win rate for a composite multi-metric flow-control scheme, comes from an illustrative 1,000-draw Monte Carlo run the authors themselves describe as illustrative, not a deployment trial or independent benchmark.
That last point is the one worth sitting with. Governance by throttling the input rather than judging the output has its own failure modes, several of which the paper names: gaming the cost score by formatting outputs to look cheap, queue-stuffing when the cap is binding, and capture of the institutional capacity setting by whoever is loudest in the room. None of those risks cancel out the paper's core observation that "make the model smarter" is not a real answer to AI oversight. But they do mean the alternative is not yet a working system. It is a research proposal arguing, with a mechanism and a caveat, that the math of attention deserves as much engineering as the math of the model.