An arXiv preprint (an unreviewed paper posted to a public research server) uses Koopman spectral analysis (a dynamical systems math method that reads a nonlinear system by its hidden rhythm) to read multi agent LLM (large language model) debates as
When a panel of language models argues a question, when does the argument actually end, and which models carried the room? A new arXiv preprint, "Certifying Collective Reasoning in Multi-Agent Systems via Koopman Spectral Analysis", reads the debate as a single dynamical system and proposes a math method that answers both questions before the discussion is over.
Koopman spectral analysis is a way of taking a complex, nonlinear system, here a group of LLMs exchanging messages on a graph, and reading its hidden rhythm. The resulting spectrum yields three machine-checkable certificates. A sub-dominant eigenvalue, called λ₂, fixes the debate's intrinsic timescale and produces a convergence deadline that can be computed before the run finishes. Its eigenvector names the coherent factions the collective split into, with the magnitude |λ₂| gating whether that post-hoc story is trustworthy. And a small set of leading spectral coordinates, eight of thirty-two, preserves the final decision at 99.7% fidelity, giving auditors a compressed message basis to inspect.
On the authors' benchmark, an attention-consensus model, the predicted deadline tracked observed convergence with a 0.93 log-log correlation and bounded it in 96% of 24 configurations. A certificate learned from 15 debates generalized to 60 of 60 held-out debates. The analysis runs in minutes on a CPU.
What the paper does not yet claim is fielded deployment. The numbers sit on the authors' own model rather than on production multi-agent systems, attribution is exact only when the spectrum itself certifies metastability, and the audit's failure modes, when the certificate does not apply, are still an open question.