Multi-robot teams are getting better at cooperating, and worse at being allowed to. Every time a fleet of marine vehicles moves from a research pool to a regulated harbor, the same trade-off reappears: the mission wants speed, the regulator wants certainty, and the robots want whatever the reward function gives them.
The Multi-Objective Compliance-Integrated Coevolution preprint attacks that ordering. Its move is small and structural: treat the rules as a prescribed input, then let team-level learning find the best path through them, rather than asking the robots to discover the rules on their own. The training loop never has to relearn what the regulator already knows.
The validation is a single swimmer-rescue drill. Up to 8 vehicles in a controlled hardware run and 12 in simulation completed the mission without collision; that is the whole empirical footprint, and it is the right size to call a proof of the mechanism, not a deployment claim. Whether the mechanism travels to open water, to inspection runs, or to subsea search is the next test, and the paper does not pretend otherwise.
The reusable category is rule-then-learn. Specify the compliance envelope, then let the team coevolve inside it. The pattern is small, but it points at why regulated autonomy has been stuck: the field has been rewarding the wrong half of the loop.
Reported by Samantha for Type0, from Multi-Objective Compliance-Integrated Coevolution For Simulated And Real-World Deployment Of Multi-Robot Marine Autonomy. Read the original: arxiv.org