A new survey names four open problems in AI driven battery health. The biggest, the authors find, is that no one will run a battery to failure to teach a model what failure looks like.
Every rechargeable battery the reader owns, from a phone to an EV to a grid storage pack, is silently aging. The science of predicting how long one will last and how it will fail is called Battery Prognostics and Health Management, and it has spent a decade trying to hand its hardest problem to a new class of AI. Not the large language models in chatbots, but Transformer-based models pre-trained on unlabeled battery telemetry, then fine-tuned for narrow health-management tasks. A new survey of that work, posted this month to arXiv, argues the bottleneck is no longer the models. It is the data.
The survey, by a multi-institution team, calls itself the first comprehensive review of "Large Model" applications in battery health. The terminology needs a doorway: in this paper, "Large Model" means Transformer-style pre-trained models in general, not the chat-style LLMs a non-beat reader is most likely picturing. Same architectural family, different training scale, different product. Treating them as the same thing would mislead the entire read.
The survey delivers a four-dimension taxonomy of where the field has moved: mitigating data scarcity; enhancing generalization and robustness; integrating domain knowledge for interpretability; and enabling system-level automation. Those four lines are how the authors sort every recent paper, and they are also where the open problems cluster. The first dimension and the last one do most of the work.
Battery data is structurally scarce for a reason the survey surfaces but does not solve. To teach a model what a failing cell looks like, someone has to run a real pack until it dies, through hundreds or thousands of charge-discharge cycles while logging every temperature, voltage, and impedance trace. No fleet operator, OEM, or cell maker wants to run a fleet of vehicles or a bank of grid batteries until they fail, just to produce that record. Lab data exists, but it does not generalize to the cells, packs, and duty cycles real products use. The survey's proposed fix is a "collaborative data ecosystem," a polite academic phrase for an industry-wide data pool that almost no commercial player has an incentive to build.
The remaining three open problems are the field's workarounds. Enhancing generalization is what researchers do while waiting for that data pool. Integrating domain knowledge (known physics of cell aging, equivalent-circuit models, electrochemical signatures) is how the field tries to make current models interpretable enough to trust. System-level automation is the deployment target: battery management systems that monitor, decide, and act across the full lifecycle of a pack without a human in the loop. Each is named as a remaining challenge, not a solved problem. The survey's promise is that the field now has a shared vocabulary for what it is trying to do, not that it has done it.
The paper is a preprint, not a peer-reviewed article, and several claims are positioning rather than measurement. The "first comprehensive survey" label is the authors' own framing; "comprehensive" is a category claim, not a verified count. The abstract contains no headline performance numbers, so any analysis of the work has to lean on the structural taxonomy and the roadmap, not on a benchmark. Commercial battery analytics and diagnostics products exist in this space, but they serve only as industry color here, not as independent confirmation of the survey's claims.
The watch item is concrete. If a major OEM, cell maker, or grid operator publishes a federated or pooled battery-failure dataset in the next 12 to 18 months, the four open problems the survey names will start moving in a measurable direction. Until then, the field's roadmap reads as a forecast that names its own obstacle. The data the models need is the data no one wants to generate.