A Weizmann Institute team's Brain Interaction Transformer (Brain IT) rebuilds images from one hour of brain scan data, where prior work needed forty, while still misreading common scenes and requiring NVIDIA's high end data center GPUs to train.
A Weizmann Institute manuscript released alongside MIT Technology Review's reporting describes a brain-scan image-reconstruction model that needs one hour of new-subject data to reach fidelity the authors compare to methods using forty hours. The team, publishing as part of MIT Technology Review's The Download newsletter, calls the system Brain-IT: a Brain Interaction Transformer that maps functionally similar voxel clusters shared across subjects into localized image features, then routes them through a structural branch built on VGG features and Deep Image Prior and a semantic branch that guides a diffusion model.
The public code and pretrained checkpoints make the work reproducible in principle, but stage-two training still demands two H200 GPUs or four H100s.
The reporting puts a hard ceiling on the "mind-reading" framing. The training set covers eight people. Reconstructions fail on common scenes: a cake becomes sandwiches, a dog becomes a goat. Independent neuroscientist Tommy Sprague raised two concerns: the high cost of scanning, and the absence of an independent replication of the approach. Dreams, imagined images, and locked-in communication appear in the paper as aspirations, not demonstrated outcomes.
The Download's Oct. 1 issue also carries a small-batteries item; details belong in their own frame.