Who is betting on what in AI, and against whom.
Rows are camps, columns are the questions. Product categories (chatbots, robots, video) cut across these; the questions are what the camps actually disagree about.
| SignalWhat does the model learn from? | WhenWhen does learning happen? | RepresentationWhat is the internal representation? | Acts inWhere does it act? | |
|---|---|---|---|---|
| Scale next-token prediction | Text, then verifiable rewards and human preference Internet-scale textVerifiable outcome rewardsHuman preference |
Offline pretrain, then freeze; memory is a retrieval layer Pretrain once, then freeze |
Tokens over a transformer Tokens over a transformer |
Software: chat, code, browsers, APIs Software |
| World models | Self-supervised prediction of video and latent state Video and sensor streams |
Offline pretraining, but built to plan at inference by rolling out imagined futures Pretrain once, then freezePlan at inference by rolling out imagined futures |
Latent embeddings (JEPA) or generated pixels (video models). These sub-camps disagree. Latent embeddings (JEPA)Generated pixels (video models) |
Simulated and physical environments; robotics and driving downstream Simulated environmentsThe physical world |
| Physics AI and AI for science | Simulation data plus the governing equations themselves, which act as a checkable reward Simulation dataGoverning equations as a checkable reward |
Offline pretraining, then optimization loops at inference: simulate, take a gradient, redesign Pretrain once, then freezeOptimisation loops at inference |
Continuous fields in 3D plus time; neural operators, resolution-invariant Continuous fields in 3D plus time |
Design and discovery loops: materials, fusion, weather, devices, drugs Design and discovery loops |
| Learn from experience | Scalar reward from interaction, plus self-generated subgoals (options) Scalar reward from interaction |
Continually, at runtime, forever. No train/deploy boundary Continually at runtime, no train/deploy boundary |
Whatever the agent builds online: learned features and options, no replay buffer Features and options built online |
Any environment, ultimately embodied. Target: 20 watts The physical worldSimulated environments |
| Physical AI and robotics | Teleop demonstrations, human video, sim-to-real, then on-robot RL Teleoperated demonstrations and human videoSimulation dataScalar reward from interaction |
Offline, with on-robot RL fine-tuning starting to matter Offline, then on-robot fine-tuning |
Mostly camp 1: a vision-language model with a diffusion or flow action head Vision-language model with an action head |
Actuators in the real world The physical world |
| Program search and other substrates | Task success on novel abstraction (ARC), free-energy minimization, symbolic consistency Task success on novel problemsFree-energy minimization |
Often at test time: search and adapt per task rather than pretrain once Search and adapt per task at test time |
Explicit programs, symbols or probabilistic models; deep nets only as a guide for search Explicit programs and symbolsProbabilistic models |
Mostly abstract problems today; some robotics via active inference Abstract problemsThe physical world |
texthuman-preferenceverifiable-rewardvideosimulationphysical-lawscalar-rewarddemonstrationstask-successfree-energyofflineinference-planninginference-optimisationon-robot-finetunecontinualtest-timetokenslatentpixelsfieldsonline-featuresvlaprogramsprobabilisticsoftwaresimulationphysicaldesign-loopsabstract