Physics-native world foundation model

The first complete end-to-end model stack.

Confidential
August 2026

TL;DR

01

Trillion-dollar infrastructure

World simulation will become the trillion-dollar infrastructure layer for physical AI.

02

Physics-native convergence

The field is converging on physics-native foundation models, but the full stack remains unsolved.

03

Generational lead

Viggle is the first to train a complete end-to-end physics-native world foundation model.

04

Compounding moat

Our proprietary stack, social-scale data flywheel, and proven scaling capability compound that lead.

Physical AI has a quadrillion-year experience gap

Physical AI can only scale through interactive learning in a world simulator.

LLM World modelexperience engine
Physical AIrobot brain + hardware Embodied AI
109 living beings in parallel
×
106 years of evolution
=
1015 years of world experience

Biological intelligence was shaped by this. Physical AI has no scalable way to acquire it.

Real robot pre-training

Not a viable path at this scale.

Accurate + efficient world simulator

The only path that scales.

The real test of a world model is generality

The endgame is a full-stack physics-native model. Others add 3D only at the output.

Tokenization Architecture Data Engine Pre-training Post-training
JST-2
World Labs / SpAItial AI
Video models

physics-native partial none

Adjacent approaches still leave world simulation unsolved.

JEPA

A learning architecture, not a complete world simulator.

Tokenization, data, and end-to-end training for world simulation remain unsolved.

Meshy / Tripo

Built to generate individual mesh objects for traditional game engines.

Their generality stops at individual objects, not physical, dynamic worlds.

Embodied AI

Real-world learning is slow, expensive, and embodiment-specific.

As hardware evolves, data loses transferability; intelligence and embodiment must be iterated together inside a world simulator.

JST unlocks a new scaling law for world simulation Video models learn pixel correlations. Physics-native foundation models learn the world's underlying structure, delivering higher accuracy while reducing cost by orders of magnitude.

Model performance vs. inference cost
Pixel-based and physics-native scaling laws for world models Cost → Model performance → pixel-based scaling law JST-2 full-stack physics-native World Labs / SpAItial AI 3D only at the output Large video models higher quality, prohibitive cost Meshy / Tripo static objects only Small video models low cost, low coherence
Pixel and reconstruction models share one cost-quality frontier. Physics-native foundation models open a new one by changing the representation, not the implementation.

JST-2 crossed the first generality threshold

Generality begins with one foundation model spanning characters, motion, and scenes.

JST-2 Meshy Tripo Cartwheel Hunyuan Marble SpAItial AI Character Image-to-character Rigging Motion Text-to-motion Video-to-motion Scene Text-to-scene Image-to-scene

Only JST-2 covers all three primitives and all six foundational tasks.

One checkpoint, across-the-board SOTA

JST-2 beats every specialist benchmarked, with no task-specific fine-tuning.

Image-to-character vs Tripo 3.04.5CharacterFidelityFacial RealismMaterialRealismClothing &Detail Fidelity
Text-to-motion vs Hunyuan 3.04.5Dailylocomotion& postureDaily object& self-careFast combat& stuntsPerformance& danceSocial &healthinteractionSportsUpper body& handgesture
Rigging vs Meshy 3.04.5Joint AccuracyBody &OutfitAdaptabilitySpeedJST-2 45sMeshy 7mins
Video-to-motion vs Cartwheel 3.04.5Motion AccuracyGlobalTranslation&OrientationRobustness toComplex/FastMotionTemporalSmoothness

JST-2 specialist

JST-2 demo

Video placeholder

Characters, motion, and scenes from one foundation model

One architecture scales to physical intelligence

Each generation advances the same model across graphics, physics, and reasoning.

JST-1Q4 2022 – Q2 2024 JST-2Q2 2024 – Q2 2026 JST-3Q2 2026 – Q4 2027 JST-XQ4 2027 – Q1 2029 Graphics Proof Solved Solved Solved Physics Basic Simulation-ready Solved Solved Reasoning Basic Advanced Solved Unlocks A proven architecture and a data flywheel operating at social scale. One foundation model spanning characters, motion, and scenes. A scalable physics-native world simulator that learns rather than follows hand-coded rules. Physical intelligence that reasons, acts, and adapts in simulated worlds at model speed. AI analogue GPT-12018 · architecture proven GPT-22019 · first $1B investment GPT-32020 · scaling breakthrough Beyond LLMsphysical intelligence

JST-1 proved our architecture and data flywheel

50M+ users improved the model through real-world creation and correction, while the product funded the entire loop.

0 → JST-1Architecture proof · complete model stack

TokenizationArchitectureData enginePre-trainingPost-training

JST-1 → JST-2Social-scale data flywheel

MODELJST-1PRODUCTViggleUSERDATADeployPut an efficient model in millions of users’ hands.CreateCreators explore diversity no lab can predefine.CorrectReal use exposes and corrects model failures.ImproveFeedback becomes training data for the next model.
A self-sustaining AI flywheel
50M+

users

300M+

videos generated

$2M

ARR

$1.2M

total annual cost

$0

paid acquisition

JST-3: from rule-based to learned physics

The first physics-native world model to learn physical dynamics directly.

JST-2 → JST-3Proven architecture · 10× to 100× model scale

Social-scale interaction exposes physical edge casesGamified consumerproductsDiverse interactiveworldsCreation, interaction,correctionThe data engine scales physical experienceSynthetic sceneand dynamics dataReal-to-sim forreal-world diversitySelf-augmentationat model scaleJST-2JST-3Learned physics10× to 100×model scale10× data10× compute

Model-as-a-Service scales to $60M ARR

One foundation model monetized across consumer, developer, and enterprise markets.

Consumer

Products monetize distribution and expand the data flywheel.

Prosumer + developers

Subscriptions and APIs monetize creators, studios, and developers.

Enterprise

Licenses monetize high-volume platform and publisher deployments.
6× ARRin twelve monthsDec 2026 target$10M ARRDec 2027 target$60M ARRConsumerProsumer +developersEnterprise

Builders of the first physics-native world model

Four years building the full stack: tokenization, architecture, data engine, pre-training, and post-training.

Hang Chu

Hang Chu, CEO

Leading researcher in 3D generative models for 10+ years.

  • Former Principal Researcher, Autodesk AI Lab.
  • Founding member, NVIDIA Toronto AI Lab.
  • First author, Modular Codec Avatar, Meta.
  • PhD ABD, University of Toronto ML Group (led by Geoffrey Hinton); Cornell MS.
  • 2.7K+ citations; h-index 19.
Ming Liang

Ming Liang, CTO

Researcher and systems builder for 3D foundation models.

  • Founding member and staff scientist, Uber ATG.
  • Core founding member, Waabi.
  • Founding member, Apple SPG.
  • Kaggle Data Science Bowl World Champion ($1M, one of Kaggle’s largest-ever prize pools).
  • Tsinghua PhD; 12.0K+ citations; h-index 31.

Core team

Jinma
JinmaData LeadTsinghua MS; 10+ years as an ML engineer
QQ
QQModel LeadFounding member, SenseTime; 4.0K+ citations
Yun
YunChief ScientistFirst author, CVPR Best Paper Finalist (top 0.2%)
SharpRuntime LeadNVIDIA CUDA team alumnus
Yao
YaoInfra LeadFormer Senior Engineer, ByteDance
Jason
JasonProduct LeadFormer AI startup founder
Nan
NanGrowth LeadMaster’s degrees from USC and LSE
KD
KDDesign LeadParsons; formerly at Google
Renjie
RenjieTheory AdvisorProfessor, UBC
Eric
EricStrategy AdvisorFormer CTO, Ctrip; formerly at eBay
Christine
ChristineGTM AdvisorFirst marketer at Twitter
Andre
AndreLegal AdvisorPartner, Osler