Teacher-rendered Stage 1 voxel data
Open-world representation research
Language-grounded Minecraft geometry.
CLMCP distills OpenAI CLIP vision features into a voxel encoder, then exposes that semantic space to a Minecraft Agent for retrieval, planning, and navigation.
Introduced after the training-world version
Successful matched navigation episode
Choose a research surface
LIVE RETRIEVAL / AUTHORITATIVE REPLAY
DEMO 01LIVE MODEL

Text → Voxel Search
Describe a Minecraft scene in English. The frozen CLIP text tower and a selected CLMCP encoder rank the same 768-scene gallery by cosine similarity.
MODEL DROPDOWNREAL VOXELS360° WEBGL
Open retrieval surface↗
DEMO 02MATCHED EPISODE

Agent Benchmark Replay
Inspect the full native Agent trace for a GLM-5.3-Flash matched distractor-house episode where CLMCP followed a nearly identical route, entered the watchtower ground floor 7.4% earlier, and used 66.1% fewer tokens.
CROSS-VERSION ZERO-SHOTAGENT TRACEFLAGSHIP VS BASELINE
Open replay surface→