MineChoice: experimental stone-pickaxe skill classifier
This model has NOT beaten Minecraft. It is not a speedrunning model. This is an initial public research checkpoint toward that goal, trained locally on an RTX 3090. It is independent of TypeSafe/Jev and Mojang/Microsoft.
What is actually trained
A 19,592-parameter MLP (15 inputs, two 128-unit ReLU hidden layers, eight categorical outputs), trained from scratch with cross-entropy on real recorded Minecraft 1.20.4 teacher decisions. This is a numerical state-to-skill classifier, not a text model or an LLM prompt wrapper. The coded teacher is only for data collection. Learned-policy rollouts execute the neural model's choices without teacher fallback. A deterministic feasibility mask removes unavailable skills.
Training rows: 99; validation rows: 11. Training episodes: 9; validation episodes: 1. Validation imitation accuracy: 1.000. These tiny-data metrics do not establish natural-world generalization.
Actual gameplay evaluation
Held-out controlled arenas: 9/9 stone-pickaxe successes. Mean time with 180-second failure penalty: 22.69 seconds. Full-game evaluation: not performed; dragon wins demonstrated: zero.
Arena layouts use separate seeds from collection. These are layout seeds in a shared superflat world, not unseen Minecraft world-generation seeds. Server commands build arenas and provide varied starting inventories. During each episode the bot uses ordinary survival controls and crafting. No arena outcome is presented as a full survival-game completion.
Runtime
Native Node.js CPU inference p50 0.0224 ms / p95 0.0364 ms on the development machine. This excludes observation collection, navigation, and Minecraft ticks. See latency.json for the scope. No hosted inference needed. The latest live run, after caching nearby blocks, measured observation plus inference p95 0.512 ms. This excludes the initial per-episode block scan and skill execution. Raw measurements for both pre-optimization and post-optimization runs are included.
import fs from 'node:fs';
import {predict} from './src/policy.mjs';
const model = JSON.parse(fs.readFileSync('policy.json', 'utf8'));
const decision = predict({
logs: 1, planks: 0, sticks: 0, tables: 0, wooden_pickaxe: 0,
cobblestone: 0, stone_pickaxe: 0, table_near: 0, log_near: 1,
stone_near: 1, health: 20, food: 20, log_distance: 5,
stone_distance: 8, table_distance: 64
}, model);
console.log(decision); // action, confidence, probabilities
See SOURCE_README.md for collection/training and local game setup. model.pt contains state_dict/schema; use torch.load(..., weights_only=True). policy.json contains the same weights and explicit feature order/scales for portable inference. model.safetensors and config.json provide a pickle-free weight export and schema.
Data and limitations
The included numerical events were generated by this project's coded teacher and local Minecraft server. They are not human speedruns, VPT demonstrations, synthetic text labels, or Jev outputs. Only completed episodes and successful skill transitions supply training rows. Failures remain in the rollout report. Splits are by episode. Validation selects the epoch. Temperature calibration is disabled when there are fewer than 100 validation rows or five episodes; see training_report.json for its actual status. Validation metrics are not independent calibration estimates. Softmax confidence is not a guarantee of correctness or game-winning probability.
Actions: gather_log, craft_planks, craft_sticks, craft_table, place_table, craft_wooden_pickaxe, mine_stone, craft_stone_pickaxe. Supports oak/stone arena observations only. No Nether, stronghold, End, dragon, food, or combat skills. Further demonstrations, implemented skills, outcome-based optimization and held-out natural-world evaluation are required for the project objective.
MIT covers original source and original model weights; third-party game software and dependencies keep their own terms. No Minecraft binaries, textures, credentials, or account files are included.
- Downloads last month
- 17