Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
Thinking Machines Lab has launched Inkling-Small, an open weights Mixture-of-Experts mannequin with 276B complete parameters and 12B lively. That is a couple of quarter the dimensions of Inkling, which carries 975B complete and 41B lively parameters. The mannequin was educated on NVIDIA GB300 NVL72 methods. It causes natively over textual content, photos and audio. The context window reaches 1M tokens, and thinking effort is adjustable. Weights ship beneath Apache 2.0 on Hugging Face.
Is it deployable
Yes, and the quantized checkpoint is why. Per the model card, the BF16 checkpoint wants a minimum of 600 GB of aggregated VRAM. That is met by 4x NVIDIA B300 or 8x NVIDIA H200. The NVFP4 checkpoint drops that ground to 180 GB. It runs W4A4 on a single B300, which requires SM100+, or W4A16 on two H200s. Supported runtimes are SGLang, vLLM, TokenVelocity, Unsloth and Hugging Face.
That single-GPU path strikes a 276B mannequin out of frontier-lab territory. Startups can self-host on one rented B300 occasion. Mid-size enterprises with current H200 capability can serve it with out new {hardware}. Regulated sectors acquire a private-weights choice: monetary companies, healthcare operations, insurance coverage, telecom and public sector. Applicable workloads embody coding brokers, terminal automation, and doc and chart understanding. Audio widens that to call-center analytics, voice interfaces and assembly summarization.
Architecture
Inkling-Small is a 42-layer decoder-only transformer with a sparse MoE feed-forward spine. Each token routes to 6 of 256 specialists, plus 2 shared specialists lively on each token. Attention is a hybrid of native and world layers. The mannequin is encoder-free and natively multimodal. Images are divided into 40×40-pixel patches and remodeled utilizing a four-layer hMLP. Audio is represented as dMel spectrograms. Both move by a light-weight embedding layer and are processed collectively with textual content tokens. Numerics help covers BF16, MXFP8 and NVFP4. Audio enter is WAV at 16 kHz, ideally beneath two minutes. Output is textual content solely.
Inkling-Small started coaching after its bigger counterpart. That let the analysis crew revise the pre-training knowledge combine and the machine studying recipe. The analysis crew post-trained an earlier checkpoint, Inkling-Small (preview), partly utilizing on-policy distillation with Inkling because the instructor. From that checkpoint, it continued scaling agentic coding RL for 2 weeks.
Benchmark outcomes
The smaller mannequin surpasses its instructor on reasoning and agentic coding. On Humanity’s Last Exam (textual content solely) Inkling-Small scores 31.6%, forward of Inkling’s 29.7%. SWE-bench Verified is 80.2% versus 77.6%, utilizing a bash-only harness. Terminal-Bench 2.1 reaches 64.7% at finest harness. Toolathlon Verified is 54.4%, towards Inkling’s 45.5%. GPQA Diamond is 89.5%, AIME 2026 is 95.5%, and IFBench is 82.2%. ARC-AGI-2 rises to 40.1% from Inkling’s 36.5%.
SimpleQA Verified falls to 20.6% from Inkling’s 43.9%, and the AA Omniscience index drops to -9.0 from 2.1. Tau 3 Banking is 15.5% versus Inkling’s 23.7%. All evaluations ran at effort 0.99 and temperature 1.0, with a 256K max-token trajectory restrict on coding evals. External scores are sourced from Artificial Analysis, Scale AI and ARC Prize.
Multimodality, epistemics and security
Multimodal scores keep near Inkling at decrease value. MMMU Pro is 74.0%. CharXiv RQ is 77.4%, rising to 81.3% when the mannequin makes use of Python to crop, zoom and examine charts programmatically. Audio MC is 54.9%, MMAU is 77.0%, and VoiceBench is 90.1%.
On epistemics, calibration was educated with RL towards correct scoring guidelines on a big corpus of real-world forecasting questions. ForecastBench with out search offers a Brier Index of 61.3 ± 0.46, forward of Inkling’s 60.1 ± 0.54. On security, StrongREJECT is 98.4%, FORTRESS adversarial is 71.6%, and FORTRESS benign is 96.9%. Thinking Machines Lab concluded the mannequin presents no materials uplift past the present open-weight ecosystem. It recommends layering downstream moderation similar to Llama Guard on consumer-facing deployments.
Both fashions can be found on Tinker with a limited-time low cost. Text, picture and audio chat run on Tinker Playground.
Interactive explainer
The embed beneath breaks the discharge into 4 interactive components. Tab one animates how sparse routing prompts 8 of 258 specialists per layer. Tab two ranks Inkling-Small towards comparable open weights fashions on ten benchmarks. Tab three traces how the three enter modalities converge into one decoder. Tab 4 sizes the {hardware} every checkpoint format requires.
Key Takeaways
- Inkling-Small is a 276B complete, 12B lively MoE mannequin beneath Apache 2.0.
- It beats the 975B Inkling on HLE, SWE-bench Verified, Terminal-Bench 2.1 and ARC-AGI-2.
- The NVFP4 checkpoint runs on a single B300 at 180 GB aggregated VRAM.
- Factual recall regressed: SimpleQA Verified 20.6% versus Inkling’s 43.9%.
- Native textual content, picture and audio enter with a 1M token context window.
Check out the Technical details and Model weight. Also, be happy to observe us on Twitter and don’t neglect to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to associate with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so forth.? Connect with us
The submit Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model appeared first on MarkTechPost.
