webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware
webAI has launched TwIL-LM, a two-model family of formal-logic reasoners at 1.7B and 3B parameters. The 3B member, TwIL-LM3, is a merged fine-tune of SmolLM3-3B; the 1.7B member is a PEFT LoRA adapter for SmolLM2-1.7B-Instruct. Both goal autoformalization: translating English into first-order logic and checking whether or not a conclusion follows from its premises. Both run regionally, with a 1.06 GB quantized construct for the 1.7B and a 1.78 GiB Q4_K_M GGUF for the 3B. webAI’s announcement frames the discharge round beating gpt-oss-120b on 4 of 5 formal-reasoning lanes.
Is it deployable?
Partially. Non-commercial use solely, as of now.
Both checkpoints ship beneath the webAI Non-Commercial License ver. 1.0. Revenue-generating deployment requires a separate settlement with webAI.
- Company degree: any dimension. The 3B
Q4_K_MGGUF is 1.78 GiB and runs on CPU or 4 GB of VRAM. The 1.7BQ4_K_Mis 1.06 GB. - Industries: compliance and RegTech, monetary providers, healthcare and pharma, authorized and contract operations, formal-methods analysis. webAI positions native execution for environments the place knowledge can’t go away the system.
- Applications: first-order logic (FOL) translation, entailment classification over premise units, pure language to structured question, Lean formalization drafting and critique, and a verifier layer that checks a bigger mannequin’s output.
How TwIL-LM3 was constructed?
Four phases sit on prime of the bottom mannequin. LoRA supervised fine-tuning on an artificial formal-logic corpus. Checkpoint fusion, averaging intermediate SFT checkpoints in parameter area. WiSE-FT interpolation again towards the pretrained base at λ = 0.25. Then MGPO, an entropy-weighted GRPO stage run towards a programmatic verifier. The printed checkpoint is step 2071.
That λ is load-bearing: solely 1 / 4 of the fine-tuned delta is retained. A sibling arm that skipped the interpolation scored increased in-domain, at macro gate 0.515, however gave again roughly twelve factors of held-out functionality. webAI didn’t publish that arm.
Performance
webAI's announcement lists 96.4 on rule induction, 87.6 on semantic parsing, 64.6 on Lean formalization, 52.0 on exact-format answering, and 68.7 on entailment labeling.
It stories two tracks. On Track A, in-domain formal logic, TwIL-LM3 scores 0.4488 on the six-lane common and 0.4218 on the macro gate, the metric the coaching pipeline gates on. It leads each arm as much as and together with LFM2.5-8B-A1B on all six goal lanes, at 0.4218 towards 0.3757 with a 3rd of the parameters. It doesn't lead the 2 largest arms. Qwen3-8B takes the gate 0.5336 to 0.4218, however most of that's loose-match credit score; beneath strict-7 the 2 sit at 0.2093 and 0.1971. gpt-oss-120b takes the six-lane common 0.5192 to 0.4488.
Efficiency is the place the mannequin card is unambiguous. TwIL-LM3 produces the shortest generations of any arm, 482 tokens on Track B, and consequently probably the most solutions per second at 32.9 towards the 120B's 4.2.

Held-out switch
TwIL-LM3 improves in-domain by +26% relative, macro gate 0.336 to 0.422, whereas additionally gaining +0.022 on the held-out core common. The mannequin card calls it the one arm within the venture that good points on each tracks. LogicBench strikes to 0.7167 from 0.6467. GSM8K slips barely to 0.8733 from 0.8833, and IFEval regresses to 0.6433 from 0.6767.
The 1.7B is a special commerce. Its macro-primary rating is 0.361 towards 0.185 for the unadapted base. Out-of-distribution outcomes are blended: LogicBench BQA improves to 0.590 from 0.563, whereas GSM8K falls to 0.380 from 0.413 and ARC-C chain-of-thought falls to 0.463 from 0.587.
Key Takeaways
- TwIL-LM3 (3B) and TwIL-LM (1.7B) goal formal logic, each beneath a non-commercial license.
- Shipping TwIL-LM3 trails gpt-oss-120b on the six-lane common, 0.4488 to 0.5192.
- Its actual edge is effectivity: 32.9 solutions/sec from 482-token generations.
- WiSE-FT at λ = 0.25 is why in-domain good points don't collapse held-out efficiency.
Check out the Model weights and Technical details. Also, be happy to observe us on Twitter and don’t overlook to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
Need to companion with us for selling your GitHub Repo OR Hugging Face Page OR Product Release OR Webinar and so on.? Connect with us
The submit webAI Releases TwIL-LM: A 1.7B and 3B Formal-Logic Model Family for Autoformalization on Local Hardware appeared first on MarkTechPost.
