Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together
Model playing cards report high quality below server-class, full-precision circumstances. Those numbers hardly ever predict how the identical mannequin behaves on a cellphone. This week, Liquid AI launched Pipette. It is an open-source platform for benchmarking basis fashions on edge gadgets, inbuilt partnership with Artificial Analysis as an unbiased methodology validator. Pipette treats on-device habits as a property of the deployed system, not the mannequin in isolation. Its unit of measurement is a full configuration: mannequin + quantization + runtime + system. The launch dataset covers 5 on-device efficiency metrics throughout greater than 1,000 mannequin × quantization × runtime × system × context configurations, spanning 30+ fashions, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to eight,192 tokens. Initial verified outcomes come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra. The sensible declare is testable: two 350M fashions on the similar quantization on the identical cellphone retain 78.4% and 33.8% of decode throughput at 4,096 tokens.
Is it deployable?
Yes, Pipette ships as Apache 2.0 infrastructure (pipette-mgmt, pipette-clients, pipette-scores), a public outcomes dataset, a hosted dashboard, and native iOS and Android benchmark apps. Nothing is waitlisted. Publication of community-submitted outcomes continues to be in beta.
- Which corporations: Any workforce transport a mannequin onto {hardware} it doesn’t personal. Solo builders and seed-stage startups can use the dashboard and apps with out infrastructure. Mid-market product groups can run the purchasers throughout an inside system fleet. Large OEMs, chip distributors and enterprises can function the entire pipeline behind their very own firewall.
- Industries: Consumer electronics and smartphone OEMs, automotive, industrial and robotics, healthcare gadgets, monetary providers, protection — anyplace latency, privateness or connectivity forces inference onto the system.
- Applications: Model and quantization choice earlier than a dash commits; SoC and {hardware} procurement validation; regression testing when a runtime, OS or driver updates; context-length capability planning; unbiased verification of vendor efficiency claims.
What Liquid AI shipped
Liquid AI launched Pipette in partnership with Artificial Analysis, an unbiased validator that reviewed and verified the methodology. The premise is slim and helpful: on-device habits is a property of the deployed system, not of the mannequin in isolation.
The launch dataset covers 5 on-device efficiency metrics throughout greater than 1,000 mannequin × quantization × runtime × system × context configurations. It spans 30+ fashions, a number of quantization codecs, llama.cpp builds for macOS, iOS, Windows and Android, and context lengths from 256 to eight,192 tokens. Initial printed outcomes come from a MacBook Pro with M5 Max, an iPhone 17 Pro and a Galaxy S26 Ultra, with AMD Ryzen AI Max+ 395 and Radeon 8060S outcomes listed as coming quickly.
In Pipette, the unit of measurement is a deployment configuration: mannequin + quantization + runtime + system. A benchmark then defines the metric and token form, producing a latency, throughput or reminiscence consequence. Quality is tracked individually on IFBench, GPQA Diamond and MATH-500. Those high quality scores at present come from llama.cpp analysis runs on NVIDIA H100 80GB reference techniques, then get matched to on-device runs sharing the identical mannequin and quantization — a top quality quantity proven subsequent to cellphone throughput was not produced on the cellphone.
Why the deployment context adjustments the reply
Four printed comparisons present how far a configuration can transfer a call:
- Context scaling can diverge at an identical parameter counts. At Q4_K_M on Galaxy S26 Ultra, Granite-4.0-H-350M retains 78.4% of its decode throughput from 256 to 4,096 enter tokens, whereas Granite-4.0-350M retains solely 33.8%.
- Sparse activation buys velocity, not reminiscence. At 2,048 enter tokens on the identical cellphone, LFM2.5-8B-A1B decodes 2.4x sooner than Qwen3.5-4B and 2.6x sooner than Ministral-3-3B-Instruct-2512. It prompts 1.5B of 8.5B parameters per token, but nonetheless peaks at 5.29 GiB as a result of all skilled weights occupy reminiscence.
- Speed and high quality don’t co-locate. On iPhone 17 Pro at Q4_K_M, MiniCPM5-1B completes a 2,048-in / 256-out workload in 3.47 seconds versus 4.12 seconds for LFM2.5-1.2B-Instruct, a 15.8% discount in elapsed time. On the identical artifacts, LFM scores 9.0 factors larger on MATH-500.
- Near-identical system profiles can cover task-level reversals. At Q4_K_M and 2,048 enter tokens on M5 Max, Granite-4.1-8B and Ministral-3-8B-Instruct-2512 differ by 2.4% in decode throughput and 1.2% in peak RAM. Granite leads IFBench by 7.3 factors; Ministral leads GPQA Diamond by 14.0 factors.
How the measurements are produced
Performance runs observe a published methodology: fastened token shapes, grasping decoding, a discarded warm-up, 5 measured repetitions and readiness gating. Before every timed repetition, a platform-specific test verifies thermal and load circumstances; failing runs will not be printed. Evaluations use a separate protocol with deterministic, model-blind scoring, and pipette-scores by no means sees technology provenance. Every submission information benchmark model, token form, mannequin artifact, quantization, runtime model and settings, and system {hardware} and OS.
Interactive explainer
Key Takeaways
- Pipette benchmarks configurations, not fashions: mannequin + quantization + runtime + system.
- Apache 2.0 stack, 1,000+ configurations, 30+ fashions, three verified gadgets at launch.
- Quality evals run on H100 references and are matched to on-device efficiency, not measured on-device.
- Identical parameter counts can differ 78.4% vs 33.8% in context-scaling retention.
Check out the Technical Details and Leaderboard. Also, be at liberty to observe us on Twitter and don’t overlook to affix our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well.
The put up Liquid AI Open-Sources Pipette: A Reproducible Benchmarking Suite That Measures On-Device Models, Quantization, Runtime and Hardware Together appeared first on MarkTechPost.
