1 link tagged with all of: benchmarks + model-compression + llm
Click any tag below to further narrow down your results
Links
PrismML’s Bonsai 8B trains a large language model with 1-bit weights from scratch, squeezing 8.2 billion parameters into just 1.15 GB. In benchmarks it ties or outperforms FP16 models like Llama 3.1 and runs at real-time speeds on phones, shifting the size-performance trade-off.
- Bonsai 8B packs 8.2B parameters into 1.15GB using native 1-bit weights, yet scores 70.5 average vs Llama 3.1's 67.1 (16GB FP16), even hitting 88.0 on GSM8K vs Llama's 76.6.
- It runs at 44 tokens/sec on an iPhone 17 Pro Max, making full on-device 8B-scale LLMs feasible without cloud infrastructure.
- Its "intelligence density" (score/GB) hits 1.062 versus Qwen's 0.098 and Llama's 0.084 — over 10x more capability per byte.
- Trade-offs appear in code generation (57.9 vs Qwen's 79.9 on HumanEval+) and multi-step reasoning (MuSR: 64.3 vs 70.0), showing 1-bit precision struggles with complex logical chains.