Ultra‑low‑bit language models can cut storage and memory bandwidth, yet a nominal “1.58‑bit” label does not fully capture the stored representation, retained capability, or runtime behavior. This work performs an end‑to‑end post‑training conversion of the instruction‑tuned 4B‑parameter Qwen model using KOTMS rotation, E2M‑ATQ ternarization, and GPTQ‑style error compensation from TWLA. The conversion is weight‑only; activations stay at 16‑bit precision, so ILA‑AMP is omitted. We evaluate effective bit accounting, task capability retention, perplexity, calibration sensitivity, checkpoint composition, and deployment behavior. The final quantized linear layers use 1.641 effective bits per weight, covering 81.62% of the parameters. Across ten capability benchmarks, accuracy drops from 64.5% to 54.7%, with uneven degradation: BoolQ retains 84.6% of teacher performance, while ARC‑Challenge falls to 43.8%. Perplexity rises from 13.639 to 18.748 on WikiText‑2, 24.700 to 31.992 on PTB, and 19.831 to 28.966 on C4. A subsequent packing step preserves ternary planes and scales, shrinking model size from 8.29 GiB to 3.96 GiB with virtually unchanged perplexity. A third‑party packing attempt was lossy and excluded from the primary artifact claim. The packed artifact has not been benchmarked end‑to‑end for task accuracy or generation throughput. A preliminary Triton GEMV microbenchmark shows a 4.6× slowdown compared to FP16 cuBLAS on one tested shape. Hence we do not claim that compression alone yields faster inference.
Review