Running a language model on edge hardware enables private, low‑latency reasoning without a network connection, but the small models that fit on such devices are often unreliable on tasks they should excel at, such as arithmetic, algebra, and formal logic. We argue that much of this unreliability is avoidable. Many queries that appear to require reasoning are actually structurally deterministic and admit fast, exact symbolic solutions. Forcing a probabilistic model to approximate them sacrifices accuracy and energy for little gain.
We therefore introduce a neurosymbolic router that first classifies each incoming query and then dispatches it to the cheapest solver that guarantees correctness: deterministic engines handle structured tasks, while the small language model (SLM) tackles open‑ended word problems. The routing logic is not hand‑coded; instead we learn a deterministic finite automaton (DFA) using the L* grammatical inference algorithm. During learning the SLM serves as a membership oracle and labeled data act as an equivalence oracle.
Evaluation on a Raspberry Pi 4B (8 GB RAM, no GPU) used 100 unseen prompts from DeepMind Mathematics, GSM8K, and RuleTaker. The learned router achieved 100% routing accuracy and 98.3% overall accuracy with a 512‑token reasoning budget (93.3% on word problems), far surpassing the strongest baseline Program‑of‑Thought (72.0%) and a tool‑calling agent with the same solvers (58.7%). Because formatted queries never reach the model, the router answers them in 1‑11 ms; in its 30‑token configuration it runs 8.8× faster and 2.8× more energy‑efficient than Program‑of‑Thought.
Core techniques
- Automatic construction of a DFA via the L* algorithm for deterministic query classification.
- Use of the SLM as a membership oracle and labeled data for equivalence checking.
- Hybrid reasoning that combines symbolic solvers with a small language model to maximize efficiency on resource‑constrained devices.
Review