AI tutoring can markedly improve learning outcomes in developing regions such as Vietnam, yet the two obvious approaches fall short. Cloud assistants like ChatGPT route sensitive student data to foreign servers, violating data‑sovereignty regulations such as Vietnam's Decree 53, and their Western‑centric pre‑training corpora are not aligned with the national textbook curriculum, resulting in fragmented local knowledge and frequent hallucinations.
Self‑hosting an open‑source model keeps data on‑premise but hits a two‑fold wall: post‑training quantization methods (AWQ, GPTQ) shrink the static weight footprint, yet the dynamic KV cache and pre‑fill latency of long tutoring contexts still cause out‑of‑memory failures and slow responses on consumer GPUs, while the model continues to hallucinate on region‑specific material.
DeepEdu‑v1 is an AI‑tutoring system for Vietnamese education built on the SCALE (Self‑improving Context‑Aware Learning Engine) framework, introducing two innovations. First, a long‑context inference engine amortizes token selection from per‑sub‑chunk to per‑cluster granularity, issuing 7.7× fewer retrieval calls during long‑context retrieval. Compared with a state‑of‑the‑art selective‑attention baseline, this cuts pre‑fill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self‑improving agentic layer continuously curates a verified playbook from past interactions instead of fine‑tuning, designed to progressively reduce reliance on dominant‑language priors as trustworthy local knowledge accumulates.
In its deployed configuration, DeepEdu achieves nearly a 2× TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per‑track gains on financial‑reasoning and interactive‑agent benchmarks.
Review