Tool observations dominate the context of software‑engineering agents, making long interaction histories expensive to keep. Existing context‑compression techniques often discard information needed for later actions, while adapting agents to soft‑token representations can alter their original behavior. To cut context size while preserving action‑critical information and behavioral fidelity, we introduce LOHA (Latent Observations, Hard Actions) and ACD (Anchored Context Distillation).\ \ LOHA compresses older tool observations into soft tokens but keeps the agent’s own turns and the last $K$ observations in raw text. This provides compact access to historical data and exact reference to recent content.\ \ ACD enables the model to read this mixed representation by distilling the base model’s full‑text predictions into the latent view, while anchoring its behavior on plain‑text inputs to the same base model to limit drift.\ \ On the SWE‑bench Verified benchmark, with $K=3$ the context per call drops by 43% for Qwen3‑4B and 57% for SWE‑Master‑4B‑RL; resolve rates become 12.1% and 21.8% versus 14.5% and 27.5% for the uncompressed bases. A single‑run recency sweep shows $K=8$ reaches 14.4% and 23.0%, indicating larger windows generally favor task performance over compression.\ \ Under a 32K‑token limit, Qwen3‑4B with $K=3$ resolves 21.1% of a 199‑instance subset, compared to 11.1% for the same adapted agent using full text. In concurrent single‑GPU serving, the compressed agent achieves 1.9× the instance throughput of its full‑text counterpart.\ \ Review: The LOHA‑ACD combination dramatically reduces context size without sacrificing critical behavior, offering substantial efficiency gains especially in resource‑constrained deployment scenarios.