Whether an information‑extraction pipeline should operate on page images or on parsed text depends on the document, and the answer flips across the layout spectrum. Under a constraint that excludes (closed‑source) cloud services, we study this trade‑off by processing privacy‑sensitive documents on‑premise with small (≤ 8 B‑parameter) text‑only and vision‑language models, evaluating both accuracy and energy. The design space spans input representation, model family, and inference configuration. Benchmarks use the near‑plain‑text Kleister‑NDA contracts and the layout‑rich VRDU forms. Results show that batching is the dominant energy lever, cutting per‑page energy by 38‑85% without accuracy loss. FP8 quantisation saves 27‑32% energy for single‑request inference, but once batching is applied the saving drops below 1 mWh per page (≈ 9‑19%). Pre‑processing then dominates the remaining energy: neural OCR consumes 17× more energy per page than classical OCR and never reaches the Pareto frontier. The winning representation flips with document type: vision‑language models excel on layout‑rich documents, while small text‑only models paired with a cheap parser dominate on near‑plain‑text documents; in both cases they are more accurate and cheaper than any vision‑language configuration. From these findings we derive concrete guidelines for energy‑efficient, privacy‑compliant local information extraction.
Review