NeFut Logo NeFut
中 Admin Login

[CS.AI] Smart Content Ingestion for Generative AI Workloads

Published at: 2026-10-07 22:00 Last updated: 2026-10-08 01:25
#algorithm #AI #Machine Learning

Smart content ingestion for generative AI workloads treats extraction as a first‑class lifecycle stage. In classic machine learning the task, data representation, labels and model architecture are tightly coupled, so data preparation is narrow, schema‑bound and visible. Generative AI decouples the model from any single task; a single foundation model serves open‑ended downstream tasks while enterprise knowledge appears in PDFs, presentations, spreadsheets, scanned forms, tables, diagrams and other mixed‑layout files that embed textual, visual, geometric and structural signals. A language model or retriever cannot reason reliably over mis‑represented information, making precise front‑end extraction essential.

The presented production‑ready system comprises selective OCR routing, a scarcity‑first curation engine, a reference‑based extraction scorer that measures character, word and table‑structure accuracy, a deterministic structure‑aware parent‑child chunker, and a read‑only retrieval evaluator that generates grounded questions per page and reports Hit@k, mean reciprocal rank and latency. On a corpus of 180 documents the best extractor achieves a composite score of 97.4 (character error rate 0.13 %, table similarity 0.995). The chunker reaches hit@1 68.6 %, hit@10 92.8 % and MRR 0.77 over 25,050 generated questions. We distill three design principles—structure before semantics, never mutate what you measure, budget your labels—and position measured content extraction as the perception layer of enterprise agentic systems. Review

Original Source: https://arxiv.org/abs/2610.07091

[h] Back to Home