We examine whether human‑readable text is required for fine‑tuning large language models (LLMs). The proposed Desired-Update‑Aligned Synthetic Data (DASA) method uses activation‑gradient feedback from a frozen reference model to directly optimize continuous synthetic input embeddings, bypassing source‑text reconstruction and linguistic fluency. DASA aims to generate embeddings that yield useful adaptation updates; these embeddings are fed straight into downstream fine‑tuning, with discrete token projections used only for qualitative inspection.
Experiments span six models from the Llama and Qwen families (1B‑32B parameters) and six benchmarks covering knowledge retrieval, mathematical reasoning, code generation, and commonsense reasoning. Under matched LoRA adaptation settings, DASA matches or exceeds the performance of natural‑language source data in most configurations and outperforms GRADMM in the majority of comparisons. Additional studies on general‑domain and task‑specific source data confirm its robustness. Overall, DASA delivers a $3.6$‑$4.9\times$ speedup over GRADMM while using comparable peak GPU memory.
Review: By synthesizing directly in embedding space and omitting the text generation step, DASA shows that human‑readable text is not a prerequisite for effective LLM fine‑tuning, offering a more efficient and resource‑friendly adaptation pathway.