High‑quality synthetic data is essential for post‑training large language models to capture the wide range of expert strategies and decisions that appear in conversations. Prompting a model directly or conditioning it only on a target scenario typically yields low‑diversity outputs that collapse onto a few dominant modes.
We introduce a generative‑flow‑network (GFlowNet) based pipeline for producing diverse synthetic dialogues. The key idea is to train a GFlowNet to sample latent conversation structures while modeling important interaction features—such as confusion‑episode dynamics and the balance of scaffolding directives—with a Gaussian‑mixture density. This makes the sampling probability of each expert strategy proportional to its prevalence in the training set.
Experiments on two structurally distinct domains—tutoring and emotional‑support dialogues—show that GFlowNet‑generated data achieve a better trade‑off among fidelity, mode coverage and authenticity, and do not copy training examples. Compared with reinforcement‑learning and end‑to‑end LLM baselines, classifiers trained on the synthetic GFlowNet conversations deliver stronger signals on three downstream outcome‑prediction tasks. Review