NeFut Logo NeFut
中 Admin Login

[CS.AI] Overview of the Pistis Technical Report

Published at: 2026-09-25 22:00 Last updated: 2026-09-28 00:49
#AI #Machine Learning #LLM

We introduce the Pistis model family, comprising a 27B‑parameter multimodal LLM built on Qwen3.6 and a 9B‑parameter counterpart built on Qwen3.5. Both models are trained using a general, scalable post‑training framework. The framework first establishes a strong foundation via large‑scale multimodal supervised fine‑tuning (SFT).

On top of this foundation we propose Interleaved Distillation and Reinforcement Learning (IDRL), a novel paradigm that tightly integrates on‑policy distillation and reinforcement learning within a single training loop. By alternating between the two objectives rather than optimizing them in isolation or via a static joint loss, IDRL achieves more effective knowledge transfer, greater optimization stability, and finer credit assignment for long‑horizon agentic trajectories, thereby improving performance while mitigating typical capability trade‑offs.

At both model scales the framework yields two specialized variants: Pistis‑Thinking, which enhances deep multimodal reasoning, and Pistis‑Agentic, which incorporates agentic trajectory data to support long‑horizon planning, iterative reasoning, and tool use. Pistis‑Agentic is especially strong in multimodal search. Both scales outperform their respective base models.

Beyond parameter optimization, we introduce Pistis‑Auto‑Harnessing (PAH), a system‑level method that automatically improves the inference harness through iterative optimization. Experiments demonstrate that PAH boosts model performance without updating model parameters or increasing the interaction budget.

Review

Original Source: https://arxiv.org/abs/2609.28554

[h] Back to Home