Back to research

What Matters for Latent Reasoning with Flow Matching

FLaRe: Flow-based Latent Reasoning

Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

arXiv preprint, October 2026

Abstract

FLaRe studies how language models can reason with continuous latent states and generate only their final answers. It combines a compact representation of symbolic reasoning, question-conditioned flow matching, and training on verified model-generated thoughts.

The work evaluates whether these thoughts improve answers, support varied reasoning paths, admit faithful explanations, benefit from additional computation, and reduce inference cost. Experiments on arithmetic tasks examine the training choices needed to make latent reasoning effective.

Key contributions

  1. A compact reasoning space

    Encode symbolic reasoning into latent codes that can also be decoded as natural-language explanations.

  2. Training that matches inference

    Train the answer reader on generated thoughts, then refine the model through verified self-training.

  3. Five evaluation criteria

    Probe usefulness, diversity, explainability, refinement and efficiency.

  4. Accuracy and latency

    On GSM8K, two-step FLaRe reaches 57.4% accuracy at 68 ms, versus 59.3% at 267 ms for explicit CoT. Stage 2; median latency on one RTX 3090, bf16, one question at a time.

Citation

If you use this work, please cite:

@article{ouali2026what,
  title={What Matters for Latent Reasoning with Flow Matching},
  author={Ouali, Yassine and Bulat, Adrian and Tzimiropoulos, Georgios},
  journal={arXiv preprint arXiv:2610.06666},
  year={2026},
  eprint={2610.06666},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  url={https://arxiv.org/abs/2610.06666}
}