What Matters for Latent Reasoning with Flow Matching
FLaRe: Flow-based Latent Reasoning
arXiv preprint, October 2026
Abstract
FLaRe studies how language models can reason with continuous latent states and generate only their final answers. It combines a compact representation of symbolic reasoning, question-conditioned flow matching, and training on verified model-generated thoughts.
The work evaluates whether these thoughts improve answers, support varied reasoning paths, admit faithful explanations, benefit from additional computation, and reduce inference cost. Experiments on arithmetic tasks examine the training choices needed to make latent reasoning effective.
Key contributions
-
A compact reasoning space
Encode symbolic reasoning into latent codes that can also be decoded as natural-language explanations.
-
Training that matches inference
Train the answer reader on generated thoughts, then refine the model through verified self-training.
-
Five evaluation criteria
Probe usefulness, diversity, explainability, refinement and efficiency.
-
Accuracy and latency
On GSM8K, two-step FLaRe reaches 57.4% accuracy at 68 ms, versus 59.3% at 267 ms for explicit CoT. Stage 2; median latency on one RTX 3090, bf16, one question at a time.
Citation
If you use this work, please cite:
@article{ouali2026what,
title={What Matters for Latent Reasoning with Flow Matching},
author={Ouali, Yassine and Bulat, Adrian and Tzimiropoulos, Georgios},
journal={arXiv preprint arXiv:2610.06666},
year={2026},
eprint={2610.06666},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2610.06666}
}