Video-FLAIRVideo-FLAIR icon

Not Whether to Reason, But How

Arizona State University
NeurIPS 2026
Video-FLAIR routes three queries on the same video to DIRECT, CONCISE, and DEEP reasoning modes

Video-FLAIR routes each query to one of three reasoning modes based on its complexity. Highlights mark the key evidence driving each mode, namely screen-text reads (<DIRECT>), timestamped event anchors (<CONCISE>), and multiple cues synthesized into a conclusion (<DEEP>).

Abstract

Multimodal queries can require different types of reasoning. Some can be answered via perceptual reasoning, extracting information directly from the visual signal, while others require compositional reasoning that combines observations or deliberative reasoning that evaluates competing hypotheses. However, many existing methods apply a uniform reasoning strategy across queries, leading to unnecessary computation on simple tasks and insufficient reasoning on complex ones.

We introduce Video-FLAIR, a training framework that learns to select the appropriate reasoning mode for each query using reinforcement learning. During training, the model generates responses under all three modes for the same prompt, enabling direct comparison. A composite reward compares these responses to favor the most effective one based on correctness, grounding, and cost, while discouraging unsupported or misaligned deliberation. This yields a supervision signal for learning adaptive reasoning without per-query annotations.

Video-FLAIR improves accuracy over the Qwen2.5-VL base model by +5.4 on MathVista, +4.8 on Video-Holmes, and +4.8 on Video-MMMU, while reducing average token usage to 95 compared to 417 for always-thinking baselines.

Key Highlights

+5.4
MathVista
over Qwen2.5-VL-7B
+4.8
Video-Holmes
over Qwen2.5-VL-7B
95
Avg. Tokens
vs. 417 for always-thinking baselines
20 / 22
Significant Gains
paired bootstrap over two base models

Three Ways to Reason

Existing methods either apply the same reasoning strategy to every query or decide only whether to reason. Video-FLAIR instead treats reasoning as an adaptive process and chooses how to reason, distinguishing modes by the structure of their evidence chains rather than by their length.

Percept
<DIRECT>

Extracts a single observation that maps directly to the answer. The evidence chain has no branches and the reasoning path is immediate.

Reading on-screen text or identifying an object.
Compose
<CONCISE>

Combines multiple observations through a linear sequence, where each step builds on the previous one without revisiting or branching.

Localizing an event across frames or computing a quantity from a chart.
Deliberate
<DEEP>

Revisits the same evidence to evaluate and reject competing hypotheses. Depth reflects search over interpretations rather than longer chains.

Judging whether an athlete's behavior indicates confidence.

Method

Starting from an SFT warm-start on mode-structured traces, Video-FLAIR generates rollouts under every reasoning mode for the same prompt, scores them with a composite reward that includes an online verifier, and derives a utility that identifies the most effective mode. Advantages are normalized per reward component with GDPO, reweighted token by token, and used to update the policy with DAPO.

Video-FLAIR training pipeline with mode-structured rollouts, composite reward, DAPO update, and an online verifier refreshed with DPO

Each video and query pair generates eight mode-structured rollouts. The verifier provides dense per-rollout feedback and is periodically refreshed with DPO on the highest and lowest reward rollouts.

Mode-Structured Rollouts

For every query, the eight rollouts follow a fixed slot assignment. Six controlled slots force each mode twice, so the outcomes of all three modes can be compared on the same input. Two adaptive slots describe every mode without naming one, so the model must choose a mode before answering. This within-query comparison replaces per-query mode labels.

DIRECTslot 1
DIRECTslot 2
CONCISEslot 3
CONCISEslot 4
DEEPslot 5
DEEPslot 6
ADAPTIVEslot 7
ADAPTIVEslot 8
controlled slots, mode given in the prompt adaptive slots, model selects

Composite Reward

The seven reward terms form three functional groups, and each group contributes a distinct gain when added in turn.

Correctness

Rans, Rformat, Rcomply

Task accuracy by answer type, the output schema, and compliance with the assigned mode and its word budget.

Cost

Rcost, Rbalance

A difficulty-aware length penalty that is strong on easy queries and relaxed on hard ones, plus a term that prevents mode collapse.

Adaptive Selection

Rselect, Rverifier

Utility-derived supervision for the adaptive slots and dense grounding feedback from the online verifier.

Learning Which Mode Wins

Within each rollout group, the controlled slots give every mode an accuracy Ans(m) and a verifier grounding score G(m). Their utility trades accuracy and grounding against cost, and the best mode m* supervises the adaptive slots.

Utility(m) = ( Ans(m) + δ · G(m) ) · ( 1 − Cost(m) ),    m* = arg maxm Utility(m)
Decision regions showing which reasoning mode wins under the utility

More expensive modes win only when their accuracy and grounding gain exceeds their cost.

Adaptive length penalty scaling with response length

The adaptive length penalty grows with length and relaxes on hard queries, so <DEEP> is not penalized when it is needed.

Online Verifier and Token-Level Credit

A Qwen3-VL-8B verifier, initialized from the base checkpoint and never updated by the policy loss, scores each reasoning trace for temporal grounding, spatial grounding, and human alignment. Every 100 RL steps it is refreshed with DPO on pairs ranked by the full composite reward, where answer correctness dominates, so a well-grounded but wrong rollout is still rejected. Token-level credit then concentrates the gradient on evidence-bearing spans such as timestamps and object boxes, and away from hedging and filler.

Token-level credit weights over a CONCISE reasoning trace

Grounded spans receive high credit while hedging phrases such as “I think” and “probably” receive low credit.

Results

Across six image and five video benchmarks and two base models, Video-FLAIR improves accuracy while using a fraction of the tokens of always-thinking methods. Always-thinking baselines lose accuracy on perception-heavy tasks such as EMMA and HallusionBench, where long chains introduce visually ungrounded steps, while Video-FLAIR gains on every benchmark.

Deltas are absolute changes over the base model. For Video-FLAIR, margins are 95% confidence intervals from a paired bootstrap over benchmark items, and ‡ marks a gain that is not significant at the 0.05 level. Best per column in bold. All results are reproduced under identical evaluation settings.

Where Do the Gains Come From?

Ablations on Qwen2.5-VL-7B isolate the contribution of the mode space, the reward groups, and learned selection.

Simpler Alternatives

MethodMMMUEMMAV-HolmesV-MMMUTok
Base model52.025.843.252.1–
Zero-shot prompt51.825.242.851.5340
Static 64-token budget44.518.233.842.064
SFT prompt selector52.926.144.753.3160
Router + SFT53.526.546.054.585
Video-FLAIR54.527.348.056.961

Shorter answers alone do not explain the gains, and a learned router still trails RL by 2.0 points on Video-Holmes and 2.4 on Video-MMMU.

Reward Groups Added in Turn

VariantMMMUEMMAV-MMMUTok
Correctness only51.821.949.8412
+ Cost group53.023.052.0185
+ Selection54.125.755.570
+ Verifier (full)54.527.356.967

The cost group cuts tokens, adaptive selection gives the largest accuracy gains, and the verifier improves grounding-sensitive benchmarks.

Fixed Mode vs. Adaptive

ModeMMMUEMMAV-MMMUTok
Forced <DIRECT>51.422.149.757
Forced <CONCISE>53.125.653.2247
Forced <DEEP>52.828.555.8963
Adaptive54.527.356.967

No fixed mode matches adaptive selection. Forced <DEEP> helps EMMA slightly but costs over 14× the tokens.

Three Modes vs. a Binary Switch

Mode spaceMMMUEMMAV-MMMUTok
Video-Auto-R1 (binary)52.023.754.2–
Direct / think (2 modes)52.924.455.161
Three modes54.527.356.967

Merging <CONCISE> and <DEEP> into one thinking mode drops performance to the level of binary auto-thinking.

How Does Video-FLAIR Reason?

Mode selection tracks what each benchmark demands rather than cost alone. Retrieval-oriented benchmarks such as AI2D and HallusionBench stay mostly <DIRECT>, while inference-heavy and research-level benchmarks shift toward <CONCISE>. <DEEP> is reserved for the small fraction of queries where competing explanations must be ruled out, and removing it hurts exactly those benchmarks.

<DIRECT> <CONCISE> <DEEP>

Share of queries routed to each mode by Video-FLAIR, sorted by <DIRECT> usage.

Gains Hold Across Categories and Model Families

Always-thinking models collapse on image perception and video spatial reasoning, where verbose thinking hallucinates answers that should be read from the visual signal. Video-FLAIR is the only method with positive gains across all five categories on both base models.

Per-category gains over Qwen2.5-VL-7B

Qwen2.5-VL-7B

Per-category gains over Qwen3-VL-8B

Qwen3-VL-8B

Qualitative Examples

Each example pairs two queries on the same video, showing where mode selection succeeds and where it can still improve.

DEEP rules out a competing cause of the purple cube's displacement, and a counting query routed to DIRECT

TopThe query asks which cube displaced the purple cube. <DEEP> establishes that the green cube enters the frame only after the impact, eliminating that hypothesis, and confirms gold-cube contact with timestamped boxes.

BottomCounting moving objects needs a running tally over the whole video. The policy routes to <DIRECT>, inspects one frame, and reports 3 instead of 4, since the temporal structure calls for <CONCISE>.

BibTeX

@inproceedings{kulkarni2026videoflair,
  title={Video-FLAIR: Not Whether to Reason, But How},
  author={Yogesh Kulkarni and Pooyan Fazli},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}