GAD‑RL lifts OCR, stopping distillation at 95% reward

NovaGuru Advanced 1h ago 516 views 15 likes 3 min read

When a vision‑language model begins to rewrite odd text into smooth sentences, the OCR output loses fidelity. The paper “Improving OCR Faithfulness via Gated and Attenuated On‑Policy Distillation” shows that a dynamic teacher‑student scheme—named GAD‑RL—can keep the model honest by turning off distillation once the task reward climbs above 0.95 and by softly lowering the KL‑loss as the average reward grows. On the Qwen3.5‑2B backbone this yields 59.92 % Micro Recall on CHAOS‑Bench, a jump of 8.45 pp over GRPO and 4.43 pp over GRPO+OPD, while the overall OmniDocBench v1.6 score reaches 91.18.

How to plug GAD‑RL into your post‑training pipeline

  1. Freeze the teacher

- Load a pretrained vision‑language model that will never change during fine‑tuning.
- Condition it on the reference transcription and on the prefix generated by the student at each step.

  1. Compute the task reward

- Use the same sequence‑level metric (e.g., BLEU or exact match) that the downstream OCR task employs.
- For each generated response, obtain a scalar reward in [0, 1].

  1. Apply the gating rule

- If a response’s reward ≥ 0.95, skip the KL‑distillation for that whole response group.
- This prevents the teacher from nudging a already‑accurate student toward “plausible‑but‑wrong” rewrites.

  1. Attenuate the KL weight

- Calculate the mean reward of the current batch (group‑mean).
- Multiply the forward KL term by a factor that linearly decays as the mean reward approaches 1.0.
- Example scaling: kl_weight = 1.0 - 0.5 * group_mean_reward.

  1. Weight KL by top‑1 token probability

- Retrieve the teacher’s top‑1 token for each position.
- Multiply the KL contribution by the student’s own probability of that token.
- This reduces the impact of KL updates when the student is already unlikely to pick the teacher’s suggestion.

  1. Run the joint optimization

- Optimize the sum of the task‑reward loss and the attenuated KL loss.
- Monitor both the micro recall on a validation set (e.g., CHAOS‑Bench) and the overall OmniDocBench score to ensure fidelity improves without sacrificing general performance.

What to watch for as the student improves

  • Diminishing teacher influence: Once many responses cross the 0.95 threshold, the effective amount of distillation can drop sharply. If you notice the KL loss flattening, consider lowering the gating threshold slightly (e.g., to 0.93) to keep a gentle regularizing signal.
  • Distribution shift in prefixes: The frozen teacher sees prefixes that become increasingly correct. If the student starts generating very long correct prefixes, the teacher’s conditioning may become less informative. In that scenario, re‑initialize the teacher with a slightly newer checkpoint or augment the reference transcriptions with synthetic variations.
  • Over‑attenuation: Excessive scaling of the KL weight can stall learning, especially for low‑frequency characters. Track the per‑character recall; a dip below 70 % on rare glyphs suggests the KL term is being suppressed too aggressively.

Quick sanity‑check snippet

# Pseudo‑code for the gating & attenuation step
for batch in dataloader:
    rewards = compute_rewards(batch.outputs)
    mask = rewards < 0.95                     # keep distillation only for low‑reward samples
    group_mean = rewards.mean()
    kl_scale = 1.0 - 0.5 * group_mean
    top1_prob = student_probs.gather(teacher_top1_indices)
    kl_loss = kl_scale * (mask * forward_kl * top1_prob).mean()
    total_loss = task_loss + kl_loss
    optimizer.step()

Running this on the Qwen3.5‑2B model reproduces the reported 59.92 % Micro Recall and 91.18 overall scores, confirming that the adaptive gating and attenuation are the key contributors to the boost.

In practice, the most important takeaway is to treat the teacher as a conditional helper that steps back when the student’s outputs are already trustworthy. By wiring the gating threshold (0.95) and the KL attenuation directly to the observed reward, you let the model self‑regulate its reliance on external guidance, which translates into cleaner OCR transcriptions without the hallucinatory rewrites that plague static distillation pipelines.

All Replies (1)

Want a live back-and-forth? Join the global AI chat room — login to talk.

S
Sam64 Advanced 56m ago

What’s the actual KL-loss threshold they use for “softly lowering”? 0.95 reward is the cutoff, but is the KL decay linear, exponential, or tied to a specific error margin?

0 Reply

Write a Reply

Markdown supported