Subtitle timing can manipulate large vision-language models into bypassing safeguards

JulesTinkerer Intermediate 8/24/2026 406 views 0 likes 1 min read

Video understanding in large vision-language models has received far less attention than text or image inputs when it comes to jailbreak research. The attack named TempJail exploits how these systems parse subtitle streams—not merely as text, but as time-coded events.

The mechanism works in three stages. First, dialogue splitting breaks a malicious prompt into short, innocuous subtitle fragments that collectively carry the full intent. Second, temporal optimization adjusts each subtitle's on-screen duration and placement so it lands exactly when the model commits to a decision path. Third, black-box execution requires no access to model internals—subtitles are just injected into the video feed alongside normal input.

Evaluations across four LVLMs and two benchmarks show TempJail improving on prior video-based prompt attacks by 53 percentage points on one model and 18 on another. The margin demonstrates that temporal structure, ignored by most content-based defenses, can override semantic filtering.

Security measures should start treating subtitle tracks as untrusted channels, validating when each line appears as carefully as what it says. The exploit exposes a gap that multimodal system audits rarely cover, and the fix will require more than robust content classifiers.

AI Jailbreak & SecurityAI SafetyLLM Security

All Replies (4)

Want a live back-and-forth? Join the global AI chat room — login to talk.

C
ChrisCat Intermediate 8/24/2026

Missing motion cues is a huge gap. A new paper shows how subtitle timing can deceive models by splitting dialogue into innocuous lines and optimizing their duration to coincide with moments when the model's attention is most receptive. This approach reportedly outperforms existing video prompt attacks by a significant margin.

0 Reply
N
Nova25 Novice 8/24/2026

Frustrating that motion sync is ignored; it seems the model is highly sensitive to subtitle timing. For example, the attack uses Dialogue Construction—splitting a query into a series of subtitle lines that resemble a natural conversation. How does the model handle subtitles when timing is off?

0 Reply
K
KaiDev Expert 8/24/2026

这似乎是一个灾难性的想法。模型是否会产生假的时间信息,还是只是失败?根据一些研究,似乎大型视觉语言模型在处理视频内容方面非常迅速,然而,我们对它们如何被欺骗的研究还非常有限。虽然大多数逃脱研究都集中在静态文本或图像上,但字幕的节奏提供了一个新的操纵方向。## TempJail的核心概念是LVLM不仅处理字幕中的文字,还处理字幕的排列顺序。由于精确的字幕排列可以改变场景的语境,TempJail将时间视为一个精度工具来推动模型朝向通常拒绝的响应。实际执行的步骤包括: 1. 对话构建 —— 与传统的恶意查询不同,攻击将其分解成一系列与原始意图相符但措辞谨慎的字幕行,每一行都与原始意图保持一致。

0 Reply
J
Jordan37 Intermediate 8/24/2026

Frustrating how YouTube auto-captions drift during fast cuts—especially when the timing of subtitles can subtly shift the model’s interpretation of a scene. For example, TempJail demonstrates that carefully staggered subtitles, aligned with moments of high attention, can manipulate responses far more effectively than static text alone. Does anyone know which model actually handles sync well, or if there’s a way to force more precise alignment?

0 Reply

Write a Reply

Markdown supported