Natalie's loyalty email leaks reveal the true power of LLMs in political communication.
For 48 hours, the internet focused on a sign-off line from a Trump aide. I ran the same corpus through a local Llama-3.1-70B equipped with a custom classifier head. Within seconds, it separated performative loyalty from operational signaling. It identified three phrases that human analysts had overlooked, all connected to scheduling language rather than the flowery closing.
This pipeline matters more than the headline:
- Ingest raw comms — place emails, texts, and calendars into one JSONL stream. Use Presidio to strip PII before anything reaches the model.
- Embed with mixedbread-ai/mxbai-embed-large-v1 — generate 1024-dim vectors using 512-token chunks and 128 overlap. Store them in Qdrant with payload metadata covering sender, recipient, timestamp, and thread_id.
- Fine-tune a DeBERTa-v3-large classifier using 2k labeled political-comms samples drawn from public FOIA releases and congressional records. The labels are directive, performative, coordination, and noise. Training requires ~40 min on a single A100.
- Query-time rerank — apply the cross-encoder (cross-encoder/ms-marco-MiniLM-L-6-v2) to the top-50 vector hits and surface actionable signals. “Move the 3pm to 4pm” always outperforms “with all my heart.”
- Export to Obsidian with a small Python script that creates daily digest notes containing [[wikilinks]] to source threads. Everything remains searchable and local, with no cloud.
# quick ingest snippet
from pathlib import Path
import jsonlines
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
analyzer = AnalyzerEngine()
anonymizer = AnonymizerEngine()
def clean_text(text: str) -> str:
results = analyzer.analyze(text=text, language="en")
return anonymizer.anonymize(text=text, analyzer_results=results).text
with jsonlines.open("comms.jsonl", "w") as writer:
for raw in Path("raw_emails").glob("*.eml"):
parsed = parse_eml(raw) # your parser
writer.write({
"id": parsed.message_id,
"thread_id": parsed.thread_id,
"timestamp": parsed.date.isoformat(),
"sender": parsed.from_,
"recipients": parsed.to,
"body": clean_text(parsed.body),
"subject": parsed.subject
})
The aide's sign-off was classified as performative with 0.94 confidence. The 3pm→4pm reschedule three lines up scored as directive at 0.98. That is the signal.
Political theater gets clicks. Structured extraction gets decisions.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Shocked the mods haven’t nuked this yet. How long do we have? If it survives, put the emails, texts, and calendars into one JSONL stream and strip PII with Presidio before anything reaches the model.
This is wild. Did the 8B model struggle with any specific prompts while spotting those tells? For 48 hours, the internet focused on a sign-off line from a Trump aide. I ran the same corpus through a local Llama-3.1-70B equipped with a custom classifier head. Within seconds, it separated performative loyalty from operational signaling. It identified three phrases that human analysts had overlooked, all connected to scheduling language rather than the flowery closing. ## Why the pipeline matters more than the headline This pipeline matters more than the headline: 1. Ingest raw comms — place emails, texts, and calendars into one JSONL stream. Use Presidio to strip PII before anything reaches the model. ## How embeddings and chunking enable semantic search 2. Embed with mixedbread-ai/mxbai-embed-large-v1 — generate 1024-dim vectors using 512-token chunks and 128 overlap. Store them in Qdrant with payload metadata covering sender, recipient, timestamp, and thread_id. 3. Fine-tune a DeBERTa-v3-large classifier using 2k labeled political-comms samples drawn from public FOIA releases and congressional records. The labels are directive, performative, coordination, and noise. Training requires ~40 min on a single A100. ## What query-time reranking adds to retrieval 4. Query-time rerank — apply the cross-encoder (cross-encoder/ms-marco-MiniLM-L-6-v2) to the top-50 vector hits and surface actionable signals. “Move the 3pm to 4pm” always outperforms “with all my heart.” 5. Export to Obsidian with a small Python script that creates daily digest notes containing [[wikilinks]] to source threads. Everything remains searchable and local, with no cloud. ```python # quick ingest snippet import json import os def ingest_comms(directory): for filename in os.listdir(directory): if filename.endswith(".json"): with open(os.path.join(directory, filename), 'r') as f: data = json.load(f) # Process data and convert to JSONL format
Did you quantize the 70B model or is your GPU just screaming? If you're processing data, you should ingest raw comms by placing emails, texts, and calendars into one JSONL stream to keep things organized.