I’m building a taste engine that reads fanfict beyond familiar tags
DeepSeek-V3 is currently outperforming Claude 3.5 Sonnet for one particular job: tagging and categorizing enormous fanfiction datasets with rigorous logic. When it comes to generating prose, however, Sonnet remains stronger. Over the past few weekends, I’ve been building a “taste engine”—a recommendation system that looks beyond labels such as “Slow Burn” and “Enemies to Lovers” to analyze a fic’s vibe and narrative pacing, then identify similar reads.
Noise is the central challenge with fanfic datasets. Tags are user-generated and wildly inconsistent. To address that problem, I benchmarked GPT-4o, Claude 3.5 Sonnet, and DeepSeek-V3 on “narrative fingerprinting.” I gave each model 2,000-word excerpts and asked it to identify latent emotional arcs and pacing metrics, such as “internal monologue vs. dialogue ratio.”
DeepSeek-V3 was the clear winner during extraction. Its precision with structured data is almost frightening. GPT-4o often invented “thematic depth” that did not exist, while DeepSeek returned raw, honest metrics. When I requested a JSON representation of the tension curve, DeepSeek mapped the beats accurately.
I rely on Claude 3.5 Sonnet for the “taste” portion. Claude has a linguistic sensitivity the others lack when distinguishing why one writing style feels “angsty” rather than “depressing.” It recognizes subtext and prose rhythm. Given a user’s favorite works, Claude can describe that taste profile with haunting accuracy; GPT-4o’s responses read more like book reports.
GPT-4o has become my “middle manager.” It cleans data quickly and reliably, but it lacks both DeepSeek’s surgical precision and Claude’s poetic intuition. In my tests, GPT-4o produced the most “generic” recommendations, favoring the most popular fics in each category instead of the closest stylistic matches.
My implementation uses a hybrid pipeline. DeepSeek vectorizes the “structural” elements of each story, while Claude generates the “aesthetic” embeddings.
This prompt logic tests “vibe” consistency across models:
Analyze the provided text. Ignore the plot. Instead, evaluate the prose on a scale of 1-10 for the following:
- Lexical Density (complexity of vocabulary)
- Emotional Temperature (cold/detached vs warm/intimate)
- Pacing Velocity (how quickly the scene moves relative to the clock)
Output strictly in JSON.
The performance gap becomes clearest with long-form content. DeepSeek’s context window feels more “stable” during long-range dependency checks, preserving the tone established in the first chapter by the tenth.
DeepSeek-V3:
Pros: Unbeatable structured data extraction, high logic consistency, and significantly cheaper bulk processing.
Cons: Occasionally too literal and unable to capture the “soul” of the prose.
Claude 3.5 Sonnet:
Pros: Superior understanding of nuance, style, and human emotion; feels like a real reader.
Cons: Strict rate limits can kill a project’s momentum; more prone to “refusing” certain edgy fanfic tropes if the system prompt isn’t tight.
GPT-4o:
Pros: Fast, great API stability, and a decent all-rounder.
Cons: “Average” output that gravitates toward the center of the bell curve, making recommendations feel bland.
For a system that must genuinely discern art or style, GPT-4o should not be the sole engine. Real “taste” emerges from the difference between DeepSeek’s structural analysis and Claude’s stylistic intuition.
All Replies (2)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Intrigued by the logic jump. I wonder if using an LLM-based classifier for 10k+ tags actually scales without hallucinating. DeepSeek-V3 was the clear winner during extraction, so I now feed its JSON tension curves into a K‑means clustering step to group fics by pacing and test scalability.

I want to try this tonight. Does it handle the 404 errors better than GPT-4o when scraping AO3? I’m curious if you’ve benchmarked how DeepSeek-V3 handles noise in fanfic datasets—specifically, its precision with structured data when identifying latent emotional arcs and pacing metrics like “internal monologue vs. dialogue ratio.”