AI Is Making Hacker News an Impossible Place to Find Signal
The signal-to-noise ratio on Hacker News has fallen sharply over the past few years, with the AI boom appearing to be the main cause. It was once the gold standard for uncovering the internet’s “hidden gems”: deep technical explorations, obscure whitepapers, and authentic engineering debates. Today, its front page is often filled with AI-generated summaries, effortless “top 10 tools” lists, and marketing material dressed up as technical writing but reading like a press release.
Why manual browsing fails now
There is simply too much content to browse manually without overlooking genuine breakthroughs. Unless you refresh every ten minutes, the most valuable discussions disappear beneath an endless pile of generic AI hype. The frustrating part is that the community still exists, but its discovery system is breaking down.
As a result, I have changed how I read the feed. Rather than visiting the main page, I now depend on curated aggregators and targeted search queries to locate the “long-form” writing that genuinely matters. I look for indicators of quality, such as a high comment-to-upvote ratio, which often suggests real discussion rather than a link that merely went viral.
How an LLM agent filters noise
For anyone building a sustainable AI workflow for gathering information, I have found that a custom LLM agent can help filter RSS feeds. Of course, there is irony in using AI to eliminate the noise that AI created. A basic setup pulls the HN API and sends the model a prompt asking it to classify posts by “technical depth” versus “marketing hype.”
Building a custom Python filter
If you want to build a similar filter from scratch, a simple Python script can query the API and retain posts containing keywords associated with high-quality engineering material, such as “implementation,” “latency,” “benchmark,” or “kernel,” while disregarding broad terms such as “revolutionary” or “game-changing.”
import requests
def get_top_stories():
top_stories_url = "https://hacker-news.firebaseio.com/v0/topstories.json"
stories = requests.get(top_stories_url).json()
# Filter for high-signal keywords to avoid AI fluff
high_signal_keywords = ['benchmark', 'implementation', 'deep dive', 'rfc', 'architecture']
filtered_stories = []
for story_id in stories[:50]: # Check top 50
item = requests.get(f"https://hacker-news.firebaseio.com/v0/item/{story_id}.json").json()
if any(word in item.get('title', '').lower() for word in high_signal_keywords):
filtered_stories.append(item)
return filtered_stories
The central difficulty is that “interesting stuff” is rapidly becoming a needle in a haystack. We may be approaching a point where finding high-quality human insight requires actively removing synthetic noise. That raises the possibility that general-purpose aggregators are nearing the end of their usefulness, with small, invite-only curated mailing lists making a comeback.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
This is exhausting. How many of these generic summaries are actually bots now?
Frustrating that I have to sort by "new" just to find real people. Anyone else doing this?
I can't tell the writing styles apart anymore. Is the AI output becoming too smooth to spot?