惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

量子位
GbyAI
GbyAI
V
Vulnerabilities – Threatpost
阮一峰的网络日志
阮一峰的网络日志
罗磊的独立博客
Recorded Future
Recorded Future
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
博客园 - 司徒正美
Y
Y Combinator Blog
Microsoft Security Blog
Microsoft Security Blog
美团技术团队
博客园 - Franky
Blog — PlanetScale
Blog — PlanetScale
B
Blog RSS Feed
V
Visual Studio Blog
Martin Fowler
Martin Fowler
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
博客园_首页
C
Cybersecurity and Infrastructure Security Agency CISA
博客园 - 叶小钗
AWS News Blog
AWS News Blog
Project Zero
Project Zero
T
Threat Research - Cisco Blogs
V
V2EX
F
Fortinet All Blogs
The GitHub Blog
The GitHub Blog
Latest news
Latest news
N
News and Events Feed by Topic
The Last Watchdog
The Last Watchdog
T
Threatpost
L
Lohrmann on Cybersecurity
小众软件
小众软件
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
博客园 - 聂微东
Engineering at Meta
Engineering at Meta
爱范儿
爱范儿
Google Online Security Blog
Google Online Security Blog
Forbes - Security
Forbes - Security
Attack and Defense Labs
Attack and Defense Labs
The Register - Security
The Register - Security
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
H
Help Net Security
Security Latest
Security Latest
Recent Announcements
Recent Announcements
C
Check Point Blog
B
Blog
Google DeepMind News
Google DeepMind News
K
Kaspersky official blog
I
InfoQ

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Building a News Aggregator Without an Engagement Algorithm
Bryan Leonar · 2026-05-04 · via DEV Community

Building a News Aggregator Without an Engagement Algorithm

I have been building a project called WeSearch:

https://wesearch.press

It is a free news aggregator that pulls from hundreds of sources, keeps discovery mostly chronological, adds source/bias context where available, preserves permanent daily archives, and allows anonymous discussion on stories.

The project started from a simple frustration:

Most news discovery products are either too personalized, too paywalled, too noisy, too opaque, or too socially distorted.

I wanted something closer to this:

  • a wide source feed
  • no account required
  • no paywall
  • no tracking
  • chronological discovery
  • source context
  • permanent archives
  • anonymous discussion
  • less algorithmic manipulation

That sounds simple, but once you start building it, the hard part is not fetching headlines.

The hard part is trust.


The problem with modern news discovery

There are several existing models for news discovery.

Google News is broad, but opaque. You get a feed, but you do not always know why certain stories are ranked or why certain sources are emphasized.

Reddit and X are fast, but socially distorted. Stories become memes, outrage cycles, or identity signals before they become information.

RSS readers are powerful, but require setup and source selection. They are great for people who already know what they want to follow. They are less useful for broad public discovery.

Ground News, AllSides, and similar products are useful because they introduce comparison and bias context, but some of the most useful features are often gated behind subscriptions or limited interfaces.

Hacker News is extremely high signal for technical and startup-related topics, but it is not a general-purpose news aggregator.

So the question I kept coming back to was:

What would a news aggregator look like if it tried to be less addictive, less opaque, and more useful for comparing coverage?

That is the question behind WeSearch.


Why chronological discovery still matters

A lot of modern feeds are optimized around engagement.

That usually means the system decides what you should see based on some mixture of clicks, dwell time, reactions, shares, prior behavior, and predicted interest.

That can be useful, but it creates a problem:

The feed stops being a window into what is happening and becomes a mirror of what the system thinks will keep you engaged.

For news, that is dangerous.

A chronological feed is not perfect. It can be noisy. It can be overwhelming. It can miss importance. But it has one major advantage:

It is legible.

You can understand why something appears.

It appeared because it was published or discovered recently.

That does not solve ranking, source quality, duplication, or bias. But it gives the user a clean baseline. From that baseline, you can add filtering, clustering, search, source context, and archive views without turning the whole thing into a black box.

That is why WeSearch leans chronological first.


Source context is useful, but bias labels are not enough

One of the obvious features for a news aggregator is source labeling.

People want to know:

  • where the article came from
  • whether the outlet has a known political tendency
  • whether the source is reliable
  • whether the article is reporting, opinion, analysis, or commentary
  • how other outlets are covering the same event

But a simple left / center / right label is dangerously incomplete.

Two articles can both come from “left” sources and still be completely different in quality.

One may be careful reporting with primary sources.

Another may be mostly emotional framing.

The same is true for “right” sources.

And “center” does not always mean “truthful” or “neutral.” Sometimes it means careful. Sometimes it means bland. Sometimes it means institutionally cautious. Sometimes it means avoiding claims that should actually be made.

So the long-term goal should not be:

Put a political label next to every article and call it solved.

The better goal is:

Show source tendency, article framing, sourcing depth, factual density, tone, and coverage asymmetry separately.

That is much harder, but it is also much more honest.


The difference between source bias and article framing

This distinction matters.

Source bias is about the outlet over time.

For example:

  • What stories does it usually emphasize?
  • What language does it tend to use?
  • Which political or institutional assumptions does it carry?
  • What audience does it appear to serve?
  • How often does it correct mistakes?
  • How close is it to primary-source material?

Article framing is about one specific article.

For example:

  • What facts does the headline emphasize?
  • What facts are buried?
  • What words carry emotional weight?
  • Who is quoted?
  • Who is ignored?
  • Is the piece written as reporting, analysis, advocacy, or outrage?
  • Does it separate claims from interpretation?

A serious news aggregator should not collapse those into one score.

An outlet can have a general bias while still publishing a fair article.

A generally reliable outlet can still publish a weak or misleading article.

A low-reputation source can sometimes surface a real story before institutions do.

That is why the interface needs to preserve nuance.


Permanent daily archives

One design choice I care about is permanent daily archives.

A normal feed disappears as it updates. Yesterday’s information gets buried. Last week’s framing is hard to reconstruct. The user sees the present feed, but not the shape of coverage over time.

Permanent daily archives solve part of that.

Each day becomes a stable page.

That makes it easier to answer questions like:

  • What was being covered on a specific day?
  • Which stories dominated?
  • Which topics disappeared quickly?
  • Which sources covered an event early?
  • How did the language around a story change?
  • What did the news environment look like before later context emerged?

This is useful for users, but it is also useful structurally.

A news aggregator should not only be a live feed. It should become a public memory layer.


Anonymous discussion: useful or dangerous?

WeSearch currently allows anonymous discussion.

That decision is controversial.

The upside is obvious:

People can comment without creating an account, building a profile, or turning every opinion into part of a permanent identity graph.

That lowers friction.

It also makes the product feel less like a social network and more like a public annotation layer.

But anonymity has risks:

  • spam
  • abuse
  • low-quality comments
  • astroturfing
  • drive-by political noise
  • reduced accountability
  • lower trust

The challenge is designing anonymous discussion so it does not become anonymous garbage.

Some possible approaches:

  • rate-limit comments
  • add lightweight moderation
  • separate “questions” from “opinions”
  • let users mark comments as useful, misleading, or low-effort
  • encourage source-backed replies
  • show discussion quality signals instead of identity signals
  • avoid follower counts and personality-driven posting

The key design question is whether discussion should be social or analytical.

For a news product, I think discussion should be closer to annotation than performance.


The trust problem

A news aggregator has a harder trust problem than most products.

If you build a todo app, users ask:

Does it work?

If you build a news aggregator, users ask:

Why should I trust what this thing chooses to show me?

That means the product needs visible trust signals.

Not fake authority. Real transparency.

Examples:

  • source list
  • source policy
  • correction policy
  • ranking methodology
  • bias-label methodology
  • explanation of what is automated
  • explanation of what is human-reviewed
  • clear distinction between source labels and article labels
  • visible date/time metadata
  • no pretending that the system is perfectly objective

The worst thing a news product can do is imply neutrality while hiding all the decisions that shape what people see.

A better approach is to expose the machinery.


What I would avoid

If I were designing a serious news comparison system, I would avoid a few traps.

1. Do not pretend one bias score explains an article

A single label can help orient the user, but it should not be the whole analysis.

Bias is multi-dimensional.

2. Do not over-personalize the feed

Personalization is convenient, but it quietly narrows perception.

For news, user control is better than hidden behavioral targeting.

3. Do not hide the source list

If a product claims to aggregate many sources, users should be able to see what those sources are.

4. Do not turn discussion into another social network

Follower mechanics, clout loops, and identity performance can damage the informational value of a news product.

5. Do not index thousands of empty pages

This is more of a technical SEO point, but it matters.

If a site creates source pages, tag pages, archive pages, and story pages, it needs to avoid exposing too many thin or empty URLs. Search engines and users both interpret that as low quality.


What I am still figuring out

The project is still early, and several hard questions are unresolved.

Story clustering

When ten outlets cover the same event, should those articles be grouped together automatically?

Probably yes.

But clustering can go wrong. Similar headlines do not always mean identical stories. Different angles may deserve separation.

Source weighting

Should a more reliable source receive stronger visibility?

Probably yes.

But if weighting is too aggressive, the system becomes another hidden ranking engine.

Bias display

Should bias labels be visible immediately, or should users first see the article/source and then open a deeper comparison panel?

I am not sure yet.

Immediate labels are useful, but they can also prime users before they read.

Anonymous discussion

Should anonymous comments be central to the product, or should they be secondary to source comparison?

This is still an open product question.

Search vs feed vs comparison

A news aggregator can become several different products:

  • live feed
  • searchable archive
  • RSS replacement
  • media-bias comparison tool
  • anonymous news discussion layer
  • research tool

Trying to be all of them at once can make the product confusing.

The hard part is choosing the primary job.


The current direction

Right now, I think the strongest direction is:

A chronological news aggregator with source context, permanent archives, and lightweight anonymous discussion.

Then, over time, add stronger comparison features:

  • related coverage clusters
  • source diversity views
  • article-level framing analysis
  • factuality/source-depth indicators
  • topic timelines
  • left/right/center coverage maps
  • correction and update tracking
  • “what is missing?” indicators

The product should not just answer:

What happened?

It should also help answer:

Who is covering it?
How are they framing it?
What context is missing?
Which claims are confirmed?
Which parts are interpretation?
How did coverage change over time?

That is where a news aggregator can become more than a headline feed.


Why I think this matters

The internet does not have an information shortage.

It has a context shortage.

There are endless headlines, feeds, posts, clips, takes, screenshots, and reactions.

But it is still hard to see the shape of coverage across sources.

It is hard to know which parts of a story are factual, which parts are framing, and which parts are omission.

It is hard to compare coverage without manually opening ten tabs.

It is hard to discuss news without the conversation becoming identity performance.

That is the space I am trying to explore with WeSearch.

Not a perfect truth machine.

Not another engagement feed.

Not another paywalled dashboard.

Just a clearer way to scan, compare, archive, and discuss what is being published.

The site is here:

https://wesearch.press

It is still rough in places, but the core structure is live.

I would be interested in criticism from people who care about search, RSS, journalism, media bias, recommendation systems, moderation, or information retrieval.

The question I keep coming back to is:

What would a news aggregator need to show before you would actually trust it?