惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 叶小钗
O
OpenAI News
V
V2EX
大猫的无限游戏
大猫的无限游戏
博客园 - 聂微东
S
Schneier on Security
C
CXSECURITY Database RSS Feed - CXSecurity.com
小众软件
小众软件
L
LINUX DO - 热门话题
C
Cybersecurity and Infrastructure Security Agency CISA
博客园 - Franky
Security Latest
Security Latest
S
SegmentFault 最新的问题
Project Zero
Project Zero
Spread Privacy
Spread Privacy
K
Kaspersky official blog
J
Java Code Geeks
V
Vulnerabilities – Threatpost
C
Cisco Blogs
C
CERT Recently Published Vulnerability Notes
月光博客
月光博客
T
The Exploit Database - CXSecurity.com
L
Lohrmann on Cybersecurity
人人都是产品经理
人人都是产品经理
博客园 - 三生石上(FineUI控件)
Scott Helme
Scott Helme
WordPress大学
WordPress大学
量子位
T
Threat Research - Cisco Blogs
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
宝玉的分享
宝玉的分享
Hugging Face - Blog
Hugging Face - Blog
AWS News Blog
AWS News Blog
Help Net Security
Help Net Security
Application and Cybersecurity Blog
Application and Cybersecurity Blog
Simon Willison's Weblog
Simon Willison's Weblog
S
Secure Thoughts
博客园 - 【当耐特】
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
V
Visual Studio Blog
Last Week in AI
Last Week in AI
T
Tailwind CSS Blog
腾讯CDC
Cyberwarzone
Cyberwarzone
IT之家
IT之家
GbyAI
GbyAI
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
云风的 BLOG
云风的 BLOG
T
Troy Hunt's Blog
D
Docker

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
An AI Interface for Research Papers
Justin Ross · 2026-05-22 · via Hacker News - Newest: "AI"

I have a new working paper out (with Whitney Afonso and Denvil Duncan) that does something I haven’t seen before in a research paper: a Model Context Protocol (MCP) that provides a structured way of interacting with the paper via a Large Language Model. I’m going to show you first what I did, and then tell you what I think the opportunities are available at present for scientific research papers.

Without specifics, our paper introduces a pair of randomized treatments (Priority and Performance) to elicit people’s stated preferences. We do all the normal things of a research paper: collect data about the survey respondents, show various tables, figures, and regression results. For regression results, we show you our “preferred” main specifications and selected robustness checks.

What is new is that I’ve built a MCP to allow users to interact with the data using natural language to run other regressions we either didn’t show in the paper or didn’t even think of. It also allows you to create new plots of the data, subset it for other balancing tests, and a bunch of other things (see the model notes for more).

To get a sense of this, suppose you go to the paper’s Github, download the data and server files, and follow the steps to connect the MCP to your favorite LLM. The MCP is built knowing what we have done in the paper, and with the flexibility to do new things.

You can now ask your LLM questions about the data and code. For example, “What’s the mean budget allocation by treatment arm in the practitioners sample?”

Here is a new regression specification that we did not think of until after I built the MCP, what if we interacted one of the treatment variables with the dummy variable on respondent sex? Here:

Now you know, and so do I.

Good research papers do a lot of important work. A paper states the research question. It forces the author into a coherent argument. It explains the institutional context, identification strategy, data, results, and limitations. It gives readers a canonical version of what the author believes has been shown, and what questions are reasonable to ask of the evidence.

Tyler Cowen recently asked whether AI will kill the research paper. Instead of reading a fixed PDF, a reader might press a button and update the paper with new data, rerun it under five alternative specifications, or turn it into a “meta-paper” that answers nearly any question about the subject.

I think something like that may eventually happen, but I worry it loses advantages offered by the research paper. Would literatures where motivated reasoning is strong (e.g. maybe minimum wages?) really improve here? I don’t know. In any case, a more immediate opportunity feasible right now is building a MCP as I depicted above. I think it offers efficient opportunities in replication and investigation, while retaining many of the advantages of the research paper design.

Empirical papers sit on top of a large apparatus of raw data, cleaned data, code, model choices, robustness checks, judgment calls, false starts, and dozens or hundreds of alternative specifications that never make it into the published tables.

Even when authors post replication packages, using them usually requires a reader to know the programming language, have compatible software and version control, understand the directory structure, install dependencies, decode the author’s conventions, and infer how the published tables relate to the files.

That is an interface problem: there are high costs to engaging with the research paper.

A MCP is an open standard that allows AI applications to connect to external data sources, tools, and workflows. The official MCP documentation describes it as something like a USB-C port for AI applications: a standardized way for an AI assistant to communicate with databases, files, APIs, and specialized tools rather than relying only on whatever is pasted into a chat window.

For research papers, this means the paper can come with a companion server. The server exposes the underlying data and analysis pipeline in a structured way. The reader does not need to learn the author’s coding conventions before asking a substantive question. The reader asks in ordinary language, and the AI assistant translates that question into verified analytical operations.

In my new paper, the reader can ask via their LLM “Replicate Table 1 for the mTurk sample.”

Or:

“Does the framing effect differ by respondent gender?”

The AI assistant then routes the request to the appropriate tool in the MCP server, executes the analysis, and returns the result. The local server we have built already exposes more than twenty tools across data discovery, summary statistics, balance tests, fixed-effects regressions, and publication-quality figures.

This is not the same as asking ChatGPT to “analyze the paper.” That is too loose or burn tokens predicting what algorithm for regression should be used based on your request. The point of an MCP is that the model is not improvising empirical analysis from memory or from a vague summary of the study. It is calling structured tools that the author has built, documented, and constrained.

The best near-term version of AI-assisted research is a carefully designed research environment that lets readers ask better questions of the actual replication materials. This does not eliminate the need for traditional research skill.

MCP’s create new demands of researchers that complement the existing traditional skills. To build the server means anticipating the appropriate questions readers might ask and choosing the right level of flexibility. Too little flexibility and the MCP is just a prettier replication archive. Too much flexibility and we are back to specification fishing with a friendlier user interface. MCP lets readers ask those questions without first figuring out why “if Lebron==0” (to borrow from our paper example) is an important restriction on the main results.

But the technical reviewer also gains something. If she wants to know whether I used “feols”, “reghdfe”, “lm”, or some custom function buried in a scripts folder, she can ask directly. If I have built the capability into the server, she can inspect alternative model implementations I never directly considered and see whether they matter.

This also changes what authors need to become good at. In the existing replication model, the author’s job is mostly to post enough code that a determined specialist can reproduce the paper. In the MCP model, the author has to anticipate genres of reader questions: descriptive statistics, figures, causal assumptions, robustness checks, subgroup analyses, and plausible extensions. Which brings me to the reviewer process….

Anyone who has published empirical work knows the current system.

A referee reads the paper and asks for ten thousand minor twists and clarifying exercises: add this control, drop that subgroup, cluster differently, try this alternative algorithm, show this balance test, redefine the outcome, move this figure, add that appendix table, yada yada yada.

Some of these requests are genuinely insightful. Some uncover fragility. Some help the referee understand what the author actually did. Others are make-work dressed up as rigor.

But all of them move through an absurdly slow loop: reviewer asks, author reruns, author writes a response, editor waits, reviewer rereads, and the process repeats. This takes months. Sometimes it takes years.

MCP-literate authors and reviewers could improve that process. The author provides the paper and the MCP. The reviewer, instead of asking the author to run every minor variation, connects to the paper’s MCP server and explores many of these questions directly.

If the reviewer wants to know whether the result is sensitive to a different control set, the server can allow that within a pre-specified menu. If the reviewer wants to inspect subgroup patterns, the server can generate them if the author has prepared that functionality. The reviewer should be able to get fast feedback on the petty queries without a round trip through the editorial process, and focus their feedback on whether their queries can get answers from the server.

This matters because AI is already creating a likely imbalance in academic publishing: author productivity outpacing reviewer productivity.

MCP-style replication servers are probably a partial, not full, offset to that problem. They let reviewers answer more of their own minor questions immediately, and reserve the editorial process for the questions that actually require judgment.

The immediate future I’m imagining is that the research paper remains the canonical argument. The replication package is the canonical archive. The MCP server becomes the canonical interface.

The printed article once solved a distribution problem: how to circulate an argument. The PDF solved a storage and access problem: how to make that argument easy to obtain. The replication archive solved a transparency problem: how to let others inspect the machinery behind the argument.

The MCP can solve an interface problem we didn’t even previously recognize. A research paper can not only inform on what the authors found, but also let them ask disciplined questions around the evidence presented.

The MCP in the budgeting paper is a first, hopefully rudimentary effort. It is a local server: you download it from Github (data and code) and follow the instructions to connect it to your preferred LLM. I have never had this go smoothly across several different computers in testing, but luckily LLMs are great at troubleshooting themselves.

The next version I’m working on will (hopefully) be a much simpler url link that you could connect in your LLM without downloading data or messing with your .settings folder. I have a tiny grant proposal out at IU for a summer project to set up a dedicated campus server to pilot letting people stash their MCPs for replications. This would hopefully help us develop some best practices, and serve as a proof of concept for a larger MCP repository for scientific research. Stay tuned.

My essays may be republished online or in print under Creative Commons license CC BY-NC-ND 4.0. I ask that you edit only for style or to shorten, provide proper attribution and link to this substack.

Discussion about this post

Ready for more?