惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

月光博客
月光博客
Stack Overflow Blog
Stack Overflow Blog
L
LangChain Blog
Jina AI
Jina AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
雷峰网
雷峰网
T
Tailwind CSS Blog
MongoDB | Blog
MongoDB | Blog
博客园 - 【当耐特】
博客园 - 聂微东
V
Visual Studio Blog
博客园_首页
Engineering at Meta
Engineering at Meta
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
The Cloudflare Blog
人人都是产品经理
人人都是产品经理
Apple Machine Learning Research
Apple Machine Learning Research
阮一峰的网络日志
阮一峰的网络日志
Microsoft Security Blog
Microsoft Security Blog
GbyAI
GbyAI
F
Fortinet All Blogs
C
Check Point Blog
罗磊的独立博客
H
Hackread – Cybersecurity News, Data Breaches, AI and More

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
AI should earn its keep: Introducing the AI Productivity ...
nadis · 2026-06-05 · via Hacker News - Newest: "AI"

By Scott Wu06.04.26

Companies are spending more on AI than ever, but most of them can’t tell you what they’re getting for it. Dashboards show activity metrics like tokens consumed and lines of code generated, but none of them actually answer the question: how much value is the business actually getting out?

The industry needs to move from maximizing usage metrics to maximizing outcomes — and right now, there’s no good standard for measuring that. AI vendors should be the ones to provide it.

We built an AI estimator that measures the productive engineering output Devin is providing to enterprise customers. We validated our estimator against engineers’ assessment of the time it would have taken to do the same work on their own.

The results made us confident enough to offer a guarantee to our enterprise customers: if Devin delivers less engineering value than you’re paying for, Cognition will fund your usage up to $10M until it does. We're calling it the AI Productivity Guarantee, and we hope other AI companies will move in a similar direction.

How it works

An agent reviews each completed Devin session and estimates two things:

  1. Did this session result in useful output?
  2. If so: how long would a human engineer have taken to produce the same work?

We measure in hours of productive output because lines of code don’t correspond to effort: a critical bug that takes hours to investigate might be a two-line fix. The estimator agent has access to the user’s prompt, the PR if one exists, every action Devin took, and codebase context from DeepWiki. If the session resulted in unmerged PRs or was classified as otherwise unproductive, the output is considered not useful. We assembled a dataset of human time estimates from users at our enterprise customers for validation. See the technical details of our methodology.

Validation and limitations

We asked a set of users across our enterprise customers how long their Devin tasks would have taken by hand. No single estimate is perfect, but across many tasks with varying complexity, the highs and lows average out.

This produces an estimate of engineering productivity from agents — hours of useful output. It does not replace measuring ROI, which requires deeper context on the business value of each task. At Cognition, our customer-facing teams collaborate directly with enterprises to understand the full ROI impact of their agent deployments. This estimator provides a baseline by measuring productive output. We plan to keep iterating and publishing what we learn.

The AI Productivity Guarantee

We built Cognition around delivering real engineering value. Devin is model-independent — we use the right model for each task, helping customers optimize price performance. Devin has fine-grained controls to manage spend and steer users towards more productive prompts already. Our teams embed directly into customer accounts: identifying high-value projects, pair-programming with engineers on their backlog, running enablement workshops on productively managing fleets of agents, and measuring outcomes.

Because of these features, our engagement model, and our review of historical productivity data, we’re now confident enough to put a financial commitment behind Devin's productivity in enterprise deployments. Engineering hours are converted to dollar value using a standard global rate and compared against each customer’s actual consumption near the end of their annual contract. If the value falls short, we issue credits up to $10M.

Every AI vendor should be able to tell their customers what they're getting for their money. We'd like to see more of the industry move in this direction. If you're interested in learning more about the AI Productivity Guarantee, contact us here. Existing customers can reach out to their account team.