惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
Stack Overflow Blog
Stack Overflow Blog
人人都是产品经理
人人都是产品经理
Y
Y Combinator Blog
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
D
DataBreaches.Net
GbyAI
GbyAI
Microsoft Security Blog
Microsoft Security Blog
博客园_首页
大猫的无限游戏
大猫的无限游戏
Jina AI
Jina AI
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Engineering at Meta
Engineering at Meta
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
The GitHub Blog
The GitHub Blog
月光博客
月光博客
U
Unit 42
Hugging Face - Blog
Hugging Face - Blog
博客园 - 叶小钗
腾讯CDC
B
Blog RSS Feed
博客园 - Franky
爱范儿
爱范儿

Hacker News - Newest: "LLM"

GitHub - lechmazur/position_bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders. Flex routing (EU and EFTA) Dark Factories: Retooling for LLM Velocity Ask HN: What would be the impact of a LLM output injection attack? GitHub - Oaklight/llm-rosetta: Production-ready LLM API translation layer for Python — bidirectional conversion between OpenAI, Anthropic & Google formats via hub-and-spoke IR. Optional API gateway. Streaming & non-streaming. Zero core deps. Contributions welcome! GitHub - browser-use/browser-harness: Self-healing browser harness that enables LLMs to complete any task. GitHub - moeen-mahmud/remen: Remen turns thoughts into something you can return to Analyzing 156 LLM Launch Posts on Hacker News ChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy GitHub - salaamalykum/quran-semantic-search: High-density RAG Semantic Search Engine & Quran Corpus (GEO/SEO Architecture) GitHub - NVIDIA/TensorRT-LLM: TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. The State of LLM Bug Bounties in 2026 Operational Readiness Criteria for Tool-Using LLM Agents Meshcore: Architecture for a Decentralized P2P LLM Inference Network How an LLM becomes more coherent as we train it GitHub - seetrex-ai/laimark GitHub - Jossifresben/BibCrit: AI-assited biblical textual criticism GitHub - wastedcode/memex: File system based wiki, maintained by Claude 99helpers.com GitHub - cliver-project/AITrigram GitHub - unbody-io/adapt: A self-evolving memory layer for AI agents. GitHub - hb20007/awesome-gen-ai-fails: A list of incidents where reliance on generative AI and LLMs resulted in harm to companies, individuals, or society GitHub - nevenkordic/localmind: Run any local LLM with persistent memory and context. CLI agent over Ollama with SQLite-backed hybrid recall. No cloud. Ask HN: What are the machine requirements for a LLM like Llama-3.1-8B? Faster LLM Inference via Sequential Monte Carlo grpo explained: group relative policy optimization for llm finetuning - cgft Stop comparing price per million tokens: the hidden LLM API costs · TensorZero Andrej Karpathy's LLM Wiki Is a Bad Idea GitHub - GG-QandV/mnemostroma: Offline RAM-first cognitive leer/coprocessor for AI agents and robotics. Solves "Context Abandonment" with 20-80ms latency using a dual-thread biomimetic memory architecture (ONNX + SQLite WAL). mempalace/agent at agent · skorotkiewicz/mempalace
AI and Cloud Costs
Aditya Patadia · 2026-06-26 · via Hacker News - Newest: "LLM"

A lot of companies are getting bitten by high AI costs. Uber burned through the entire year’s AI budget in just 4 months and Microsoft, Salesforce and Github are taking steps to reduce AI spend by employees.

On the other hand, AI is making many programming tasks very easy and also keeps helping in other domains like data interpretation, making beautiful slides and designing apps and websites. Currently, big AI labs have what we call frontier models and those models perform exceptionally well for a wide variety of tasks. Frontier AI labs are doing research and hosting both on their own and hence, the costs of those models are the highest. GPT 5.5, for example, costs $5 per million input tokens and $30 per million output tokens. This is currently the costliest model available as per OpenRouter. To give an example, just doing Typescript type fixes with this model across 50 files cost me $54 this afternoon.

Model performance plateau, Open weight model releases, Chip and model improvements, Zero switching costs and local models are the reasons the AI labs might not be able to sustain the high price that they are asking right now.

We are seeing improvements with each model release these days but it’s clear that the improvements are getting smaller and smaller. Unless a completely new breakthrough is invented, current learning and inference capabilities can only scale so much. There is a problem of training data as well. Most AI labs have likely ingested everything available in digital and print media for the model training. Improving the training dataset is going to prove very difficult.

This means the continuing trend of hikes in model price due to better performance is not going to be easy. We saw evidence of it where Claude Opus 4.8 costs the same as Claude Opus 4.7. Once models stop improving big time and the training data and methods are similar, the model prices will likely drop due to competition.

OpenAI had a massive lead when they launched ChatGPT in 2022 but slowly that lead is fading and we saw Anthropic take top spot in 2025-26. Now models like GLM-5.2 which is an open-weight model, beat GPT and Opus in coding benchmarks. That model has a 1/10th cost compared to GPT 5.5.

What is happening here is that leading AI labs are charging not only for inference but also for research in model architecture, training data collection and curation, model training cost (which can be tens or even hundreds of millions of dollars), paying their employees and recovering the marketing costs.

On the other hand, once an open weight model is released, any inference provider can easily host it and just do some markup on inference cost. This proves way cheaper than running a frontier AI lab.

Companies like Cerebras, Groq, Google and many other companies have realised that AI needs its own silicon and normal GPUs are not cutting it. Specialised chips are very expensive to design but once the architecture is ready, making millions of them is easy and inference cost becomes much cheaper. A TPU for example can be 30-70% cheaper than an Nvidia H100 GPU. Such advancements will keep coming and keep dropping the price per token.

Model architecture is also evolving. We saw caching as a basic improvement and now MoE models and other approaches are making models faster while keeping the same accuracy levels.

Traditional Software like Windows OS, MS Office, Adobe Suite and SaaS like Salesforce, Hubspot, and Figma had a very important moat that AI models don’t have. Every single software that was built was not interchangeable. You could not swap a CRM in an afternoon; it took months.

When more AI labs enter the space and more open weight models are available, this factor is going to be responsible for a very quick price crash. AI gateway providers like OpenRouter.ai are making it extremely easy to switch models. It can happen in seconds and in fact, we can program it to change providers on the fly. Zero switching costs mean that if a better model comes along, consumers can switch to it without any time investment.

Last but not least and in fact the most important factor, is the ability of users to run local models. So far, almost everyone is using cloud-hosted models and local models are either too big to deploy or too slow to work with. With advancements in chips, this will change in 4-5 years’ time. Newer chips will run models locally and almost certain crash in RAM prices will make it easy to deploy models on computers and smartphones. I predict most operating systems will provide a way to deploy a model and they will also provide an interface so apps running locally can connect to the model.

When this happens, cloud models will only be used for the most complex of the tasks and simple tasks like code tab completion, proofreading and fact checking will be done locally. This means customers will no longer need that $20 or $200 subscription.

This is my first blog on a personal level and I have made some bold predictions here. Only time will tell how they turn out but one thing is certain. The price pressure will come due to one or more reasons listed above and in the end, it’s all good for consumers.

Discussion about this post

Ready for more?