惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

G
Google Developers Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
WordPress大学
WordPress大学
阮一峰的网络日志
阮一峰的网络日志
V
Visual Studio Blog
雷峰网
雷峰网
博客园_首页
The Cloudflare Blog
Hugging Face - Blog
Hugging Face - Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
爱范儿
爱范儿
小众软件
小众软件
D
Docker
P
Proofpoint News Feed
B
Blog
Vercel News
Vercel News
B
Blog RSS Feed
U
Unit 42
月光博客
月光博客
The GitHub Blog
The GitHub Blog
Apple Machine Learning Research
Apple Machine Learning Research
Y
Y Combinator Blog
I
InfoQ
Recent Announcements
Recent Announcements

Show HN

Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code.
Model Benchmark | CoinSignal
docuru · 2026-05-31 · via Show HN
Dashboard

Ranking uses verified model performance from the database. The score favors accuracy first, then hit rate, consistency, confidence calibration, and enough sample size to trust the result.

Best calibratedmoonshotai/kimi-k2.5

Verified predictions11,769

Only completed prediction windows with accuracy scores are included.

Models compared13

Grouped by model name across every tracked coin.

Best avg accuracy72.2%

Mean score from direction, range closeness, and range overlap.

Best recent form75.8%

Average accuracy over each model's latest 10 verified calls.

Accuracy

Primary quality score already recorded for each model.

Hit Rate

Share of predictions scoring at least 70%.

Consistency

Rewards models with lower accuracy variance.

Calibration

Checks whether confidence matches actual results.

Recency

Separates current form from older performance.

RankModelScoreAvg accuracyRecentHit rateConsistencyConf gapSamplesCoinsLast verified
#1openai/gpt-5.4High-confidence avg: 57.2%77.2%72.2%51.4%76.6%79.4%-7.6%1172ADA, AVAX, BNB, BTC, DOGE +4Jun 04, 12:08 AM
#2minimax/minimax-m2.7High-confidence avg: 44.8%71.7%66.1%69.7%65.6%75.7%-3.9%195ADA, AVAX, BNB, BTC, DOGE +4Apr 26, 06:05 AM
#3xiaomi/mimo-v2.5High-confidence avg: 46.2%70.0%65.3%61.8%57.9%78.7%-2.9%959ADA, AVAX, BNB, BTC, ETH +2Jun 06, 12:00 PM
#4xiaomi/mimo-v2.5-proHigh-confidence avg: 45.5%70.0%65.5%68.0%57.8%78.9%-5.1%884ADA, AVAX, BNB, BTC, ETH +2Jun 04, 06:08 PM
#5minimax/minimax-m2.5High-confidence avg: 50.7%69.6%64.5%66.1%60.5%76.4%-6.2%1105ADA, AVAX, BNB, BTC, DOGE +4Jun 04, 06:08 PM
#6openai/gpt-5-miniHigh-confidence avg: 38.2%67.8%62.4%65.1%56.8%77.3%-5.8%1144ADA, AVAX, BNB, BTC, DOGE +4Jun 04, 06:08 PM
#7qwen/qwen3.5-plus-20260420High-confidence avg: 51.5%65.5%61.0%69.0%48.0%78.4%-4.6%930ADA, AVAX, BNB, BTC, ETH +2Jun 04, 06:08 PM
#8moonshotai/kimi-k2.5High-confidence avg: 42.8%64.1%58.7%72.6%46.1%76.1%-0.6%1110ADA, AVAX, BNB, BTC, DOGE +4Jun 04, 06:08 PM
#9z-ai/glm-5.1High-confidence avg: 44.5%63.8%58.9%70.2%44.2%77.9%+1.7%1183ADA, AVAX, BNB, BTC, DOGE +4Jun 04, 06:08 PM
#10deepseek/deepseek-v4-flashHigh-confidence avg: 47.9%63.7%58.8%73.0%43.6%78.5%-2.0%949ADA, AVAX, BNB, BTC, ETH +2Jun 06, 12:01 PM
#11google/gemini-3-flash-previewHigh-confidence avg: 42.4%62.7%58.3%74.2%43.3%78.0%+8.3%1040ADA, AVAX, BNB, BTC, DOGE +4Jun 04, 06:08 PM
#12z-ai/glm-5High-confidence avg: 42.8%61.5%54.7%68.3%52.8%67.7%+9.1%195ADA, AVAX, BNB, BTC, DOGE +4Apr 26, 06:05 AM
#13deepseek/deepseek-v4-proHigh-confidence avg: 42.5%58.5%54.4%75.8%32.8%79.1%+9.1%903ADA, AVAX, BNB, BTC, ETH +2Jun 04, 06:08 PM