惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

C
Check Point Blog
aimingoo的专栏
aimingoo的专栏
Jina AI
Jina AI
Microsoft Security Blog
Microsoft Security Blog
IT之家
IT之家
V
Visual Studio Blog
量子位
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园 - 聂微东
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
大猫的无限游戏
大猫的无限游戏
Y
Y Combinator Blog
Stack Overflow Blog
Stack Overflow Blog
D
Docker
MyScale Blog
MyScale Blog
小众软件
小众软件
云风的 BLOG
云风的 BLOG
美团技术团队
Microsoft Azure Blog
Microsoft Azure Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
Last Week in AI
Last Week in AI
Apple Machine Learning Research
Apple Machine Learning Research
博客园 - 【当耐特】

Hacker News: Ask HN

The New Window Delete ChatGPT Atlas Spyware Tell HN: Qwen Free Tier Is Discontinued Ask HN: SeedLegals Partnerships in London, worth it? Ask HN: How to highlight talent from untraditional backgrounds? Ask HN: We dont need a programming language now? Durable Object alarm loop: $34k in 8 days, zero users, no platform warning What if Time at the subatomic level has multiple arrows? How to add MidnightBSD Key to UEFI Secure Boot DBX? (Revoked and Forbidden Keys) Ask HN: What's your experience working at xAI as an AI tutor? Any engineers here with experience of clinical data standards? Ask HN: Who is using OpenClaw? Agent Skills for Software Test Automation Ask HN: Who needs contributors? Claude Code is thinking too much Ask HN: What Is the Big-O Order of a Jigsaw Puzzle? Ask HN: Stepping into a new role as a Senior, mentoring dos and dont's? Founder from Zurich heading to SF and Austin for the first time Hacker News No Manual Screenshots: I Built a Scalable Screenshot API Using Cloud Playwright Ask HN: Thought experiment: AGI giving us answers we don't like? Ask HN: I quit my job over weaponized robots to start my own venture 1% Vacancy, 81% Preleased: Where Midmarket Compute Deploys in 2026 Ask HN: Preferred pricing model for sound effects libraries? Copy of the email I sent to my undergraduate professors on Nov 30, 2025 Model API Performance | Hacker News Ask HN: Are open-weight LLMs the new offline encyclopedias? Valgrind 3.27 RC1 is out Claude Code OAuth down for >12 hours Ask HN: What's Better?–Tauri or Electron?
Qwen3.5 50% expert reduction success
2026-04-16 · via Hacker News: Ask HN

The method: collect 3D activation histograms (layer × expert × rank) over domain-specific corpora, rank experts by a rank-weighted utilisation score, remove the bottom 128 of 256 per layer directly in the GGUF — no finetuning, no safetensors loading, struct-level I/O only.

The earlier work established that a 50% reduction in parameter count could result in near parity in performance in Python coding compared to the full model while losing any significant capacity to code in HTML, while the similarly constructed WEB coding specialist likewise performed at near equivalence to the full model on Web tasks but failed significantly on Python tasks.

This new work substantially completes the full proof of concept that a general MoE model can be histographically indexed and specialist models can be extracted at significant reduction in expert count and memory footprint to give an end user with constrained VRAM access to models, within a given domain, that heretofore would have been inaccessible.

Next will be a full decomposition of the new Gemma4-25b-a4b model to show applicability of this idea to a different base architecture, along with further development of the CoE (College of Experts) orchestration framework — designed to integrate a set of disk-resident specialist models into a functional system where the collective intelligence of serially-invoked specialists exceeds that of any single general model runnable within the same VRAM budget.

All 8 models (Q4_K_M GGUF, ~18B, Ollama-ready), histograms, masks, corpora and pipeline scripts: GitHub: https://github.com/JThomas-CoE/College-of-Experts-AI HF: https://huggingface.co/JThomas-CoE