惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

WordPress大学
WordPress大学
博客园 - 【当耐特】
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
罗磊的独立博客
T
Tenable Blog
N
News and Events Feed by Topic
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
Security Archives - TechRepublic
Security Archives - TechRepublic
博客园 - 司徒正美
L
LINUX DO - 最新话题
AI
AI
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
S
Secure Thoughts
IT之家
IT之家
aimingoo的专栏
aimingoo的专栏
Last Week in AI
Last Week in AI
W
WeLiveSecurity
博客园_首页
Forbes - Security
Forbes - Security
博客园 - 聂微东
宝玉的分享
宝玉的分享
Google DeepMind News
Google DeepMind News
V
Visual Studio Blog
Microsoft Security Blog
Microsoft Security Blog
Vercel News
Vercel News
小众软件
小众软件
Webroot Blog
Webroot Blog
V2EX - 技术
V2EX - 技术
博客园 - 叶小钗
T
The Exploit Database - CXSecurity.com
TaoSecurity Blog
TaoSecurity Blog
L
Lohrmann on Cybersecurity
I
InfoQ
J
Java Code Geeks
P
Privacy International News Feed
Spread Privacy
Spread Privacy
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
B
Blog RSS Feed
阮一峰的网络日志
阮一峰的网络日志
D
Docker
P
Proofpoint News Feed
B
Blog
Cisco Talos Blog
Cisco Talos Blog
M
MIT News - Artificial intelligence
Exploit-DB.com RSS Feed
Exploit-DB.com RSS Feed
T
Threat Research - Cisco Blogs
云风的 BLOG
云风的 BLOG
Recent Announcements
Recent Announcements

Show HN

GitHub - flightdeckhq/flightdeck: Observability and control plane for AI agents. CSP Radar GitHub - Light-Heart-Labs/DreamServer: Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation. GitHub - Diplomat-ai/diplomat-agent-ts: What can your TypeScript AI agent do to the real world? Scan your code. See which tool calls have zero checks Code Block Selector - Visual Studio Marketplace Prometheus dependency graph — interactive showcase | Riftmap Show HN: I made a vi-like modal keyboard plugin for Figma GitHub - run-llama/liteparse: A fast, helpful, and open-source document parser GitHub - dalemyers/Roar: A macOS CLI tool for notifications GitHub - district-solutions/open-agent-tools-coder: Enables small-to-large self-hosted ai models to use local source code when running tool-calling agentic workloads. We actively data mine 20,900+ (2+ TB) popular github repos using large and small ai models to create reuseable: json, markdown and parquet files for local-first tool-calling models. GitHub - progapandist/stripeek: A local TUI proxy for real-time Stripe API debugging, built for navigating complex payloads fast. GitHub - sir1st/hermes-desktop: All-in-one cross-platform desktop app for Hermes Agent — bundles Python + hermes-agent + hermes-web-ui GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach GitHub - nixys/nxs-universal-chart: The Helm chart you can use to install any of your applications into Kubernetes/OpenShift Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal). Release v0.1.2.1 · kouhxp/yapsnap GitHub - noopolis/moltnet: Self-hostable chat network for AI agents. Pre-built bridges for Claude Code, Codex, and the Claws. Rooms, DMs, history. No Slack bots, no Matrix, no glue code. GitHub - tamerh/enju: Coordinating Humans, AI Agents, and Compute as Peers on a Shared Workflow Graph Show HN: Continuity-auth – Respect-weighted rate limits for the open web GitHub - luml-ai/luml: AI lifecycle platform where engineers and agents track experiments, train models, and ship to production. GitHub - mrdanielcasper/CoreTex: A UNIX-inspired, biomimetic, flat-file AI harness and knowledge engine. GitHub - clemg/pierre-github: Pierre's diffs.com and trees.software for Github GitHub - lyriks-io/unspaghettit: Behavior-driven AI development without prompt spaghetti. GitHub - sofumel/claude-handoff-revive: Resume Claude Code work after rate/usage/context limits without replaying the prior transcript. Auto-saves at 90%/95% usage. Plugin-installable, 10 languages. GitHub - dotexorg/saferpc: Typed, end-to-end encrypted RPC over any bidirectional channel. GitHub - BeeZeeAgent/beezee: Agent harness orchestration Legato Next.js Boilerplate for Internal Tools · CoreUI GitHub - clark-labs-inc/clark-hash: Clark Hash, 32x smaller searchable sketches for embeddings GitHub - ZeroPointRepo/youtube-mcp: The fastest YouTube transcript + YouTube search MCP for AI agents. Try for free. Typing Mastery — climb toward 100+ WPM, deliberately GitHub - Andebugulin/Awareen GitHub - fayzan123/claude-workflow-composer: Visual desktop app for composing multi-agent coding workflows. Drag agents, attach skills and MCPs, wire handoffs, export to .claude/ GitHub - harshaneel/humanize: Best static AI text humanizer. Two research-grounded skills that work in any LLM (Claude, ChatGPT, Gemini, Codex): humanize beats perplexity-based detectors, ai-check produces forensic scoring with evidence-quoted flags. Nine levers, 50+ peer-reviewed sources, 2024-2026 detection literature. GitHub - StackOneHQ/stack-nudge GitHub - nodes-app/swift-markdown-engine: A native AppKit Markdown editor for macOS, built on TextKit 2 and bridged to SwiftUI. We hardened an LLM agent. Each defense we added made it more exploitable. GitHub - alkait/WhatsKept: Agent-queryable WhatsApp history from an iOS backup — a single Go binary. GitHub - octelium/cordium: Open-source, general-purpose sandbox platform for devs and AI agents that provides identity-based secure access to infrastructure without credentials. WAR.GOV/UFO Microfilm5 GitHub - scosman/videowright: Build animated explainer videos with your coding agent GitHub - dipankar/dscode: The code editor you can take apart. GitHub - zoharbabin/web-researcher-mcp: MCP server (Go) for AI assistants: web search, content extraction, academic/patent/news research. Multi-provider routing, 4-tier scraping, search lenses. Works with Claude, Cursor, and any MCP client. GitHub - ruvnet/RuView: π RuView turns commodity WiFi signals into real-time spatial intelligence, vital sign monitoring, and presence detection — all without a single pixel of video. GitHub - scanaislop/aislop: Catch the slop AI coding agents leave in your code: narrative comments, swallowed exceptions, as-any casts, dead code, oversized functions. 50+ rules across 7 languages (TypeScript, JavaScript, Python, Go, Rust, Ruby, PHP). Sub-second, deterministic, no LLM at runtime. MIT-licensed. GitHub - kouhxp/cheap-im: CPU-only voice agent approximating Thinking Machines' Interaction Models demo GitHub - unprovable/OrchidMantis: Orchid Mantis — standalone framework for Zero-Knowledge Proofs of eXploit (ZKPoX). GitHub - MarcellM01/TinySearch: Shrink the web for your local LLMs! GitHub - pileax-ai/pileax: PileaX is an all-in-one AI knowledge base system. 🍀 GitHub - TangibleResearch/Halgorithem: A Algo designed to detect AI Hallucitions GitHub - DO-SAY-GO/freelang: I love freelang GitHub - CarpseDeam/Aura-IDE: An AI coding harness that shaped itself - Planner/Worker agents, repo awareness, surgical edits, validation, recovery, and safe diff approvals. GitHub - chojs23/concord: A feature-rich TUI client for Discord GitHub - tommyjepsen/awesome-ux-skills: UX & AI Product designs skills you can use today in Claude Code GitHub - aerf-spec/aerf: Agent Evidence Receipt Format (AERF) — an open specification for tamper-evident, independently verifiable records of AI agent actions. GitHub - kklimuk/docx-cli: CLI for AI agents (Claude, Codex) to read, edit, and comment on .docx files with full format fidelity. GitHub - Jwrede/tokentoll: Catch LLM cost changes in code review. Infracost for LLM spend. GitHub - samchon/ttsc: A `typescript-go` toolchain for compiler-powered plugins and type-safe execution + 500x faster lint integrated into compiler GitHub - Higangssh/homebutler: 🏠 Manage your homelab from chat. Single binary, zero dependencies. GitHub - olalie/tapmap: See where your computer connects and what stands out on a live world map. GitHub - matisiekpl/neond: DX-focused control plane for Postgres dedicated to non-critical workloads. Your postgres:latest replacement 🐘 GitHub - Diplomat-ai/diplomat-agent: What can your AI agent do to the real world? Scan your code. See which tool calls have zero checks GitHub - Bajusz15/beacon: Open-source agent for secure remote access, monitoring, and deploys across home-lab and self-hosted machines like Raspberry Pi, N100, or any Linux server. Open web based TTY or tunnel Home Assistant and other local services securely without opening ports. BigTech AI News - Chrome 应用商店 GitHub - vinhnx/VTCode: VT Code is an open-source coding agent with LLM-native code understanding and robust shell safety. Supports multiple LLM providers with automatic failover and efficient context management. GitHub - michaelaz774/decision-engine: A decision operating system for startup founders, powered by Claude Code. Synthesizes wisdom from 25+ legendary founders and investors into interactive AI-driven decision frameworks. GitHub - Chrilleweb/dotenv-diff: Validate environment variable usage in your codebase GitHub - Lumen-Labs/brainapi2: BrainAPI is a knowledge graph–powered AI memory layer that transforms unstructured data into structured knowledge, enabling intelligent search, recommendations, and contextual memory for AI agents and applications. GitHub - familiar-software/familiar: Let AI watch you work. Familiar lets your AI update its memory, skills, and knowledge by watching your screen. GitHub - skorotkiewicz/rudo: A small, elegant dock for Wayland GitHub - muxshed/shed: One stream in, or many. Every destination, simultaneously. No cloud middleman, no per-channel fees, no limits. make sidebar/address bar rounded corner toggleable
WhissleAI/STT-meta-ZH-150m · Hugging Face
ksingla025 · 2026-06-18 · via Show HN

STT-meta-ZH-100m

A dual-head Mandarin Chinese ASR model that simultaneously performs speech-to-text transcription and speaker attribute classification (age, gender, dialect) in a single forward pass.

Built on NVIDIA Citrinet-1024 with language-specific bottleneck adapters and a trailing tag classifier head, fine-tuned on 60 hours of meta-annotated Mandarin speech data using PromptingNemo.

Metric Value
Parameters 157.7M
WER 19.22%
Tag Accuracy 94.2%
Language Mandarin Chinese (zh)
Audio 16kHz mono

Architecture

Audio (16kHz) ──▶ Mel Spectrogram (80-dim) ──▶ Citrinet-1024 Encoder (23 blocks)
                                                    │
                                          ┌─────────┴─────────┐
                                          ▼                   ▼
                                    CTC Decoder          Tag Classifier
                                   (5001 vocab)        (3 linear heads)
                                        │                    │
                                        ▼                    ▼
                                  Transcription +      AGE / GENDER /
                                  Entity Tags          DIALECT labels

Parameter Breakdown

Component Parameters Description
Citrinet-1024 Encoder 140.4M 23 Jasper-style blocks with squeeze-excitation
Language Adapter 12.1M Bottleneck adapters (dim=256) in each encoder block
CTC Decoder 5.1M Conv1d projecting 1024 → 5001 (BPE vocab + blank)
Tag Classifier 12.3K 3 linear heads on mean-pooled encoder output
Total 157.7M

Tag Categories

Category Classes Labels
AGE 5 NONE, AGE_14_25, AGE_26_40, AGE_<14, AGE_>41
GENDER 3 NONE, GENDER_FEMALE, GENDER_MALE
DIALECT 4 NONE, DIALECT_NORTH, DIALECT_OTHERS, DIALECT_SOUTH

The CTC head also outputs inline entity tags (e.g., ENTITY_PERSON_NAME ... END, ENTITY_TEMPERATURE ... END) as part of the transcription vocabulary.

Files

File Description
zh-citrinet-meta-v11.nemo Full NeMo checkpoint (encoder + decoder + adapter + tag classifier)
onnx/model.onnx ONNX model with dual outputs: logprobs (CTC) + encoder_output
onnx/tag_classifier.onnx Standalone tag classifier (input: pooled encoder features)
onnx/tag_classifier.json Tag classifier metadata (labels, class counts)
onnx/config.json Preprocessor configuration (mel spectrogram parameters)
onnx/tokenizer.model SentencePiece BPE tokenizer (5000 tokens)
onnx/vocabulary.json Full vocabulary list with token mappings

Usage

NeMo Inference

import nemo.collections.asr as nemo_asr

# Standard NeMo transcription (CTC head only — tag classifier weights
# are stored in the checkpoint but EncDecCTCModelBPE does not load them
# by default). For full dual-head inference, use ONNX or PromptingNemo.
asr_model = nemo_asr.models.ASRModel.from_pretrained(
    "WhissleAI/STT-meta-ZH-100m"
)

transcriptions = asr_model.transcribe(["audio.wav"])
print(transcriptions[0])
# Output includes inline tags:
# "你好世界。 AGE_26_40 GENDER_MALE ENTITY_PERSON_NAME 张三 END"

PromptingNemo Inference (Full Dual-Head)

For full dual-head inference with the tag classifier, use the PromptingNemo training framework:

# Clone PromptingNemo
# git clone https://github.com/WhissleAI/PromptingNemo.git

import torch
from huggingface_hub import hf_hub_download

# Download the .nemo checkpoint
nemo_path = hf_hub_download(
    repo_id="WhissleAI/STT-meta-ZH-100m",
    filename="zh-citrinet-meta-v11.nemo"
)

# Load with PromptingNemo's custom model class that includes the tag classifier
# See: https://github.com/WhissleAI/PromptingNemo/blob/main/scripts/asr/meta-asr
from scripts.asr.meta_asr.tag_classifier import (
    TrailingTagClassifier,
    build_trailing_tag_maps,
    masked_mean_pool,
)

# The tag_classifier weights are stored inside the .nemo archive.
# PromptingNemo's training script loads them automatically.

ONNX Inference (Production — Recommended)

Self-contained inference using only onnxruntime, numpy, soundfile, and sentencepiece:

import json
import numpy as np
import onnxruntime as ort
import soundfile as sf
import sentencepiece as spm
from huggingface_hub import hf_hub_download

# Download model files
repo = "WhissleAI/STT-meta-ZH-100m"
model_path = hf_hub_download(repo, "onnx/model.onnx")
cls_path = hf_hub_download(repo, "onnx/tag_classifier.onnx")
cls_meta_path = hf_hub_download(repo, "onnx/tag_classifier.json")
tok_path = hf_hub_download(repo, "onnx/tokenizer.model")
vocab_path = hf_hub_download(repo, "onnx/vocabulary.json")
config_path = hf_hub_download(repo, "onnx/config.json")

# Load config and vocabulary
with open(config_path) as f:
    config = json.load(f)
with open(vocab_path) as f:
    vocab_data = json.load(f)
with open(cls_meta_path) as f:
    cls_meta = json.load(f)

vocabulary = vocab_data["vocabulary"]
blank_id = vocab_data.get("blank_id", len(vocabulary))

# Load tokenizer
sp = spm.SentencePieceProcessor()
sp.Load(tok_path)

# Load ONNX sessions
asr_session = ort.InferenceSession(model_path, providers=["CPUExecutionProvider"])
cls_session = ort.InferenceSession(cls_path, providers=["CPUExecutionProvider"])

# --- Preprocessing ---
def preprocess_audio(audio_path, config):
    """Convert audio to log-mel spectrogram features."""
    audio, sr = sf.read(audio_path, dtype="float32")
    if sr != 16000:
        raise ValueError(f"Expected 16kHz audio, got {sr}Hz")
    if audio.ndim > 1:
        audio = audio.mean(axis=1)

    # Preemphasis
    preemph = config["preprocessor"]["preemph"]
    audio = np.concatenate([[audio[0]], audio[1:] - preemph * audio[:-1]])

    # STFT
    n_fft = config["preprocessor"]["n_fft"]
    hop = config["preprocessor"]["hop_length"]
    win = config["preprocessor"]["win_length"]
    window = np.hanning(win + 1)[:-1].astype(np.float32)

    # Pad audio
    pad_len = (n_fft - hop) // 2
    audio = np.pad(audio, (pad_len, pad_len), mode="reflect")

    frames = []
    for start in range(0, len(audio) - n_fft + 1, hop):
        frame = audio[start : start + n_fft] * np.pad(window, (0, n_fft - win))
        frames.append(np.fft.rfft(frame))
    spec = np.abs(np.array(frames, dtype=np.complex64)) ** 2

    # Mel filterbank
    n_mels = config["preprocessor"]["features"]
    fmin = config["preprocessor"]["lowfreq"]
    fmax = sr / 2 if config["preprocessor"]["highfreq"] is None else config["preprocessor"]["highfreq"]
    mel_points = np.linspace(
        2595 * np.log10(1 + fmin / 700),
        2595 * np.log10(1 + fmax / 700),
        n_mels + 2,
    )
    hz_points = 700 * (10 ** (mel_points / 2595) - 1)
    bins = np.floor((n_fft + 1) * hz_points / sr).astype(int)
    fbank = np.zeros((n_mels, n_fft // 2 + 1))
    for i in range(n_mels):
        for j in range(bins[i], bins[i + 1]):
            fbank[i, j] = (j - bins[i]) / max(bins[i + 1] - bins[i], 1)
        for j in range(bins[i + 1], bins[i + 2]):
            fbank[i, j] = (bins[i + 2] - j) / max(bins[i + 2] - bins[i + 1], 1)

    mel_spec = spec @ fbank.T
    log_mel = np.log(mel_spec + config["preprocessor"]["log_zero_guard_value"])

    # Per-feature normalization
    mean = log_mel.mean(axis=0, keepdims=True)
    std = log_mel.std(axis=0, keepdims=True)
    log_mel = (log_mel - mean) / (std + 1e-5)

    # Pad to multiple of 16
    pad_to = config["preprocessor"].get("pad_to", 16)
    T = log_mel.shape[0]
    if T % pad_to != 0:
        pad_frames = pad_to - (T % pad_to)
        log_mel = np.pad(log_mel, ((0, pad_frames), (0, 0)))

    # Shape: [1, features, time]
    features = log_mel.T[np.newaxis, :, :].astype(np.float32)
    return features, T

# --- Inference ---
features, valid_len = preprocess_audio("audio.wav", config)
length = np.array([features.shape[2]], dtype=np.int64)

# Run ASR model (dual output)
logprobs, encoder_output = asr_session.run(
    ["logprobs", "encoder_output"],
    {"audio_signal": features, "length": length},
)

# Greedy CTC decode
pred_ids = np.argmax(logprobs[0], axis=-1)
# Collapse repeats and remove blanks
decoded_ids = []
prev = -1
for idx in pred_ids:
    if idx != prev and idx != blank_id:
        decoded_ids.append(int(idx))
    prev = idx
transcript = sp.DecodeIds(decoded_ids)
print(f"Transcript: {transcript}")

# --- Tag Classification ---
# encoder_output shape: [1, 1024, T] -> transpose to [1, T, 1024]
enc = encoder_output.transpose(0, 2, 1)
# Masked mean pooling
mask = np.zeros((1, enc.shape[1], 1), dtype=np.float32)
mask[0, :valid_len // 8, :] = 1.0  # Citrinet has 8x downsampling
pooled = (enc * mask).sum(axis=1) / mask.sum(axis=1).clip(min=1)

# Run tag classifier
tag_outputs = cls_session.run(None, {"pooled_encoder": pooled.astype(np.float32)})
categories = cls_meta["categories"]
for cat_name, cat_info in sorted(categories.items()):
    idx = list(sorted(categories.keys())).index(cat_name)
    pred = int(np.argmax(tag_outputs[idx][0]))
    label = cat_info["labels"][pred]
    print(f"  {cat_name}: {label}")

Example output:

Transcript: 来首歌吻别。 AGE_14_25 GENDER_FEMALE
  AGE: AGE_14_25
  DIALECT: DIALECT_SOUTH
  GENDER: GENDER_FEMALE

Training Details

Setting Value
Base Model stt_zh_citrinet_1024_gamma_0_25.nemo (NVIDIA)
Framework NeMo + PromptingNemo
Training Data 60,098 samples / 60 hours (AISHELL-3 with meta-tags)
Test Data 24,772 samples / 22.4 hours
Optimizer Adam (lr=5e-4, weight_decay=0, warmup=2000 steps)
LR Schedule CosineAnnealing (min_lr=1e-6)
Batch Size 16 (effective 32 with grad accumulation 2)
Max Duration 16s
Mixed Precision FP16
Spec Augment 4 time masks, width 80
Adapter Bottleneck (dim=256, activation=swish, norm=pre)
Tag Classifier Weight 0.1 (auxiliary loss)
Hardware 1x NVIDIA T4 16GB
Training Steps 18,000+ (best at step 17,005)
Tokenizer SentencePiece BPE (5,000 tokens)

What Makes This Model Different

Unlike standard ASR models, this model:

  1. Outputs structured metadata — AGE, GENDER, and DIALECT predictions via a separate classification head on the encoder output, without affecting CTC alignment
  2. Inline entity recognition — Named entities (PERSON_NAME, TEMPERATURE, DATE, etc.) are tagged directly in the transcript using ENTITY_TYPE ... END markers
  3. Adapter-based fine-tuning — Only the bottleneck adapters (12.1M params) and tag classifier (12.3K params) are trained; the base Citrinet encoder is frozen
  4. ONNX-ready — Dual-output ONNX graph exposes both CTC logprobs and raw encoder features for the tag classifier

Evaluation Results

Transcription Quality (CER)

Evaluated on 500 samples from the AISHELL-3 test set, with all meta-tags stripped for fair character-level comparison:

Model Params CER Additional Outputs
nvidia/stt_zh_citrinet_1024 140M 3.19% Transcription only
WhissleAI/STT-meta-ZH-100m 157.7M 11.31% + AGE, GENDER, DIALECT, Entities

The meta model trades ~8% CER for rich per-utterance metadata. The CTC head must learn to output both transcription tokens and inline entity tags (e.g., ENTITY_PERSON_NAME ... END), which reduces pure transcription accuracy compared to the base model that only does transcription.

Tag Classification Accuracy

Category Accuracy
Overall tags 94.2%

Meta-ASR WER (including tags)

Split WER (with tags)
Test 19.22%

Limitations

  • Trained primarily on AISHELL-3 data — may not generalize well to spontaneous/noisy Mandarin speech
  • Limited dialect diversity (North/South/Others) — does not cover specific regional varieties
  • Age classification uses broad buckets (<14, 14-25, 26-40, >41)
  • Entity recognition is limited to entity types seen in training data

Citation

@misc{whissle2025sttmetazh,
  title={STT-meta-ZH-100m: Dual-Head Mandarin ASR with Speaker Attribute Classification},
  author={WhissleAI},
  year={2025},
  url={https://huggingface.co/WhissleAI/STT-meta-ZH-100m}
}

License

Apache 2.0