惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Cisco Talos Blog
Cisco Talos Blog
量子位
小众软件
小众软件
Microsoft Azure Blog
Microsoft Azure Blog
V
Visual Studio Blog
I
InfoQ
Jina AI
Jina AI
The Cloudflare Blog
Recorded Future
Recorded Future
Recent Announcements
Recent Announcements
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
G
Google Developers Blog
Stack Overflow Blog
Stack Overflow Blog
阮一峰的网络日志
阮一峰的网络日志
Microsoft Security Blog
Microsoft Security Blog
美团技术团队
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
Martin Fowler
Martin Fowler
T
Tailwind CSS Blog
博客园 - Franky
酷 壳 – CoolShell
酷 壳 – CoolShell
F
Fortinet All Blogs
WordPress大学
WordPress大学
P
Proofpoint News Feed
D
DataBreaches.Net
爱范儿
爱范儿
雷峰网
雷峰网
D
Docker
B
Blog
Engineering at Meta
Engineering at Meta
腾讯CDC
N
Netflix TechBlog - Medium
C
Check Point Blog
博客园 - 【当耐特】
Apple Machine Learning Research
Apple Machine Learning Research
T
Tenable Blog
GbyAI
GbyAI
Security Archives - TechRepublic
Security Archives - TechRepublic
博客园 - 三生石上(FineUI控件)
T
The Blog of Author Tim Ferriss
博客园 - 聂微东
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
SecWiki News
SecWiki News
S
Security @ Cisco Blogs
S
Security Affairs
V
V2EX
Application and Cybersecurity Blog
Application and Cybersecurity Blog
云风的 BLOG
云风的 BLOG
C
CERT Recently Published Vulnerability Notes
Y
Y Combinator Blog

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
I built a deep learning framework in Rust from scratch — Part 3: the road to crates.io
Pavel · 2026-04-25 · via DEV Community

In Part 1 I argued why a graph-based DL framework in pure Rust was a project worth doing.

In Part 2 I wrote the GPU backend on wgpu and figured out how to make TransformerBlocktrain on it. Both posts ended with the same honest admission: the code was fine, but the project wasn't ready for other humans.

This post is about closing that gap. Six phases of work, a v0.2.0 → v0.3.1
bump, and a crate that now looks like something you'd actually reach for.

Here's the plan I committed to at the start:

Phase 1: cleanup and consistency — get to 0 warnings.
Phase 2: API reliability — declarative layer API so users don't
hand-manage HashMap<String, Shape>.
Phase 3: GPU completeness — every CPU op should have a WGSL twin.
Phase 5: ecosystem — RoPE done properly, Slice/Concat primitives, CI.
Phase 6: pre-release polish — fmt, clippy, docs.rs, CNN example.

(Phase 4 — performance — is intentionally deferred to v0.5. I'll explain.)

Phase 1 — 121 warnings to zero

The first thing I did was run cargo build --release --all-targets. The
output scrolled for a full screen:

warning: `rustyasg` (bin "rustyasg") generated 121 warnings (14 duplicates)
warning: `rustyasg` (lib) generated 21 warnings

Enter fullscreen mode Exit fullscreen mode

142 warnings feels like somebody stopped caring somewhere, but what I
actually found was amusing: src/main.rs started with

mod analysis;
mod asg;
mod autograd;
mod gui_viewer;
// ... seven more lines

Enter fullscreen mode Exit fullscreen mode

That's the binary recompiling the entire library as a separate crate. So
every pub struct in nn/ that main.rs didn't touch became "never used" —
and there are a lot of them. Fix: replace those mod declarations with
use rustyasg::*. ~100 false-positive warnings gone in one diff.

The rest was actual cleanup: deprecated rand::thread_rngrand::rng,
unused imports, a let minus_one = lit_scalar(-1.0) that was never read, and
the rotate_half function in RoPE that turned out to be a stub that added
cos as a bias
. I couldn't fix it at this point (needed Slice/Concat ops)
but I added a big doc-comment calling it out so nobody would accidentally ship
it to production:

/// **STUB IMPLEMENTATION.** Full RoPE requires Slice/Concat operations
/// to split head_dim into pairs. Current code just adds `cos` as a bias —
/// this is **not** mathematically correct RoPE and must be fixed before
/// production use.
fn rotate_half(&self, x: &Tensor, _seq_offset: usize) -> Tensor { ... }

Enter fullscreen mode Exit fullscreen mode

At the end of Phase 1 cargo build --release --all-targets was clean and I
had one open wound in RoPE to close later.

Phase 2 — the declarative layer API

This was the change I was most worried about because it was a breaking
change
for every v0.2 user. Look at the main.rs I inherited:

let layer1 = Linear::new(&ctx, "layer1");
// ...
let mut initial_shapes = HashMap::new();
initial_shapes.insert("x".to_string(), (vec![4, 2], DType::F32));
initial_shapes.insert("layer1.weights".to_string(), (vec![2, 8], DType::F32));
initial_shapes.insert("layer1.bias".to_string(),    (vec![1, 8], DType::F32));
// ...
if name.contains("w_q.weights") || name.contains("w_k.weights") || ... {
    shape = vec![embed_dim, embed_dim];
} else if name.contains("linear1.weights") {
    shape = vec![embed_dim, ff_hidden_dim];
}

Enter fullscreen mode Exit fullscreen mode

That's string-matching to figure out parameter shapes. When I saw
name.contains("w_q") in the binary I knew what Phase 2 had to be.

The fix: layers should own their shape information. Every nn::*
constructor takes dimensions, and the layer self-registers with
GraphContext:

pub fn new(ctx: &Rc<RefCell<GraphContext>>, name: &str, in_f: usize, out_f: usize) -> Self {
    let weights = Tensor::new_parameter_with_shape(
        ctx,
        &format!("{name}.weights"),
        vec![in_f, out_f],
        Initializer::XavierUniform,
    );
    // bias too...
}

Enter fullscreen mode Exit fullscreen mode

Behind the scenes GraphContext got a new parameter_meta: HashMap<String,
ParameterMeta>
where ParameterMeta carries shape, dtype, and an
Initializer. Then two helper methods close the loop:

ctx.build_shape_map(&input_shapes)     // → feeds ShapeInference
ctx.init_parameters(&mut runtime_data) // → samples weights

Enter fullscreen mode Exit fullscreen mode

I also wrote nn::init with nine standard initializers — Zeros, Ones,
Constant, Uniform, Normal, Xavier (uniform/normal), Kaiming (uniform/normal) —
because if layers are going to pick initializers on the user's behalf, the
defaults should be good defaults. Xavier-uniform for Linear, Kaiming-uniform
for Conv2d, Ones/Zeros for LayerNorm gamma/beta, Normal(0, 0.02) for
embeddings (GPT-2 conventions).

The user-facing difference is shocking. main.rs used to be 270 lines with
a hand-rolled 50-line shape-dispatch block. After Phase 2 it's 225 lines,
and the shape dispatch is one line:

ShapeInference::run_with_context(&mut graph, &ctx.borrow(), &input_shapes)?;

Enter fullscreen mode Exit fullscreen mode

The XOR example dropped from 275 lines to 190 lines with no loss of clarity.
The string-matching in the binary is gone entirely.

Obviously this broke every single example, every single test. I rewrote them
all. That's what v0.3.0 is: one coherent breaking change, one SemVer bump.

Phase 3 — GPU completeness

At the start of Phase 3, if you ran cargo run --release -- --gpu on the
TransformerBlock demo, you got:

thread 'main' panicked:
  UnimplementedOperation("node type not supported on GPU:
                          LayerNorm { input: 0, gamma: 10, beta: 11, eps: 1e-5 }")

Enter fullscreen mode Exit fullscreen mode

The README had been saying "✅ GPU backend" for months. It was lying by
omission. LayerNorm is not a composite op on GPU — it's a specialized
NodeType — and nobody had written the WGSL shaders.

So I wrote them. Four shaders for LayerNorm alone:

  • LayerNorm (forward): one worker per row.
  • LayerNormBackward (∂L/∂x): the formula is dx = inv_std * (dy·γ - mean(dy·γ) - x_norm·mean(dy·γ·x_norm)) — two reduction passes per row before the final output.
  • LayerNormGradGamma: parallelize over columns, each worker scans all rows.
  • LayerNormGradBeta: simple column-wise sum.

I added a helper dispatch_rowwise next to the existing dispatch_shader,
because LayerNorm's "one thread = one row, iterate the reduction axis
internally" pattern keeps showing up (and it did — in EmbeddingGrad later).

Then came the avalanche. Each took roughly the same pattern:

Operation Shaders added
Conv2dBackwardInput, Conv2dBackwardWeight 2
MaxPool2d, MaxUnpool2d (backward) 2
AvgPool2d, AvgUnpool2d, AdaptiveAvgPool2d 3
Embedding, EmbeddingGrad 2
ConvTranspose2d 1 (+ bias shader)

Each with a matching CPU-vs-GPU parity test. The trickiest was
MaxUnpool2d: naïvely you'd scatter from grad_output back to the positions
of max values, but WGSL doesn't have atomic f32, so concurrent scatter is a
data race. I worked around it by parallelizing over input positions — for
each (n, c, ih, iw), find which windows cover it, recompute argmax in each
of those windows, and accumulate only if we're the argmax. O(kH²·kW²) per
element, which is fine for typical 2×2 / 3×3 kernels.

After Phase 3:

Epoch  1, Loss: 9.347217
Epoch  2, Loss: 1.233081
...
Epoch 15, Loss: 0.000002
--- TRAINING COMPLETE in 576ms ---

Enter fullscreen mode Exit fullscreen mode

TransformerBlock trains on GPU end-to-end. 42 parity tests (now 46) verify
every GPU op matches the CPU reference to 1e-5.

Phase 5 — closing the RoPE wound, and ecosystem

The stub RoPE from Phase 1 had been eating at me. Fixing it required adding
three new primitives:

  • NodeType::Slice { input, axis, start, end } — the obvious building block.
  • NodeType::Concat { inputs, axis } — the dual.
  • NodeType::SliceBackward { grad_output, axis, start, full_size } — for the gradient of Slice. Concretely: zero-pad grad_output back to the original shape.

Full coverage means: NodeType + shape inference + CPU impl + GPU WGSL +
autograd. For Concat I punted on a pure-GPU implementation (would need a
multi-input kernel with dynamic strides) — the GPU path reads the inputs
back to CPU, concatenates via ndarray, and re-uploads. It's slow, but
correct, and RoPE only concatenates a few tensors of moderate size.

With Slice/Concat in place, rotate_half becomes the textbook split-half:

let x1 = x.slice(3, 0, half_dim);
let x2 = x.slice(3, half_dim, self.head_dim);

let rot1 = &(&x1 * &cos_tensor) - &(&x2 * &sin_tensor);
let rot2 = &(&x1 * &sin_tensor) + &(&x2 * &cos_tensor);

rot1.concat(&[&rot2], 3)

Enter fullscreen mode Exit fullscreen mode

Mathematically correct. End-to-end differentiable. Zero stubs.

Also in Phase 5: GitHub Actions CI, CHANGELOG.md in Keep-a-Changelog
format, and CONTRIBUTING.md documenting every place you have to edit when
adding a new NodeType (six, it turns out).

Phase 6 — the polish that makes or breaks a release

This is the unglamorous phase, but it's the one that separates "published a
crate" from "published a crate that people use."

cargo fmt --all -- --check: 28 files reformatted.

cargo clippy --all-targets -- -D warnings: 33 warnings in lib alone
when I started. Most were mechanical (div_ceil, assign_op), and
cargo clippy --fix --allow-dirty chewed through them. For three of them —
too_many_arguments, type_complexity, should_implement_trait — the
clippy suggestion was worse than the original, so I allow'd them at crate
level with a comment explaining why. The library is now clippy-clean under
-D warnings, and so are the tests, binary, and examples (modulo one
#![allow(clippy::if_same_then_else)] in the MNIST example where the
"identical blocks" are part of an 0–9 pattern-generation lookup).

Strict rustdoc: RUSTDOCFLAGS="-D rustdoc::broken_intra_doc_links" caught
ten broken links, all of the form [N, C_in, H, W] in doc-comments where
I'd written tensor shape notation inside markdown link syntax. Fix: wrap
shape notations in backticks: `[N, C_in, H, W]`.

Cargo.toml for docs.rs:

[package.metadata.docs.rs]
all-features = true
rustdoc-args = ["--cfg", "docsrs"]

[profile.release]
opt-level = 3
lto = "thin"
codegen-units = 1
strip = "debuginfo"

exclude = ["logo.png", "target/*", ".github/*", "*.log"]

Enter fullscreen mode Exit fullscreen mode

The exclude is important — the published crate is 120 KB instead of 1.3 MB.
Users don't need the logo to compile.

A real CNN example. Up to this point every example was a MLP, which made
all the Conv2d work feel theoretical. I wrote examples/cnn_classifier.rs:

Input [N, 1, 8, 8]
  → Conv2d(1→8, 3×3, pad=1) → ReLU → AvgPool2d(2×2)    → [N, 8, 4, 4]
  → Conv2d(8→16, 3×3, pad=1) → ReLU → AdaptiveAvgPool2d → [N, 16, 1, 1]
  → reshape → Linear(16→8) → ReLU → Linear(8→3)         → [N, 3]

Enter fullscreen mode Exit fullscreen mode

Trained with Adam on a 3-class synthetic dataset. Converges to 100% test
accuracy in under a second
. First real exercise of Conv2dBackwardInput,
Conv2dBackwardWeight, AvgUnpool2d, AdaptiveAvgPool2d as an actual
training loop.

The last boss: CI on three platforms

After all of that, I pushed to GitHub and got this from Actions:

  • cargo fmt (Ubuntu)
  • cargo clippy (-D warnings) (Ubuntu) — exit code 101
  • cargo doc (Ubuntu)
  • test (ubuntu-latest)
  • test (macos-latest)
  • test (windows-latest) — exit code 1

The clippy one was predictable in retrospect: my local rustc was 1.89.0,
but dtolnay/rust-toolchain@stable on CI was grabbing whatever stable
pointed at on the morning of the build. New rustc, new clippy lints, new
-D warnings failures. Fix: pin the toolchain via env var:

env:
  RUST_TOOLCHAIN: "1.89.0"

Enter fullscreen mode Exit fullscreen mode

And reference @master with the pinned version in every job. Now CI is
deterministic.

The Windows failure was more interesting. Two tests passed on Ubuntu and
macOS but failed on Windows:

#[test]
fn test_save_load_checkpoint() {
    let path = "test_checkpoint_dir";
    save_checkpoint(path, &checkpoint).unwrap();
    // ...
    fs::remove_dir_all(path).ok();
}

Enter fullscreen mode Exit fullscreen mode

cargo test runs tests in parallel by default. Two tests opening the same
relative path
in the same process's working directory is a race.
Windows' filesystem is stricter than Linux about concurrent deletion — Linux
will usually let you remove a directory while another handle is open, Windows
tells you to go away. The test passed on Ubuntu because the race was benign;
it failed on Windows because it wasn't.

Fix: temp directory with a unique suffix.

let path = std::env::temp_dir().join(format!(
    "rustyasg_ckpt_{}_{}",
    std::process::id(),
    std::time::SystemTime::now()
        .duration_since(std::time::UNIX_EPOCH)
        .unwrap()
        .as_nanos()
));

Enter fullscreen mode Exit fullscreen mode

Plus --test-threads=1 in CI as defence-in-depth. Now all three platforms
pass.

Going international: two READMEs

Until last week the README was in Russian, because I'm Russian and I never
expected anyone else to read it. That's the kind of default that quietly
keeps a project from getting discovered.

Fix: make the primary README.md English — docs.rs, crates.io, GitHub's
front page. Then mirror it as README.ru.md for Russian readers. Both
reference each other in the first lines.

Same treatment for every file in src/. I delegated the code translation
to a subagent with explicit instructions — preserve formatting, preserve
identifiers, only rewrite strings and comments — and verified with

grep -rP "[\x{0400}-\x{04FF}]" --include="*.rs" .

Enter fullscreen mode Exit fullscreen mode

The only Cyrillic left in the repo is in README.ru.md (intentional) and
a stale target/package/0.2.0/ artifact (not published, not tracked).

docs.rs docs are now professional English. That matters more than I
expected it would.

The numbers

Before After
Warnings 142 0
Tests 85 + 26 + 8 = 119 87 + 46 + 8 = 141
GPU ops supported ~25 basic + Conv2d fwd +11 (LayerNorm, Conv2d bwd, pool, embedding, ConvTranspose2d, Slice/Concat/SliceBackward)
Lines in main.rs 270 225
String-matching in binary yes zero
Layer constructors Linear::new(ctx, name) Linear::new(ctx, name, in, out)
RoPE stub (+ cos as bias) correct split-half
CI none 4 jobs, 3 OSes, strict fmt/clippy/doc
README Russian only English + Russian mirror
Published crate size 1.3 MB (with logo) ~120 KB
TransformerBlock loss (15 epochs) 0.000002 on GPU
CNN example accuracy (none existed) 100% on 3-class synthetic

Publishing

I'm tagging this one v0.3.1 and pushing to crates.io. The command list is
underwhelming given how much work led up to it:

# Ensure everything is clean.
cargo fmt --all -- --check
cargo clippy --release --all-targets -- -D warnings
cargo test --release --lib --tests --test grad_check -- --test-threads=1
RUSTDOCFLAGS="-D rustdoc::broken_intra_doc_links" cargo doc --lib --no-deps

# Commit + tag.
git add -A
git commit -m "Release v0.3.1"
git tag -a v0.3.1 -m "v0.3.1"
git push origin master --tags

# Dry-run — builds the crate archive and compiles it in isolation.
# Catches anything that works on your machine but would fail from scratch.
cargo publish --dry-run

# Ship.
cargo publish

Enter fullscreen mode Exit fullscreen mode

Then docs.rs picks it up automatically and rebuilds documentation at
https://docs.rs/rustyasg/0.3.1.

What's next (v0.5 — performance)

I deliberately deferred Phase 4. The reason is boring and correct: you can't
optimise what you haven't measured, and you shouldn't measure what isn't
correct yet. v0.3.1 is correct. v0.5 will be the performance release:

  • GPU buffer pool. Currently every training step allocates fresh buffers for every intermediate tensor. An arena that reuses allocations should be a big win — but I want to benchmark before and after with criterion so I can quote real numbers instead of vibes.
  • Kernel fusion. MatMul + Bias + Activation is a three-kernel dance that should be one WGSL shader. Detection + code generation for the fusion pass is a chunk of compiler work.
  • Mixed precision (f16) with loss scaling.
  • Inference-only mode. Skip autograd graph construction when the user just wants predictions. The graph-to-graph design makes this almost free — just don't build the gradient graph.
  • Tiny GPT example. Needs causal masking + proper multi-batch, which needs a working inference path first.

That's v0.5. And then v1.0 adds ONNX export, WebAssembly, a real model zoo,
and a much-improved visualiser.

Reflection

The thing I keep relearning: the distance between "code that works" and
"code someone else can use" is enormous and almost entirely about things
nobody celebrates. Renaming warnings. Writing CONTRIBUTING.md. Choosing a
temp path instead of a hardcoded one. Pinning a toolchain version.
Translating a README.

If you're thinking about open-sourcing a project, my honest advice: do Phase 6 first.
Get the strict CI green, get docs.rs building, get the README
to be the first thing you actually want someone to read. Then write the
code. Your future self will thank you — and so will everybody who finds the
crate on a search.

Code is at https://github.com/Xzdes/RustyAsg. Crate is at
https://crates.io/crates/rustyasg. Issues and PRs welcome.

Part 4 will be the performance push. See you there.