惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
宝玉的分享
宝玉的分享
Jina AI
Jina AI
Martin Fowler
Martin Fowler
W
WeLiveSecurity
V
Vulnerabilities – Threatpost
GbyAI
GbyAI
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
P
Privacy International News Feed
D
DataBreaches.Net
Security Archives - TechRepublic
Security Archives - TechRepublic
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Scott Helme
Scott Helme
U
Unit 42
Hacker News - Newest:
Hacker News - Newest: "LLM"
Google DeepMind News
Google DeepMind News
酷 壳 – CoolShell
酷 壳 – CoolShell
Vercel News
Vercel News
I
InfoQ
N
Netflix TechBlog - Medium
罗磊的独立博客
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
IT之家
IT之家
雷峰网
雷峰网
Hugging Face - Blog
Hugging Face - Blog
T
Tailwind CSS Blog
S
Securelist
S
Schneier on Security
T
Troy Hunt's Blog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
C
Cisco Blogs
H
Hacker News: Front Page
Spread Privacy
Spread Privacy
T
Tenable Blog
博客园 - Franky
Apple Machine Learning Research
Apple Machine Learning Research
Recent Commits to openclaw:main
Recent Commits to openclaw:main
D
Darknet – Hacking Tools, Hacker News & Cyber Security
博客园_首页
量子位
AI
AI
L
LINUX DO - 最新话题
J
Java Code Geeks
小众软件
小众软件
爱范儿
爱范儿
月光博客
月光博客
H
Help Net Security
aimingoo的专栏
aimingoo的专栏

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
NovelPilot: A Novel Writing Agent Powered by Gemma 4
Doraking · 2026-05-23 · via DEV Community

This is a submission for the Gemma 4 Challenge: Build with Gemma 4

Most AI story generators work like this:

prompt in → wall of text out

That is useful, but it does not feel like a real writing process.

When people write fiction, they do not only generate paragraphs. They plan the premise, design characters, build the world, structure the plot, manage foreshadowing, write scenes, edit style, check continuity, and prepare the final piece for readers.

So I built NovelPilot.

NovelPilot is a Gemma 4-powered AI writing room that turns one prompt into a complete story creation pipeline.

One prompt goes in.

Nine agents start working.

A finished story comes out.


What I built

NovelPilot is a web app that helps users create short fiction through a structured multi-agent workflow.

The user starts with a simple prompt, such as:

Write a melancholic sci-fi mystery set in modern Tokyo. A graduate student who lost his memory investigates a disappearance in a quantum computing lab.

Then NovelPilot launches a sequence of specialized AI agents:

  1. Premise Architect
  2. Character Director
  3. World Builder
  4. Plot Strategist
  5. Chapter Architect
  6. Prose Writer
  7. Style Editor
  8. Continuity Detective
  9. Publisher Agent

Each agent performs a specific part of the writing process.

The result is not just a generated story. It is a full creative package:

  • Story concept
  • Character profiles
  • Worldbuilding notes
  • Plot structure
  • Chapter outline
  • Chapter 1 draft
  • Style editor report
  • Foreshadowing tracker
  • Continuity detective report
  • Title ideas
  • Publication summary
  • Browser reading mode
  • Polished PDF export

NovelPilot is designed to demonstrate Gemma 4 as a multi-agent creative reasoning engine, not just a text completion model.


Demo

Live demo: https://novelpilot.vercel.app

How to try it:

  1. Open the live demo.
  2. Click Run Judge Demo.
  3. Watch the nine-agent pipeline complete.
  4. Read the finished novel in the browser.
  5. Review the Foreshadowing Tracker and Continuity Detective.
  6. Download the final story as a polished PDF.

The Judge Demo works without an API key, so reviewers can test the full experience immediately.

For live generation, NovelPilot supports Gemma 4 through a provider abstraction, with OpenRouter as the recommended provider.


Sample prompt and output

Here is the sample prompt I used to test NovelPilot.

The protagonist is Ren Kanzaki, a 24-year-old graduate student working in a quantum computing laboratory. A few days ago, he lost part of his memory. He cannot remember what he was researching, why his professor suddenly disappeared, or why his own name appears in an old experimental log.

The story begins on a rainy night in Tokyo. Ren enters the university research building after midnight and finds an old experiment log hidden inside a locked drawer. On the final page, he sees the sentence:

“Ren Kanzaki will be removed from the observation target as of today.”

The story should focus on quiet tension, memory gaps, emotional unease, and the unsettling atmosphere of the laboratory. Avoid flashy action. Let the mystery emerge through scenery, silence, dialogue, and small contradictions.

Main theme:
If memories disappear, can a person still remain the same self?

Main characters:
- Ren Kanzaki: A graduate student who lost part of his memory. Calm and intelligent, but emotionally repressed.
- Mio Shiraishi: Ren’s labmate. She knows something about Ren’s memory loss but refuses to tell him the truth.
- Professor Kuon: The missing professor. He was researching quantum memory transfer.
- Associate Professor Kurosaki: The person currently managing the laboratory. He seems helpful, but some of his statements contradict the records.

Tone:
Intellectual, quiet, melancholic, slightly literary, and mysterious.

Enter fullscreen mode Exit fullscreen mode

  • Language: en
  • Genre: sci-fi
  • Tone: melancholic
  • Target Length: short-story

I also exported the generated story as a polished PDF.

Sample output PDF: Download the generated novel PDF

This PDF was generated directly from NovelPilot’s finished reader view.


Code

GitHub repo: https://github.com/dorakingx/novelpilot

Tech stack:

  • Next.js App Router
  • TypeScript
  • Tailwind CSS
  • shadcn/ui-style components
  • Gemma 4 provider abstraction
  • OpenRouter-compatible live mode
  • Mock mode for the zero-setup judge demo
  • Browser-based polished PDF export
  • Vercel deployment

The app has two main modes:

Mode Purpose
Demo / Mock Mode Lets judges try the full workflow without an API key
Live Mode Uses Gemma 4 through the configured provider

The provider layer is intentionally isolated in lib/gemma.ts, so the model provider can be changed without rewriting the app.


How I used Gemma 4

Gemma 4 is the reasoning engine behind the multi-agent writing pipeline.

NovelPilot uses Gemma 4 for:

  • structured story concept generation
  • character design
  • worldbuilding
  • plot planning
  • chapter outlining
  • prose drafting
  • style editing
  • foreshadowing tracking
  • continuity auditing
  • publisher copy generation

Each agent receives the accumulated story bible and previous structured outputs.

This means Gemma 4 is not just generating paragraphs. It acts as the structural memory and reasoning layer for the whole novel creation process.

The important design decision was to make every agent return structured data whenever possible. That allows the UI to render the model output as real product features: timelines, cards, reports, trackers, reader views, and exports.


Why I chose this Gemma 4 model

For the live version, NovelPilot is designed to use a Gemma 4 model through OpenRouter.

I chose this approach because the app needs strong reasoning and structured generation across multiple steps. The model must follow JSON schemas, preserve context from earlier agents, and reason about story structure, character consistency, and foreshadowing.

NovelPilot focuses especially on:

  • long-context creative reasoning
  • structured JSON generation
  • story memory across multiple steps
  • continuity checking
  • literary planning and drafting

Gemma 4 is a good fit because the project is not only asking the model to write a paragraph. It asks the model to behave as a coordinated writing room.


What makes NovelPilot different

Most AI writing tools generate text.

NovelPilot generates a writing process.

The user does not only receive a draft. They see how the story is built:

Prompt
  ↓
Premise
  ↓
Characters
  ↓
World
  ↓
Plot
  ↓
Chapter outline
  ↓
Draft
  ↓
Style edit
  ↓
Continuity audit
  ↓
Publisher package
  ↓
Reader view
  ↓
PDF export

Enter fullscreen mode Exit fullscreen mode

This makes the output easier to inspect, revise, and trust.


Key feature: Foreshadowing Tracker

One of my favorite parts is the Foreshadowing Tracker.

Instead of only writing a draft, NovelPilot tracks story threads like this:

{
  "item": "The cracked silver watch",
  "introducedIn": "Chapter 1",
  "status": "unresolved",
  "suggestedPayoff": "It reveals the exact time the protagonist's memory was overwritten.",
  "payoffChapter": "Chapter 3",
  "emotionalPurpose": "Connects guilt, identity, and lost time."
}

Enter fullscreen mode Exit fullscreen mode

This makes the output more useful for writers.

It also shows why a structured model workflow matters. The app is not only asking Gemma 4 to write prose. It is asking Gemma 4 to reason about narrative structure.


Key feature: Continuity Detective

The Continuity Detective checks the generated story for structural problems.

It returns issues with:

  • category
  • severity
  • evidence
  • suggested fix

Example structure:

{
  "category": "foreshadowing",
  "severity": "high",
  "issue": "The experiment log is introduced as important but has no planned payoff.",
  "evidence": "The log appears in Chapter 1 and is referenced in the outline, but no chapter resolves its origin.",
  "suggestedFix": "Reveal in the final chapter that the log was written by an earlier version of the protagonist."
}

Enter fullscreen mode Exit fullscreen mode

This was important to me because many AI writing tools can generate plausible fiction, but fewer tools help the user understand whether the story actually holds together.


Final reader experience

After all agents finish, NovelPilot automatically transitions into a Completed Novel Reader.

The user can read the finished story directly in the browser.

They can also go back to the Agent Workspace to inspect:

  • agent outputs
  • story bible
  • foreshadowing tracker
  • continuity report
  • publisher package

The final reader is not a one-way screen. Users can freely move between the production workflow and the finished novel.


PDF export

I also added polished PDF export.

Instead of relying on the browser’s default print layout, NovelPilot generates a designed A4-style manuscript PDF.

The PDF includes:

  • cover page
  • novel title
  • metadata
  • chapter title
  • formatted manuscript body
  • optional story notes

This makes the app feel closer to a complete writing product, not just a demo.


UI/UX design

I wanted the app to feel like an AI creative studio.

The flow has three stages:

1. Prompt Launcher

The first screen is focused.

The user only sees:

  • prompt input
  • language
  • genre
  • tone
  • target length
  • Generate Story
  • Run Judge Demo

This keeps the experience simple.

2. Agent Workspace

After generation starts, the app transitions into the agent workspace.

This screen shows:

  • active agent timeline
  • story bible
  • foreshadowing tracker
  • manuscript preview
  • continuity detective
  • export tools

3. Completed Novel Reader

When all agents finish, the app opens the final reading screen.

The user can read the story, download a PDF, or go back to review the agent outputs.


Technical architecture

The core architecture is simple:

app/page.tsx
  Main app phase control:
  launcher → workspace → reader

lib/useStoryProject.ts
  Client-side orchestration of the pipeline

app/api/generate-agent/route.ts
  Runs one agent per request

lib/gemma.ts
  Provider abstraction for Gemma 4 / OpenRouter / mock mode

lib/prompts.ts
  Prompt templates for each writing agent

lib/agents.ts
  Merges structured agent outputs into the Story Bible

lib/types.ts
  Shared TypeScript types

components/
  Prompt launcher, agent workspace, reader, trackers, reports, export panels

Enter fullscreen mode Exit fullscreen mode

The app uses a state-first architecture because this is a hackathon project. I intentionally avoided authentication, databases, and user accounts so the core experience stays fast and easy to judge.


Agent workflow

Here is the high-level pipeline:

User Prompt
  ↓
Premise Architect
  ↓
Character Director
  ↓
World Builder
  ↓
Plot Strategist
  ↓
Chapter Architect
  ↓
Prose Writer
  ↓
Style Editor
  ↓
Continuity Detective
  ↓
Publisher Agent
  ↓
Completed Novel Reader + PDF Export

Enter fullscreen mode Exit fullscreen mode

Each step builds on the previous one.

For example, the Character Director does not work from the original prompt alone. It receives the premise and theme created by the Premise Architect.

The Plot Strategist receives the concept, characters, and worldbuilding.

The Continuity Detective receives the story bible, chapter outline, draft, and previous reports.

This makes the app feel like an actual production pipeline rather than a single model call.


What I learned

The biggest lesson was that structured outputs are more powerful than plain prose outputs for creative tools.

A single prose response is hard to inspect.

But structured outputs can become:

  • timelines
  • cards
  • story bibles
  • trackers
  • reports
  • reader views
  • exports

I also learned that judge experience matters.

That is why I added Run Judge Demo. Reviewers can experience the full product without configuring an API key.

Another lesson was that a creative AI product should not end at “generation complete.” It should end with something the user can actually consume. That is why I added the final reader and PDF export.


Challenges

The biggest challenge was balancing autonomy and control.

If the app is too automatic, it feels like the user has no creative role.

If the app asks for too much input, it stops feeling agentic.

So I designed NovelPilot around this principle:

The AI agents do the heavy lifting, but the user can always review, regenerate, edit, read, and export.

Another challenge was making the final output feel complete. The Completed Novel Reader and PDF export helped turn the generated draft into something closer to a finished product.


What’s next

I would like to add:

  • full multi-chapter generation
  • persistent projects
  • local storage
  • streaming agent output
  • genre-specific prompt packs
  • vertical Japanese reading mode
  • richer PDF themes
  • user-editable story bible
  • side-by-side draft revision

Final thoughts

NovelPilot is my attempt to make AI fiction generation feel less like a chatbot and more like a writing room.

The core idea is simple:

One prompt. Nine agents. A complete story pipeline.

Gemma 4 is the reasoning engine behind the process. It plans, writes, edits, tracks foreshadowing, checks continuity, and packages the final story.

That is what makes NovelPilot more than a story generator.

It is an AI-powered novel production studio.