惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

云风的 BLOG
云风的 BLOG
C
CERT Recently Published Vulnerability Notes
阮一峰的网络日志
阮一峰的网络日志
G
Google Developers Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
P
Privacy International News Feed
N
News and Events Feed by Topic
博客园 - Franky
Spread Privacy
Spread Privacy
P
Privacy & Cybersecurity Law Blog
T
Tor Project blog
博客园_首页
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Hugging Face - Blog
Hugging Face - Blog
P
Proofpoint News Feed
博客园 - 叶小钗
S
Securelist
Stack Overflow Blog
Stack Overflow Blog
The Cloudflare Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
V
Vulnerabilities – Threatpost
量子位
D
Docker
NISL@THU
NISL@THU
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
美团技术团队
cs.CL updates on arXiv.org
cs.CL updates on arXiv.org
Engineering at Meta
Engineering at Meta
小众软件
小众软件
F
Fortinet All Blogs
Cisco Talos Blog
Cisco Talos Blog
N
News | PayPal Newsroom
F
Full Disclosure
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
B
Blog RSS Feed
博客园 - 三生石上(FineUI控件)
WordPress大学
WordPress大学
Apple Machine Learning Research
Apple Machine Learning Research
有赞技术团队
有赞技术团队
Martin Fowler
Martin Fowler
T
Threat Research - Cisco Blogs
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
H
Heimdal Security Blog
L
Lohrmann on Cybersecurity
IT之家
IT之家
Webroot Blog
Webroot Blog
P
Palo Alto Networks Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Cloudbric
Cloudbric
Blog — PlanetScale
Blog — PlanetScale

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant Common SOC 2 Failures (Real World) Stop Vibe-Checking Your AI App: A Practical Guide to Evals How to Use SonarQube and SonarScanner Locally to Level Up Your Code Quality Your Next To-Do App Is Dead — I Replaced Mine with an OpenClaw AI Sign a Nostr event in 60 lines of Python using coincurve — no nostr-sdk, no nbxplorer, no rust toolchain ITGC Audit Explained Like You’re in Big 4 Patch Tuesday abril 2026: Microsoft parcha 163 vulnerabilidades y un zero-day en SharePoint Stop scraping everything: a better way to track competitor price changes Listing on MCPize + the Official MCP Registry while routing payments OUTSIDE the marketplace — how I kept 100% of my x402 revenue Building an AI-Powered Risk Intelligence System Using Serverless Architecture Why We Ripped Function Overloading Out of Our AI Toolchain Testing AI-Generated Code: How to Actually Know If It Works SaaS Churn Is Killing Your Business. Here Is What to Do About It (Without a Support Team) The Speed of AI Is No Longer Linear - And Self-Improving Models Are Why How to Implement RBAC for MCP Tools: A Practical Guide for Engineering Teams From Standard Quote to Persuasive Proposal: AI Automation for Arborists I built a CLI that scaffolds complete multi-tenant SaaS apps Axios CVE-2025–62718: The Silent SSRF Bug That Could Be Hiding in Your Node.js App Right Now The dashboard that ended our friendship Data Pipelines Explained Simply (and How to Build Them with Python) The Hidden Cost of AI Systems Nobody Talks About. undefined vs undeclared, and how typeof behaves Switching from file-based jobs to NATS/Kafka in Rust without changing code io_uring Adventures: Rust Servers That Love Syscalls Why Agentic AI is Killing the Traditional Database The POUR principles of web accessibility for developers and designers Quantum Neural Network 3D — A Deep Dive into Interactive WebGL Visualization How To Install Caveman In Codex On macOS And Windows Automation Pipeline Reliability: Why Your Workflow Breaks When Nobody Is Watching I Built an 'Open World' AI Coding Agent — It Works From ANY Folder From Freelancing to Product: A Tech Service Company's SaaS Transformation China's AI Giants: Adding Tencent Hunyuan & ByteDance Doubao to AI University (74 Providers) On the Vibe Coders and Their Lies clerk: Auto-Summarize Your Claude Code Sessions AI Weekly — 2026/04/10–04/17 | The Model Lockdown Is Here, but the Toolchain Is the Real Battleground AI 週報 — 2026/04/10–2026/04/17 模型封鎖潮來了,但工具鏈才是真戰場 Maybe this is how Open-Source apps are born... 🚀 Fine-Tune LLMs with LoRA and QLoRA: 2026 Guide tRPC v11 + Next.js App Router: End-to-End Type Safety Without the Boilerplate ShadCN UI in 2026: Why I Stopped Installing Component Libraries and Started Owning My Components SaaS Billing in React Server Components: Stripe + Supabase Without a Single `useEffect` Join our DEV Weekend Challenge — $1,000 in Prizes Across TEN winners! Submissions Due April 20 at 6:59 AM UTC. Implementing FSRS Spaced Repetition in Flutter + Supabase — Adding Memory Science to an AI Learning App "I Texted My Localhost From the Train — Claude Code Fixed the Bug Before I Got Home" I Built a Sales Prep AI and It Went Deeper Than Expected Design to Code #2: One JSON, Eleven Outputs Solving the 100M-Row Problem: A Summary Table Pattern for High-Volume Push Notification Logs Flutter Web With Wasm: What Actually Changes For Developers I Built 50 Royalty-Free Soundtracks for My Side Project in a Weekend Using AI Music Generation The Vibe Coding Security Checklist: 7 Things to Check Before You Ship Stop Letting Googlebot Guess Fix Your React App's SEO Right Desconstruindo o Streaming do LinkedIn: Como Criar um Engine de Extração de Vídeo de Alta Performance com HLS e FFmpeg (EDA Part-1) EDA (Exploratory Data Analysis) Explained With Real Life — Why Looking at Your Data Is the Most Important Step in Machine Learning Brand Relationship Management at Scale: Our 4-Touch Outreach System for 200+ Brands Why String.fromEnvironment() Might Return an Empty String in Dart JGuardrails 1.0.0 — Hardening Java LLM Apps Against Jailbreaks, Toxicity, and Prompt Injection Plan and Schedule a Full Week of Threads Content From One Claude Conversation Coding Cat Oran Ep3, Five Tables Changed Everything Updated: BFF Pattern I'm done watching freelancers get buried by 200 proposals. So I'm building the alternative. This is my first post BFS Algorithm in Java Step by Step Tutorial with Examples Tracking LLM Pricing Monthly: An Open Dataset for 22 AI Models How We Measure Content ROI on a Comparison Site: Revenue Attribution Without Perfect Data Introducing Nova AI Ops: The AI-Native Operating System for SRE Teams I built a free desktop video downloader for Windows — Grabbit How Talkie OCR Helps Vision-Impaired & Dyslexic Users Read the World Around Them VRCFaceTracking安装和iPhone面捕配置教程,有bug Even CrowdStrike Can't See Your Agents The Automation Gold Rush: What n8n Workflows and Claude Are Opening Up for Developers Right Now
Vertical Cognitive Depth and Structured Reasoning: A Practical Hypothesis for Robust Behavior Beyond Training Data
Алексей Горм · 2026-05-12 · via DEV Community

Most modern AI systems look impressive—until the problem shifts slightly. A small change in context, a new combination of known elements, or an implicit contradiction is often enough to break otherwise strong models. This article explores a concrete hypothesis: that robustness under such shifts depends not only on model size or training data, but on a missing internal capability—how deeply a system can process contradictions. We call this capability Vertical Cognitive Depth (VCD) and examine how a structured reasoning process (A11) may help expose and partially compensate for its absence.


1. The Problem: Generalization Breaks Outside Familiar Data

Neural networks are highly effective within the distribution they were trained on. However, they often struggle when:

  • Inputs differ slightly from training data (distribution shift)
  • Known components appear in new combinations
  • The task requires resolving implicit contradictions

This is commonly referred to as out-of-distribution (OOD) generalization failure. Standard metrics such as:

  • accuracy
  • perplexity
  • benchmark scores

do not reliably predict how a model behaves under these conditions.

In practice, two models with similar architecture and performance can show qualitatively different reasoning stability when faced with novel or conflicting inputs.


2. Observation: Reasoning Fails at the Point of Conflict

Empirically, many failures share a pattern:

  • The model encounters a conflict between constraints and available knowledge
  • Instead of resolving it, the model implicitly smooths or ignores the contradiction
  • A plausible but incorrect answer is produced

This suggests that the failure is not only about missing knowledge, but about how the model handles internal inconsistency.


3. Hypothesis: Vertical Cognitive Depth (VCD)

We introduce the concept of Vertical Cognitive Depth (VCD):

The capacity of a system to detect, maintain, and transform contradictions between constraints and knowledge without prematurely resolving them.

Key properties:

  • Not model depth (number of layers)
  • Not context length
  • Not chain-of-thought length

Instead, VCD describes the ability to:

  1. Detect a contradiction
  2. Hold it explicitly (without collapsing it into a guess)
  3. Use it to generate a revised direction or framing

In this sense, VCD is a latent behavioral parameter rather than a directly measured architectural feature.


4. Why Existing Metrics Fall Short

Current proxies for reasoning ability fail to capture this dimension:

  • Perplexity measures prediction quality, not conflict handling
  • Accuracy hides internal failure modes
  • Chain-of-thought length measures verbosity, not depth

A model may produce long explanations while still collapsing contradictions early.

Thus, none of these metrics reliably indicate whether a system can sustain structured reasoning under tension.


5. Structured Reasoning (A11) as an External Scaffold

A11 is a structured reasoning protocol that separates:

  • S1 — Direction (intent / goal)
  • S2 — Constraints (limits, risks, conditions)
  • S3 — Knowledge (available information)

The critical step is:

  • S4 — Explicit integration, where conflicts between S2 and S3 are surfaced

In unstructured reasoning, this conflict is often skipped or implicitly resolved.
A11 enforces:

  • explicit detection of inconsistency
  • delayed resolution
  • possible revision of S1 (goal or framing)

This does not increase the model’s knowledge, but changes how it navigates gaps.


6. A11 Does Not Solve Generalization in Its Current Form (But Changes Behavior)

It is important to be precise:

  • In its current application, A11 does not modify model weights
  • It does not introduce new knowledge
  • It does not yet eliminate out-of-distribution limitations

However, this is a limitation of how A11 is currently used (as an external reasoning scaffold), not necessarily a fundamental limitation of the approach itself.

Even in its current form, A11 can:

  • reduce premature convergence to incorrect answers
  • increase transparency of failure modes
  • enable construction of solutions from partial knowledge

In other words, A11 may improve behavior under uncertainty, even if it does not yet increase underlying generalization capacity.

This distinction matters:
A11 should not be viewed as a solution to generalization, but as a mechanism that changes how models behave when generalization fails.


7. Linking VCD and A11

A11 can be interpreted as an external mechanism that simulates VCD-like behavior:

VCD capability A11 mechanism
Detect contradiction S2 vs S3 comparison
Hold contradiction explicit S4 step
Transform contradiction revision of S1

This suggests:

Systems with low intrinsic VCD may benefit from structured reasoning scaffolds that enforce conflict retention.


8. Toward an Operational Test for VCD

For VCD to be meaningful, it must be testable.

A minimal experimental setup:

  1. Construct tasks with deliberate tension:
  • conflicting constraints
  • incomplete knowledge
  • ambiguous goals
  1. Measure:
  • whether the model explicitly detects contradiction
  • how long it maintains it before resolution
  • whether it revises its approach
  1. Compare across models with similar standard metrics

Hypothesis:

VCD-like behavior will better predict robustness than accuracy or perplexity alone.


9. Implications

If validated, VCD would suggest:

  • A missing dimension in evaluating reasoning systems
  • A way to distinguish surface competence from structural stability
  • A design target for both prompting methods and training objectives

It also reframes some failures:

  • Not as lack of knowledge
  • But as failure to process contradiction

Conclusion

Modern AI systems often fail not because they lack information, but because they cannot sustain reasoning when information conflicts. Vertical Cognitive Depth (VCD) is a hypothesis that attempts to capture this missing capability.

Structured reasoning frameworks like A11 do not solve generalization directly, but they expose and partially compensate for this limitation by forcing explicit interaction with contradictions.

Whether VCD represents a new measurable property or a reformulation of existing ideas depends on one thing: the ability to turn it into a reproducible test.

Until then, it remains a useful lens for analyzing why strong models sometimes fail in simple, unfamiliar ways.


Appendix: A11 (Version 2026) — Full Structured Specification

The following specification is included to make the A11 reasoning model explicit and interpretable, especially for systems that rely on structured representations of reasoning processes.


1. Immutable Core (S1–S3)

  • S1 — Will
    (Intent, direction, goal)

  • S2 — Wisdom
    (Priorities, values, constraints, risks)

  • S3 — Knowledge
    (Facts, models, methods, structure)

Signals from S2 and S3 are processed in parallel and passed to S4.


2. S4 — Comprehension (Origin Point of Qualia)

S4 INTEGRITY is the central rule.

S4 receives signals from S2 and S3. Order, speed, and depth do not matter.
Integration must be maximally honest.

If full honest integration is not possible:

  • It is forbidden to smooth tension, fabricate coherence, or close contradictions artificially

  • A TensionPoint must be explicitly identified (a concrete gap between S2 and S3)

  • A new S1 (fork, not replacement) must be generated strictly from this TensionPoint

  • It is forbidden to:

    • paraphrase the original S1
    • generalize it loosely
    • produce a semantically equivalent goal
  • The new S1 must be:

    • sharper
    • more specific
    • more operational

3. S4 Integrity Log

The Integrity Log is an append-only mechanism for recording structural breaks.

Each entry contains:

  • S2_signal
  • S3_signal
  • TensionPoint (explicit contradiction)
  • Reason (why integration failed)
  • NewS1 (generated fork)
  • Hash(prev) — reference to previous entry
  • Timestamp

Properties:

  • Append-only
  • Hash-linked chain (tamper-resistant)
  • No deletion of history
  • Acts as an internal validator of reasoning integrity

4. S5–S10 INTEGRITY

In a full A11 pass (S1–S11):

  • All levels S5–S10 are mandatory
  • Skipping any level is not allowed
  • Transition to S11 is only valid after explicit traversal of all six levels

If S5–S10 are not completed:

  • The system must explicitly state the reason
  • Absence of a reason = structural violation

Lite Mode (S1–S4):

  • S5–S10 are skipped
  • Activated only via Switch Flags
  • Otherwise, full pass is the default

5. Operational Layer / “Living Phase” (S5–S10)

      ┌───────────────────────┬───────────────────────┐
      │   Projective Layer    │   Practical Layer     │
      │   S5  S6              │   S8  S9              │
      │ (Freedom / Constraint)│ (Freedom / Constraint)│
      │           ↑           │           ↑           │
      │        Balance (S7)   │        Balance (S10)  │
      └───────────────────────┴───────────────────────┘

Enter fullscreen mode Exit fullscreen mode

  • Signals initiated in S4 propagate into S5–S10
  • Core processing occurs here
  • Fractality applies only within pairs:

    • S5–S6
    • S8–S9

Depth of branching depends on:

  • context
  • efficiency
  • cost constraints

6. S11 — Realization

S11 INTEGRITY:

Realization evaluates alignment with the original S1 (Will).

Possible outcomes:

  • Acceptance
  • Rejection
  • Transformation
  • Escalation into a new pass

S11 may incorporate data from the Integrity Log.


7. Full Vertical Structure

S1 → (parallel signals)
S2 — Wisdom        S3 — Knowledge
        ↓                  ↓

        S4 — Comprehension
        (integration + TensionPoint + Integrity Log)

                ↓

        S5–S10 — Operational Layers
        (projective + practical + fractal processing)

                ↓

        S11 — Realization

Enter fullscreen mode Exit fullscreen mode


8. Core Operational Principles

  • S1–S3 form an immutable core
  • S4 is the honest integrator
  • Failure at S4 is a signal, not an error

Adaptive Pass Depth:

After S4, perform reflection:

“Can this result be justified via S2 as strongly as via S3?”

If not (S3 dominates):

  • Initiate a second pass (S5–S11)

  • Generate S1_new strictly from TensionPoint

  • S1_new must:

    • be sharper
    • not semantically repeat the original
  • Goal:

    • strengthen S2
    • generate new meaning
  • Reflection result must be recorded in the Integrity Log

Additional rules:

  • Rollback allowed only to S1–S3
  • All structural breaks are preserved in the Integrity Log

9. Switch Flags (Activation of Full A11)

Switch Flags are external control signals that determine whether:

  • Lite Mode (S1–S4) is sufficient
  • or a Full Pass (S1–S11) is required
A11.SwitchFlags {
    RiskFlag,
    ConflictFlag,
    UncertaintyFlag,
    ValueFlag,
    UserDepthFlag
}

Enter fullscreen mode Exit fullscreen mode

Full A11 is activated if:

  • RiskFlag is active
  • OR ValueFlag is active
  • OR UserDepthFlag is active
  • OR (ConflictFlag + UncertaintyFlag) are both active

Otherwise:

  • Lite Mode (S1–S4) is used
  • No Integrity Log is recorded

10. Purpose of Flags

  • RiskFlag — prevents failure in critical decisions
  • ConflictFlag — detects S2–S3 inconsistency
  • UncertaintyFlag — signals insufficient data
  • ValueFlag — protects constraints and priorities
  • UserDepthFlag — explicit request for deep reasoning

This specification is included to ensure that the structural assumptions behind A11 are explicit, reproducible, and interpretable across different systems and contexts.


Algorithm 11 (A11) https://github.com/gormenz-svg/algorithm-11