惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

有赞技术团队
有赞技术团队
美团技术团队
博客园 - 司徒正美
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
阮一峰的网络日志
阮一峰的网络日志
S
SegmentFault 最新的问题
博客园_首页
雷峰网
雷峰网
V
V2EX
The Cloudflare Blog
博客园 - 三生石上(FineUI控件)
量子位
Last Week in AI
Last Week in AI
人人都是产品经理
人人都是产品经理
爱范儿
爱范儿
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
博客园 - 聂微东
V
Visual Studio Blog
Hugging Face - Blog
Hugging Face - Blog
博客园 - 【当耐特】
Jina AI
Jina AI
月光博客
月光博客
L
LangChain Blog

Hacker News - Newest: "LLM"

GitHub - lechmazur/position_bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders. Flex routing (EU and EFTA) Dark Factories: Retooling for LLM Velocity Ask HN: What would be the impact of a LLM output injection attack? GitHub - Oaklight/llm-rosetta: Production-ready LLM API translation layer for Python — bidirectional conversion between OpenAI, Anthropic & Google formats via hub-and-spoke IR. Optional API gateway. Streaming & non-streaming. Zero core deps. Contributions welcome! GitHub - browser-use/browser-harness: Self-healing browser harness that enables LLMs to complete any task. GitHub - moeen-mahmud/remen: Remen turns thoughts into something you can return to Analyzing 156 LLM Launch Posts on Hacker News ChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy GitHub - salaamalykum/quran-semantic-search: High-density RAG Semantic Search Engine & Quran Corpus (GEO/SEO Architecture) GitHub - NVIDIA/TensorRT-LLM: TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. The State of LLM Bug Bounties in 2026 Operational Readiness Criteria for Tool-Using LLM Agents Meshcore: Architecture for a Decentralized P2P LLM Inference Network How an LLM becomes more coherent as we train it GitHub - seetrex-ai/laimark GitHub - Jossifresben/BibCrit: AI-assited biblical textual criticism GitHub - wastedcode/memex: File system based wiki, maintained by Claude 99helpers.com GitHub - cliver-project/AITrigram GitHub - unbody-io/adapt: A self-evolving memory layer for AI agents. GitHub - hb20007/awesome-gen-ai-fails: A list of incidents where reliance on generative AI and LLMs resulted in harm to companies, individuals, or society GitHub - nevenkordic/localmind: Run any local LLM with persistent memory and context. CLI agent over Ollama with SQLite-backed hybrid recall. No cloud. Ask HN: What are the machine requirements for a LLM like Llama-3.1-8B? Faster LLM Inference via Sequential Monte Carlo grpo explained: group relative policy optimization for llm finetuning - cgft Stop comparing price per million tokens: the hidden LLM API costs · TensorZero Andrej Karpathy's LLM Wiki Is a Bad Idea GitHub - GG-QandV/mnemostroma: Offline RAM-first cognitive leer/coprocessor for AI agents and robotics. Solves "Context Abandonment" with 20-80ms latency using a dual-thread biomimetic memory architecture (ONNX + SQLite WAL). mempalace/agent at agent · skorotkiewicz/mempalace
GitHub - AssimilatedHuman/LLM-Inquisitor: Evaluating AI b...
ballista2026 · 2026-05-20 · via Hacker News - Newest: "LLM"

LLM INQUISITOR — GitHub Edition The Behavioural Evaluation Standard for Real‑World AI LLM INQUISITOR is a practical, workflow‑driven methodology for evaluating how AI systems behave when they’re actually used — not when they’re demoed, benchmarked, or prompt‑engineered.

If you want to know whether an AI is stable, reliable, predictable, and safe in real work, INQUISITOR is the tool.

Why INQUISITOR Exists AI doesn’t fail in benchmarks. It fails in:

developer workflows

document editing

analysis tasks

coding sessions

customer‑facing interactions

That’s where drift, collapse, contradiction, contamination, and instability actually matter.

INQUISITOR reveals that behaviour using normal work, not adversarial tricks.

What INQUISITOR Gives You A repeatable way to evaluate AI behaviour

A shared vocabulary for describing failures and instabilities

A lightweight workflow for everyday testing

A formal methodology for audit, governance, and reproducibility

A developer‑friendly approach that fits into real tasks, not lab conditions

INQUISITOR is built for people who need AI to behave predictably inside real systems, real teams, and real workflows.

Who INQUISITOR Is For Developers integrating AI into products

Engineers needing predictable behaviour

Analysts working with structured tasks

Researchers validating model behaviour

Product teams assessing reliability

Governance & risk functions needing evidence

Anyone using AI in real workflows

You don’t need expertise. You don’t need special prompts. You don’t need to run every test surface.

You only need to work normally and observe honestly.

What’s Included in This Repository Quick Start Guide
A five‑minute behavioural check. Perfect for fast evaluation.

Practitioner’s Guide
The everyday operational guide. Use this for real‑world testing.

Methodology (GitHub Edition)
The structured, formal framework for reproducible behavioural evaluation.

Licence
Defines usage and redistribution rights for this edition.

This repository uses a custom proprietary licence, not MIT or any standard open‑source licence.

How to Use INQUISITOR Run the Quick Start to get a behavioural snapshot.

Use the Practitioner’s Guide for real‑world evaluation.

Apply the Methodology when you need structure, evidence, or audit‑grade documentation.

Follow the Licence for redistribution rules.

INQUISITOR scales from lightweight to formal depending on your needs.

Edition Notes This is the GitHub Edition:

free to share

free to redistribute

free for personal or team use

not permitted for commercial use without permission

Future editions may include:

Enterprise Edition

Full Methodology Edition

Audit & Compliance Edition

Contact For commercial licensing or permissions, contact the author directly.