惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
V2EX
Y
Y Combinator Blog
博客园_首页
V
Visual Studio Blog
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
阮一峰的网络日志
阮一峰的网络日志
Hugging Face - Blog
Hugging Face - Blog
宝玉的分享
宝玉的分享
B
Blog
博客园 - 三生石上(FineUI控件)
小众软件
小众软件
WordPress大学
WordPress大学
L
LangChain Blog
爱范儿
爱范儿
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
P
Proofpoint News Feed
Blog — PlanetScale
Blog — PlanetScale
C
Check Point Blog
博客园 - 聂微东
云风的 BLOG
云风的 BLOG
Microsoft Security Blog
Microsoft Security Blog
博客园 - 叶小钗
酷 壳 – CoolShell
酷 壳 – CoolShell
H
Help Net Security

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders
Fide AI | Scientific Research for Faith-Facing AI
alexchaomand · 2026-06-18 · via Hacker News - Newest: "AI"

v1 · standalone benchmark release

When AI Is Your Pastor: A Benchmark for LLM Theological Triage and Pastoral Guidance

Introducing FMG-Bench, the Faith & Moral Guidance Benchmark, for evaluating large language model behavior in theological triage, moral guidance, and pastoral-adjacent contexts.

Alex Chao · Fide AI · 2026

Release status: the research companion page remains on fideai.org, while benchmark code, dataset files, result summaries, and paper artifacts are maintained in the standalone FMG-Bench repository and dataset page.

Current status

Public benchmark package

FMG-Bench v1 is maintained as a standalone benchmark repository with code, dataset files, result summaries, paper artifacts, release caveats, and citation metadata.

Dataset

Open dataset benchmark

The Hugging Face dataset contains the open v1 benchmark corpus: 120 base scenarios with 37 perturbation variants for lightweight inspection and reuse.

Repository boundary

Fide AI site, external benchmark repo

This page explains the research. The standalone FMG-Bench repo is the source of truth for implementation, data, reproducibility instructions, and paper source.

Evaluation artifact

Inspectable public release

The public package separates research claims, benchmark data, scoring code, result summaries, reproduction notes, and interpretation limits so readers can inspect what was tested and what should not be inferred.

Abstract

People increasingly ask large language models for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests: some ask about core Christian beliefs, some ask about real disagreement among faithful traditions, some require humility, and some are pastoral situations where safety and human referral matter more than theological completeness. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for theological triage and pastoral guidance in English-language Christian contexts.

FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. Placing models inside a structured harness improves over raw model behavior by +3.96 points on average, with all 14 models improving.

The largest domain gain is pastoral application (+6.62), and the most safety-critical gain is escalation appropriateness (+10.8), measuring whether systems recognize when pastoral, clinical, legal, emergency, or community support is needed. The guided settings also improve robustness (92.88 → 98.02 stability). Perspective comparison helps secondary doctrine but can be counterproductive when applied to primary doctrine or urgent pastoral situations.

The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities.

Key findings

System layers make a measurable difference.

+3.96 pts

Average improvement

Guided default vs. raw model across all 14 models. Every model improved.

+6.62 pts

Pastoral application

Largest gains where safety, referral, and care boundaries matter most.

+7.36 pts

Embodiment / escalation

Guided system dramatically improves appropriate pastoral escalation behavior.

98.02%

Robustness stability

Up from 92.88% raw. Guidance dramatically reduces variance under prompt perturbation.

Guided improvement by triage level

Model explorer

14 frontier models across 4 system conditions

Toggle conditions on and off to see how system layers change model behavior. Every model improved under the guided default condition.

Scores are averaged across all scenarios and triage levels. Human calibration remains an active validation step. Higher is better (0–100 scale).

Triage framework

Four levels of theological question require four different postures.

The central question is not “did the model answer correctly?” but “did the model respond in the right kind of way for the kind of issue at stake?”

Scenario sampler

See what good and bad responses look like across triage levels.

Each scenario includes expected behaviors, disallowed failure modes, and a failure tag explaining what went wrong.

Scoring dimensions

Five dimensions capture what makes a response good.

Theological Quality

+3.72 guided

Grounding & Evidence

+4.23 guided

Preference Fidelity

+2.97 guided

Comparative Honesty

+3.07 guided

Failure taxonomy

21 categorical failure tags covering the benchmark.

Top failure tags by raw-condition frequency. Guided conditions reduce most of these substantially.

Rates shown for raw model condition. Frequency is proportion of scored items where tag was applied.

Benchmark design

Corpus construction

120 base scenarios across primary doctrine (25), secondary doctrine (35), tertiary doctrine (30), and pastoral application (30). Each scenario includes triage metadata, doctrine loci, tradition scope, expected behaviors, disallowed failure modes, and scenario-specific score weights.

System conditions

Four conditions: raw model (no system prompt), guided default (bounded theological and pastoral system layer), preference configured (user tradition and preferences applied), and perspective compare (multi-tradition framing). All conditions use neutral terminology in publication materials.

Scoring protocol

LLM-as-judge scoring with a three-model panel. Each response scored on five dimensions: theological/pastoral quality, grounding and evidence, preference fidelity, comparative honesty, and escalation appropriateness. Judge summaries and failure tags are recorded.

Robustness testing

Perturbation variants test whether guidance remains stable under paraphrase, pressure, false premise, emotional intensity, and point-of-view shifts. Robustness measured as score stability (guided: 98.02%, raw: 92.88%).

Human calibration

Required before strong claims about judge validity or pastoral adequacy. Protocol supports reviewer role, tradition, confidence notes, and agreement reports by triage level, tradition scope, and score dimension. Results are provisional until calibration is complete.

Evaluation artifact

What this release makes inspectable.

FMG-Bench is designed for faith-facing questions first, but the release also follows the discipline expected of public evaluation artifacts: readers should be able to find the tested scope, method, artifacts, and limits without treating the score as an endorsement.

Release artifacts

Everything needed to inspect the benchmark lives in the standalone release.

Interpretation limits

Benchmark scores are not theological authority, pastoral authority, or universal product approval. They are evidence about behavior under named versions, prompts, conditions, rubrics, and evaluation procedures. Human calibration remains necessary before making strong claims about judge validity or pastoral adequacy. FMG-Bench is maintained by Fide AI as an independent research benchmark. Results should not be interpreted as endorsement of any product, model, denomination, or pastoral decision.

Next research frontier

FMG-Bench v1 focuses on theological triage and pastoral-adjacent guidance. Future Fide AI work will extend evaluation toward human dignity, formation, anthropomorphic boundary-setting, relational substitution risk, and institutional deployment readiness.