ๆƒฏๆ€ง่šๅˆ ้ซ˜ๆ•ˆ่ฟฝ่ธชๅ’Œ้˜…่ฏปไฝ ๆ„Ÿๅ…ด่ถฃ็š„ๅšๅฎขใ€ๆ–ฐ้—ปใ€็ง‘ๆŠ€่ต„่ฎฏ
้˜…่ฏปๅŽŸๆ–‡ ๅœจๆƒฏๆ€ง่šๅˆไธญๆ‰“ๅผ€

ๆŽจ่่ฎข้˜…ๆบ

The GitHub Blog
The GitHub Blog
A
About on SuperTechFans
The Cloudflare Blog
G
Google Developers Blog
ๅš
ๅšๅฎขๅ›ญ_้ฆ–้กต
Martin Fowler
Martin Fowler
Apple Machine Learning Research
Apple Machine Learning Research
L
LangChain Blog
D
Docker
C
Check Point Blog
T
Tailwind CSS Blog
ๅš
ๅšๅฎขๅ›ญ - ๅธๅพ’ๆญฃ็พŽ
ๅฅ‡ๅฎขSolidotโ€“ไผ ้€’ๆœ€ๆ–ฐ็ง‘ๆŠ€ๆƒ…ๆŠฅ
ๅฅ‡ๅฎขSolidotโ€“ไผ ้€’ๆœ€ๆ–ฐ็ง‘ๆŠ€ๆƒ…ๆŠฅ
Hugging Face - Blog
Hugging Face - Blog
Microsoft Security Blog
Microsoft Security Blog
V
V2EX
ๅš
ๅšๅฎขๅ›ญ - ๅถๅฐ้’—
T
The Blog of Author Tim Ferriss
้…ท ๅฃณ โ€“ CoolShell
้…ท ๅฃณ โ€“ CoolShell
ITไน‹ๅฎถ
ITไน‹ๅฎถ
M
MIT News - Artificial intelligence
Microsoft Azure Blog
Microsoft Azure Blog
ๅš
ๅšๅฎขๅ›ญ - ใ€ๅฝ“่€็‰นใ€‘
GbyAI
GbyAI

Hacker News - Newest: "LLM"

GitHub - lechmazur/position_bias: A benchmark for testing whether LLM judges keep the same preference when two lightly edited versions of the same story are shown in opposite orders. Flex routing (EU and EFTA) Dark Factories: Retooling for LLM Velocity Ask HN: What would be the impact of a LLM output injection attack? GitHub - AronDaron/dataset-generator: No-code desktop app for generating high-quality synthetic datasets to fine-tune LLMs โ€” plan-then-execute pipeline, LLM-as-judge, HuggingFace upload. GitHub - Oaklight/llm-rosetta: Production-ready LLM API translation layer for Python โ€” bidirectional conversion between OpenAI, Anthropic & Google formats via hub-and-spoke IR. Optional API gateway. Streaming & non-streaming. Zero core deps. Contributions welcome! GitHub - browser-use/browser-harness: Self-healing browser harness that enables LLMs to complete any task. GitHub - moeen-mahmud/remen: Remen turns thoughts into something you can return to Analyzing 156 LLM Launch Posts on Hacker News ChatGPT vs Gemini vs Claude: The Best LLM Subscription You Should Buy GitHub - salaamalykum/quran-semantic-search: High-density RAG Semantic Search Engine & Quran Corpus (GEO/SEO Architecture) GitHub - NVIDIA/TensorRT-LLM: TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way. The State of LLM Bug Bounties in 2026 Operational Readiness Criteria for Tool-Using LLM Agents Meshcore: Architecture for a Decentralized P2P LLM Inference Network How an LLM becomes more coherent as we train it GitHub - seetrex-ai/laimark GitHub - Jossifresben/BibCrit: AI-assited biblical textual criticism GitHub - wastedcode/memex: File system based wiki, maintained by Claude 99helpers.com GitHub - cliver-project/AITrigram GitHub - unbody-io/adapt: A self-evolving memory layer for AI agents. GitHub - hb20007/awesome-gen-ai-fails: A list of incidents where reliance on generative AI and LLMs resulted in harm to companies, individuals, or society GitHub - nevenkordic/localmind: Run any local LLM with persistent memory and context. CLI agent over Ollama with SQLite-backed hybrid recall. No cloud. Ask HN: What are the machine requirements for a LLM like Llama-3.1-8B? Faster LLM Inference via Sequential Monte Carlo grpo explained: group relative policy optimization for llm finetuning - cgft Stop comparing price per million tokens: the hidden LLM API costs ยท TensorZero Andrej Karpathy's LLM Wiki Is a Bad Idea GitHub - GG-QandV/mnemostroma: Offline RAM-first cognitive leer/coprocessor for AI agents and robotics. Solves "Context Abandonment" with 20-80ms latency using a dual-thread biomimetic memory architecture (ONNX + SQLite WAL).
GitHub - stef41/lmscan: ๐Ÿ” Detect AI-generated text and fi...
2026-04-11 ยท via Hacker News - Newest: "LLM"

Detect AI-generated text. Fingerprint which LLM wrote it. Open-source GPTZero alternative.

PyPI Downloads License Python CI Tests OpenSSF Scorecard

GPTZero charges $15/month. Originality.ai charges per scan. Turnitin locks you into institutional contracts.

lmscan is free, open-source, works offline, and tells you which model wrote the text.

demo

$ lmscan "In today's rapidly evolving digital landscape, it's important
to note that artificial intelligence has become a pivotal force in
transforming how we navigate the complexities of modern life..."

๐Ÿ” lmscan v0.1.0 โ€” AI Text Forensics
โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•

  Verdict:     ๐Ÿค– Likely AI (77% confidence)
  Words:       184
  Sentences:   10
  Scanned in 0.01s

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚ Feature                    โ”‚ Value    โ”‚ Signal             โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚ Burstiness                 โ”‚ 0.07     โ”‚ ๐Ÿ”ด Very low (AI)    โ”‚
โ”‚ Sentence length variance   โ”‚ 0.27     โ”‚ ๐ŸŸก Below average    โ”‚
โ”‚ Slop word density          โ”‚ 20.7%    โ”‚ ๐Ÿ”ด High (AI)        โ”‚
โ”‚ Transition word ratio      โ”‚ 2.2%     โ”‚ ๐ŸŸก Elevated         โ”‚
โ”‚ Readability consistency    โ”‚ 0.00     โ”‚ ๐Ÿ”ด Very low (AI)    โ”‚
โ”‚ ...                        โ”‚          โ”‚                     โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ”Ž Model Attribution
  1. GPT-4 / ChatGPT    62% โ€” "delve", "tapestry", "beacon", "landscape" (ร—2), +19 more
  2. Claude (Anthropic)  13% โ€” "robust", "nuanced", "comprehensive"
  3. Gemini (Google)      9% โ€” "furthermore", "additionally"

โš ๏ธ  Flags
  โ€ข Very low burstiness (0.07) โ€” AI text is more uniform in complexity
  โ€ข High slop word density (20.7%) โ€” contains known AI vocabulary markers

Install

pip install lmscan

Zero dependencies. Works with Python 3.9+. No API keys. No internet. No GPU.

Usage

# Scan text directly
lmscan "Your text here..."

# Scan a file
lmscan document.txt

# Pipe from stdin
cat essay.txt | lmscan -

# JSON output (for scripts and CI)
lmscan document.txt --format json

# Per-sentence breakdown
lmscan document.txt --sentences

# CI gate: fail if AI probability > 50%
lmscan submission.txt --threshold 0.5

Python API

from lmscan import scan

result = scan("Text to analyze...")

print(f"AI probability: {result.ai_probability:.0%}")
print(f"Verdict: {result.verdict}")
print(f"Confidence: {result.confidence}")

# Which model wrote it?
for model in result.model_attribution:
    print(f"  {model.model}: {model.confidence:.0%}")
    for evidence in model.evidence[:3]:
        print(f"    โ†’ {evidence}")

# Per-sentence analysis
for sentence in result.sentence_scores:
    if sentence.ai_probability > 0.7:
        print(f"  ๐Ÿค– {sentence.text[:60]}... ({sentence.ai_probability:.0%})")

Scan entire directories

from lmscan import scan_file
import glob

for path in glob.glob("submissions/*.txt"):
    result = scan_file(path)
    print(f"{path}: {result.verdict} ({result.ai_probability:.0%})")

How It Works

lmscan uses 12 statistical features derived from computational linguistics research to distinguish AI-generated text from human writing:

Feature What it measures AI signal
Burstiness Variance in sentence complexity AI text is unusually uniform
Sentence length variance How much sentence lengths vary AI produces uniform lengths
Vocabulary richness Type-token ratio (Yule's K corrected) AI reuses words more
Hapax legomena ratio Fraction of words appearing once AI has fewer unique words
Zipf deviation How word frequencies follow Zipf's law AI deviates from natural distribution
Readability consistency Flesch-Kincaid variance across paragraphs AI maintains constant readability
Bigram/trigram repetition Repeated word pairs and triples AI repeats phrase structures
Transition word ratio "however", "moreover", "furthermore"... AI overuses transitions
Slop word density Known AI vocabulary markers "delve", "tapestry", "beacon"...
Punctuation entropy Diversity of punctuation usage AI is more predictable

Each feature produces a signal via sigmoid transformation. The weighted combination produces the final AI probability.

Model Fingerprinting

lmscan includes vocabulary fingerprints for 5 major LLM families:

Model Distinctive markers
GPT-4 / ChatGPT "delve", "tapestry", "landscape", "leverage", "multifaceted", "it's important to note"
Claude (Anthropic) "certainly", "I'd be happy to", "straightforward", "I should note"
Gemini (Google) "crucial", "here's a breakdown", "keep in mind"
Llama / Meta "awesome", "fantastic", "hope this helps"
Mistral / Mixtral "indeed", "moreover", "hence", "noteworthy"

Attribution uses weighted vocabulary matching, phrase detection, and hedging pattern analysis.

Accuracy & Limitations

What lmscan is good at:

  • Detecting text with strong AI stylistic patterns
  • Identifying which model family generated text
  • Scanning at scale (thousands of documents) with zero cost
  • Providing explainable evidence (not a black box)

What lmscan cannot do:

  • Detect AI text that has been manually edited or paraphrased
  • Work reliably on very short text (<50 words)
  • Detect AI text in non-English languages (English-only for now)
  • Replace human judgment โ€” use as a signal, not a verdict

This is statistical analysis, not a neural classifier. It detects stylistic patterns, not watermarks. It works best on unedited LLM output and degrades gracefully on edited text.

CI Integration

GitHub Actions

- name: AI Content Check
  run: |
    pip install lmscan
    lmscan submission.txt --threshold 0.7 --format json

Pre-commit

repos:
  - repo: https://github.com/stef41/lmscan
    rev: v0.1.0
    hooks:
      - id: lmscan
        args: ["--threshold", "0.7"]

Research Background

lmscan's approach is informed by published research on AI text detection:

  • DetectGPT (Mitchell et al., 2023) โ€” perturbation-based detection using log probability curvature
  • GLTR (Gehrmann et al., 2019) โ€” statistical visualization of token predictions
  • Binoculars (Hans et al., 2024) โ€” cross-model perplexity comparison
  • Zipf's Law in NLP โ€” word frequency distributions differ between human and AI text
  • Stylometry โ€” decades of authorship attribution research applied to AI forensics

lmscan takes the statistical intuitions from these papers and implements them as lightweight, dependency-free heuristics that work without requiring a reference language model.

FAQ

Q: Is this as accurate as GPTZero? A: GPTZero uses neural classifiers trained on labeled data. lmscan uses statistical heuristics. GPTZero is more accurate on edge cases; lmscan is free, offline, and explainable. Use both if accuracy matters.

Q: Can students use this to evade AI detection? A: lmscan shows which features trigger detection, which could help someone understand why text reads as AI-generated. This is by design โ€” understanding AI writing patterns makes everyone a better writer. The same information is available in published research papers.

Q: Does it work on non-English text? A: Currently English-only. The slop word lists and transition word lists are English-specific. Statistical features (entropy, burstiness) work across languages but haven't been calibrated.

Q: Does it phone home? A: No. Zero network requests. No telemetry. No API keys. Everything runs locally.

Q: How is model attribution possible without running the model? A: Each LLM family has characteristic vocabulary biases. GPT-4 loves "delve" and "tapestry". Claude says "I'd be happy to". These are statistical fingerprints โ€” not guaranteed attribution, but strong signals.

See Also

License

Apache-2.0