惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

J
Java Code Geeks
G
Google Developers Blog
Blog — PlanetScale
Blog — PlanetScale
U
Unit 42
A
About on SuperTechFans
Vercel News
Vercel News
B
Blog
Martin Fowler
Martin Fowler
MyScale Blog
MyScale Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
腾讯CDC
D
Docker
V
Visual Studio Blog
博客园 - 叶小钗
The Cloudflare Blog
Jina AI
Jina AI
B
Blog RSS Feed
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
WordPress大学
WordPress大学
T
Tailwind CSS Blog
MongoDB | Blog
MongoDB | Blog
D
DataBreaches.Net
月光博客
月光博客
大猫的无限游戏
大猫的无限游戏

DEV Community

Authentication Security Deep Dive: From Brute Force to Salted Hashing (With Java Examples) Why AI Systems Don’t Fail — They Drift Spilling beans for how i learn for exam😁"Reinforcement Learning Cheat Sheet" I Replaced Chrome with Safari for AI Browser Automation. Here's What Broke (and What Finally Worked) How Python Borrows Other People's Work The $40 Architecture: Processing 1 Billion API Requests with 99.99% Uptime Vibe Coding: A Workflow Guide (From Zero to SaaS) Most webhook security guides protect the wrong side. The scary part is delivery. Headless CMS for TanStack Start: Build a Blog with Cosmic EU Age Verification App "Hacked in 2 Minutes" — What Actually Happened Comfy Cloud’s delete function does not actually remove files Running AI Models on GPU Cloud Servers: A Beginner Guide Event-driven media intelligence with AWS Step Functions and Bedrock I scored 500 AI prompts across 8 quality dimensions — here's what broke How to Call Google Gemini API from Next.js (Free Tier, No Backend Needed) The Portal Protocol: Reclaiming Human Connection in the Age of AI How to Fix Your Team's Scattered Knowledge Problem With a Self-Hosted Forum Intro to tc Cloud Functors: A Graph-First Mental Model for the Modern Cloud Designing Multi-Tenant Backends With Both Ownership and Team Access I Built a Neumorphic CSS Library with 77+ Components — Here's What I Learned PostgreSQL Performance Optimization: Why Connection Pooling Is Critical at Scale Cómo construí un SaaS multi-rubro para gestionar expensas en Argentina con FastAPI + Vue 3 🚀 I Built an Ethical Hacking Scanner Tool – Open Source Project I Replaced /usage and /context in Claude Code With a Single Statusline A Pythonic Way to Handle Emails (IMAP/SMTP) with Auto-Discovery and AI-Ready Design I Collected 8.9 Million Polymarket Price Points — Here's What I Found About How Markets Really Move EcoTrack AI — Carbon Footprint Tracker & Dashboard Everyone's Using AI. No One Agrees How. 5 self-hosted ebook managers worth trying in 2026 Building Your First AI Agent with LangChain: From Chatbot to Autonomous Assistant
From Variant CSV to Review-Ready Report: A Python Workflo...
Oluwagbade Odimayo · 2026-06-15 · via DEV Community
Cover image for From Variant CSV to Review-Ready Report: A Python Workflow With Docker and GitHub Actions

Oluwagbade Odimayo

Variant prioritisation often starts with a table.

But a table alone does not answer the most important question:

Which variants deserve closer review, and why?

The ClinVar Variant Prioritisation Workflow was built to answer that question with transparent scoring, validation, reporting, Docker, and CI.

Repository:

GitHub

Tech Stack

Python
pandas
Pydantic
matplotlib
pytest
Make
Docker
GitHub Actions
mamba

What the Workflow Does

The workflow takes a curated inherited-disease variant dataset and ranks variants using transparent evidence rules.

Each variant receives:

priority score out of 100
priority tier
ranked output
review recommendation

Dataset Fields

The curated dataset includes:

variant_id
gene
chromosome
position
reference
alternate
consequence
clinvar_significance
review_status
allele_frequency
inheritance
phenotype_match_score
computational_score
disease_area

Validation Layer

Before scoring, the workflow checks:

required columns
valid allele frequency values
valid phenotype match score range
valid computational score range
record schema consistency

Pydantic is used for schema validation.

This prevents the scoring logic from running on malformed records.

Scoring Framework

The score is out of 100:

ClinVar-style significance: 30
Review status: 15
Variant consequence: 20
Allele frequency rarity: 15
Phenotype match: 20

Priority tiers:

>= 80   high_priority
60-79   moderate_priority
40-59   low_priority
< 40    minimal_priority

This is not a clinical diagnostic score. It is a transparent prioritisation score for review.

Example Result

Top ranked variants from the current dataset:

Rank Variant Gene Consequence Score
1 VAR010 DMD stop_gained 99
2 VAR001 BRCA1 stop_gained 98
3 VAR014 FBN1 splice_donor_variant 96
4 VAR019 MLH1 splice_acceptor_variant 95
5 VAR008 SCN1A frameshift_variant 94

Outputs

The pipeline generates:

results/tables/ranked_variants.csv
results/tables/top_prioritised_variants.csv
results/reports/top_variant_review_report.md
results/figures/priority_score_distribution.png
results/figures/priority_tier_counts.png
results/figures/top_gene_priority_scores.png

Makefile Commands

make test
make score
make report
make figures
make pipeline

The full pipeline loads data, validates records, scores variants, generates review outputs, and creates figures.

Docker Workflow

docker build -t clinvar-variant-prioritisation:latest .
docker run --rm clinvar-variant-prioritisation:latest make test
docker run --rm clinvar-variant-prioritisation:latest make pipeline

Docker exposed two real issues.

First, make was missing inside the image.

Second, the non-root container user could not overwrite files under /app/results.

Both were fixed in the Dockerfile.

CI Workflow

GitHub Actions runs:

pytest test suite
full pipeline
expected output file checks

The workflow was also updated to opt into the Node.js 24 runtime.

Documentation

The repository includes:

README.md
docs/methods.md
docs/limitations.md
docs/evidence_map.md
docs/reviewer_guide.md
docs/evidence/

Main Takeaway

The project demonstrates how a small variant dataset can become a reproducible scientific workflow.

It includes:

validation
transparent scoring
ranked outputs
review reporting
visual analytics
Docker reproducibility
CI
evidence tracking

The result is not a clinical diagnostic system. It is a professional bioinformatics workflow showing how variant prioritisation logic can be made transparent, reproducible, and review-ready.