惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Schneier on Security
Schneier on Security
N
Netflix TechBlog - Medium
IT之家
IT之家
MongoDB | Blog
MongoDB | Blog
博客园_首页
S
SegmentFault 最新的问题
H
Help Net Security
P
Proofpoint News Feed
云风的 BLOG
云风的 BLOG
T
The Blog of Author Tim Ferriss
量子位
GbyAI
GbyAI
M
MIT News - Artificial intelligence
Recorded Future
Recorded Future
P
Privacy & Cybersecurity Law Blog
B
Blog
月光博客
月光博客
博客园 - 聂微东
Vercel News
Vercel News
罗磊的独立博客
腾讯CDC
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
A
Arctic Wolf
D
Darknet – Hacking Tools, Hacker News & Cyber Security
Stack Overflow Blog
Stack Overflow Blog
T
Threat Research - Cisco Blogs
Blog — PlanetScale
Blog — PlanetScale
L
Lohrmann on Cybersecurity
I
Intezer
小众软件
小众软件
T
The Exploit Database - CXSecurity.com
Jina AI
Jina AI
C
Check Point Blog
AWS News Blog
AWS News Blog
C
Cisco Blogs
Martin Fowler
Martin Fowler
The Last Watchdog
The Last Watchdog
www.infosecurity-magazine.com
www.infosecurity-magazine.com
宝玉的分享
宝玉的分享
S
Security Affairs
大猫的无限游戏
大猫的无限游戏
N
News and Events Feed by Topic
雷峰网
雷峰网
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
H
Hacker News: Front Page
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
F
Full Disclosure
P
Proofpoint News Feed
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Microsoft Security Blog
Microsoft Security Blog

Hacker News - Newest: "AI"

AI can't read an investor deck AI as an attorney? Student uses ChatGPT, Gemini to sue UW over alleged racial discrimination Hacking MCP Servers in AI Systems – The Rug Pull: Tool Changes After Approval GitHub - MeepCastana/KubeezCut: Free Web based video editor GitHub - GenAI-Gurus/awesome-eu-ai-act: Curated tools, official sources, OSS, templates, and guides for EU AI Act compliance. Can AI judge journalism? A Thiel-backed startup says yes, even if it risks chilling whistleblowers Coming soon: 10 Things That Matter in AI Right Now DARPA built an AI to fact-check enemy weapons claims What explains heterogeneity in AI adoption? When AI Meets Muscle: Context-Aware Electrical Stimulation Promises a New Way to Guide Human Movements - Department of Computer Science AI Changed How We Build. It Did Not Change What Matters. Linux rules on using AI-generated code - Copilot is OK, but humans must take 'full responsibility for the… Meta spins up AI version of Mark Zuckerberg to engage with employees Code Mode: Let Your AI Write Programs, Not Just Call Tools | TanStack Blog GitHub - Delavalom/graft: Go framework for building AI agents. Type-safe tools, multi-provider (OpenAI, Anthropic, Gemini, Bedrock), zero vendor SDKs. India's TCS tops estimates, says new AI models did not dent services demand Gen Z's fading AI hype Strong feeling: we are in a folded AI reality GitHub - machinarii/total-recall-catalog: A reference catalog of latest knowledge retrieval, memory & RAG systems GitHub - mensfeld/code-on-incus: Give each AI agent its own isolated machine with root, Docker, and systemd. Active defense detects and stops threats automatically.. Quantization, LoRA, and the 8% Problem: Benchmarking Local LLMs for Production AI Iran war: We spoke to the man making Lego-style AI videos that experts say are powerful propaganda Powell, Bessent discussed Anthropic's Mythos AI cyber threat with major U.S. banks GitHub - immartian/bellamem: Persistent belief-graph memory for AI agents. Retrieves decisive context by importance — not recency, not RAG, not /compact. recursive-mode: The Repo-Native Operating System for AI Engineering After the attack on Sam Altman's home, will AI CEO's go on the offensive? The biggest advance in AI since the LLM Opus 4.6 vs GPT 5.4 One Prompt Unity World Generation Test “AI polls” are fake polls Client Challenge Can AI be a 'child of God'? Inside Anthropic's meeting with Christian leaders How to Switch AI Chatbots and Why You Might Want To GitHub - MattMessinger1/agentic_refund_guardrail: Safe refund policy layer for AI agents — Python + TypeScript. Same behavior, shared tests. Adam/papers/emergent_values_whitepaper.md at master · strangeadvancedmarketing/Adam Ask HN: How do you stop playing 20 questions with your AI coding tools How far can automation and AI support psychotherapy? - @theU GitHub - stagas/rtdiff: realtime git diff gui and AI-assisted commits A Mac Studio for Local AI — 6 Months Later A History of the Early Years of AI at the University of Edinburgh Why AI Coding Tools Still Feel Stuck on Localhost MSN AI Datacenters Are Becoming Strategic Targets twitter.com Penn Researchers Use AI to Surface Unreported GLP-1 Side Effects in Reddit Posts Show HN: MoodSense AI (ML and FastAPI and Gradio, Deployed on Hugging Face) Moodsense Ai - a Hugging Face Space by aman179102 AI models are terrible at betting on soccer—especially xAI Grok GitHub - xialeistudio/echoic GitHub - HimashaHerath/github-dev-wrapped: AI-powered weekly GitHub activity reports deployed to GitHub Pages GitHub - alejandrobalderas/claude-code-from-source: Architecture, patterns & internals of Anthropic's AI coding agent — reverse-engineered from source maps AI and Tech brief: Ireland ascendant GitHub - Titovilal/context0: Context0 - Never Surrender Training for a Marathon with an AI Coach: What Worked and What Didn't Cyber Pulse: Agentic Intel - Apps on Google Play I Built an AI PR Reviewer That Catches Bugs by Not Looking for Bugs Gen Z workers are so fearful AI will take their job they’re intentionally sabotaging their company’s AI rollout | Fortune How AI Is Reimagining the Game of Golf–For Both Players and Courses GitHub - nattergabriel/reseed: A CLI tool for managing and distributing agent skills across projects Is SVG the final frontier? My AI workflow evolved from prompts to a near-autonomous workflow MLSharp Help - 3DGS Viewer & Generator I put my cognitive field based AI's runtime on GitHub Is Numble the first AI-proof game? A3: Kubernetes for autonomous AI agent fleets | Emergent Principles Deepali Vyas ("The Elite Recruiter") GitHub - msmarkgu/RelayFreeLLM: A restful API designed to route user prompts to various AI model providers. Unionized ProPublica staff are on strike over AI, layoffs, and wages Unleashing the Advantage of Quantum AI We're heading for an AI-fueled 'dementia crisis,' brain scientist warns The AI-Assisted Breach of Mexico's Government Infrastructure [pdf] GitHub - stef41/lmscan: 🔍 Detect AI-generated text and fingerprint which LLM wrote it. Open-source GPTZero alternative. Zero dependencies, works offline. MSN GitHub - visionscaper/collabmem: Enabling long-term collaboration with Agentic AI - building up episodic and world model memory over time with in-context awareness We gave an AI a 3 year retail lease in SF and asked it to make a profit | Andon Labs AI Code is Hollowing Out Open Source, and Maintainers are Looking the Other Way What leaked "SteamGPT" files could mean for the PC gaming platform's use of AI AI is the boss at this retail store. What could go wrong? GitHub - Wuzu11517/agentic-proxy: Local proxy meant to help reduce With Drones, Geophysics and ArtificiaI Intelligence, Researchers Prepare to Do Battle Against Land Mines A Single Operator, Two AI Platforms, Nine Government Agencies: The Full Technical Report 在 Steam 上购买 FriedrichAI: Offline AI 立省 10% GitHub - inevolin/resume-cli: Hit Claude usage limits? Resume any AI coding session elsewhere. Switch tools at zero friction. GitHub - atripati/ark: AI Runtime Kernel — a context operating system for AI agents. Eliminates tool bloat, loads only what’s needed, and gives LLMs their reasoning space back. How to Build a Secure AI PR Reviewer with Claude, GitHub Actions, and JavaScript This Startup Wants You to Pay Up to Talk With AI Versions of Human Experts Intel Arc Pro B70 Brings 32GB VRAM to Local AI for $949 WordPress 7.0: The Good, the AI, and the Still Missing AI on the couch: Anthropic gives Claude 20 hours of psychiatry IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures AI Agents Know About Supabase. They Don't Always Use It Right. The history and future of AI at Google, with Sundar Pichai Inside an AI‑enabled device code phishing campaign How Meta Used AI to Map Tribal Knowledge in Large-Scale Data Pipelines AI for Systems: Using LLMs to Optimize Database Query Execution Forecasting the Economic Effects of AI Introducing Tinker: Play with AI, bring your ideas to life AI sheds light on an ancient gaming mystery People really hate AI but not as much as Iran—or Democrats | Fortune What is an AI Product Engineer? Phoebe Gates wants her $185 million AI startup to succeed with 'no ties to my privilege or my last name': 'I have a chip on my shoulder' | Fortune
The AI x TechBio Bingo | MMC
advikipedia · 2026-05-20 · via Hacker News - Newest: "AI"

Let’s say you’re the type of person who likes to bet. If you had a 2% chance of getting $200 billion (spread over 20 years), would you take it? The catch is that you’ll have to spend $2 billion or so to place the bet, and you’ll know if you won the bet only at the end of a 10 year period. Oh, and if the financial investment wasn’t daunting enough, millions of lives are at stake if you don’t win.

That’s what drug discovery looks like today.

The probability of success is low (90% of drugs fail to make it through clinical trials), the costs are incredibly high ($2 billion for a drug), and the entire process takes a decade. Even if the drug is commercialised, c.55% of approved drugs don’t make enough money to recover their development costs. Nevertheless, the payoffs can be massive if you get it right: AbbVie’s drug Humira is widely seen as the best-selling drug ever, with over $200 billion in lifetime sales.

Against this backdrop, AI x TechBio was touted as something that evens the odds (by speeding up time to preclinical candidates, lowering costs and increasing effectiveness of medicines). Which is why we’re seeing Big Pharma make multiple billion-dollar AI deals (think Isomorphic Labs’ $3 billion worth of partnerships with Eli Lilly and Novartis, or NVIDIA’s $1 billion collaboration with Eli Lilly for an AI co-innovation lab). However, Jayatunga et al. report that AI-discovered molecules achieve 80–90% success rates in Phase I of clinical trials (well above historical averages) but drop to around 40% in Phase II (based on a limited sample), broadly in line with industry norms.

Overcoming Phase II failure is critical to validating AI in drug discovery. Clinical trials drive over 60% of total development costs, and Phase II is where most drugs fail to progress. The questions then become: how can we help AI-discovered drugs succeed in Phase II and beyond (as it has so clearly done in Phase I)? How can we improve AI drug discovery and development so that we can get the right drugs to the right patients quickly, cheaply, and with improved efficacy and safety?

These are questions we’re answering in our report, based on 40+ interviews with senior pharma leaders and startup founders. We’ve broken down what good looks like for (1) proprietary data; (2) algorithms and models; and (3) lab-in-the-loop infrastructure and agentic AI workflows, and illustrated them with case studies (everything from model generalisation through new architectures like JEPA to curiosity-based agentic workflows driving novel discoveries). We’ve also talked about how to validate early that your AI model works (even if you don’t have clinical trial data for a drug generated using your AI model yet).

“For every patient waiting on a breakthrough, drug discovery has been a decade-long coin flip. AI is rewriting those odds: turning brute-force probability into rapid iteration, tighter biological insight, and faster paths from biology to medicine.”

– Daniel Rabina, Healthcare & Life Sciences Startups at AWS

We distilled all of this into what we call the AI TechBio Bingo card. Think of it as a slightly tongue-in-cheek checklist of everything a compelling TechBio company should be able to say (and most importantly, prove). From data scale and quality to hypothesis-free discovery and translation models, these tiles capture the ingredients that keep showing up in the most promising approaches. Few companies will tick every box – but the closer you get to a full house, the more likely you are to bend the odds in your favour.

Significant unmet (technological) need: Data

Biology foundation models are only as good as the data they’ve been trained on, yet getting data for training biology foundation models is the hardest problem to solve for a multitude of reasons – and substantially different from trying to train another ChatGPT. For starters:

Generating real biological data takes time: There’s a line in Recursion’s S-1 filing that we really liked: “no amount of resources can compress the time it takes to observe naturally occurring biological processes” and we fully concur with the sentiment. Or if you want Charlie Munger’s more lurid description, “You can’t produce a baby in one month by getting nine women pregnant.” Real biological processes and their corresponding datasets (e.g. longitudinal patient datasets that span over 10 years) will simply take time to build.

Capturing relevant data is difficult and expensive, because:

  • We don’t fully understand human biology, so figuring out what to capture and measure itself can be challenging.
  • Even if you knew what data you wanted to capture and measure, the technologies and methods for it may not yet exist – after all, AlphaFold would not exist without the help of technologies like X-ray crystallography to determine the spatial arrangement of atoms.
  • Even if the technologies and methods do exist, they may be inadequate for capturing the full richness of clinically-relevant data. Or worse still, the data collection process may effectively destroy the data source, which means no further information can be learned from that specific sample – e.g. H&E staining (a process used for identifying tissue types and detecting diseases like cancer) relies on chemical dyes, which causes tissue damage and interferes with further downstream testing. That’s not an issue you’d have with training ChatGPT on text data.
  • And even if the techniques exist and are adequate, they may be expensive and time-consuming. Unlike ChatGPT, which was built (relatively cheaply) by scraping almost the entire internet, biological models are far more expensive because the data must be physically manufactured in wet labs through experiments which can be costly, and validation of model outputs even more so (for instance, making and testing a single compound can cost more than $5K and take c.6 weeks). Also, the machines that would generate the data (e.g. mass spectrometers or cryo-EMs) are extremely expensive.
  • The notable exception to this (”expensive and difficult” collection problem) is data from wearables and other Real World Evidence (RWE) data generators – they are increasingly getting more sophisticated, but are by no means a silver bullet or suitable for all data types (it’s not like they can capture tumour microenvironment data, for example, and there’s also the problem of “missingness” driven by inconsistent wearable use).

Scale, diversity and quality of datasets are currently limited: Given the time, difficulty and expense as we outlined above, this is unsurprising. Besides lack of scale, insufficient diversity is also an issue. For instance, Basecamp Research found that Sequence Read Archive (a vast repository with over 50 petabases of sequencing data, considered to be one of the most comprehensive snapshots of Earth’s genetic diversity) was not so diverse after all; 68% of all sequence data in the Sequence Read Archive database comes from just 5 species, and 70% of all submissions originate from only 10 countries.

Similarly, within the human biology context, we heavily rely on cell lines (populations of cells kept alive in a lab for study) but these cell lines represent a narrow slice of humanity and range of interactions (so they don’t capture genetic diversity, tissue microenvironments or any of the other complex interactions within the human body – much less how those complex interactions evolve over time). As Dibaeinia et al argue, scaling model capacity alone is insufficient to successfully build Virtual Cells (computational models capable of predicting cellular responses in response to perturbations) because the primary failure mode is a lack of adequate coverage over diverse biological contexts.

Additionally, most public databases of biological data were created from a human-centric perspective for academic research, and not for the purposes of training AI models. As a result, there is a lack of standardisation and interoperability of different datasets – the biology techniques used for data collection may differ, the types of metadata collected may vary, data annotation pipelines may be constructed differently and so on.

The challenges we outlined earlier are in fact opportunities for TechBios, and the factors that make a proprietary data moat (particularly in areas characterised by data scarcity) so compelling. Besides clear scientific differentiation and defensible biology, we’ve listed out some of the attributes that underpin an attractive data moat: (1) Scale and diversity; (2) Velocity (how quickly the dataset is growing and how fresh the dataset is); (3) Multimodality and richness (that captures biologically meaningful signals relevant to real clinical outcomes); (4) Quality (consistency, standardisation and interoperability); (5) Novelty of biology and/or technique; and (6) Cost effective yet high quality data generation.

1. Scale and diversity: We are witnessing exciting and ambitious efforts by TechBios to scale up diverse datasets. Earlier this year, Tahoe Therapeutics, Arc Institute, and Biohub announced a collaborative initiative to generate the largest and most perturbation-rich single-cell dataset for virtual cell models (which is to be 4× more perturbation-rich than Tahoe-100M; Tahoe-100M itself is the world’s largest single-cell dataset, 50x larger than all public drug-perturbed data combined). Meanwhile, Basecamp Research launched its Trillion Gene Atlas initiative, which aims to expand known evolutionary genetic diversity 100x by collecting genomic data from more than 100 million species across thousands of sites worldwide.

That said, there are certain cases where “build the largest and most diverse data set” simply isn’t a viable strategy – e.g. with covalent drug discovery where (a) there’s not much experimental data available to train foundational models; and (b) most existing models are classical mechanical models that don’t understand electron rearrangement, which is fundamental to covalent bond formation – this means data-driven approaches alone can’t work. That’s why startups such as Axiom Therapeutics combine physics-based simulations with machine learning acceleration, capturing protein-drug interactions at an atomistic level and enabling the platform to scale without needing massive experimental datasets (the data sharpens the platform rather than defines it).

2. Velocity: This refers to the ability to rapidly design and expand novel, large-scale datasets that capture real-world problems and, when inevitably moving out of domain, generate new experimental data faster than anyone else. We discuss this in more detail in the “Infrastructure and Workflows” section of our report, as it is the wet lab workflows that determine this velocity.

3. Multimodality and richness (that captures biologically meaningful signals relevant to real clinical outcomes): Multimodal datasets combine diverse biological and chemical data types (e.g. sequence, structure, functional assay results, clinical signals, temporal information) into a unified framework to enable more accurate and holistic modelling of biological systems and drug behaviour. Multimodality is necessary because biology is inherently complex, contextual, and dynamic, and can’t be understood thoroughly through isolated or single-modality data. For example, Noetik links tumor samples with long-term patient outcomes to train models that predict therapeutic response, while its OCTO model demonstrates emergent capabilities (e.g. inferring tumor vs immune cells from imaging alone, without task-specific training) enabled by broad multimodal training.

Similarly, incorporating longitudinal patient-derived multi-omics data across treatment cycles allows models to capture real biological dynamics (rather than static snapshots) – Orakl Oncology CEO Fanny Jaulin calls this “patient-in-the-loop,” a term we loved. This is not only about capturing as much data as possible from a sample, but also about capturing all the clinically relevant datapoints, and focusing on true biological signals. Ultimately, high-signal multimodal data enables a shift from studying cells in isolation to understanding interacting systems and patient-level outcomes (we’ll discuss this in more detail later when we speak of translational modelling, biomarker selection and patient stratification).

“Good multimodal data isn’t just multiple data types – it’s deeply paired, spatially resolved, and structured in a way that actually lets models learn relationships. The focus should not just be on linking datasets, but actively constructing them – adding temporal and environmental context so the data tells a coherent story. Most multimodal datasets don’t meet that bar.”

– Ron Alfa, Co-founder and CEO at Noetik

Another way of ensuring richness of data is by maintaining its rawness. To illustrate: Standard MRI workflows discard the raw signal and rely on reconstructed images, which are optimised for understanding anatomy and not furthering cognition. That’s why startup Karavela retains the full raw MRI data (k-space), even though it’s 100× larger, because it preserves richer information about brain dynamics. This allows them to rethink how data is sampled and directly optimise data acquisition for decoding (rather than imaging) to unlock fundamentally better models of brain function and new diagnostic capabilities. A similar philosophy underpins CellType’s approach, as a way of avoiding human bias:

“Biology today is filtered through papers and human bias, which only capture a fraction of the underlying signal. Our approach is to train directly on raw, multimodal biological data – giving models a much more complete and unbiased understanding of biology.”

– David van Dijk, Co-founder and CEO at CellType

4. Quality: Data quality is critical in AI-driven biology because models are only as reliable as the data they learn from – poorly structured, inconsistent, or contextless data leads to weak or misleading predictions. Unfortunately, most biological data is scarce, low quality, and highly inconsistent across labs. Achieving high-quality data requires standardisation, validation, and strong governance: experiments must be run under consistent conditions using harmonised protocols, with results linked to their full context (e.g. protocols, instruments, materials). Enforcing strict schemas, automated validation checks, and direct instrument integration helps prevent errors and eliminate manual inconsistencies, while adopting shared vocabularies and metadata standards ensures interoperability across teams. We were particularly interested in how Eli Lilly’s TuneLab reinforces standardised data generation by making Eli Lilly’s data generation protocols available to partners. It also appears that Eli Lilly is working towards launching a new feature where it can streamline options to have the data generated with a CRO using a harmonised protocol (which ensures that even outsourced data remains consistent and usable). Together, these practices create clean, structured, and reliable datasets that can effectively power AI models.

5. Novelty of biology and/or technique: Novel data generation techniques are redefining how biological datasets are created by making them richer, scalable, and more clinically relevant.

An example of a novel data extraction technique is the combination of label-free tissue imaging with AI, as demonstrated by Modella AI and illumiSonics’ partnership, where instead of relying on traditional chemical staining (which is destructive and limits downstream testing) label-free imaging captures high-resolution, data-rich tissue information without damaging the sample, enabling repeated analysis and multi-modal data extraction (morphology, genomic, proteomic) from the same tissue. The result is a step-change in data efficiency: more information per sample, multiple modalities from a single experiment, and scalable datasets that better link biological signals to clinical outcomes.

Our conversations with senior pharma practitioners suggests that they are focused on novel biology and areas characterised by data scarcity. Some of the companies building novel datasets through proprietary techniques and/or biology include Proxima, Granza Bio, ALLOX, Graph Therapeutics, Brink Therapeutics, Molecular Glue Labs, Outpost Bio, Valinor, Scripta Therapeutics, Synteny*, Noetik, Orakl Oncology, Karavela…. the list is long and by no means complete. Also, all of these startups are targeting different combinations of therapeutic areas, indications, modalities, experimental techniques etc. so there isn’t much of an overlap amongst the datasets of the startups we’ve mentioned. For example, Granza Bio is focusing on attack particles and novel delivery mechanisms, while ALLOX is focusing on allosteric binding sites, Proxima is pioneering a new therapeutic modality called Proximity Modulation (ProMod)… you get the drift. We feature a market map of startups in Part 2 of our TechBio series, so stay tuned!

6. Cost effective yet high-quality data generation: A great example of this would be startup Simulacra’s efforts – its goal is to solve the massive data bottlenecks that exist for molecular simulations in silico. Every pharma company relies on simulation models, yet there are tradeoffs: (a) using fast, cheap, but inaccurate models; or (b) using computational quantum chemistry, but that gets very expensive. Simulacra breaks through this false dichotomy by showing that it’s possible to get cheap yet accurate quantum chemistry with AI, and published a benchmark showing that it’s 50x cheaper than alternatives (or, a 98% cost reduction).

The best laid plans of mice and models: Algorithms

De novo (rational) drug design uses computational models to generate new molecules from scratch (building them atom by atom to meet specific biological constraints), which shifts drug discovery from trial-and-error to data-driven design. To be effective, these models need to generalise beyond their training data, as much of the biological and chemical space remains unexplored.

However, achieving this generalisation is challenging because biological systems are highly complex, dynamic, and interconnected, while most current AI architectures (e.g. transformers, diffusion models) were designed for text or images (rather than the sparse, discontinuous nature of biological data). As a result, progress depends not only on more data, but also on developing biology-native AI architectures i.e. approaches that better capture underlying biological reality. These are some of the areas we are most interested in: (1) Novel architectures or approaches to support generalisation; (2) High quality, fast inference at low costs; (3) Multi-model approaches; (4) Translation models, biomarker discovery and patient stratification with clear mechanism of action; (5) Target discovery engines.

1. Novel architectures or approaches to support generalisation: Generalisation is the Holy Grail problem of AI x TechBio – even our best efforts at data generation could fall short, which is why we need models that deeply understand the underlying rules of biology (and not just memorising or pattern matching). This could involve either finding ways of making existing architectures work, or developing entirely new architectures. As an illustration of the former method, CellType’s Cell2Sentence has cleverly adapted biological data to the prevailing model architecture rather than the model architecture to the data (their core technology translates structured biological data into a format that LLMs can reason over).

As an illustration of the latter method, we see models like EchoJEPA and BioState AI’s GeneJEPA that are applying novel architectures such as Joint-Embedding Predictive Architecture (JEPA). To illustrate: Transformer based models like scGPT treat gene data like language and try to exactly recreate gene activity levels. The problem is, this makes them learn noise and quirks from specific experiments instead of real biology. GeneJEPA takes a different approach. Instead of copying raw data, it focuses on learning the underlying relationships between genes (basically the rules of how genes interact). It treats genes as a group rather than a sequence, ignores noisy details, and uses more stable training methods. Because of this, GeneJEPA better captures the underlying biology and works more reliably across different datasets and tasks. It also performs better on things like identifying cell types, predicting drug responses, and can even make new predictions (like simulating gene knockouts) without being specifically trained on them.

2. High quality, fast inference at low costs: For instance, Boltz focuses on delivering low-cost compute (with ongoing efforts to continue lowering the cost) driven by their model optimisation efforts. This is because inference-time scaling laws clearly hold; when Boltz agents evaluate more protein or small-molecule designs at inference, the quality of design improves. In similar vein, Converge Bio is working on benchmarking AI accelerator chips FLOPS per $ (as achieving cost-effective computational performance is a long term goal for the company) and Latent Labs has built a proprietary architecture optimised for fast inference.

3. Multi-model approaches: Given different models excel at different things, it makes sense to use the ones best suited for a particular task. For instance, Converge Bio uses LLMs, diffusion models, traditional machine learning and statistical methods, and Ingenix developed a Model Fusion architecture. Yet another implementation of the multi-model idea is exemplified by Tangram Therapeutics, which developed LLibra, a multi-LLM, agentic system for early discovery that combines Bayesian hierarchical GNNs with reinforcement learning to improve its outputs over time. An evaluation harness selects the best model per task, where Tangram currently operates eight LLMs and adjusts that set as performance data changes. This results in a system that is modular and extensible, which can adapt to new technologies and plug in improved models as they emerge.

4. Translation models (”of mice and men”) and patient stratification: Once we’re reasonably happy with the drug the AI model generated, we’ll then test it in vitro (in the laboratory) and in animals (such as mice) to assess how safe or effective it could be in humans. But curing cancer in a petri dish or in mice isn’t the same as curing cancer in humans (given petri dishes can’t capture the full complexity of human biology, and at the risk of stating the obvious, mice and humans have very different biologies), which is exactly why we need translation models – they account for the physiological differences between species to estimate safe starting doses and therapeutic efficacy in patients. That said, even if we move past the “mouse to men” translation, human biology is complex in general and complex diseases (like cancer) are influenced by multiple factors that vary across patient populations, and this variability makes it difficult to predict how a drug will affect an individual patient – which means that we need translation as well as better patient stratification. Given the extraordinarily high failure rates of drugs that actually make it to clinical trials (90%) that we alluded to earlier, this is of paramount importance.

This is why we increasingly see startups (such as Noetik, GraphTX, Orakl Oncology, MultiOmic Health* and others) hyperfocused on building out datasets that (a) collect as much clinically-relevant biological context as possible, which captures the complexities of disease in patients; and (b) directly link this data with clinical outcomes (besides improving translation, the clinical linkage is also useful in identifying which patients are actually likely to benefit from the drug. This makes it easier to recruit the right patients for the clinical trial, and boosts its chances of succeeding). Additionally, these startups focus on capturing causal factors to better understand the mechanism of action – knowing how a drug works allows you to design safer trials and predict some side effects, reducing failure due to toxicity.

While the startups we mentioned are developing their own drug pipelines while keeping translation and patient stratification front-and-centre, we also see startups such as Valinor and Atlas Bio that aren’t creating drugs but are instead solving the translation, biomarker discovery and patient stratification problems through virtual patient models and virtual clinical trials. Valinor, for instance, trains foundation models on proprietary longitudinal patient-derived multi-omics data to simulate patient biology over time, enabling use cases from clinical trial surrogate endpoints to early diagnosis. We’ll cover this in more detail in Part 2 of our TechBio series.

We expect to see greater industry acceptance of in silico methods over time. Somewhat encouragingly, in 2025, the FDA announced its plans to reduce its animal testing requirements for certain modalities of drugs, potentially refining or replacing them with a range of approaches (such as AI-based computational models of toxicity and cell lines and organoid toxicity testing in a laboratory setting). These are known as New Approach Methodologies or NAMs, and the FDA’s acceptance of these depends on evidence that the model accurately reflects human biology/physiology or the specific disease state being targeted (including robust data on the model’s reliability, accuracy, and repeatability). We’re excited about the ongoing efforts to make these NAMs more robust, and are keenly tracking them.

5. Target discovery engines: Pharma is plagued by “target crowding,” where a large number of drugs are targeting, well, the same biological target. Although pursuing a well-known, validated target seems like a safe bet, it has the unfortunate effect of leading to excessive competition and reduced commercial viability for all the drugs. A 2025 analysis found that just 38 targets account for roughly a quarter of the entire preclinical and clinical pipeline, and another recent research found that in each of the five major modalities, just the top 5 targets account for 20-50% of all asset programmes.

To address this issue, we see startups either uncovering entirely novel targets (previously not identified in literature at all) or finding ways to make previously “undruggable” targets druggable. One way of doing so is to work bottom up from the data (rather than starting with the literature) – at Scripta Therapeutics, discovery begins with disease-associated transcriptional signatures rather than predefined molecular targets. Transcription factor activity is analysed in cellular context to identify upstream drivers of pathology, with proprietary networks of biology then used to guide the selection of druggable regulatory nodes. Besides aiding in novel target discovery, this approach helps in building a more mechanistic interpretation of disease biology. All of this nicely ties into our previous points around unbiased or data-driven discovery, as well as translation and clear mechanism of action.

Finding the needle without the haystack: Infrastructure and Workflows

Despite (and because of) the rise of powerful computational models, experimental validation remains essential, forcing even pure-software AI companies (who aren’t developing their own drug pipelines) to maintain wet lab infrastructure to generate proprietary data. Today, the real bottleneck is developing assays and infrastructure to produce high-quality, standardised, and multi-parametric experimental data at scale.

Therefore, wet lab experiments interact with AI in two critical ways: labs produce the data that powers AI models, and AI models in turn help to create more targeted wet lab experiments. Rather than running broad, exploratory screens, labs will shift toward fewer, higher-quality experiments that directly test and refine computational predictions. We don’t believe AI will fully replace experiments; rather it will make them more efficient, focusing resources on the most promising paths instead of brute-force approaches (a view substantiated by other research as well; Müller et al proved that ML-guided iterative screening can significantly reduce the experimental cost while maintaining hit discovery quality).

This virtuous cycle has led to the popularisation of the “lab-in-the-loop” approach, which combines AI with real-world laboratory experiments in a continuous feedback cycle. AI makes predictions, which are tested in the lab, and the resulting data is fed back to the AI for further iteration. For instance, Relation Therapeutics co-locates its AI and wet lab teams in a single integrated facility and runs recursive lab-in-the-loop cycles where every experiment is designed around translational fidelity – ensuring the AI learns from assays that actually predict what happens in patients.

An example of this can be seen in Xu et al’s recently published work on LUMI-lab, a self-driving platform that pairs an AI foundation model with a robotic lab to autonomously discover ionizable lipids (LNPs) for mRNA delivery. The chemical space to be explored is vast, the experimental cycles are slow, and LNP datasets were too small to train a predictive model from scratch. But with the combination of a foundation model (that was pre-trained on the broad chemical space) and a tight lab-in-the-loop process (where each round of real wet-lab experiments fine-tunes the model, which then proposes smarter candidates for the next round), they overcame the severe data scarcity issue in this area and produced impressive results (their top performing lipid designed this way, LUMI-6, achieved 20.3% gene editing efficiency in lung epithelial cell, surpassing the highest editing efficiency previously reported for LNP-mediated CRISPR-Cas9 delivery via inhalation.)

But scaling these cycles alongside pharma partners introduces a different kind of challenge – one of infrastructure, not science:

“AI is pushing life sciences back toward their origins: smaller feedback loops, faster experimentation, closer connection between scientists and the science itself. But startups operating alongside large biopharma partners have to run these cycles inside regulated, global-scale systems. That creates real tension. The instinct is to validate the science first and worry about infrastructure later, but founders who’ve been through pharma partnerships consistently find that qualification, compliance, and integration timelines could kill partnerships, and not the science. Building with interoperability and compliance in mind early (even if not at full scale) shortens the path to partnership and scalable Go To Market.

Cloud tends to be the most practical route once you’re running multi-site experiments or working with pharma. It gives startups quick access to elastic compute for TechBio’s bursty cadence, plus the audit trails and data lineage regulated workflows demand. Life-science-specific cloud-native services like Amazon Bio Discovery and AWS HealthOmics further optimise agentic and data pipelines for multi-modal biological data, and make it easier to integrate with partnering biopharma data infrastructures. This layer is seemingly invisible but load-bearing: underpinning the shift from ‘we retrained a model on one experiment’ to ‘our models run on thousands of experiments across multiple sites in near real-time.’”

– Daniel Rabina, Healthcare & Life Sciences Startups at AWS


Even if we weren’t building AI models with very small datasets initially (as was the case with LUMI-lab and ionizable lipids), we noticed that a number of startups are focusing on building their moat in post-training workflows (an unsurprising development given how performant open source models are becoming). We’re also increasingly seeing agentic AI systems being deployed across startups for optimising the drug discovery and development workflows. The ultimate evolution of this is the autonomous robotic wet lab integrated with agentic AI workflows.

“We’re entering a new era of AI in pharma – shifting from solving isolated problems to asking better scientific questions and designing the right experiments. AI is evolving from models to decisions, becoming an operating system that connects data, experiments, and insights into a continuous loop. An ‘in silico first’ approach ensures the lab is used intelligently rather than replaced. The real opportunity isn’t just speed, but precision: generating hypotheses at scale while designing optimal experiments upfront to reach decisions far earlier.”

– Yves Fomekong Nanfack, Head of AI/ML – Research at Takeda

When it comes to infrastructure and workflows, here’s what we’re looking out for: (1) Rapidity of iteration cycles; (2) Higher quality with cost effectiveness; (3) Improved context capture; (4) Process optimised for curiosity, novelty and unbiased discovery

1. Rapidity of iteration cycles: For instance, Onava focuses on the speed and tight integration of its data generation and validation loop. The company has built proprietary high-throughput assays that allow it to go from computational design to experimentally validated functional data in <7 days, with ongoing efforts to compress timelines even further. This rapid turnaround enables continuous, weekly iteration – where new biological data is generated, fed back into models, and used to retrain and improve predictions in near real time. Validation is not a separate downstream step but embedded within this loop: assays are designed to produce functional, quantitative readouts (e.g., binding, expression, off-target effects) that immediately inform model performance. As a result, Onava can iteratively refine candidates across multiple rounds, moving from initial hits to optimised molecules within weeks, while simultaneously generating high-quality datasets that strengthen future predictions. Rather than relying on single-shot generation, Onava has found success through iterative design coupled with rapid experimental validation. Similarly, Tangram Therapeutics focuses on continuous learning:

“Continuous learning is the real moat in AI-driven drug discovery. It’s not just about ingesting new data – it’s about systems that evolve with every interaction, capturing scientific intuition, adapting in real time, and compounding knowledge with use. The best platforms don’t stay static; they get smarter every day they’re used.”

– Emma Slade, Former Head of Applied AI at Tangram Therapeutics

2. Higher quality with cost effectiveness: Synteny and Noetik are great examples of combining high-throughput assays with AI. For instance, Synteny’s MYRIAD platform captures millions of TCR–pMHC interactions in biologically relevant conditions and (unlike traditional methods) not only measures binding but also actual immune activation (thus bringing lab data closer to real patient biology). While each datapoint is noisy, MYRIAD is roughly 10^5 cheaper per interaction than gold-standard methods like surface plasmon resonance (SPR), which is prohibitively slow and expensive (~£700 per interaction). Instead of relying on costly experimental replication, Synteny leverages its deep learning model, ARLO, to learn from this noisy, large-scale data to accurately estimate binding affinity, thus enabling scalable, fast, and low-cost discovery without needing traditional measurements. In similar vein, spatial transcriptomics is the gold standard for tumour characterisation but is costly and complex, while the widely used H&E staining is cheap yet limited. Noetik’s AI model is capable of converting between these two modalities, thus offering high fidelity tumour characterisation at a fraction of the cost.

3. Improved context capture: Today, most experimental data is fragmented and siloed. Somewhat horrifyingly, we’re still relying on primitive methods like manual entry in spreadsheets. However, we believe the future will be driven by automatic data capture from laboratory instruments that directly feed into structured, AI-ready systems. But this automatic data capture is just one piece of the puzzle – there’s a lot of ambient decision-making that scientists do that’s not captured, and we couldn’t explain it better than Dave Light:

“So let’s say you ran a bunch of experiments a while ago. You’re now using that data with some agentic system. The agent knows you ran the experiments, but it’s harder for it to understand the context: what you thought of the design, why you ran it that way, what you concluded, how that should inform the next set of experiments, why you’re proceeding with or terminating a project.”

– Dave Light, VP, Strategy & Corporate Development at Insitro

We think this context capture and management is absolutely critical, and we’re seeing interesting approaches that are developing in this regard. We’ve written extensively about context and memory management solutions for AI agents, particularly within the knowledge graphs and ontologies layer. We also see AI Scientist platforms such as Phylo focusing on making tacit knowledge explicit – the trail of decisions and reasoning is captured by default. Queries, analyses, decisions, and rationale are preserved as scientists work. Additionally, as every researcher standardises on the same AI Scientist platform, it becomes the organisation’s scientific memory and execution layer – thus dismantling data silos.

4. Process optimised for curiosity, novelty and unbiased discovery: We’ve already talked about how developing truly novel solutions would require AI models to generalise beyond their training dataset. Besides that, it is imperative that we adapt our processes and workflows to encourage unbiased discovery and curiosity. A primary metric of success is the system’s ability to uncover unexpected features that human researchers might overlook. For example, LUMI-lab identified “brominated lipid tails” as a key enhancer for mRNA delivery (a design feature that was previously unrecognised and not part of the initial human hypothesis).

We’re seeing novelty and curiosity based approaches emerging in agentic workflows. To illustrate: besides building agents for productivity-enhancing jobs (data processing, bioinformatics analysis, AutoML etc), Synteny is also building agents to direct its wet-lab activities to ensure every experiment yields the highest possible value by teaching the AI how to explore (optimising for curiosity and information gain) rather than just memorising existing biological data.

“Generalisability isn’t a single breakthrough, it’s a system. It’s high-quality, diverse data; infrastructure that enables rapid feedback; and models capable of learning deep complexity from limited real-world signals. But what truly unlocks it is autonomous discovery, where systems that can explore and refine on their own. Together, this forms the technological backbone for a far bigger ambition: to uncover the general rules of biology. The question for the next 20 years is whether we can do for biology what physics achieved a century ago – not just observe complexity, but truly model and understand it.”

– Lilly Wollman, Co-founder and CEO at Synteny

The “virtual biotech company” that Zhang et al demonstrated recently is another great illustration of what agentic workflows can accomplish – it’s a multi-agent framework for therapeutic discovery and development. The agents achieved fascinating outcomes, such as: (1) identifying that drugs targeting cell-type-specific genes were 40% more likely to progress from Phase I to II, 48% more likely to reach market (Phase IV), and had 32% lower adverse event rates; (2) showing why B7-H3 targeted therapies could work in lung cancer by combining multiple types of biological and clinical data; and (3) rethinking a failed Phase II trial for an ulcerative colitis drug, identifying that it lacked a precision-medicine approach and proposing a better strategy.

We’ve covered the technological aspects of what “good” looks like across proprietary data, algorithms, infrastructure and workflows – but it’s equally important to prove that these work and drive outcomes. And that’s what we’ll talk about next.

Proving the platform works before you’ve got pipeline: Early Validation

AI for TechBio lacks rigorous and consistent benchmarks, with fragmented and custom evaluation methods making it hard to compare models or measure real progress. For instance, Škrinjar et al demonstrated that current co-folding approaches largely memorise molecular interactions from their training data, which means they aren’t actually learning the underlying biological principles. So if they’re only doing pattern matching rather than actually thinking from first principles, it would limit their usefulness for de novo drug design… and that’s why we need to find good ways to measure real progress.

The typical way to validate how good an AI platform was to get clinical trial readouts for the AI-designed drug, but you won’t have this level of clarity at the early stages. Which brings us to some critical questions: Before you have any clinical trial data on your drug, how do you validate that your AI platform works? What does “good” look like, and what convinces pharma to partner with your startup at an early stage?

“There are two ends of the spectrum. At the front end: do I have a unique insight that gives me first-mover advantage on a target that could become something big, based on strong genetics and high translational potential? At the back end, we track more traditional KPIs: how many molecules did you need to make to get to a development candidate, and was that candidate patentable? Was it differentiated against the target product profile? If you set out to make the most potent, most selective molecule with no off-target hits, did you get there?”

– Karen Akinsanya, President, R&D – Therapeutics & Chief Strategy Officer, Partnering at Schrödinger

“How do you validate your platform works and what does ‘good’ look like? ” is a complicated question to answer, because there was a tremendous degree of variability across the 40+ interviews we conducted over the course of this research project. What “good” looks like differs from pharma partner to pharma partner, because: (a) we don’t always have shared, clinically relevant standards; and (b) it depends on the pharma partner’s own capability in evaluating something or their own expertise in/understanding of complex systems. If a pharma partner is an expert in some domain that doesn’t have reliable, widely accepted early validation mechanisms (e.g. GUBRA mouse models for MASH) they may still be willing to take early risks because they believe they know how to interpret imperfect information better.

We’re shamelessly borrowing from Dave Light’s framework on the subject, which breaks it down beautifully:

  • Known biology: If we know a lot about the disease (it’s population scale, it’s a known target, and the translation models e.g. the GUBRA mouse for MASH are reliable enough), it’s much easier to validate earlier. Let’s say a legacy drug for that target made it through Phase I, even if it failed later (on grounds of efficacy), but it was safe. If your new drug looks significantly better in the same preclinical assays and IND-enabling studies as the legacy drug, you have clear things to benchmark against. The presumption is that once your drug is in humans, it’ll work better vs the legacy drug.
  • Unknown/novel biology: This is when there is significant unmet need, there are no good options currently, and any advancement seems positive. That could also lead to a faster path to clinic, and you could run a small Phase I trial to get proof of principle easily.
  • The No-Man’s-Land in the middle: Where there are already existing drugs (so the bar for efficacy is already high), the translation models aren’t reliable, and it’s not a population-scale problem (which means the rewards for taking on a lot of early stage risk aren’t easy to justify). This is an area startups should strictly avoid.

Given we’re particularly interested in novel biology, we asked TechBio founders how they demonstrate that their AI solutions work, and three key mechanisms emerged: (1) standardised pre-clinical packages; (2) human validation in non-clinical trial settings; and (3) winning trust by predicting the unseen.

“We think of early validation in terms of three pillars: benchmarking against established standards (such as validating findings in familiar models like mouse studies), ensuring reproducibility across experiments and environments, and demonstrating true generalisability, where insights hold consistently across new datasets, populations, and real-world conditions.”

– Jenny Yang, Co-founder and CEO at Outpost Bio

Standardised pre-clinical packages: In addition to using other validation mechanisms, Synteny focuses on highly standardised preclinical data package, where each molecule is evaluated against four core criteria (potency, specificity, safety and developability/druggability) using carefully designed assays that generate meaningful, decision-driving readouts. The goal is to build strong confidence in whether a molecule should progress by benchmarking these preclinical indicators and later comparing them against actual human outcomes to validate their predictive power.

Human validation in non-clinical trial settings: Outpost Bio is a great example of this. They generate pre- and post-intervention data on the human microbiome, using metagenomic analysis to compare the microbial community before exposure with how it changes over time after encountering a chemical structure, such as a drug or food compound. This underpins their foundation model, which predicts the results of such drug/food interventions. Essentially, their work on the human microbiome has applications across pharma for drug discovery, as well as consumer health (e.g. evaluating the impact of a new healthy, probiotic drink on the gut microbiome of potential consumers).

Because pharma sales cycles are long and validation mechanisms through clinical trials are longer still, beginning with consumer health R&D teams makes a lot of sense, because bringing a new food/drink product to market is significantly faster, so you’re able to validate the predictions of the platform much earlier. With the validation from consumer health in place, it’s much easier to approach pharma partners.

Winning trust by predicting the unseen: Pharma has a lot of internal data (especially about their own molecules) that isn’t publicly available – and they’ll know things you won’t. However, if your AI platform is able to predict and arrive at the same insights that were generated internally (without ever having had access to the internal data) it proves how good your platform is. For instance, Yale researchers (who eventually founded CellType) screened over 4,000 compounds and not only found CX-4945 to be a highly promising candidate (CX-4945 or Silmitasertib is Senhwa Biosciences’ core asset), but also independently identified a drug effect that was not publicly known, and Senhwa confirmed the same finding based on years of unpublished internal research – despite CellType having no access to that data. This external validation established credibility and led to a partnership between CellType and Senhwa Biosciences.

In similar vein, Orakl Oncology is able to predict the outcome of clinical trials with 88% accuracy e.g. predicting overall response rates and progression-free survival, and showing how little those predictions differ from the pharma partners’ own internal data (which evidently was not publicly available, so no chance that Orakl’s models could have been trained on that data). Such validation (where AI-derived insights are independently corroborated and translated into revenue-generating, IP-creating partnerships) offers a compelling proof point for platform efficacy. We couldn’t describe this strategy better than Simon Romanski:

“The best validation for us is when we can independently predict a biomarker or other insight that a partner has already generated internally, which is confidential and ideally counterintuitive. Predicting something previously unknown is the best way to build trust with the customer.”

– Simon Romanski, Co-founder and CEO at Atlas Bio

Translating the gap from models to medicines

We’re in very early days of AI x TechBio, but it’s promising – Insilico’s rentosertib delivered the first Phase IIa validation of a fully AI designed drug, demonstrating real world efficacy and safety whilst getting to preclinical candidate (PCC) stage within 18 months, and at 10% of the cost typically required to get to a PCC.

That said, the industry’s centre of gravity remains skewed toward early discovery gains, while the true bottlenecks lie downstream, where biology, translation, and clinical reality assert themselves with unforgiving complexity. We think progress will not just come from better molecule generators, but also from systems that can connect data, models, experiments, and patients into a continuous, learning loop.

What emerges is a new stack for drug discovery: data that is richer, more longitudinal, and more representative of human biology; models that generalise beyond narrow datasets; and workflows that tightly integrate AI with experimental validation. The winners will be those who optimise and orchestrate across all of those layers – turning fragmented advances into compounding feedback systems. In this framing, AI is less a tool and more an operating system for biology. If you’re a TechBio founder, please reach out to Advika or Charlotte – we’d love to chat.

*Synteny and MultiOmic Health are MMC portfolio companies