BioTender AI Biology Infrastructure Atlas

Mapping the infrastructure powering the AI biology ecosystem
v0.2 · Curated by BioTender (biotender.online) · Collection date 2026-07-21 · 11 layers (L0–L10) incl. AI Models & Tools · Every field provenance-tagged (VERIFIED / DERIVED / SUBJECTIVE / UNVERIFIED)
313Entities
802Relations
11Layers
40AI Models
38Tools
219Open-source

Data Explorer

Search and filter all 313 infrastructure entities across 11 layers (L0–L10, now including AI Models and Bioinformatics Tools). Every card opens a full provenance-tagged record: capabilities, access, security, pricing, sources, and relationships. Open-vs-commercial is a first-class filter.

AI Biology Stack

The infrastructure stack, bottom-up: biological DataComputeModel distributionSoftwareWorkflowAgentExperimentalEnterpriseAI ModelsTools, with cross-cutting standards & organizations alongside. Click any entity to explore its dependencies.

Infrastructure Registries

Four curated registries of the building blocks of an AI-biology stack: Compute providers, Model distribution & models, Agent infrastructure, and Workflow engines. Each row is a real, sourced entity — click for its full provenance record. Use the search box to filter within a registry.

Infrastructure Dependency Graph

A hand-rolled force-directed graph of 802 documented relationships across 21 relationship types (depends_on, integrates_with, supports, hosts, hosts_model_repository, trains, runs, implemented_by, validated_by, develops, developed_by, and more — each colour-coded). Drag to pan, scroll to zoom, click a node to highlight its neighbourhood. Every edge carries an evidence note, confidence (H/M/L), and last-verified date.

nodes · edges shown · drag background to pan · scroll to zoom · click node to focus

Compare Infrastructure

Put 2–4 infrastructures side by side across 15 dimensions — purpose, scale, openness, commercial use, API/SDK, self-hosting, compute, data handling, security, pricing, maintenance, community, biology coverage, and enterprise signals. Cells stay honest: unverifiable dimensions show "Unverified / N/A".

Presets:

Stack Generator — What Infrastructure Do I Need?

Pick your task and your needs; a rule-driven engine assembles a recommended, evidence-backed stack from real, sourced infrastructure, then expands it along documented dependency edges (models pull in the frameworks they depend on, the datasets that train them, and the registries that host them). Fully transparent — every pick and every derived link is a click away from its provenance.

1 · Task
2 · Needs (optional, multi-select)

Build Your AI Biology Stack

Start from a reference preset — Small Lab (lean, mostly-free) or Enterprise (governed, private-data) — then toggle components on/off. Each slot is filled with real, sourced infrastructure; the summary tallies open-source vs commercial.

Preset

Infrastructure Timeline

Founding / first-public-release year of each infrastructure, as stated by its sources. Color marks the layer. No dates are fabricated — anything without a verifiable date sits in a separate "undated" lane.

Organization View

Infrastructure grouped by the organization that builds and maintains it — from consortia (wwPDB, GA4GH) and public institutes (EMBL-EBI, NCBI, Broad) to hyperscalers and startups. .

Infrastructure Health & Data Quality

Operational status (Active / Maintained / Deprecated / Archived / Unavailable) plus honest uncertainty tracking: which fields could not be verified from primary sources. This is a transparency ledger, not a scorecard of quality.