About BioRoute
Literature-grounded enzyme discovery for researchers and scientists.
BioRoute is a literature-grounded enzyme discovery platform. It helps scientists find the right enzyme for their project in minutes instead of weeks of manual literature searching. Every recommendation is backed by evidence from peer-reviewed papers and curated databases — nothing is generated, everything is cited.
Describe what you need in plain English — a substrate you want to degrade, a reaction you want to catalyze, a pathway you want to engineer — and BioRoute returns a ranked list of enzyme candidates with kinetic parameters, experimental protocols, and full provenance for every claim.
The problem we solve
Enzyme selection is one of the most time-consuming steps in biochemistry and biotechnology projects. The information scientists need is scattered across dozens of disconnected resources, and assembling it by hand can take days or weeks.
A typical manual workflow requires you to:
- Search across multiple disconnected databases — UniProt, BRENDA, KEGG, and others — each with its own query language and data format.
- Read dozens of papers to find experimental conditions, expression hosts, and purification protocols.
- Cross-reference kinetic data from BRENDA with sequence annotations in UniProt and pathway context from KEGG.
- Compare candidates by hand with no systematic framework for weighing evidence quality, data completeness, or experimental relevance.
- Often miss the best enzyme because it is buried in a paper you did not find, or listed under an alternative name you did not search.
BioRoute does all of this automatically in minutes. It searches every major enzyme database and 214M+ scientific papers simultaneously, merges the results, and delivers a ranked, evidence-tagged comparison you can act on immediately.
How it works — the full pipeline
Every BioRoute search passes through an 8-stage pipeline. Each stage is fully automated and completes in seconds.
- 1
Query understanding (NLU)
Users describe what they need in plain English. BioRoute's AI parses this into a structured scientific query using tool_use structured output — no fragile regex JSON parsing — identifying the target substrate, desired reaction, organism preferences, and environmental constraints such as temperature range, pH, or cofactor requirements.
- 2
Multi-database enzyme discovery (9 sources in parallel)
BioRoute simultaneously queries 16+ live data sources to cast the widest possible net for candidate enzymes:
KEGG
Metabolic pathways, reactions, and enzyme search
UniProt
570K+ reviewed entries with sequences and computed protein properties
ExPASy
ENZYME nomenclature — EC number classification and cross-links
BRENDA
Curated kinetics for 5,400+ EC entries (bundled cache)
PubTator 3.0
Pre-extracted chemical, gene, and species entity tags
Semantic Scholar
214M+ papers — bulk search, abstracts, TLDR summaries
OpenAlex
Scholarly metadata, open-access PDF links, license signals
NCBI
Gene and protein records, taxonomy, and cross-database links
STRING
Protein interaction partners with evidence channels
IntAct
Experimentally validated protein–protein interactions
ChEMBL
Inhibitor and bioactivity data — IC50, Ki, mechanism
Rhea
Curated reaction equations and balanced stoichiometry by EC
InterPro
Protein domain architecture and family signatures
ChEBI
Substrate chemistry — SMILES, InChI, formula, synonyms
PDB / AlphaFold
Crystal structures and AI-predicted 3D models
Europe PMC
Full-text search across European life-science literature
Unpaywall
Open-access PDF resolution for full-text retrieval
CORE
Open-access aggregator — 200M+ metadata records
- 2.5
Entity resolution
Deduplicates and merges enzymes discovered from multiple sources. Resolves synonyms, alternative names, and cross-references into canonical identities so the same enzyme found in KEGG, UniProt, and a Semantic Scholar paper is recognized as a single candidate, not three separate entries.
- 2.6
Relevance gate (BM25 + categorical + LLM rerank)
Filters low-confidence candidates before expensive downstream aggregation. A three-layer gate — BM25 text scoring, categorical feature matching, and an LLM reranker — ensures only genuinely relevant enzymes proceed to the data-intensive stages that follow.
- 3
Literature search (214M+ papers)
For each candidate enzyme, BioRoute searches Semantic Scholar's corpus of 214M+ papers to find relevant research. It uses multi-strategy search with relevance scoring to surface the most informative papers — not just the most cited ones — for each specific candidate.
- 4
Data aggregation (kinetics, interactions, domains from 7+ sources)
For each candidate, aggregates kinetic parameters (Km, kcat, Vmax), protein properties (molecular weight, isoelectric point, stability), interaction partners, domain architecture, and experimental insights from the literature into a unified profile. Every data point is tagged with its source — whether a curated database entry or a specific paper — along with a confidence level and a full citation.
- 5
8-dimension ranking
Candidates are scored across 8 weighted dimensions. This produces a transparent, evidence-weighted ranking where every score is explainable and traceable.
Query match0.20Constraint match0.10Identity confidence0.15Literature quality0.15Experimental evidence0.15Source diversity0.10Protocol availability0.025Data completeness0.025 - 6
Extended Data Sheet generation (9 practical fields)
For each ranked candidate, BioRoute generates a structured data sheet covering 9 fields that bench scientists actually need: thermostability, solubility, cofactors, substrate scope, scale-up considerations, safety, storage stability, immobilization compatibility, and known inhibitors. Each field is grounded in database evidence and literature citations.
- 7
Protocol extraction (user-triggered, verbatim from PDFs)
When a user requests protocols for a candidate, BioRoute fetches full-text PDFs and extracts verbatim experimental methods, expression conditions, purification steps, and assay protocols directly from the papers. These are not paraphrased or generated — they are the actual methods sections from published research.
What makes BioRoute different
vs. UniProt / BRENDA / KEGG (individual databases)
- Those are databases, not discovery tools. You need to already know what you are looking for.
- BioRoute searches all of them simultaneously and cross-references the results into a single unified view.
- BioRoute ranks and compares candidates with transparent scoring — databases just list entries.
vs. BLAST / sequence search tools
- BLAST finds sequence homologs. BioRoute finds functionally relevant enzymes for your specific application.
- BioRoute incorporates literature evidence, kinetic data, and experimental protocols — not just sequence similarity.
vs. manual literature review
- BioRoute searches 214M+ papers in seconds, not weeks.
- Extracts structured data from papers automatically.
- Provides reproducible, transparent scoring instead of subjective judgement.
vs. ChatGPT / general AI
- BioRoute never hallucinates enzyme properties. Every claim is traced to a specific database entry or paper.
- Results include real kinetic parameters, real EC numbers, and real PDB structures — not plausible-sounding fiction.
- AI is used for query understanding and text extraction, not for generating scientific claims. All LLM calls use tool_use structured output with schema validation — no regex-based JSON parsing.
Key capabilities
Cross-candidate Intelligence
Groups enzymes by family, compares domain architectures, and detects underexplored patterns across the entire result set — surfacing insights that are invisible when reviewing candidates one at a time.
Extended Data Sheets
Nine practical fields per candidate — thermostability, solubility, cofactors, substrate scope, scale-up, safety, stability, immobilization, and inhibitors — each grounded in cited evidence.
Protocol Extraction
Verbatim experimental methods extracted from full-text PDFs, with consensus protocols built across multiple papers for each candidate.
Evidence Transparency
Every score and every data point shows its provenance. Users can trace any claim back to its source paper or database entry.
Kinetic Parameters
Real Km, kcat, kcat/Km, and Vmax values from BRENDA, with substrate specificity and source organism context for each measurement.
3D Structure Viewer
Integrated Mol* viewer showing PDB crystal structures or AlphaFold predicted structures directly in the results interface.
Relevance Gate
A three-layer filter — BM25 text scoring, categorical feature matching, and LLM reranking — eliminates low-confidence candidates before expensive data aggregation begins.
Export
Results exportable as PDF, Word, or Markdown for lab notebooks, grant applications, and publications.
Who it's for
BioRoute is built for anyone who needs to identify, compare, or characterize enzymes as part of their research or engineering work:
- Biochemistry and biotechnology researchers
- Enzyme engineers designing or optimizing biocatalysts
- Bioremediation scientists evaluating degradation pathways
- Industrial biotechnology teams scaling enzymatic processes
- Pharmaceutical researchers exploring enzymatic drug targets
- Students and educators in biochemistry and molecular biology
Our commitment to scientific honesty
BioRoute is a discovery and literature-aggregation tool, not an oracle. We hold ourselves to the same standard of transparency that scientists expect from their own work.
- Every data point is labeled by confidence level: supported, inferred, or low confidence. You always know how strong the evidence is.
- Results require expert review before wet-lab use. BioRoute accelerates the search — it does not replace scientific judgement.
- We distinguish between curated database evidence and text-mined claims, so you can weigh each appropriately.
- We never overclaim. If evidence is thin, we say so. If a kinetic parameter comes from a single paper, that context is visible.
Security and trust
BioRoute is built with defense-in-depth protections to ensure data integrity and safe operation:
- SSRF protection prevents internal network probing through any user-supplied URLs or external API callbacks.
- Prompt injection hardening defends against adversarial inputs designed to manipulate the AI pipeline.
- Every LLM output is schema-validated before it reaches the application — malformed or unexpected responses are rejected, not silently passed through.