Explore small molecule libraries
Search a large library against a protein target by scoring only a fraction of it, choosing each molecule from everything scored so far.
When to explore instead of screen
Section titled “When to explore instead of screen”A library screen scores every molecule you give it. That is the right tool when the library is small enough to afford end to end.
Exploration is for the opposite case: the library is far larger than the budget you want to spend on
it. You submit the whole library and a budget — how many molecules to actually score — and the run
spends that budget where it is most likely to pay off.
Each molecule is chosen using every result that came before it, so the run concentrates on the regions of the library that are scoring well. For the same number of scored molecules, that recovers several times more of the library’s best-scoring compounds than picking the same number blindly.
Start with start(), then poll retrieve() for status and phase and page list_results() as
molecules are scored. Exploration has no one-call run() helper: a run is long-lived and resumable,
so you drive it yourself.
import osfrom boltz_api import Boltz
client = Boltz( base_url="https://api.boltz.bio", api_key=os.environ["BOLTZ_API_KEY"])
target = { "entities": [{"type": "protein", "value": "MKTIIALSYIFCLVFA", "chain_ids": ["A"]}],} # see Input format for pocket_residues, reference_ligands, constraints, …
exploration = client.small_molecule.explore.start( target=target, library={ "format": "csv", "source": {"type": "url", "url": "https://example.com/library.csv"}, "smiles_column": "smiles", }, budget=7000,)print(exploration.id) # sm_exp_...Write your request body to small-molecule-explore.json (see Input format), then:
export BOLTZ_BASE_URL="https://api.boltz.bio"EXPLORE_ID=$( boltz-api --format raw small-molecule:explore start \ --input @json://./small-molecule-explore.json | jq -r '.id')
boltz-api small-molecule:explore retrieve --id "$EXPLORE_ID"import Boltz from "boltz-api";
const client = new Boltz({ baseURL: "https://api.boltz.bio", apiKey: process.env["BOLTZ_API_KEY"] });
const target = { entities: [{ type: "protein", value: "MKTIIALSYIFCLVFA", chain_ids: ["A"] }],}; // see Input format for pocket_residues, reference_ligands, constraints, …
const exploration = await client.smallMolecule.explore.start({ target, library: { format: "csv", source: { type: "url", url: "https://example.com/library.csv" }, smiles_column: "smiles", }, budget: 7000,});Input format
Section titled “Input format”An exploration takes a target to score against, a file-backed library to explore, and a budget capping how much of it gets scored.
{
"target": {
"entities": [{ "type": "protein", "value": "MKTIIALSYIFCLVFA", "chain_ids": ["A"] }]
# see Target for pocket_residues, reference_ligands, constraints, …
},
"library": {
"format": "csv", # csv | tsv
"source": { "type": "url", "url": "https://example.com/library.csv" },
"smiles_column": "smiles", # defaults to "smiles"
"id_column": "catalog_id" # optional; falls back to "id", then to the row index
},
"budget": 7000, # how many molecules to score; must not exceed the accepted library size
"molecule_filters": {
# optional, and off by default here — see Molecular filters
"boltz_smarts_catalog_filter_level": "recommended"
}
}| Field | Required | What it is | Link |
|---|---|---|---|
target | Yes | The protein and binding pocket to score against. | Target |
library | Yes | The CSV or TSV library to explore, passed by reference. | Library |
budget | Yes | How many molecules to actually score. | Budget |
molecule_filters | No | Which molecules are eligible to be scored. | Molecular filters |
Target (target)
Section titled “Target (target)”The target is the protein you’re exploring against, and takes the same shape as a
library screen: list its entities (protein
chains only), then optionally point the pipeline at the binding pocket with pocket_residues,
reference_ligands, constraints, and bonds. Omit the pocket hints and the pipeline auto-detects
the pocket.
Library (library)
Section titled “Library (library)”The library is a CSV or TSV file passed by reference, not inline. format is csv or tsv,
and source is a url or base64 file source.
- A URL source may use the full file limit: 375 MiB, up to 5,000,000 data records.
- A base64 source is additionally bound by the API’s 50 MiB request-body limit, so prefer a URL for anything large.
The file must be UTF-8 and may contain only the columns you select; column order does not matter.
smiles_column defaults to smiles. id_column names the column carrying your own identifiers.
When omitted, a distinct id column is used if present, and otherwise IDs are generated from
zero-based data-record indexes. Whatever lands there comes back as external_id on the matching
result, so you can correlate results to your input library. IDs are limited to 1,024 UTF-8 bytes.
Budget (budget)
Section titled “Budget (budget)”budget is the number of molecules to score, and it is the unit of work and of billing, not the
library size. It must not exceed the accepted library size, which is what remains after invalid
molecules are dropped and duplicates are merged, so it can be smaller than the row count you
submitted.
Molecular filters (molecule_filters)
Section titled “Molecular filters (molecule_filters)”Filters take the same shape as a library screen, with one difference that matters:
Watching a run
Section titled “Watching a run”Exploration spends real time before it scores anything: fetching and canonicalizing the library,
then building the neighbor graph that selection reads. On a large library that is most of the wall
clock, and every count stays zero throughout, so progress.phase tells you which stage the run is
actually in rather than leaving that stretch indistinguishable from a stalled run.
| Phase | What is happening |
|---|---|
preparing_library | The submitted file is being fetched, validated, and de-duplicated. |
building_graph | The neighbor graph and target inputs are being prepared. |
scoring | Molecules are being selected and scored. Results are arriving. |
Phases only move forward, and a resumed run does not repeat one it has already finished.
import time
# Poll for status and phase. library_size is absent until the library is prepared.while exploration.status not in ("succeeded", "failed", "stopped"): time.sleep(30) exploration = client.small_molecule.explore.retrieve(exploration.id) p = exploration.progress if p.phase == "scoring": print(f"scoring: {p.num_molecules_scored}/{p.total_molecules_to_score} of {p.library_size}") else: print(f"{p.phase}: no molecules scored yet")
# Page through scored molecules; use external_id to correlate back to your library.results = list(client.small_molecule.explore.list_results(exploration.id))results.sort(key=lambda r: r.metrics.binding_confidence, reverse=True)for r in results[:5]: print(f"{r.id} ext={r.external_id} bind={r.metrics.binding_confidence:.2f} {r.smiles}")
# Stop early once you've collected enough; everything already scored stays available.client.small_molecule.explore.stop(exploration.id)
# Resume later from the surrogate as it stood; no library or graph work is repeated.client.small_molecule.explore.resume(exploration.id)export BOLTZ_BASE_URL="https://api.boltz.bio"boltz-api small-molecule:explore retrieve --id "$EXPLORE_ID" # status, phase, and progressboltz-api small-molecule:explore list-results --id "$EXPLORE_ID" # scored molecules so farboltz-api small-molecule:explore stop --id "$EXPLORE_ID" # stop early; partial results stayboltz-api small-molecule:explore resume --id "$EXPLORE_ID" # continue without redoing prep// Poll for status and phase. library_size is absent until the library is prepared.let run = exploration;while (!["succeeded", "failed", "stopped"].includes(run.status)) { await new Promise((r) => setTimeout(r, 30000)); run = await client.smallMolecule.explore.retrieve(run.id); const p = run.progress; console.log( p.phase === "scoring" ? `scoring: ${p.num_molecules_scored}/${p.total_molecules_to_score} of ${p.library_size}` : `${p.phase}: no molecules scored yet`, );}
// Stream scored molecules as they arrive; use external_id to correlate back to your library.for await (const result of client.smallMolecule.explore.listResults(run.id)) { console.log(`${result.id} ext=${result.external_id} bind=${result.metrics.binding_confidence}`);}
// Stop early; everything already scored stays available. Resume picks up without redoing prep.await client.smallMolecule.explore.stop(run.id);await client.smallMolecule.explore.resume(run.id);num_molecules_failed does not consume budget: a molecule that fails to score is replaced by
another selection, so a run still completes at its full budget.
Output format
Section titled “Output format”Results are written as each molecule finishes, so they can be paged while the run is still going and remain retrievable if it fails partway. They use the same shape a library screen produces, so a comparison between the two needs no reconciling.
This is what list_results() streams:
{
"data": [
{
"id": "pres_8f3a2b", # unique result ID
"external_id": "catalog-00417", # the id_column value for this molecule, if any
"created_at": "2026-02-25T13:03:40Z",
"smiles": "CC(=O)OC1=CC=CC=C1C(=O)O", # the scored molecule
"metrics": {
"binding_confidence": 0.94, # 0–1; confidence protein binding occurs; 0.7+ high-confidence
"optimization_score": 0.53, # 0–1; ranks relative binding strength for lead optimization
"structure_confidence": 0.95, # 0–1; confidence in the predicted structure
"iptm": 0.91, # 0–1; interface predicted TM-score
"ptm": 0.92, # 0–1; global predicted TM-score
"complex_plddt": 0.95, # 0–1; pLDDT across the full complex
"complex_iplddt": 0.88 # 0–1; interface pLDDT
},
"artifacts": {
# short-lived presigned download URLs; check url_expires_at and download promptly
"structure": { # predicted bound structure (.cif); may be null until ready
"url": "https://.../structure.cif",
"url_expires_at": "2026-02-25T14:03:40Z"
},
"archive": { # full result archive (.tar.gz): structure, metrics.json, and pae.npz
"url": "https://.../archive.tar.gz",
"url_expires_at": "2026-02-25T14:03:40Z"
}
},
"warnings": [] # optional quality warnings for this result, if any
}
# ...more results on this page
],
"has_more": True, # true if more pages remain
"first_id": "pres_8f3a2b", # ID of the first item; pass as before_id for the previous page
"last_id": "pres_4ab7e0" # ID of the last item; pass as after_id for the next page
}The run object tracks status, phase, and progress. It’s what retrieve() returns:
{
"id": "sm_exp_8f3a2b",
"status": "running", # pending | running | succeeded | failed | stopped
"progress": {
"phase": "building_graph", # preparing_library | building_graph | scoring
"library_size": 134795, # distinct molecules accepted; absent while the library is prepared
"total_molecules_to_score": 7000, # the budget you requested
"num_molecules_scored": 0, # scored and available to download so far
"num_molecules_failed": 0, # terminal failures; these do not consume budget
"rejection_summary": {
"filtered_count": 0, # removed by molecule_filters
"invalid_count": 1, # rejected as invalid input
"duplicate_count": 2800 # rows that collapsed onto a molecule already in the library
},
"latest_result_id": None # ID of the most recently scored result, once scoring starts
},
"error": None, # { code, message } once status is "failed"
"pipeline": "boltzmol",
"pipeline_version": "1.0",
"livemode": True, # false for runs created with a test key
"workspace_id": "ws_3a2b",
"created_at": "2026-02-25T12:00:00Z",
"started_at": "2026-02-25T12:00:05Z",
"completed_at": None, # set when the exploration finishes
"stopped_at": None, # set if you stop the exploration early
"data_deleted_at": None # set once the run's data is deleted
# "input" echoes the request you submitted (null after data deletion)
}Metrics
Section titled “Metrics”Exploration produces the same metrics a library screen does. See
Metrics for the full table; binding_confidence
and optimization_score are the two you will usually rank on.
Status values
Section titled “Status values”| Status | Meaning |
|---|---|
pending | The exploration is queued and has not started yet. |
running | The exploration is preparing the library, building the graph, or scoring. Check progress.phase. |
succeeded | The full budget was scored. |
failed | The exploration encountered an error. Check the error field. |
stopped | The exploration was stopped early. Partial results are available, and it can be resumed. |