---
title: Explore small molecule libraries | Boltz API Docs
description: Search a large library against a protein target by scoring only a fraction of it, choosing each molecule from everything scored so far.
---

## When to explore instead of screen

A [library screen](/docs/api/guides/small-molecule-library-screen/index.md) scores every molecule you give it. That is the right tool when the library is small enough to afford end to end.

Exploration is for the opposite case: the library is far larger than the budget you want to spend on it. You submit the whole library and a `budget` — how many molecules to actually score — and the run spends that budget where it is most likely to pay off.

**Each molecule is chosen using every result that came before it**, so the run concentrates on the regions of the library that are scoring well. For the same number of scored molecules, that recovers several times more of the library’s best-scoring compounds than picking the same number blindly.

Scoring around **7% of the library** is where the advantage over random selection is clearest. A smaller budget still works — it simply leaves the run less scored history to choose from.

## Run

Start with `start()`, then poll `retrieve()` for status and phase and page `list_results()` as molecules are scored. Exploration has no one-call `run()` helper: a run is long-lived and resumable, so you drive it yourself.

- [Python](#tab-panel-0-0)
- [CLI](#tab-panel-0-1)
- [TypeScript](#tab-panel-0-2)

```
import os
from boltz_api import Boltz


client = Boltz(
    base_url="https://api.boltz.bio",
    api_key=os.environ["BOLTZ_API_KEY"])


target = {
    "entities": [{"type": "protein", "value": "MKTIIALSYIFCLVFA", "chain_ids": ["A"]}],
}  # see Input format for pocket_residues, reference_ligands, constraints, …


exploration = client.small_molecule.explore.start(
    target=target,
    library={
        "format": "csv",
        "source": {"type": "url", "url": "https://example.com/library.csv"},
        "smiles_column": "smiles",
    },
    budget=7000,
)
print(exploration.id)  # sm_exp_...
```

Write your request body to `small-molecule-explore.json` (see [Input format](#input-format)), then:

Terminal window

```
export BOLTZ_BASE_URL="https://api.boltz.bio"
EXPLORE_ID=$(
  boltz-api --format raw small-molecule:explore start \
    --input @json://./small-molecule-explore.json | jq -r '.id'
)


boltz-api small-molecule:explore retrieve --id "$EXPLORE_ID"
```

```
import Boltz from "boltz-api";


const client = new Boltz({
  baseURL: "https://api.boltz.bio",
  apiKey: process.env["BOLTZ_API_KEY"] });


const target = {
  entities: [{ type: "protein", value: "MKTIIALSYIFCLVFA", chain_ids: ["A"] }],
}; // see Input format for pocket_residues, reference_ligands, constraints, …


const exploration = await client.smallMolecule.explore.start({
  target,
  library: {
    format: "csv",
    source: { type: "url", url: "https://example.com/library.csv" },
    smiles_column: "smiles",
  },
  budget: 7000,
});
```

## Input format

An exploration takes a **target** to score against, a file-backed **library** to explore, and a **budget** capping how much of it gets scored.

Copy

```
{
  "target": {
    "entities": [{ "type": "protein", "value": "MKTIIALSYIFCLVFA", "chain_ids": ["A"] }]
    # see Target for pocket_residues, reference_ligands, constraints, …
  },
  "library": {
    "format": "csv", # csv | tsv
    "source": { "type": "url", "url": "https://example.com/library.csv" },
    "smiles_column": "smiles", # defaults to "smiles"
    "id_column": "catalog_id" # optional; falls back to "id", then to the row index
  },
  "budget": 7000, # how many molecules to score; must not exceed the accepted library size
  "molecule_filters": {
    # optional, and off by default here — see Molecular filters
    "boltz_smarts_catalog_filter_level": "recommended"
  }
}
```

| Field              | Required | What it is                                              | Link                                                     |
| ------------------ | -------- | ------------------------------------------------------- | -------------------------------------------------------- |
| `target`           | Yes      | The protein and binding pocket to score against.        | [Target](#target-target)                                 |
| `library`          | Yes      | The CSV or TSV library to explore, passed by reference. | [Library](#library-library)                              |
| `budget`           | Yes      | How many molecules to actually score.                   | [Budget](#budget-budget)                                 |
| `molecule_filters` | No       | Which molecules are eligible to be scored.              | [Molecular filters](#molecular-filters-molecule_filters) |

### Target (`target`)

The target is the protein you’re exploring against, and takes the same shape as a [library screen](/docs/api/guides/small-molecule-library-screen#target-target/index.md): list its **entities** (protein chains only), then optionally point the pipeline at the binding pocket with `pocket_residues`, `reference_ligands`, `constraints`, and `bonds`. Omit the pocket hints and the pipeline auto-detects the pocket.

### Library (`library`)

The library is a CSV or TSV file passed **by reference**, not inline. `format` is `csv` or `tsv`, and `source` is a `url` or `base64` file source.

- A **URL** source may use the full file limit: 375 MiB, up to 5,000,000 data records.
- A **base64** source is additionally bound by the API’s 50 MiB request-body limit, so prefer a URL for anything large.

The file must be UTF-8 and may contain only the columns you select; column order does not matter. `smiles_column` defaults to `smiles`. `id_column` names the column carrying your own identifiers. When omitted, a distinct `id` column is used if present, and otherwise IDs are generated from zero-based data-record indexes. Whatever lands there comes back as `external_id` on the matching result, so you can correlate results to your input library. IDs are limited to 1,024 UTF-8 bytes.

A URL source is fetched once, asynchronously, after the run is accepted, and re-fetched from the start if that fetch is retried. The URL must stay valid well past the moment you submit it. A presigned URL that expires in minutes will fail the run.

### Budget (`budget`)

`budget` is the number of molecules to score, and it is the unit of work and of billing, not the library size. It must not exceed the **accepted** library size, which is what remains after invalid molecules are dropped and duplicates are merged, so it can be smaller than the row count you submitted.

### Molecular filters (`molecule_filters`)

Filters take the same shape as a [library screen](/docs/api/guides/small-molecule-library-screen#molecular-filters-molecule_filters/index.md), with one difference that matters:

Built-in SMARTS filtering is **disabled** unless you set `molecule_filters.boltz_smarts_catalog_filter_level` explicitly. A screen defaults it to `recommended`; an exploration does not, because you are exploring a library you curated. Set it to `recommended` or higher if you want Boltz’s structural-alert filtering applied.

## Watching a run

Exploration spends real time before it scores anything: fetching and canonicalizing the library, then building the neighbor graph that selection reads. On a large library that is most of the wall clock, and every count stays zero throughout, so `progress.phase` tells you which stage the run is actually in rather than leaving that stretch indistinguishable from a stalled run.

| Phase               | What is happening                                                  |
| ------------------- | ------------------------------------------------------------------ |
| `preparing_library` | The submitted file is being fetched, validated, and de-duplicated. |
| `building_graph`    | The neighbor graph and target inputs are being prepared.           |
| `scoring`           | Molecules are being selected and scored. Results are arriving.     |

Phases only move forward, and a resumed run does not repeat one it has already finished.

- [Python](#tab-panel-1-0)
- [CLI](#tab-panel-1-1)
- [TypeScript](#tab-panel-1-2)

```
import time


# Poll for status and phase. library_size is absent until the library is prepared.
while exploration.status not in ("succeeded", "failed", "stopped"):
    time.sleep(30)
    exploration = client.small_molecule.explore.retrieve(exploration.id)
    p = exploration.progress
    if p.phase == "scoring":
        print(f"scoring: {p.num_molecules_scored}/{p.total_molecules_to_score} of {p.library_size}")
    else:
        print(f"{p.phase}: no molecules scored yet")


# Page through scored molecules; use external_id to correlate back to your library.
results = list(client.small_molecule.explore.list_results(exploration.id))
results.sort(key=lambda r: r.metrics.binding_confidence, reverse=True)
for r in results[:5]:
    print(f"{r.id}  ext={r.external_id}  bind={r.metrics.binding_confidence:.2f}  {r.smiles}")


# Stop early once you've collected enough; everything already scored stays available.
client.small_molecule.explore.stop(exploration.id)


# Resume later from the surrogate as it stood; no library or graph work is repeated.
client.small_molecule.explore.resume(exploration.id)
```

Terminal window

```
export BOLTZ_BASE_URL="https://api.boltz.bio"
boltz-api small-molecule:explore retrieve --id "$EXPLORE_ID"      # status, phase, and progress
boltz-api small-molecule:explore list-results --id "$EXPLORE_ID"  # scored molecules so far
boltz-api small-molecule:explore stop --id "$EXPLORE_ID"          # stop early; partial results stay
boltz-api small-molecule:explore resume --id "$EXPLORE_ID"        # continue without redoing prep
```

```
// Poll for status and phase. library_size is absent until the library is prepared.
let run = exploration;
while (!["succeeded", "failed", "stopped"].includes(run.status)) {
  await new Promise((r) => setTimeout(r, 30000));
  run = await client.smallMolecule.explore.retrieve(run.id);
  const p = run.progress;
  console.log(
    p.phase === "scoring"
      ? `scoring: ${p.num_molecules_scored}/${p.total_molecules_to_score} of ${p.library_size}`
      : `${p.phase}: no molecules scored yet`,
  );
}


// Stream scored molecules as they arrive; use external_id to correlate back to your library.
for await (const result of client.smallMolecule.explore.listResults(run.id)) {
  console.log(`${result.id}  ext=${result.external_id}  bind=${result.metrics.binding_confidence}`);
}


// Stop early; everything already scored stays available. Resume picks up without redoing prep.
await client.smallMolecule.explore.stop(run.id);
await client.smallMolecule.explore.resume(run.id);
```

`num_molecules_failed` does not consume budget: a molecule that fails to score is replaced by another selection, so a run still completes at its full budget.

## Output format

Results are written as each molecule finishes, so they can be paged while the run is still going and remain retrievable if it fails partway. They use the same shape a library screen produces, so a comparison between the two needs no reconciling.

This is what `list_results()` streams:

Copy

```
{
  "data": [
    {
      "id": "pres_8f3a2b", # unique result ID
      "external_id": "catalog-00417", # the id_column value for this molecule, if any
      "created_at": "2026-02-25T13:03:40Z",
      "smiles": "CC(=O)OC1=CC=CC=C1C(=O)O", # the scored molecule
      "metrics": {
        "binding_confidence": 0.94, # 0–1; confidence protein binding occurs; 0.7+ high-confidence
        "optimization_score": 0.53, # 0–1; ranks relative binding strength for lead optimization
        "structure_confidence": 0.95, # 0–1; confidence in the predicted structure
        "iptm": 0.91, # 0–1; interface predicted TM-score
        "ptm": 0.92, # 0–1; global predicted TM-score
        "complex_plddt": 0.95, # 0–1; pLDDT across the full complex
        "complex_iplddt": 0.88 # 0–1; interface pLDDT
      },
      "artifacts": {
        # short-lived presigned download URLs; check url_expires_at and download promptly
        "structure": { # predicted bound structure (.cif); may be null until ready
          "url": "https://.../structure.cif",
          "url_expires_at": "2026-02-25T14:03:40Z"
        },
        "archive": { # full result archive (.tar.gz): structure, metrics.json, and pae.npz
          "url": "https://.../archive.tar.gz",
          "url_expires_at": "2026-02-25T14:03:40Z"
        }
      },
      "warnings": [] # optional quality warnings for this result, if any
    }
    # ...more results on this page
  ],
  "has_more": True, # true if more pages remain
  "first_id": "pres_8f3a2b", # ID of the first item; pass as before_id for the previous page
  "last_id": "pres_4ab7e0" # ID of the last item; pass as after_id for the next page
}
```

The run object tracks status, phase, and progress. It’s what `retrieve()` returns:

Copy

```
{
  "id": "sm_exp_8f3a2b",
  "status": "running", # pending | running | succeeded | failed | stopped
  "progress": {
    "phase": "building_graph", # preparing_library | building_graph | scoring
    "library_size": 134795, # distinct molecules accepted; absent while the library is prepared
    "total_molecules_to_score": 7000, # the budget you requested
    "num_molecules_scored": 0, # scored and available to download so far
    "num_molecules_failed": 0, # terminal failures; these do not consume budget
    "rejection_summary": {
      "filtered_count": 0, # removed by molecule_filters
      "invalid_count": 1, # rejected as invalid input
      "duplicate_count": 2800 # rows that collapsed onto a molecule already in the library
    },
    "latest_result_id": None # ID of the most recently scored result, once scoring starts
  },
  "error": None, # { code, message } once status is "failed"
  "pipeline": "boltzmol",
  "pipeline_version": "1.0",
  "livemode": True, # false for runs created with a test key
  "workspace_id": "ws_3a2b",
  "created_at": "2026-02-25T12:00:00Z",
  "started_at": "2026-02-25T12:00:05Z",
  "completed_at": None, # set when the exploration finishes
  "stopped_at": None, # set if you stop the exploration early
  "data_deleted_at": None # set once the run's data is deleted
  # "input" echoes the request you submitted (null after data deletion)
}
```

Download URLs expire. Check `url_expires_at` and download promptly. Once a URL expires, a new one can be generated until the data is deleted. By default data is retained for 7 days; see [Data Retention](/docs/api/guides/data-retention/index.md).

Metric values and molecules shown in this guide are illustrative.

## Metrics

Exploration produces the same metrics a library screen does. See [Metrics](/docs/api/guides/small-molecule-library-screen#metrics/index.md) for the full table; `binding_confidence` and `optimization_score` are the two you will usually rank on.

## Status values

| Status      | Meaning                                                                                           |
| ----------- | ------------------------------------------------------------------------------------------------- |
| `pending`   | The exploration is queued and has not started yet.                                                |
| `running`   | The exploration is preparing the library, building the graph, or scoring. Check `progress.phase`. |
| `succeeded` | The full budget was scored.                                                                       |
| `failed`    | The exploration encountered an error. Check the `error` field.                                    |
| `stopped`   | The exploration was stopped early. Partial results are available, and it can be resumed.          |
