✳ REVERSE IMAGE SEARCH FOR PROTEINS
A picture.
A world of
structure.
Find the proteins behind the picture.
One image. Over 800,000 molecular views.

PDB 1UBQ · 76 residues
01 / THE SEARCH
What are you looking at?
Upload a ribbon diagram or a crop from a paper.
Drop a protein image here
or
PNG, JPG OR WEBP · UP TO 10 MBUse a close crop of one protein, with minimal labels.
Your image is not saved.
Use in Codex ↗Example proteins
TRY AN EXAMPLE ↓↳ Training examples: try the search, not a test of unseen proteins.
YOUR SEARCH / STRUCTURAL CANDIDATES
A closer look.
| RANK | STRUCTURE | PROTEIN / ORGANISM | SIMILARITY | Actions |
|---|
Previews show RCSB biological assembly 1; the 3D viewer shows deposited coordinates. These can differ from the object in your image. Visual similarity alone does not establish identity, homology or function.
02 / FROM COORDINATES TO CONNECTIONS
Many views.
One way to find them.
From a 2D image to candidate 3D structures.


224 × 224 pixels
14 × 14 pixels each
Actual image pixels; the model learns the features.
Start with known structures.
We take experimentally determined protein structures from the Protein Data Bank and prepare asymmetric units and biological assemblies for rendering.
03 / A COLLECTION OF PERSPECTIVES
Built to look
beyond one view.
One image, compared with a collection of molecular views.
Try your own image ↗Model & scoring
DINOv2-L with a trained retrieval head and four fine-tuned transformer blocks. Each PDB is scored by its five best views, using a rule selected on validation data.
04 / PUTTING IT TO THE TEST
Learning to see
the structure.
Finding the same PDB from a different rendered view.
Loading benchmark results…
Exact-PDB retrieval across all 38,833 indexed PDBs, with the query views excluded. These are rendered-image results; paper figures are evaluated below.
Smaller-gallery comparison & protocol
Every method uses the same queries, reference images and mean-of-five scoring rule.
| Method | First result | Top 10 |
|---|
Test PDBs use the existing sequence-component split at 40% sequence identity. Silhouette scaling is fitted on reference images only. Equal scores are ordered by PDB ID. Existing checkpoints and the validation-selected scoring rule were fixed before this rerun; no retraining or test-set tuning.
Download measurements & provenance ↓5,174 held-out query views; 831,225 reference images after exclusions. Fixed model and mean-of-five scoring. Full results & provenance ↓
05 / IN THE WILD
From the paper.
Back to the protein.
Retrieval from real paper figures, using the live collection.
EXPLORATORY BENCHMARK · 20 SEPTEMBER 2026
Candidate recall at K
A hit means at least one listed source PDB appears in the first K results.
Loading frozen benchmark measurements…
Exploratory results. 407 crops have unconfirmed panel-to-PDB mappings, so candidate recall can overestimate exact matches. The all-crop rate counts 362 missing-reference cases as misses.
Results & protocol
| K | Hits | Indexed source · n=117 | All crops · n=479 |
|---|
DINOv2-L ft_b searched 836,399 reference views across 38,833 PDB entries. Approved crops were standardized to 512-pixel square masters, then resized to 224 × 224 for inference. Entries are ranked by the mean similarity of their five best views. The checkpoint and scoring rule were fixed; no training or tuning on these queries.
This is a development diagnostic, not a blind test of unseen proteins. Source PDBs can belong to training, validation or test reference splits; multiple crops can come from the same paper. The smaller confirmed, single-source subset has 16 indexed crops: 4/16 hits at top 10 and 7/16 at top 50. No structural or biological-relatedness grading is applied.
Download results & provenance ↓FIVE TOP-10 CANDIDATE HITS
A source PDB found.
Paper crop → retrieved reference view.
THREE FAILURE CASES
Where retrieval falls short.
Two ranking misses and one missing reference. The right image is the first result.
Selected examples, not a representative sample. A different PDB can still be structurally related.