Find research datasets, and know what you may do with them

ORBIT searches about 73 million research datasets through the DOI registry that Zenodo, figshare, Dryad, PANGAEA and most institutional archives all mint their identifiers through. Every result says whether you are allowed to reuse it, links out to the repository by DOI, and can be exported to a reference manager. The whole page is free and takes no credit.

Being able to download it is not the same as being allowed to use it

This is the distinction the tool exists for. A dataset search that told you only where the files are would be answering the easy half of the question. ORBIT shows every result in one of three states, and the third is the one that matters most.

Openly licensed means the record states a licence permitting reuse — CC0, CC BY, MIT and the like. Licence restricted means a licence is stated and carries limits worth reading, such as non-commercial. Licence not stated means the record names no licence at all, which is common and is not permission: in most jurisdictions the default is all rights reserved. ORBIT will not present a blank field as a green light.

Where the datasets come from

The DOI registry research repositories mint through, read via its own API. That is why one search reaches Zenodo, figshare, Dryad, PANGAEA, Dataverse and a long tail of institutional archives at once, with no arrangement needed with any of them.

ORBIT does not crawl repository pages, and it does not host, mirror or redistribute anybody's data. Every result links out to the source by DOI — the identifier that still resolves after a repository reorganises its site, which for a dataset citation is the whole point.

Two things ORBIT cleans up, and says it is cleaning up

Repositories mint a separate DOI for every version of a dataset and index the parent record beside them, so raw results repeat themselves — nine of twenty-five rows in a live measurement. ORBIT collapses versions into one row and tells you how many exist, grouping them by the identifiers repositories publish rather than by matching titles, because two different datasets can share a name.

Metadata quality varies enormously. Some records carry a full abstract, named creators and a licence; some carry a title and nothing else. ORBIT marks the sparse ones and names what is missing rather than presenting every hit as equally trustworthy. A thin record can describe excellent data — you simply cannot tell, and you should know that before you spend an afternoon on it.

Questions about dataset discovery

Is this really free? What is the catch?

There is no catch and no paid tier behind it. Searching, filtering, following the links and exporting are all free and take no credit, and unlike the literature tool there is no paid analysis pass on this page at all. The registry ORBIT reads is free to query, so there is nothing here worth charging you to recover.

If I can download a dataset, can I use it in my paper?

Not necessarily, and this is the single most important thing on this page. Being able to download data is not the same as having permission to reuse it. ORBIT shows every result in one of three states: openly licensed, licence restricted, or licence not stated. That third state is common, and it does not mean the data is free to use — it means nobody has said. In most jurisdictions the default is all rights reserved, so a dataset with no stated licence is one to ask the repository about before you build on it.

Where do the datasets come from?

The DOI registry that research repositories mint their identifiers through, which is why one search covers Zenodo, figshare, Dryad, PANGAEA, Dataverse and a great many institutional archives at once, without ORBIT needing an arrangement with any of them. About 73 million dataset DOIs. Read through the registry's own API — ORBIT does not crawl repository pages, and it does not host, mirror or redistribute anybody's data. Every result links out to the source.

Why do some results say the record is sparse?

Because it is, and hiding that would be dishonest in the direction that costs you time. Metadata quality varies enormously between repositories: some records carry a full abstract, named creators and a licence, and some carry a title and nothing else. ORBIT marks the sparse ones and says exactly what is missing — no description, no named creator, no year — rather than presenting every hit as equally trustworthy. A thinly described dataset can still be excellent; you just cannot tell from the record, and you should know that before you spend an afternoon on it.

I searched and saw the same dataset several times. Does ORBIT fix that?

Yes. Repositories mint a separate DOI for every version of a dataset, and both the versions and the parent record are indexed, so raw results repeat themselves — in one live measurement, nine of twenty-five rows were duplicates. ORBIT collapses them into a single row and tells you how many versions exist, so you can still go and find an older one if a paper cited it. It groups them using the identifiers the repositories publish, never by matching titles, because two genuinely different datasets can share a name and merging those would hide data from you.

Can I get the results into my reference manager?

Yes, in four formats and all free: RIS, BibTeX, CSV and a plain Markdown list. The RIS and BibTeX exports mark each entry as a dataset rather than as a journal article — the wrong type here is not cosmetic, because it silently produces a wrong citation in your bibliography. The exports also carry the version and the licence state, so what you read back in six months says the same thing the screen said today.

Does ORBIT keep a record of what I search for?

No. Results are cached so that two people searching the same topic share one answer, but the cache holds no user id and cannot be traced back to anyone — the key is a hash of the search itself. There is deliberately no search history. What a researcher is looking for is a map of work they have not published yet.