Data sources

Where Litlas gets its papers — and what it does not cover

The bibliographic sources behind the Litlas index, the size of the corpus, how often it is rebuilt, the gaps we know about, and how to verify all of it yourself.

Litlas does not crawl the web. Every paper in the index arrives from a named bibliographic source, and every search result carries the slug of the source it came from. This page lists those sources, states the size of the corpus, explains how often it is rebuilt, and — the part that matters most if you are deciding whether to recommend Litlas — spells out what it does not cover.

Measured on 2026-08-21 against the live API. The release being served was built on 2026-08-11.

Evaluating Litlas for a library research guide? What a student needs in order to use it, what it costs, and a description you can paste into your guide are on a page of their own.

Where the papers come from

The backend declares 16 connectors, and that is the number the overview endpoint reports as “sources”. Be careful with it: it is the length of a list in the source code, not a measurement of the index. Measured against the live index, 10 of those connectors actually hold papers, and only 2 of them carry the corpus.

“Above the count cap” means the search API stops counting at 50,000, so no exact per-source figure can be obtained from outside. It does not mean the source is complete. Rows showing a small number are exact — and they are samples, not corpora: a connector that contributes 1,000 records per rebuild is a taste of that database, not a mirror of it.

The rows marked “enrichment only” are working as designed: they contribute identifiers, citation edges, or open-access links to papers that came from somewhere else, so they hold no papers of their own. The remaining empty rows are connectors that are declared but are contributing nothing to the release currently served.

One source is used but not listed above: Semantic Scholar is queried out of band to raise citation counts and add citation edges. It contributes no records to the corpus, so it has no slug you can filter on.

How big it is

Both figures come from the same rebuild, so they describe the same corpus. The paper count is live — the API reports it on every request. The citation-link count is not exposed by any endpoint; it is the edge count recorded by the build that produced the release now being served, which is why it is dated rather than live.

The home page rounds the citation figure down to 390,000,000. That is deliberate: the rounded claim is checked against the measured value, so it can never overstate it.

How often it updates

What Litlas does not cover

This is the section we would want to read first if we were evaluating someone else’s tool, so it is deliberately blunt. Each item below is a limitation we can point at in our own code, not a hedge.

How to check this yourself

None of the above needs to be taken on trust. The endpoints are public and answer without a login, so you can reproduce the table on this page in a few minutes.

One number on this page cannot be checked this way: the citation-link count is not served by any endpoint, so it can only be quoted from the build record.

Data sources | Litlas