SYNTOLOGY The Modern Ontology for AI Home Explorer Atlas Code Methodology Attribution About Developers
Attribution and upstream licences

Some of what the graph shows you came from somewhere else.

Every edge in the Syntology graph carries a provenance string naming where it came from, and one of those strings points at a body of work this project did not produce: the archived Papers with Code dataset. That data is published under a licence which asks, in return for the freedom to use it, that the source be named and the licence carried forward. This page is where that is done. It is linked from every surface that serves the material.

Papers with Code (archived)

CC BY-SA 4.0 · snapshot of 28 July 2025

Syntology's paper-to-code links are derived in part from the Papers with Code dataset, created by the Papers with Code project (paperswithcode.com, operated by Meta AI until the site was shut down in July 2025) and preserved after the shutdown by the pwc-archive community mirror on Hugging Face. The specific material used is the dataset pwc-archive/links-between-paper-and-code — 300,161 rows of (paper, repository) associations frozen at the last public snapshot of the data, retrieved 28 July 2025.

That material is licensed under the Creative Commons Attribution-ShareAlike 4.0 International Public License (CC BY-SA 4.0). The full text of that licence is at creativecommons.org/licenses/by-sa/4.0/legalcode; a plain-language summary is at creativecommons.org/licenses/by-sa/4.0/. Syntology is not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror; naming them here is attribution, not endorsement.

What we changed

The material has been modified. It was not reproduced as published. Rows were filtered to papers this graph already held, reshaped into typed graph edges rather than table rows, joined against code links this project extracted independently from paper full text, and re-keyed onto our own node identifiers. The upstream fields is_official, framework, mentioned_in_paper and mentioned_in_github are carried on the edge under a pwc_ prefix, deliberately kept in their own namespace so that upstream assertions are never blended with, or allowed to overwrite, what our own extraction found. We are not aware of prior modifications by the pwc-archive mirror beyond the format conversion its dataset card describes.

Where it appears

Anywhere the graph shows you that a paper has code. Concretely: the repository list and official / mentioned in paper / mentioned in github labels in the Atlas detail panel; the code-coverage colouring on the Atlas map; the repositories returned by the graph API and the MCP tools; and the paper-to-code counts quoted on our own pages. Each such edge carries the provenance string external:paperswithcode_snapshot_2025-07-28, so you can always ask the graph itself which of the two sources an answer came from.

How much, measured

As of 13 September 2026: 125,485 of 159,573 live IMPLEMENTED_BY edges (78.6%) carry the Papers with Code provenance string, and 128,842 carry at least one pwc_ field — the extra 3,357 are links our own extraction found first and which the archive then corroborated, where our provenance is preserved untouched. The remaining 34,085 edges come from Syntology's own extraction of GitHub URLs out of paper full text (deterministic:regex_extraction). These figures are re-derivable: groundwork/49_attribution_compliance/audit_pwc_surfaces.py.

Also held from the same archive, under the same licence

The archived leaderboards — 155,456 rows over 21,951 tables. Fetched and measured, not loaded into the graph and not served. Named here anyway, so that this page describes what we hold and not only what we show.

The datasets and methods catalogues

Papers with Code's own benchmark and method catalogues, used to decide whether a name our extractor produced refers to a benchmark the field actually shares. A small number of Dataset nodes carry external:paperswithcode_catalogue as their provenance for this reason.

Other upstream sources

named for transparency; each used under its own terms

This page exists for the Papers with Code archive, whose licence carries an explicit attribution condition. The sources below are named so that a reader can see where material in the graph comes from, and each is used under the terms its own publisher sets.

arXivPreprint metadata and full text, from which almost everything downstream is extracted.
Semantic Scholar (S2)Identity resolution and the majority of resolved citation edges.
OpenReviewPeer reviews, stored one node per reviewer with rating and confidence kept raw.
GitHubRepository metadata and source. Source files carry their own repository's licence, which is why the code index is not redistributed.
OpenAlexSupplementary bibliographic records.
CrossrefDOI resolution and venue backfill.
Hugging FaceHosts the pwc-archive mirror above, and supplies the task taxonomy behind Task nodes.
GROBIDThe open-source PDF-to-TEI extractor every paper passes through.
If we have this wrong

If you maintain one of the sources above and this page misnames you, under-credits you, or omits something your licence asks for, write to media@syntology.ai and it will be corrected. Attribution that is wrong is worse than attribution that is late, and we would rather hear it from you than not hear it.

The Modern Ontology for AITermsPrivacyMethodologyEthosmedia@syntology.aiLast reviewed 13 September 2026