CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding
CodeGraph maps 167 million source files into a Wikidata-grounded graph of code concepts.
The paper describes a pipeline that uses a code-specialized LLM to extract concepts such as algorithms, paradigms, design patterns, and application domains from source files. It links those concepts to Wikidata with a three-stage process, then rolls up parent categories into the graph. Applied to the Stack-Edu corpus, CodeGraph contains about 158 million nodes and roughly 1 billion typed edges across 14 programming languages. The authors say it is the first known large-scale open-taxonomy knowledge graph for source code. HF Daily Papers' note
The paper describes a pipeline that uses a code-specialized LLM to extract concepts such as algorithms, paradigms, design patterns, and application domains from source files. It links those concepts to Wikidata with a three-stage process, then rolls up parent categories into the graph. Applied to the Stack-Edu corpus, CodeGraph contains about 158 million nodes and roughly 1 billion typed edges across 14 programming languages. The authors say it is the first known large-scale open-taxonomy knowledge graph for source code. HF Daily Papers' note
score 4