Gauvain Bourgne, Jean-Gabriel Ganascia, Evangelia Zve
We have not lifted any functions out of this paper's repositories yet, so there is nothing we have run. The repositories linked to it are listed below.
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Some documents that embedding-based topic models initially classify as noise later become founding members of emerging topics. At publication time, however, they appear as scattered points in embedding space and are difficult to distinguish from ordinary noise without the benefit of hindsight. We study whether such anticipatory outliers can be predicted prospectively, using only information available when a document first appears. We derive labels from the subsequent trajectories of outlier documents, distinguishing those that anticipate new topics from those that reinforce existing topics or remain isolated, and estimate label confidence through agreement across multiple embedding models. On two French news corpora, anticipatory outliers prove predictable at publication time. Under cross-validation, $F_1$ rises from about 0.77 over the full eligible population to above 0.90 on high-consensus subsets, and remains at 0.76-0.80 under a strictly chronological evaluation. Predictive performance is driven mainly by geometric features capturing each outlier's position in embedding space.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2609.29183")
get_code_for_paper("2609.29183")
have("2609.29183")
The run record, dated, one paper per request, free:
curl https://syntology.ai/api/ran/2609.29183.json
A badge for a README (the split and the date, never a ratio):
[](https://syntology.ai/paper/2609.29183)
Connect an agent — have() is free.