Tong Zhang, Mathieu Salzmann, Haoqi Wang
We lifted 1 functions out of this paper's own repositories and ran 1 of them in a sandbox. "Ran" means the function executed on a synthesized input and returned a value. It is not a reproduction of the paper's results.
| Repository | Role | Ran |
|---|---|---|
| haoqiwang/sinder | — | 1 of 1 |
| Function | Status | Where it lives |
|---|---|---|
| SVDLinearAddition | Ran | haoqiwang/sinder/sinder/repair.py code served (permissive licence) · get_code("9fa3073b349c0a06") |
Some links come from the archived Papers with Code dataset (CC BY-SA 4.0): attribution and licence.
Vision Transformer models trained on large-scale datasets, although effective, often exhibit artifacts in the patch token they extract. While such defects can be alleviated by re-training the entire model with additional classification tokens, the underlying reasons for the presence of these tokens remain unclear. In this paper, we conduct a thorough investigation of this phenomenon, combining theoretical analysis with empirical observations. Our findings reveal that these artifacts originate from the pre-trained network itself, specifically stemming from the leading left singular vector of the network's weights. Furthermore, to mitigate these defects, we propose a novel fine-tuning smooth regularization that rectifies structural deficiencies using only a small dataset, thereby avoiding the need for complete re-training. We validate our method on various downstream tasks, including unsupervised segmentation, classification, supervised segmentation, and depth estimation, demonstrating its effectiveness in improving model performance. Codes and checkpoints are available at https://github.com/haoqiwang/sinder.
The same record, over MCP at https://syntology.ai/mcp:
get_harvested_code_for_paper("2407.16826")
get_code_for_paper("2407.16826")
have("2407.16826")
Connect an agent — have() is free.