Loading the SOTA2 catalog…
SOTA2 Research · papers
Find papers, implementations, and the benchmark evidence behind state-of-the-art AI systems.
| CLIP-RL: Surgical Scene Segmentation Using Contrastive Language-Vision Pretraining & Reinforcement Learning |
|---|
| Fatmaelzahraa Ali Ahmed, Muhammad Arsalan, Abdulaziz Al-Ali |
| 2025 |
| arxiv 2507.04317 |
| Predicting the spatial distribution and demographics of commercial swine farms in the United States | Felipe E. Sanchez, Thomas A. Lake, Jason A. Galvis | 2025 | arxiv 2511.00132 |
|---|
| TrueCity: Real and Simulated Urban Data for Cross-Domain 3D Scene Understanding | Duc Nguyen, Yan-Ling Lai, Qilin Zhang | 2025 | arxiv 2511.07007 |
|---|
| DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation | Yan Gong, Jianli Lu, Yongsheng Gao | 2025 | arxiv 2511.13047 |
|---|
| UNIV: Unified Foundation Model for Infrared and Visible Modalities | Fangyuan Mao, Shuo Wang, Jilin Mei | 2025 | arxiv 2509.15642 |
|---|
| EGSA-PT:Edge-Guided Spatial Attention with Progressive Training for Monocular Depth Estimation and Segmentation of Transparent Objects | Gbenga Omotara, Ramy Farag, Seyed Mohamad Ali Tousi | 2025 | arxiv 2511.14970 |
|---|
| AdaptFly: Prompt-Guided Adaptation of Foundation Models for Low-Altitude UAV Networks | Jiao Chen, Haoyi Wang, Jianhua Tang | 2025 | arxiv 2511.11720 |
|---|
| Phys4DGen: Physics-Compliant 4D Generation with Multi-Material Composition Perception | Jiajing Lin, Zhenzhong Wang, Dejun Xu | 2024 | arxiv 2411.16800 |
|---|
| Temporal-Guided Visual Foundation Models for Event-Based Vision | Ruihao Xia, Junhong Cai, Luziwei Leng | 2025 | arxiv 2511.06238 |
|---|
| MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning | Thanh-Dat Truong, Christophe Bobda, Nitin Agarwal | 2025 | arxiv 2508.10133 |
|---|