| ★ | Visual7W: Grounded Question Answering in Images | Xavi Giró-i-Nieto @xavigiro-i-nieto | 12 | — | |
| 16 | Do AI Systems Have World Models? Probing Reasoning, Forecasting, and Generalization | Simons Institute for the Theory of Computing @simonsinstitute | 10.0K | Estados Unidos | 87 |
| 9 | From "Umwelt" to "World" models | Simons Institute for the Theory of Computing @simonsinstitute | 5.6K | Estados Unidos | 88 |
| 17 | 350 - Efficient Image Retrieval with Vision Transformer (ViT) and FAISS | DigitalSreeni @digitalsreeni | 2.6K | Estados Unidos | 87 |
| 19 | SketchGPT: A Sketch-based Multimodal Interface for Application-Agnostic LLM Interaction | | 138 | Estados Unidos | 87 |
| 18 | Beyond Symbols: Motion Perception Cues Enhance Dual-Task Performance w Wearable Directional Guidance | | 129 | Estados Unidos | 87 |
| 1 | Deep Image Retrieval: Learning global representations for image search | Xavi Giró-i-Nieto @xavigiro-i-nieto | 104 | — | 91 |
| 2 | Visual Question Answering (VQA) | | 72 | — | 89 |
| 8 | Hybrid Handcrafted & Deep Multi-Angle Features For Rotation-Invariant Texture-Based Image Retrieval | Computer Science & IT Conference Proceedings @computerscienceitconferenc7375 | 61 | Australia | 88 |
| 13 | NeuroSync: Intent-Aware Code-Based Problem Solving via Direct LLM Understanding Modification | | 57 | Estados Unidos | 87 |
| 4 | Uncover the Secret to RanKing Your Videos on Google! | Nex Gen AI 2.0 - ChatGPT @nexgenai2.0 | 54 | Estados Unidos | 88 |
| 12 | Using Vision-Language Models to Evaluate and Secure Mixed Reality Experiences with Maria Gorlatova | | 51 | Estados Unidos | 87 |
| 11 | GenTune: Toward Traceable Prompts to Improve Controllability of Image Refinement in Environment D... | | 36 | Estados Unidos | 87 |
| 7 | Understanding Visual Attention and Checking Behavior during Mobile Text Input | | 34 | Estados Unidos | 88 |
| 20 | Projecte d'Impacte: Assistència intel·ligent a la descripció contextualitzada de fotos per notícies | | 25 | — | 87 |
| 15 | Morae: Proactively Pausing UI Agents for User Choices | | 23 | Estados Unidos | 87 |
| 14 | "Debating with Chat Gpt: Exploring AI's Argumentative Skills" | AI Explorers @aiexplorers-iw2gz | 15 | Estados Unidos | 87 |
| 6 | Multimodal Transformers Explained: Unifying All Data Types | THE FACT FACTORY @thefactfactoryf | 14 | — | 88 |
| 5 | ViseGPT: Towards Better Alignment of LLM-generated Data Wrangling Scripts and User Prompts | | 13 | Estados Unidos | 88 |
| 10 | Analysis of short videos on TikTok for learning Portuguese as a foreign language | RevistaComunicar @revistacomunicar | 9 | España | 87 |
| 3 | Charla de José M. Saavedra: "Sketch-based Understanding in Computer Vision" | | 1 | Chile | 89 |