← 전체 목록

2026-09-16 AI 리서치 브리핑

최신 VLM, sLLM, on-device AI 논문과 연구 블로그를 한눈에 정리합니다. 중복 기사 방지를 위해 URL 기준으로 추적합니다.

총 11건 요약 자동 생성

VLM 업데이트

멀티모달 비전-언어 모델의 최신 논문과 리더보드 변화

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

Paper Hugging Face Papers

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Anatomical Grounding and Leakage-Aware Multimodal Contrastive Learning for Alzheimer's Disease Classification from Structural MRI

Paper arXiv cs.CV (recent)

Deep networks trained on structural MRI for Alzheimer's disease (AD) staging often reach reasonable accuracy while attending to anatomically irrelevant regions, and multimodal models that add clinical tables frequently rely on variables that were used to assign the diagnostic label in the first place. We study both issues with a deliberately lightweight slice-based encoder (ResNet18 with a one-layer Transformer over slices) on 1,075 baseline T1-weighted scans from ADNI-1. First, we use FastSurfer segmentations as an anatomical reference: YOLOv8 models trained on segmentation-derived labels localize Alzheimer-relevant structures with mAP_50 above 0.96, and a Grad-CAM comparison shows that the image-only classifier frequently attends to the skull, orbits and background. Second, we adapt a CLIP-style image - tabular contrastive framework and organize ADNIMERGE variables along a label-leakage spectrum. Fusion with cognitive scores yields 87.3% three-way accuracy, which we treat as a leakage-driven upper bound rather than an imaging result; fusion with regional volumes yields 73.0%. We observe that the choice of contrastive target changes what the image encoder learns: on MCI vs. CN, the image-only head reaches 52.4% when the encoder is aligned to cognitive scores and 73.8\% when aligned to volumes, although no tabular input is used at inference. Third, restricting the input to a per-subject crop of the medial temporal lobe raises image-only three-way accuracy from 58.7% to 65.1%. All results come from single runs on a small balanced test set, and we report confidence intervals and the protocol differences that prevent direct comparison with published numbers.

원문 보기

On-Device AI

디바이스 내 추론 및 엣지 최적화 동향

Stereo4DWalker: Learning 4D-aware Embodied Urban Navigation from Internet Stereo Videos

Paper arXiv cs.CV (recent)

Despite rapid progress, embodied navigation in dynamic and unstructured urban environments remains brittle. Most existing approaches directly map monocular visual inputs to actions through end-to-end pixel-to-action training, assuming that accurate spatiotemporal (4D) scene understanding will emerge implicitly. While appealing, this paradigm requires large amounts of pixel-to-action supervision that are difficult to obtain. This challenge is amplified in dynamic, unstructured settings, where robust navigation requires precise 4D scene modeling. To address these limitations, we present Stereo4DWalker, a 4D-aware embodied navigation model that leverages stereo inputs and explicitly builds structured 4D representations of geometry and motion. These 4D structures are integrated into the navigation transformer through simple yet effective 4D-conditioned attention layers. To support scalable training, we curate a large-scale stereo navigation dataset with automatically annotated actions from Internet stereo videos. Our experiments show that Stereo4DWalker surpasses state-of-the-art performance using only 1.5% of the training data, highlighting the effectiveness of explicit 4D visual modeling for data-efficient and robust urban navigation.

원문 보기

Preserving Guidance in Cost-Volume Retrieval under Extreme LiDAR Sparsity in Iterative Stereo

Paper arXiv cs.CV (recent)

While accurate LiDAR depth has been shown to improve stereo matching, high-end LiDAR remains costly and difficult to deploy at scale, motivating guidance from sparse LiDAR measurements. In this paper, we revisit how extremely sparse LiDAR can guide iterative stereo by examining its cost-volume retrieval mechanism. Our analysis shows that LiDAR guidance degrades sharply when only a few hundred points are available, and we provide a signal-processing explanation for why iterative stereo fails to fully exploit such sparse input. This insight leads to a simple yet highly effective remedy: pre-filling the initial disparity map so that cost-volume retrieval becomes more reliable under extremely sparse LiDAR. We find that pre-filling is also effective when injecting LiDAR depth into image features via early fusion, but for a fundamentally different reason, necessitating a distinct pre-filling approach. Combining both forms of guidance, our proposed Guided RAFT-Stereo (GRAFT-Stereo) achieves substantial improvements over existing LiDAR-guided stereo methods across multiple datasets.

원문 보기

AI 뉴스 & 리서치

기업/연구기관의 주요 발표와 블로그 업데이트

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation

Paper Hugging Face Papers

Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Atria Dawn: The Dawn of Agentic Superintelligence

Paper Hugging Face Papers

Atria Dawn: The Dawn of Agentic Superintelligence에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

Paper Hugging Face Papers

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Dream-RSI: Recursive Self-Improvement through Evolving Worlds

Paper Hugging Face Papers

Dream-RSI: Recursive Self-Improvement through Evolving Worlds에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

A Chosen Future Can Still Be Rewritten: Causal Writability in Video Models

Paper arXiv cs.CV (recent)

When a video model generates physically incorrect motion, did it fail to learn the correct motion, or did it learn it but fail to use it? We show the latter: the correct motion remains available inside the model and can still be made to control the generated video. We train on videos where red masses oscillate slowly and blue masses oscillate quickly, then test a red mass with fast observed motion. Even when the model generates slow motion in this conflicting case, a low-dimensional edit predicted from simple physical variables restores the correct fast motion. We call this ability causal writability. At fixed strength, we find a sharp depth boundary: the same edit changes the video before the boundary but not after it. This closure marks commitment for that write. The motion signal nevertheless remains, and a stronger downstream write can restore physical motion, while excessive gain overshoots. Early causal writability predicts which errors training later corrects: those errors are writable at more network depths than errors that persist. We reproduce both causal writability and its sharp closure in a pretrained 1.3B video model, supporting generality across model scale and training regime.

원문 보기

SAMReg: SAM-enabled Image Registration with ROI-based Correspondence

Paper arXiv cs.CV (recent)

This paper describes a new spatial correspondence representation based on paired regions-of-interest (ROIs), for medical image registration. The distinct properties of the proposed ROI-based correspondence are discussed, in the context of potential benefits in clinical applications following image registration, compared with alternative correspondence-representing approaches, such as those based on sampled displacements and spatial transformation functions. These benefits include a clear connection between learning-based image registration and segmentation, which in turn motivates two cases of image registration approaches using (pre-)trained segmentation networks. Based on the segment anything model (SAM), a vision foundation model for segmentation, we develop a new registration algorithm SAMReg, which does not require any training (or training data), gradient-based fine-tuning or prompt engineering. The proposed SAMReg models are evaluated across five real-world applications, including intra-subject registration tasks with cardiac MR and lung CT, challenging inter-subject registration scenarios with prostate MR and retinal imaging, and an additional evaluation with a non-clinical example with aerial image registration. The proposed methods outperform both intensity-based iterative algorithms and DDF-predicting learning-based networks across tested metrics including Dice and target registration errors on anatomical structures, and further demonstrates competitive performance compared to weakly-supervised registration approaches that rely on fully-segmented training data. Open source code and examples are available at: https://github.com/sqhuang0103/SAMReg.git.

원문 보기

September 15, 2026 Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train Algorithms & Theory · Data Mining & Modeling · Generative AI

News Google Research Blog

September 15, 2026 Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train Algorithms & Theory · Data Mining & Modeling · Generative AI에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

참고한 소스