← 전체 목록

2026-09-09 AI 리서치 브리핑

최신 VLM, sLLM, on-device AI 논문과 연구 블로그를 한눈에 정리합니다. 중복 기사 방지를 위해 URL 기준으로 추적합니다.

총 18건 요약 자동 생성

VLM 업데이트

멀티모달 비전-언어 모델의 최신 논문과 리더보드 변화

MindTopo reveals VLMs’ spatial reasoning abilities

News Microsoft Research Blog

MindTopo 연구를 통해 VLM(Vision-Language Models)의 공간 추론 능력이 밝혀졌습니다. 이 발견은 VLM이 시각적 정보를 이해하고 추론하는 방식에 대한 통찰력을 제공합니다.

원문 보기

sLLM 트렌드

경량화·효율화를 위한 스몰 LLM 연구

A Generalizable Feature Extractor for Alzheimer's-Related Brain MRI Tasks

Paper arXiv cs.CV (recent)

When there is not enough labeled data to properly train deep learning models, transfer learning can help. We still do not fully understand how effective it is in neuroimaging, especially for Alzheimer's disease research. It is also not clear if these transferred models can work on new datasets without being retrained for each specific task. We evaluate whether a compact, supervised pretrained model can serve as a reusable foundation model for downstream neuroimaging tasks. We freeze the 7.18 million weights of a 3D CNN previously trained for brain-age prediction, and adapt it to each task using Low-Rank Adaptation (LoRA), requiring only ~1% additional trainable parameters. We evaluate generalizability in six experiments. Adapting the model to classify cognitively normal versus Dementia on ADNI gave an AUC of 0.964 on held-out folds (Experiment #1). Applying that adapted model unchanged to OASIS-3, with no retraining, gave an AUC of 0.871 (Experiment #2). Reusing its output logit together with age and a cognitive score distinguished stable from progressing MCI with an AUC of 0.828 (Experiment #3). Adapting the same backbone to predict amyloid positivity from structural MRI gave an AUC of 0.804 (Experiment #4). Finally, the same approach estimated ICV-normalized hippocampal and white matter hypointensity volumes directly from the T1w image, with R^2 of 0.80 and 0.91 respectively, tasks normally addressed with much larger U-Net networks (Experiments #5 and #6). A compact model supervised on brain age can therefore serve as a reusable backbone, adapting to each task with ~1% additional parameters and transferring to an unseen cohort without any training. Our findings suggest that a carefully trained brain age model can serve as an effective foundation model for Alzheimer's related tasks, even under strict data constraints.

원문 보기

AI 뉴스 & 리서치

기업/연구기관의 주요 발표와 블로그 업데이트

Unlocking Lossless Speedups in LLMs via Discrete Diffusion

Paper Hugging Face Papers

Unlocking Lossless Speedups in LLMs via Discrete Diffusion에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience

Paper Hugging Face Papers

FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation

Paper Hugging Face Papers

ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Causal Foundation Models

Paper Hugging Face Papers

Causal Foundation Models에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

Paper Hugging Face Papers

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

WorldSculpt: Generating Compositional Worlds from Grounded Videos

Paper arXiv cs.CV (recent)

We study the problem of generating a compositional 3D representation of a cluttered scene containing hundreds of objects. The goal is to represent the scene as a collection of individual object meshes placed in a shared world frame, as required by downstream applications such as gaming, AR/VR, simulation, and robotics. This task is challenging in densely cluttered scenes, where objects heavily occlude one another and each view reveals only a fraction of their geometry. Geometry-based approaches typically reconstruct the scene as a single representation and leave incomplete geometry in occluded regions, while existing compositional methods with generative priors are largely limited to relatively simple scenes. We show that complex scenes with hundreds of objects can instead be generated compositionally by adapting a strong single-object 3D generative prior to multi-view observations. We instantiate this paradigm with Pixal3D, extending it with a multi-view conditioning pathway that grounds object generation in multiple posed observations. Although the model is finetuned entirely on single objects in canonical space, it generalizes to large scenes with severe occlusion without any scene-level training, demonstrating the feasibility and scalability of this paradigm. We further introduce UE-MeshyScene, a photorealistic benchmark of densely cluttered scenes with hundreds of objects, per-object annotations, and ground-truth meshes. Across single-object, controlled multi-object, and UE-MeshyScene evaluations, our method consistently outperforms prior approaches, with larger gains as scene complexity and occlusion increase. Finally, we demonstrate broader applicability by converting generated 3DGS worlds, such as Marble and HY-World 2.0, into compositional mesh scenes.

원문 보기

UniMate: One Unified Model to Animate Diverse Skeletons

Paper arXiv cs.CV (recent)

Recent advances in automatic rigging now deliver animation-ready 3D assets at scale, yet generating the motion to drive them remains a bottleneck. Existing learned animators are topology-constrained: they rely on category-specific templates or require per-skeleton fine-tuning and reference motions at inference. We present UniMate, a unified foundation model that synthesizes articulated motion for arbitrary skeletons from a rigged 3D asset and a text prompt, with no test-time optimization or per-skeleton retraining. UniMate introduces a topology-aware diffusion transformer, which integrates skeletal topology into attention via three mechanisms: (1) a graph-aware attention bias from pairwise joint relations and geodesic distances; (2) a spectral rotary position embedding generalizing RoPE to arbitrary kinematic trees via the graph Laplacian; and (3) a global topological conditioner attention-pooled from the rest-pose skeleton. We also curate UniML3D, 13,006 motion sequences spanning bipedal, quadrupedal, avian, marine, insectoid, serpentine, and articulated rigid objects with unified canonicalization and text pairing. Trained on this dataset, UniMate outperforms state-of-the-art baselines in quality, generalization, and efficiency, and supports zero-shot cross-topology transfer, in-betweening, expansion, and text-guided editing. Our project page is available at https://linzhanmou.com/unimate/.

원문 보기

From Interpretability Methods to Interpretable Models

Paper arXiv cs.CV (recent)

More than a decade in, explainable AI (XAI) for computer vision has assembled a mature toolbox: attribution, feature visualization, concept-based, and circuit-based methods. Yet almost all of the field's effort has gone into building and comparing these methods, and little into the question they were meant to answer---how interpretable are our models, and are we making progress as they evolve? We argue for shifting the field's focus from methods to models, along two complementary lines. One is already within reach: existing tools let us characterize and compare what different models represent and compute. The other is harder, and largely neglected: whether a model can actually be understood by the humans who rely on it---the independent evaluators on whom trust and certification depend, not the experts confirming what they already expect. It can only be measured, not inferred. We review why the toolbox is mature enough to support both, survey the thin body of work comparing models, draw a parallel to systems neuroscience, and close with a model-centric XAI agenda.

원문 보기

CrossDepth: Geometry-Constrained Attention for Generalizable Multi-View Surround Depth Estimation

Paper arXiv cs.CV (recent)

Reliable 3D understanding of the surrounding environment is a core requirement for autonomous driving. Multi-view surround camera rigs provide broad scene coverage, but the spatially adjacent images typically overlap only minimally. Consequently, the depth of most pixels must be inferred from monocular appearance cues. These cues can appear differently across images and may therefore be interpreted differently by the depth estimation model. We target two main sources of cross-image inconsistency: differences in camera intrinsics and the limited receptive field of each image. We address the former by conditioning the features on per-pixel camera-aware ray embeddings, enabling the network to account for camera-dependent variations in monocular cues. We address the latter by extending each pixel's context beyond its own image through cross-image attention constrained to geometrically plausible regions, derived from the calibrated rig setup. The model is trained in a fully self-supervised manner based on photometric consistency. Evaluations on DDAD and nuScenes show improved overall depth accuracy and cross-image depth consistency over state-of-the-art self-supervised methods under in-domain and cross-domain evaluation. Code is available at https://abualhanud.github.io/CrossDepthPage/.

원문 보기

September 3, 2026 Transfer learning for genomic prediction in underrepresented populations General Science · Machine Intelligence

News Google Research Blog

September 3, 2026 Transfer learning for genomic prediction in underrepresented populations General Science · Machine Intelligence에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

September 3, 2026 A connectomics milestone: Mapping the complete male fruit fly brain General Science · Health & Bioscience · Machine Intelligence · Open Source Models & Datasets

News Google Research Blog

September 3, 2026 A connectomics milestone: Mapping the complete male fruit fly brain General Science · Health & Bioscience · Machine Intelligence · Open Source Models & Datasets에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

September 1, 2026 Mapping global methane emissions from space with deep learning Climate & Sustainability · Earth AI · Machine Intelligence

News Google Research Blog

September 1, 2026 Mapping global methane emissions from space with deep learning Climate & Sustainability · Earth AI · Machine Intelligence에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

August 31, 2026 TimesFM-3: A zero-shot foundation model for multivariate forecasting Data Management · Machine Intelligence · Product

News Google Research Blog

August 31, 2026 TimesFM-3: A zero-shot foundation model for multivariate forecasting Data Management · Machine Intelligence · Product에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

August 27, 2026 Planetary prediction engine: Automating global models via Earth AI Earth AI · Generative AI · Machine Intelligence

News Google Research Blog

August 27, 2026 Planetary prediction engine: Automating global models via Earth AI Earth AI · Generative AI · Machine Intelligence에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models

News Microsoft Research Blog

GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models에 관한 최근 업데이트입니다. 자세한 내용은 원문 링크에서 확인할 수 있습니다.

원문 보기

Broadening access to Skala creates a faster path to predictive DFT

News Microsoft Research Blog

Skala에 대한 접근성을 확대하여 예측 DFT(밀도 범함수 이론)로 가는 더 빠른 경로를 만들었습니다. 이는 예측 DFT 분야의 발전을 가속화하는 데 기여할 것으로 기대됩니다.

원문 보기

참고한 소스