paper-with-me

홈 › Papers

X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar Generation

2024-05-02 · Yiwei Ma, Zhekai Lin, Jiayi Ji, Yijun Fan, Xiaoshuai Sun, Rongrong Ji

Recent advancements in automatic 3D avatar generation guided by text have made significant progress. However, existing methods have limitations such as oversaturation and low-quality output. To address these challenges, we propose X-Oscar, a progressive framework for generating high-quality animatable avatars from text prompts. It follows a sequential Geometry->Texture->Animation paradigm, simplifying optimization through step-by-step generation. To tackle oversaturation, we introduce Adaptive Variational Parameter (AVP), representing avatars as an adaptive distribution during training. Additionally, we present Avatar-aware Score Distillation Sampling (ASDS), a novel technique that incorporates avatar-aware noise into rendered images for improved generation quality during optimization. Extensive evaluations confirm the superiority of X-Oscar over existing text-to-3D and text-to-avatar approaches. Our anonymous project page: https://xmu-xiaoma666.github.io/Projects/X-Oscar/.

📄 PDF Abstract BibTeX arXiv:2405.00954

Code (0)

등록된 구현이 없습니다.

Tasks

Text to 3D

Similar Papers 제목 키워드 기반

mOSCAR: A Large-scale Multilingual and Multimodal Document-level Corpus

2024-06-13 · Matthieu Futeral, Armel Zebaze, Pedro Ortiz Suarez, Julien Abadji 외

Multimodal Large Language Models (mLLMs) are trained on a large amount of text-image data. While most mLLMs are trained on caption-like data only, Alayrac et al. [2022] showed that additionally training them on interleav…

Few-Shot LearningIn-Context Learning

OSCAR: Object Status and Contextual Awareness for Recipes to Support Non-Visual Cooking

2025-03-07 · Franklin Mingzhe Li, Kaitlyn Ng, Bin Zhu, Patrick Carrington

Following recipes while cooking is an important but difficult task for visually impaired individuals. We developed OSCAR (Object Status Context Awareness for Recipes), a novel approach that provides recipe progress track…

Object

Reading the unreadable: Creating a dataset of 19th century English newspapers using image-to-text language models

2025-02-18 · Jonathan Bourne

Oscar Wilde said, "The difference between literature and journalism is that journalism is unreadable, and literature is not read." Unfortunately, The digitally archived journalism of Oscar Wilde's 19th century often has …

Image to textOptical Character RecognitionOptical Character Recognition (OCR)Topic Classification

OScaR: The Occam's Razor for Extreme KV Cache Quantization in LLMs and Beyond

2026-05-19 · Zunhai Su, Rui Yang, Chao Zhang, Yaxiu Liu 외 arxiv

The rapid advancement toward long-context reasoning and multi-modal intelligence has made the memory footprint of the Key-Value (KV) cache a dominant memory bottleneck for efficient deployment. While the established per-…

OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization

2026-05-18 · Zhongzhu Zhou, Donglin Zhuang, Jisen Li, Ziyan Chen 외 arxiv

INT2 KV-cache quantization is attractive for long-context LLM serving, but it remains difficult to make both accurate and deployable. Simple rotations such as Hadamard transforms reduce outliers, but still degrade at INT…