paper-with-me

Papers

Data Center Audio/Video Intelligence on Device (DAVID) -- An Edge-AI Platform for Smart-Toys

2023-11-18 · Gabriel Cosache, Francisco Salgado, Cosmin Rotariu, George Sterpu, Rishabh Jain, Peter Corcoran

An overview is given of the DAVID Smart-Toy platform, one of the first Edge AI platform designs to incorporate advanced low-power data processing by neural inference models co-located with the relevant image or audio sensors. There is also on-board capability for in-device text-to-speech generation. Two alternative embodiments are presented: a smart Teddy-bear, and a roving dog-like robot. The platform offers a speech-driven user interface and can observe and interpret user actions and facial expressions via its computer vision sensor node. A particular benefit of this design is that no personally identifiable information passes beyond the neural inference nodes thus providing inbuilt compliance with data protection regulations.

📄 PDF Abstract BibTeX arXiv:2311.11030

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Similar Papers 제목 키워드 기반

Audio Surveillance: a Systematic Review

2014-09-27 · Marco Crocco, Marco Cristani, Andrea Trucco, Vittorio Murino

Despite surveillance systems are becoming increasingly ubiquitous in our living environment, automated surveillance, currently based on video sensory modality and machine intelligence, lacks most of the time the robustne…

Object Tracking

Agentic Very Long Video Understanding

2026-01-26 · Aniket Rege, Arka Sadhu, Yuliang Li, Kejie Li 외 arxiv

The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the con…

Self-supervised learning for audio-visual speaker diarization

2020-02-13 · Yifan Ding, Yong Xu, Shi-Xiong Zhang, Yahuan Cong 외

Speaker diarization, which is to find the speech segments of specific speakers, has been widely used in human-centered applications such as video conferences or human-computer interaction systems. In this paper, we propo…

Self-Supervised Learningspeaker-diarizationSpeaker DiarizationTriplet+1

Self-supervised Audio Spatialization with Correspondence Classifier

2019-05-14 · Yu-Ding Lu, Hsin-Ying Lee, Hung-Yu Tseng, Ming-Hsuan Yang

Spatial audio is an essential medium to audiences for 3D visual and auditory experience. However, the recording devices and techniques are expensive or inaccessible to the general public. In this work, we propose a self-…

TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion Models

2025-06-03 · Chetwin Low, Weimin WANG

In this paper, we present TalkingMachines -- an efficient framework that transforms pretrained video generation models into real-time, audio-driven character animators. TalkingMachines enables natural conversational expe…

DecoderKnowledge DistillationLanguage ModelingLanguage Modelling+2