paper-with-me

홈 › Papers

Real World Voice Assistant System for Cooking

2019-10-01 · WS 2019 10 · Takahiko Ito, Shintaro Inuzuka, Yoshiaki Yamada, Jun Harashima

This study presents a voice assistant system to support cooking by utilizing smart speakers in Japan. This system not only speaks the procedures written in recipes point by point but also answers the common questions from users for the specified recipes. The system applies machine comprehension techniques to millions of recipes for answering the common questions in cooking such as {``}人参はどうしたらよいですか (How should I cook carrots?){''}. Furthermore, numerous machine-learning techniques are applied to generate better responses to users.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

BIG-bench Machine LearningReading Comprehension

Similar Papers 제목 키워드 기반

GRILLBot: An Assistant for Real-World Tasks with Neural Semantic Parsing and Graph-Based Representations

2022-08-31 · Carlos Gemmell, Iain Mackie, Paul Owoicho, Federico Rossetto 외

GRILLBot is the winning system in the 2022 Alexa Prize TaskBot Challenge, moving towards the next generation of multimodal task assistants. It is a voice assistant to guide users through complex real-world tasks in the d…

Semantic Parsing

VoiceBench: Benchmarking LLM-Based Voice Assistants

2024-10-22 · Yiming Chen, Xianghu Yue, Chen Zhang, Xiaoxue Gao 외

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)BenchmarkingGeneral Knowledge+2

TWIZ-v2: The Wizard of Multimodal Conversational-Stimulus

2023-10-03 · Rafael Ferreira, Diogo Tavares, Diogo Silva, Rodrigo Valério 외

In this report, we describe the vision, challenges, and scientific contributions of the Task Wizard team, TWIZ, in the Alexa Prize TaskBot Challenge 2022. Our vision, is to build TWIZ bot as an helpful, multimodal, knowl…

Language ModelingLanguage ModellingLarge Language Model

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

2026-06-08 · Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh, Rajeev Yasarla 외 arxiv

Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large language models (LLMs) as task guidance assi…

VILT: Video Instructions Linking for Complex Tasks

2022-08-23 · Sophie Fischer, Carlos Gemmell, Iain Mackie, Jeffrey Dalton

This work addresses challenges in developing conversational assistants that support rich multimodal video interactions to accomplish real-world tasks interactively. We introduce the task of automatically linking instruct…

Retrieval