paper-with-me

Papers

HiLight: Technical Report on the Motern AI Video Language Model

2024-07-10 · Zhiting Wang, Qiangong Zhou, Kangjie Yang, Zongyang Liu, Xin Mao

This technical report presents the implementation of a state-of-the-art video encoder for video-text modal alignment and a video conversation framework called HiLight, which features dual visual towers. The work is divided into two main parts: 1.alignment of video and text modalities; 2.convenient and efficient way to interact with users. Our goal is to address the task of video comprehension in the context of billiards. The report includes a discussion of the concepts and the final solution developed during the task's implementation.

📄 PDF Abstract BibTeX arXiv:2407.07325

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Det-SAM2:Technical Report on the Self-Prompting Segmentation Framework Based on Segment Anything Model 2

2024-11-28 · Zhiting Wang, Qiangong Zhou, Zongyang Liu

Segment Anything Model 2 (SAM2) demonstrates exceptional performance in video segmentation and refinement of segmentation results. We anticipate that it can further evolve to achieve higher levels of automation for pract…

Video SegmentationVideo Semantic Segmentation

HiLight: A Hierarchical Reinforcement Learning Framework with Global Adversarial Guidance for Large-Scale Traffic Signal Control

2025-06-17 · Yaqiao Zhu, Hongkai Wen, Geyong Min, Man Luo

Efficient traffic signal control (TSC) is essential for mitigating urban congestion, yet existing reinforcement learning (RL) methods face challenges in scaling to large networks while maintaining global coordination. Ce…

Hierarchical Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1

Learning Evidence Highlighting for Frozen LLMs

2026-04-24 · Shaoang Li, Yanhang Shi, Yufei Li, Mingfu Liang 외 arxiv

Large Language Models (LLMs) can reason well, yet often miss decisive evidence when it is buried in long, noisy contexts. We introduce HiLight, an Evidence Emphasis framework that decouples evidence selection from reason…

Sequential RecommendationReinforcement LearningQuestion Answering

Pegasus-v1 Technical Report

2024-04-23 · Raehyuk Jung, Hyojun Go, Jaehyuk Yi, Jiho Jang 외

This technical report introduces Pegasus-1, a multimodal language model specialized in video content understanding and interaction through natural language. Pegasus-1 is designed to address the unique challenges posed by…

Language ModelingLanguage ModellingQuestion AnsweringVideo Question Answering+1

A CLIP-Enhanced Method for Video-Language Understanding

2021-10-14 · Guohao Li, Feng He, Zhifan Feng

This technical report summarizes our method for the Video-And-Language Understanding Evaluation (VALUE) challenge (https://value-benchmark.github.io/challenge\_2021.html). We propose a CLIP-Enhanced method to incorporate…