paper-with-me

Papers

Transcripts per million ratio: applying distribution-aware normalisation over the popular TPM method

2022-05-05 · Hilbert Lam Yuen In, Robbe Pincket

Current popular methods in literature of RNA sequencing normalisation do not account for gene length when compared across samples, whilst adjusting for count biases in the data. This creates a gap in the normalisation as bigger genes in RNA sequencing accumulate more reads due to shotgun sequencing methods. As a result, the proportions of these reads inter-sample are not properly accounted for in current normalisation methods. Alternatively, methods which account for gene length do not account for the pan-sample biases in the data by accounting for a central read average. Thus, in order to fill in the gap in the literature, we propose a novel method of Transcripts Per Million Ratio and its relatives in RNA-sequencing differential expression normalisation that can be used in different conditions, which takes into account the gene length as well as relative expression in normalisation.

📄 PDF Abstract BibTeX arXiv:2205.02844

Code (1)

Chokyotager/Ribonorma 공식 구현

Similar Papers 제목 키워드 기반

Learning ASR-Robust Contextualized Embeddings for Spoken Language Understanding

2019-09-24 · Chao-Wei Huang, Yun-Nung Chen

Employing pre-trained language models (LM) to extract contextualized word representations has achieved state-of-the-art performance on various NLP tasks. However, applying this technique to noisy transcripts generated by…

Spoken Language Understanding

Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Summaries

2025-08-18 · Kawin Mayilvaghanan, Siddhant Gupta, Ayush Kumar arxiv

Abstractive summarization is a core application in contact centers, where Large Language Models (LLMs) generate millions of summaries of call transcripts daily. Despite their apparent quality, it remains unclear whether …

Assessing the Use of Prosody in Constituency Parsing of Imperfect Transcripts

2021-06-14 · Trang Tran, Mari Ostendorf

This work explores constituency parsing on automatically recognized transcripts of conversational speech. The neural parser is based on a sentence encoder that leverages word vectors contextualized with prosodic features…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Constituency ParsingReranking+3

Predicting TED Talk Ratings from Language and Prosody

2019-05-21 · Md. Iftekhar Tanveer, Md Kamrul Hassan, Daniel Gildea, M. Ehsan Hoque

We use the largest open repository of public speaking---TED Talks---to predict the ratings of the online viewers. Our dataset contains over 2200 TED Talk transcripts (includes over 200 thousand sentences), audio features…

BIG-bench Machine Learning

WildChat: 1M ChatGPT Interaction Logs in the Wild

2024-05-02 · Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie 외

Chatbots such as GPT-4 and ChatGPT are now serving millions of users. Despite their widespread use, there remains a lack of public datasets showcasing how these tools are used by a population of users in practice. To bri…

ChatbotInstruction Following