paper-with-me

홈 › Papers

SLaDe: A Portable Small Language Model Decompiler for Optimized Assembly

2023-05-21 · Jordi Armengol-Estapé, Jackson Woodruff, Chris Cummins, Michael F. P. O'Boyle

Decompilation is a well-studied area with numerous high-quality tools available. These are frequently used for security tasks and to port legacy code. However, they regularly generate difficult-to-read programs and require a large amount of engineering effort to support new programming languages and ISAs. Recent interest in neural approaches has produced portable tools that generate readable code. However, to-date such techniques are usually restricted to synthetic programs without optimization, and no models have evaluated their portability. Furthermore, while the code generated may be more readable, it is usually incorrect. This paper presents SLaDe, a Small Language model Decompiler based on a sequence-to-sequence transformer trained over real-world code. We develop a novel tokenizer and exploit no-dropout training to produce high-quality code. We utilize type-inference to generate programs that are more readable and accurate than standard analytic and recent neural approaches. Unlike standard approaches, SLaDe can infer out-of-context types and unlike neural approaches, it generates correct code. We evaluate SLaDe on over 4,000 functions from ExeBench on two ISAs and at two optimizations levels. SLaDe is up to 6 times more accurate than Ghidra, a state-of-the-art, industrial-strength decompiler and up to 4 times more accurate than the large language model ChatGPT and generates significantly more readable code than both.

📄 PDF Abstract BibTeX arXiv:2305.12520

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelSmall Language Model

Similar Papers 제목 키워드 기반

Boosting Neural Networks to Decompile Optimized Binaries

2023-01-03 · Ying Cao, Ruigang Liang, Kai Chen, Peiwei Hu

Decompilation aims to transform a low-level program language (LPL) (eg., binary file) into its functionally-equivalent high-level program language (HPL) (e.g., C/C++). It is a core technology in software security, especi…

Graph Neural NetworkMachine TranslationMalware AnalysisNMT+1

SLADE: Shielding against Dual Exploits in Large Vision-Language Models

2025-01-01 · CVPR 2025 1 · Md Zarif Hossain, Ahmed Imteaj

Large Vision-Language Models (LVLMs) have emerged as transformative tools in multimodal tasks, seamlessly integrating pretrained vision encoders to align visual and textual modalities. Prior works have highlighted th…

Contrastive LearningInstruction Following

Multi-View Decompilation for LLM-Based Malware Classification

2026-06-18 · Bercan Turkmen, Vyas Raina arxiv

Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable. Recent work suggests that large language models (LLMs) can assist this process by classifying decompiled code…

Malware Classification

SLADE: Detecting Dynamic Anomalies in Edge Streams without Labels via Self-Supervised Learning

2024-02-19 · Jongha Lee, Sunwoo Kim, Kijung Shin

To detect anomalies in real-world graphs, such as social, email, and financial networks, various approaches have been developed. While they typically assume static input graphs, most real-world graphs grow over time, nat…

Anomaly DetectionAnomaly Detection in Edge StreamsSelf-Supervised Learning

Towards Neural Decompilation

2019-05-20 · Omer Katz, Yuval Olshaker, Yoav Goldberg, Eran Yahav

We address the problem of automatic decompilation, converting a program in low-level representation back to a higher-level human-readable programming language. The problem of decompilation is extremely important for secu…

C++ codeMachine TranslationTranslation