hangul-syllable-steganography

Overview

Hangul Syllable steganography is a custom payload-hiding technique in which raw binary data (typically a PE or .NET assembly) is encoded as a string of Korean Hangul Syllable characters. The Hangul Syllable block occupies Unicode range U+AC00–U+D7A3 (11,172 code points), more than sufficient to encode any byte value (0–255) via a simple linear mapping: byte = charcode - 0xAC00. The resulting string looks like Korean text to casual inspection, evading string-based detection and surviving transport through systems that expect Unicode text.

Encoding Scheme

Scheme Formula Unicode Range Notes
Hangul Syllable byte = charcode - 0xAC00 U+AC00–U+D7A3 11,172 code points; common in East-Asian locale contexts

Observed Implementations

  • JScript Hangul dropper (sample 7129076f): Encodes a 349,696-byte .NET assembly as 347 Hangul strings assigned to a JavaScript dictionary (Dictionary.Add(key, value)). A separate order array dictates concatenation sequence. Decoded via charcode - 0xAC00 and reconstructed in-memory via PowerShell System.Reflection.Assembly::Load. ^[/intel/analyses/7129076f2b648b20cbd7b35eb8612ba4315be053ebcb5fa852b689f1ef72deed.html]

Detection Opportunities

  • Entropy: A long string of Hangul characters with uniform distribution will have high entropy (~7.9 bits/byte), similar to encrypted data. Natural Korean text has lower entropy (~4–5 bits/byte) and non-uniform byte distribution.
  • Regex: Hunt for sequences of 50+ Hangul characters in JavaScript or text files, especially when interleaved with ASCII variable assignments or WScript.Shell calls.
  • Behavioral detection: Monitor wscript.exe spawning powershell.exe or cmd.exe with encoded commands. This is the reliable kill chain regardless of obfuscation.

Related Techniques