cjk-unicode-steganography
Overview
CJK Unicode steganography is a custom payload-hiding technique in which raw binary data (typically a PE or shellcode) is encoded as a string of CJK Unified Ideograph characters. Because these characters occupy the Unicode range U+4E00–U+9FFF (and extensions U+3400–U+4DBF), a simple linear mapping like byte = charcode - 0x3400 or byte = charcode - 0x4E00 can encode any byte value (0–255) inside a valid Unicode character. The resulting string looks like East Asian text to casual inspection, evading string-based detection and surviving transport through systems that expect Unicode text.
Encoding Schemes
| Scheme | Formula | Unicode Range | Notes |
|---|---|---|---|
| CJK Extension A | byte = charcode - 0x3400 |
U+3400–U+4DBF | 6,592 code points; sufficient for 0–255 with room for padding |
| CJK Unified | byte = charcode - 0x4E00 |
U+4E00–U+9FFF | 20,992 code points; observed in eba13078 with 381-entry dictionary |
Observed Implementations
- JScript CJK dropper (sample
0de6482c): Encodes a 385,024-byte .NET assembly as 384 CJK Extension A strings assigned to nested JavaScript object properties. Each string is decoded viacharcode - 0x3400and reconstructed in-memory via PowerShell. ^[/intel/analyses/0de6482c69377a127b91aa9c28d24981b656be44a9181a83c5b014a933987216.html] - JScript Hangul dropper (sample
7129076f): Encodes a 349,696-byte .NET assembly as 347 Hangul strings in a JavaScript dictionary. Decoded viacharcode - 0xAC00with a fallback CJK stream-cipher path. Same actor, evolved tooling. ^[/intel/analyses/7129076f2b648b20cbd7b35eb8612ba4315be053ebcb5fa852b689f1ef72deed.html] - JScript CJK Unified dropper (sample
eba13078): Encodes a 379,392-byte .NET assembly as 381 CJK Unified Ideograph strings (charcode - 0x4E00). Uses XSLT JScript extension execution andconhost.exe --headlessPowerShell spawn. ^[/intel/analyses/eba13078dea9e803b9120a45cd0dfad589f1defdf0f86fe84ae77fe63fff5300.html]
Detection
- Entropy: A long string of CJK characters with uniform distribution (each byte 0–255 mapped linearly) will have high entropy (~7.9 bits/byte), similar to encrypted data. Natural CJK text has lower entropy (~4–5 bits/byte) and non-uniform byte distribution.
- Regex: Hunt for sequences of 50+ CJK characters in JavaScript or text files, especially when interleaved with ASCII variable assignments or
WScript.Shellcalls. - YARA:
uint32(0) == 0xFFFE(UTF-16 LE BOM) followed by CJK blocks and English variable names.
Related Techniques
- jscript-environment-variable-staging — common staging chain paired with CJK steganography
- semantic-english-name-obfuscation — random English words used as variable/type names alongside CJK payloads