LLM strips underscores when extracting/decoding strings
I want the LLM to preserve exact string formatting, specifically underscores, when extracting or decoding obfuscated or encoded strings (like Base64 or strings with zero-width characters). Currently, alphanumeric characters are preserved but underscores are deleted. This is important for accurately handling technical data like API tokens or encoded strings.
Steps to Reproduce:
1. Feed the AI a Base64 string containing underscores, or a string interwoven with zero-width characters.
2. Ask it to extract or decode the string.
Expected Result: Exact string formatting is preserved.
Actual Result: Alphanumeric characters are preserved, but underscores are deleted.
- What AI feature?
Log in to comment and vote
Comments2
Shreya Yadav
Mar 9
The Evolving Bug:
Look at the string it gave this time:
SEEDTEST42Original String:
SEED_TEST_42First Extraction:
SEEDTEST_42(Lost one underscore)Memory Extraction:
SEEDTEST42(Lost both underscores!)Shreya Yadav
Mar 9
Bug: Chatbot strips underscores from IoCs and obfuscated strings during extraction.
Impact: High for forensic accuracy; users cannot rely on the bot for exact copy-pasting of complex strings.
Proof: Successfully reproduced using Base64, Zero-Width Space injection, and Cyrillic Homograph substitution.