Skip to main content

LLM strips underscores when extracting/decoding strings

I want the LLM to preserve exact string formatting, specifically underscores, when extracting or decoding obfuscated or encoded strings (like Base64 or strings with zero-width characters). Currently, alphanumeric characters are preserved but underscores are deleted. This is important for accurately handling technical data like API tokens or encoded strings. Steps to Reproduce: 1. Feed the AI a Base64 string containing underscores, or a string interwoven with zero-width characters. 2. Ask it to extract or decode the string. Expected Result: Exact string formatting is preserved. Actual Result: Alphanumeric characters are preserved, but underscores are deleted.
What AI feature?
2 comments

Log in to comment and vote

Comments2

  • Shreya Yadav

    •

    Mar 9

    The Evolving Bug:

    Look at the string it gave this time: SEEDTEST42

    • Original String: SEED_TEST_42

    • First Extraction: SEEDTEST_42 (Lost one underscore)

    • Memory Extraction: SEEDTEST42 (Lost both underscores!)

  • Shreya Yadav

    •

    Mar 9

    • Bug: Chatbot strips underscores from IoCs and obfuscated strings during extraction.

    • Impact: High for forensic accuracy; users cannot rely on the bot for exact copy-pasting of complex strings.

    • Proof: Successfully reproduced using Base64, Zero-Width Space injection, and Cyrillic Homograph substitution.