Research Context

XOR is the canonical encryption primitive in bare-metal assembly. It is symmetric — applying the same key twice returns the original byte — so the encrypt and decrypt loops are identical code. No library call, no padding, no block size: one instruction per byte. The two implementations here cover the two standard variants: a fixed single-byte key (fast, simple, weak against frequency analysis) and a rolling key that changes every iteration (stronger, same loop structure, one extra add). Both read a file, encrypt in-place, and write to stdout using only Linux syscalls.


1-XOR Properties: Why One Instruction Handles Both Directions

Three identities make XOR useful as an encryption operation:

x ^ 0 = x        (key 0 is a no-op)
x ^ x = 0        (anything XORed with itself is zero)
x ^ k ^ k = x    (applying k twice returns x — symmetry)

The third identity is the one that matters here: ciphertext = plaintext ^ key, and plaintext = ciphertext ^ key. The operation is its own inverse. This means the decrypt loop is byte-for-byte identical to the encrypt loop — run the same code twice and you are back to the original file.


2-Fixed-Key XOR: File Read, In-Place Encrypt, Write

The first implementation reads a file into a .bss buffer, iterates over every byte XORing it against 0x11, copies the result to an rbp-anchored stack buffer, then writes to stdout:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
; nasm -f elf64 xor_encrypt.asm -o xor_encrypt.o && ld xor_encrypt.o -o xor_encrypt

section .data
    file db "example.txt", 0

section .bss
    fd_no  resb 4
    buffer resb 1024

section .text
global _start
_start:
    ; rbp anchor — 512-byte zeroed frame on the stack
    sub rsp, 512
    cld
    mov rdi, rsp
    xor rax, rax
    mov rcx, 64
    rep stosq
    mov rbp, rsp

    ; sys_open("example.txt", O_RDONLY=0, 0)
    mov rax, 2
    lea rdi, [file]
    xor rsi, rsi            ; O_RDONLY = 0
    xor rdx, rdx
    syscall
    test rax, rax
    jl _error
    mov [fd_no], rax        ; save fd

    ; sys_read(fd, buffer, 1024)
    xor rax, rax            ; sys_read = 0
    mov rdi, [fd_no]
    lea rsi, [buffer]
    mov rdx, 1024
    syscall
    test rax, rax
    jle _error
    mov r14, rax            ; bytes read

    ; XOR encrypt loop — in-place over buffer
    xor rcx, rcx
    lea rdx, [buffer]       ; rdx = buffer start address

_encrypt_loop:
    xor byte [rdx], 0x11    ; XOR current byte in-place (NASM byte keyword required)
    mov [rbp + 100 + rcx], dl  ; copy to output buffer at rbp+100
    inc rcx
    inc rdx
    cmp rcx, r14
    jne _encrypt_loop

    ; sys_write(stdout=1, output_buf, bytes_read)
    mov rax, 1
    mov rdi, 1
    lea rsi, [rbp + 100]
    mov rdx, r14
    syscall

    mov rax, 60
    xor rdi, rdi
    syscall

_error:
    neg rax
    mov rdi, rax
    mov rax, 60
    syscall

The loop body is four instructions per byte: xor, mov, inc, inc, plus the cmp/jne at the end. r14 holds the byte count from sys_read, so the loop runs exactly as many times as there are bytes in the file.


3-The lea vs mov Pointer Bug

The most common mistake in this pattern — and in any buffer loop — is writing mov rdx, [buffer] instead of lea rdx, [buffer].

1
2
3
4
5
; WRONG — dereferences the label: loads the first 8 bytes of the file into rdx
mov rdx, [buffer]   ; rdx = *(buffer) — this is now file data treated as an address

; CORRECT — loads the address of buffer itself
lea rdx, [buffer]   ; rdx = &buffer

What happens at runtime with the wrong version: rdx contains the first eight bytes of example.txt interpreted as a 64-bit integer — for a text file this is typically something in the range 0x6C6C6548... (ASCII characters). The first xor byte [rdx], 0x11 writes to that address, which is not mapped. The result is SIGSEGV.

PRO TIP: The rule is simple — mov reg, [label] is a load (dereference), lea reg, [label] is an address calculation (no memory access). Whenever you need a pointer to a buffer, use lea. Whenever you need the value stored at a label (like mov rdi, [fd_no]), use mov.

The same distinction applies to lea rsi, [buffer] in the sys_read call. Both must be lea — sys_read expects a pointer to write into, not a value loaded from there.


4-Rolling Key XOR: One Extra Instruction, Much Stronger

The fixed-key version has a repeating pattern in the ciphertext — every occurrence of the same plaintext byte produces the same ciphertext byte, which makes frequency analysis straightforward. The rolling key variant fixes this by updating the key after each byte:

 1
 2
 3
 4
 5
 6
 7
 8
 9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
; nasm -f elf64 rolling_xor.asm -o rolling_xor.o && ld rolling_xor.o -o rolling_xor
; ./rolling_xor > encrypted.bin

section .data
    file    db "example.txt", 0
    xor_key db 0x11             ; starting key

section .bss
    fd_no  resb 4
    buffer resb 1024

section .text
global _start
_start:
    sub rsp, 512
    cld
    mov rdi, rsp
    xor rax, rax
    mov rcx, 64
    rep stosq
    mov rbp, rsp

    ; sys_open
    mov rax, 2
    lea rdi, [file]
    xor rsi, rsi
    xor rdx, rdx
    syscall
    test rax, rax
    jl _error
    mov [fd_no], rax

    ; sys_read
    xor rax, rax
    mov rdi, [fd_no]
    lea rsi, [buffer]
    mov rdx, 1024
    syscall
    test rax, rax
    jle _error
    mov r14, rax

    xor rcx, rcx
    lea rdx, [buffer]
    movzx r15, byte [xor_key]   ; r15b = starting key (0x11)

_encrypt_loop:
    mov al, [rdx]           ; load current byte
    xor al, r15b            ; XOR against rolling key
    mov [rdx], al           ; write back in-place
    mov [rbp + 100 + rcx], al  ; copy to output buffer

    inc rcx
    inc rdx
    add r15b, 0x11          ; update key: 0x11 > 0x22 > 0x33 > ...
                            ; r15b is 8-bit: 0xFF + 0x11 = 0x10 (auto wrap)
    cmp rcx, r14
    jne _encrypt_loop

    ; sys_write
    mov rax, 1
    mov rdi, 1
    lea rsi, [rbp + 100]
    mov rdx, r14
    syscall

    mov rax, 60
    xor rdi, rdi
    syscall

_error:
    neg rax
    mov rdi, rax
    mov rax, 60
    syscall

The only additions compared to the fixed-key version are movzx r15, byte [xor_key] before the loop and add r15b, 0x11 inside it. The key sequence produced is 0x11, 0x22, 0x33, ..., 0xFF, 0x10, 0x21, ... — a deterministic keystream that the receiver can reproduce given only the starting key.


5-8-bit Register Auto Wrap-Around

1
add r15b, 0x11   ; add to the low byte of r15 only

r15b is the lowest 8 bits of r15. An 8-bit register holds values 0x00–0xFF. When the value exceeds 0xFF, the carry is discarded — not propagated to the upper bytes of r15. This is modulo-256 arithmetic for free, no and r15, 0xFF required.

r15b = 0xEE > add 0x11 > 0xFF   (no wrap)
r15b = 0xFF > add 0x11 > 0x10   (wrapped — carry dropped)
r15b = 0xF5 > add 0x11 > 0x06   (wrapped)

PRO TIP: movzx r15, byte [xor_key] zero-extends the 8-bit value into the full 64-bit register before the loop. Without movzx, the upper bytes of r15 might contain garbage from earlier computation, and xor al, r15b still works since it only touches the low byte — but keeping r15 clean avoids confusion during debugging.


6-Decrypt: Same Code, Same Key

Because XOR is symmetric, decrypt is the encrypt loop run again with the same starting key. The fixed-key version needs no state — just run the program on the ciphertext file with the same 0x11 key. The rolling-key version requires the starting key and the same increment (0x11) — both known to the receiver.

Python verification of the rolling decrypt:

1
2
3
4
key = 0x11
for b in data:
    out.append(b ^ key)
    key = (key + 0x11) & 0xFF

The & 0xFF in Python does what the 8-bit register does in assembly — clamps the key to one byte. Run this against the output of rolling_xor and the original example.txt is recovered exactly.


Conclusion

Two programs, one core idea: a pointer, a counter, a key, and xor byte [ptr], key inside a loop. The fixed-key version establishes the structure; the rolling-key version adds one line that fundamentally changes the statistical properties of the output. The most instructive part of building both is the lea vs mov bug — it appears the first time every assembly programmer writes a buffer loop, and understanding why it segfaults ingrains the distinction between an address and a value for every subsequent program.


Coding: CFG Flattening with CMOV: Antivirus & EDR Evasion via Control Flow Obfuscation in x64 Assembly

Coding: VESQER: DPCM+RLE Hybrid Shellcode Compression in x64 Assembly

Coding: Linux x64 Assembly Syscall ABI: Registers, File Descriptors & .bss Segment


This article is written for educational purposes and security research only. The techniques described are standard computer science knowledge covered in any cryptography or systems programming curriculum. The author is not responsible for any misuse. Applying these methods to systems you do not own or have explicit written permission to test is strictly prohibited.

MITRE ATT&CK: T1027 · T1027.013