Direct answer
Deserializing untrusted pickle data is remote code execution (RCE, an attacker running arbitrary commands on a machine they were never given access to) because pickle's object format does not just describe data, it describes instructions for rebuilding an object, and one of those instructions is "call this function." An attacker who controls the bytes controls which function gets called and with what arguments, before any of your own application code ever runs. Fixing this is an architecture decision (never deserialize untrusted pickle, full stop), not a patch to pickle itself. For a penetration tester, the write-up value is in explaining precisely why this is worse than an average input-validation bug and why the compromised worker's existing access, not the bug itself, usually sets the real severity.
Structured elaboration
Why pickle deserialization is RCE by design, not by accident
Serialization turns a live in-memory object into a byte stream so it can be stored or sent elsewhere; deserialization turns those bytes back into a live object. JSON's deserializer only knows how to build passive data: strings, numbers, lists, and dictionaries. Pickle's deserializer is different: it is a small stack-based virtual machine, and its opcode set includes an instruction (historically GLOBAL, STACK_GLOBAL in newer protocol versions) that says "import this name," followed by REDUCE, which says "call the thing you just imported, using these arguments, to produce the object." That two-step "import a callable, then call it" behavior is not a bug pickle happens to have. It is how Python's __reduce__ protocol is supposed to let any object customize its own reconstruction, for example a class that needs to reopen a file handle or rebuild a C-extension object it can't just copy field-by-field.
The security problem is that pickle.loads() does not check who is asking it to import and call something. If the byte stream names any importable callable and supplies attacker-chosen arguments, the unpickler will call it, whether that callable is str (harmless) or a callable that spawns a shell (not harmless). Security researchers call a reachable dangerous callable like this a "gadget." There is no logic bug in the receiving application to find here: pickle.loads(data) on attacker-controlled data is the entire vulnerability. That is why unsafe deserialization is catalogued as its own weakness class, CWE-502 (Deserialization of Untrusted Data), and why OWASP folded insecure deserialization into the Top 10:2021 A08 category (Software and Data Integrity Failures): it is a class of finding, not a single library defect.
Why compromising the worker is rarely the end of the story
The scenario names an internal worker that consumes the pickle payload, and that detail matters more than it looks. A worker process almost always already carries whatever trust its role needs: a service credential to a database, network reachability to hosts on a private subnet, an IAM role, or a message-queue identity. RCE inside that process grants no new privilege; it hands the attacker everything the worker was already allowed to do. In kill-chain or MITRE ATT&CK terms (a public knowledge base that names common attacker tactics and techniques so findings can be compared across engagements), this bug covers Initial Access and Execution; whatever follows (credential access, lateral movement, discovery) is not a second vulnerability, it is the attacker simply exercising the worker's pre-existing reach. That is the reasoning to put in a report: the finding's blast radius is "everything that identity can already touch," which is why a "low exposure" internal worker can still justify a critical rating even when the entry point (a public-facing service accepting serialized objects) looks unremarkable on its own.
How a defender detects it
- Static review, before any code runs: Python's own
pickletools.dis() disassembles a pickle stream into its opcodes without executing them, so a reviewer or a CI check can flag a GLOBAL/STACK_GLOBAL opcode naming an unexpected module (os, subprocess, eval) in data that should only ever be application state. This is the safe way to inspect pickle traffic; never deserialize something to see what it does.
- Code review / static analysis (SAST): grep or lint for
pickle.load/pickle.loads (and pickle-backed helpers like joblib.load) reachable from any network- or user-influenced input, the same way you'd flag eval() on request data.
- Runtime telemetry: an endpoint detection and response (EDR) tool or basic process-lineage logging that flags a Python worker unexpectedly spawning a shell or child process right after consuming a queue message is a strong, if late, signal.
- Network layer: a web application firewall (WAF) or intrusion detection system (IDS) signature on suspicious opcode bytes is a weak backstop, easily varied away; treat it as defense in depth, never the primary control.
How a defender prevents it
- Don't deserialize untrusted pickle at all. This is the actual fix, not a mitigation: swap the wire format for one whose deserializer can only build declared, passive data.
- If a legacy system genuinely cannot drop pickle short term, treat origin as the control: sign the payload with a keyed-hash message authentication code (HMAC, a secret-keyed integrity check) and verify it before calling
pickle.loads, so a tampered or attacker-authored stream is rejected before it reaches the unpickler.
- Restrict what the unpickler may import by subclassing
pickle.Unpickler and overriding find_class to allow-list a small set of safe classes and reject everything else. This is fragile if the allow-list is loosened carelessly, so treat it as a compensating control, not a substitute for step 1.
- Run any deserialization that must remain on untrusted-adjacent data in a low-privilege, network-isolated worker, so that even a successful gadget call inherits as little reachability as possible. This directly shrinks the "chains into deeper access" problem described above.
Safe alternatives to pickle
| Format | Data model | Executes code on load? |
|---|
| Pickle | Arbitrary Python objects, including instructions to call constructors/functions | Yes, by design |
| JSON | Strings, numbers, booleans, lists, objects (dicts) | No |
| MessagePack | Same data model as JSON, compact binary encoding | No |
| Protocol Buffers / Avro / Thrift | Fields defined by an explicit, versioned schema | No |
All three alternatives share the property that matters: their deserializer only knows how to reconstruct data that matches a fixed, passive shape. There is no opcode in any of them that means "call an arbitrary function," so there is no gadget for an attacker to aim at.
Worked example
The safest way to see what pickle's design permits is to disassemble two streams, one plain data and one from an object with a custom __reduce__, without running anything dangerous. This uses only the standard library and a harmless callable (str) to show the mechanism:
python
import pickle
import pickletools
PROTOCOL = 4 # pinned so the opcode stream below is reproducible on any Python 3
# A normal object: just passive data (a dict of strings/numbers).
safe_obj = {"user": "alice", "count": 3}
safe_bytes = pickle.dumps(safe_obj, protocol=PROTOCOL)
print("--- opcode stream for a plain data object ---")
pickletools.dis(safe_bytes)
# An object that customizes __reduce__. The unpickler does not just
# copy fields back in: it calls whatever callable __reduce__ names,
# with whatever arguments __reduce__ supplies.
class Demo:
def __reduce__(self):
# (callable, args_tuple) -> unpickler will run callable(*args_tuple)
return (str, ("this text was produced by a real function call made during unpickling",))
demo_bytes = pickle.dumps(Demo(), protocol=PROTOCOL)
print("\n--- opcode stream for an object with a custom __reduce__ ---")
pickletools.dis(demo_bytes)
print("\n--- loading the second stream actually calls str(...) as part of rebuilding the object ---")
result = pickle.loads(demo_bytes)
print(result)
Output:
--- opcode stream for a plain data object ---
0: \x80 PROTO 4
2: \x95 FRAME 30
11: } EMPTY_DICT
12: \x94 MEMOIZE (as 0)
13: ( MARK
14: \x8c SHORT_BINUNICODE 'user'
20: \x94 MEMOIZE (as 1)
21: \x8c SHORT_BINUNICODE 'alice'
28: \x94 MEMOIZE (as 2)
29: \x8c SHORT_BINUNICODE 'count'
36: \x94 MEMOIZE (as 3)
37: K BININT1 3
39: u SETITEMS (MARK at 13)
40: . STOP
highest protocol among opcodes = 4
--- opcode stream for an object with a custom __reduce__ ---
0: \x80 PROTO 4
2: \x95 FRAME 96
11: \x8c SHORT_BINUNICODE 'builtins'
21: \x94 MEMOIZE (as 0)
22: \x8c SHORT_BINUNICODE 'str'
27: \x94 MEMOIZE (as 1)
28: \x93 STACK_GLOBAL
29: \x94 MEMOIZE (as 2)
30: \x8c SHORT_BINUNICODE 'this text was produced by a real function call made during unpickling'
101: \x94 MEMOIZE (as 3)
102: \x85 TUPLE1
103: \x94 MEMOIZE (as 4)
104: R REDUCE
105: \x94 MEMOIZE (as 5)
106: . STOP
highest protocol among opcodes = 4
--- loading the second stream actually calls str(...) as part of rebuilding the object ---
this text was produced by a real function call made during unpickling
Notice the difference: the plain-data stream never mentions a module or callable, it only pushes strings and numbers onto the stack. The __reduce__ stream contains a STACK_GLOBAL opcode naming builtins.str followed by REDUCE, meaning "import that name, then call it." The only thing separating this harmless example from a dangerous one is which name the stream imports; the calling mechanism is identical either way, which is exactly the design property that makes pickle unsafe on untrusted input. It is also literally the static-review technique from above: pickletools.dis() on a stream you have not yet loaded is how you catch a suspicious GLOBAL/STACK_GLOBAL reference before it can run.
Trade-offs & pitfalls
- A wrapper is not a fix. Wrapping a pickle payload in JSON, base64, or TLS does not change what happens once the inner bytes reach
pickle.loads. Transport security and input trust are different axes, and this bug lives entirely on the trust axis.
find_class allow-listing is a real control, but it is only as strong as its maintenance discipline. Every new type the application legitimately needs to pickle is a chance to over-broaden the allow-list back into unsafety; treat any addition to it as a security-relevant change requiring review.
- Migrating off pickle costs implicit convenience, not correctness. Pickle's appeal was serializing arbitrary Python objects (custom classes, numpy arrays, closures) with zero extra code. Schema-based formats force you to declare exactly what crosses the trust boundary, and that declaration is the fix itself: it is what removes the "arbitrary callable" primitive in the first place.
- Argue severity from reachability, not from the reflex to call every RCE a 9 or 10. Two identical pickle bugs in two different workers can carry very different real severities depending on what each worker's identity can already reach; reporting "RCE" without reasoning about the compromised principal's blast radius under-argues the finding, in either direction.