How the RC4 Algorithm Works and Why It Was Banned

RC4 is a stream cipher designed by Ron Rivest in 1987 that once dominated internet encryption, protecting everything from web traffic to wireless networks. Its appeal was straightforward: the algorithm is tiny, runs fast in software, and can work with variable-length keys. For roughly two decades it was the default cipher in SSL/TLS connections and the backbone of WEP and early WPA wireless security. That era is over. A steady accumulation of cryptographic attacks exposed serious statistical biases in RC4’s output, and in 2015 the Internet Engineering Task Force formally prohibited its use in TLS.1IETF. RFC 7465: Prohibiting RC4 Cipher Suites Understanding how RC4 works and why it failed offers a useful window into the lifecycle of a cryptographic primitive.

How RC4 Works

RC4 belongs to the stream cipher family, meaning it generates a long sequence of pseudo-random bytes (called a keystream) and combines each byte with a byte of plaintext using a simple XOR operation. The recipient, who shares the same secret key, generates the identical keystream and XORs it with the ciphertext to recover the original message. The entire algorithm revolves around a 256-byte array, usually called S, and two index pointers. There are only two phases: key scheduling and pseudo-random generation.

The Key Scheduling Algorithm (KSA) initializes the S array so that it contains every value from 0 to 255 in some scrambled order determined by the secret key. It starts with the identity permutation (0, 1, 2, … 255) and then performs 256 swap operations, each driven by a byte of the key. Once the KSA finishes, S holds a permutation that encodes the secret key’s influence. The key can be anywhere from 1 to 256 bytes long, which gave RC4 unusual flexibility compared to block ciphers that demanded fixed-size keys.

The Pseudo-Random Generation Algorithm (PRGA) then produces one keystream byte at a time. Each step increments one index pointer, uses the S array to compute a new position for the second pointer, swaps two entries in S, and outputs a byte looked up from S at a position derived from the sum of the two swapped values.2Journal of Mathematical Cryptology. A complete characterization of the evolution of RC4 pseudo random generation algorithm Because S is continually shuffled by the swaps, the keystream should look random to anyone who does not know the initial permutation. The entire generation loop is just a handful of operations per byte, which is why RC4 ran circles around block ciphers on the hardware of the late 1980s and 1990s.

Why RC4 Became So Popular

Speed was the primary selling point. RC4 requires only a few table lookups and swaps per output byte, with no multiplication or complex arithmetic. On early processors this made it dramatically faster than DES or Triple-DES in software. The algorithm’s simplicity also meant it could be implemented in very few lines of code, which mattered for embedded devices, smart cards, and firmware where memory was scarce. Hardware implementations on FPGAs have demonstrated that the cipher can be pipelined efficiently to achieve high throughput even in constrained logic environments.3Elsevier / Procedia Computer Science. Efficient FPGA Implementation of the RC4 Stream Cipher using Block RAM and Pipelining

Adoption snowballed because of network effects. When Netscape chose RC4 for SSL in the mid-1990s, every major browser and web server followed. The IEEE chose it for WEP, the original Wi-Fi encryption standard. Microsoft used it inside protocols for Office document encryption and Windows NTLM authentication. By the early 2000s, an enormous share of encrypted internet traffic was running through RC4. This ubiquity created inertia: even after researchers began publishing worrying results, operators were reluctant to switch because RC4 “worked everywhere.”

The Leaked Source Code Story

Rivest designed RC4 as a trade secret for RSA Security, not as an open standard. The algorithm was never formally published. In 1994, an anonymous post to a mailing list leaked what appeared to be the complete source code. The community quickly confirmed that it produced output identical to the commercial RC4 implementation. Because the name “RC4” remained trademarked, open-source projects began referring to compatible implementations as “ARC4” or “ARCFOUR.” This leak had an unusual side effect: it made the algorithm even more widely deployed, since anyone could now implement it freely without licensing fees, but it also opened it to much broader cryptanalytic scrutiny.

Statistical Biases in the Keystream

A perfect stream cipher would produce output indistinguishable from truly random bytes. RC4 does not meet that bar, and the gaps are measurable. The problems start in the very first bytes of output. The second byte of the RC4 keystream is twice as likely to be zero as it should be. This bias, discovered by Itsik Mantin and Adi Shamir in 2001, was the first widely recognized flaw in the PRGA output. The practical consequence is that if an attacker can collect many ciphertexts encrypted under different keys but with the same or predictable structure in the first few plaintext bytes, those early keystream biases leak information about the plaintext.

Subsequent research found that the biases are not limited to the first few bytes. Work on the so-called “Mantin biases” showed that predictable statistical patterns extend further into the keystream and can be exploited to recover plaintext from RC4-encrypted traffic. In a TLS context, where parts of the plaintext (like HTTP headers) follow a predictable structure, these biases allow an attacker to target unknown bytes such as session cookies or passwords that sit near known plaintext sequences.4PubMed Central. Analysing and exploiting the Mantin biases in RC4 The attack works by collecting a large number of ciphertexts (on the order of millions) where the same secret value is encrypted repeatedly, which is exactly what happens when a browser sends the same cookie with every request to a server.

Beyond the Mantin biases, researchers also proved that additional open biases in the RC4 keystream generator are directly relevant to TLS ciphertext recovery.5SpringerLink / Designs, Codes and Cryptography. Proving TLS-attack related open biases of RC4 Each new published bias narrowed the number of ciphertexts an attacker needed to collect and increased the range of plaintext bytes that could be recovered. By 2015, the situation was clear: RC4’s output was biased enough that real-world decryption of secrets was feasible given traffic volumes that a patient attacker could realistically capture.

The WEP Catastrophe

The most public failure of RC4 was in WEP, the Wired Equivalent Privacy protocol used to protect early Wi-Fi networks. WEP used RC4 with a 24-bit initialization vector (IV) prepended to the secret key. Because the IV was short, it inevitably repeated after a few thousand packets, and worse, the way the IV and key were concatenated created related-key conditions that leaked information about the secret key itself. Researchers showed in 2001 that by passively collecting traffic, an attacker could recover the entire WEP key in minutes. Tools to automate WEP cracking became freely available and trivially easy to use, effectively ending WEP’s credibility. The IEEE replaced WEP with WPA and later WPA2, both of which abandoned RC4 in favor of AES-based encryption.

It is worth noting that WEP’s failure was not entirely RC4’s fault. The protocol’s designers made serious mistakes in how they used the cipher: the short IV, the related-key construction, and the lack of a proper message authentication code all compounded RC4’s inherent weaknesses. A well-designed protocol can sometimes compensate for a cipher with minor statistical flaws. WEP’s protocol design, however, amplified every weakness RC4 had.

Formal Prohibition in TLS

After years of accumulating attacks, the IETF published RFC 7465 in February 2015, which flatly prohibits the use of RC4 cipher suites in TLS. The language is unambiguous: TLS clients must not include RC4 cipher suites in their connection requests, and TLS servers must not select an RC4 cipher suite even if the client offers one.6IETF. RFC 7465: Prohibiting RC4 Cipher Suites This prohibition applies to all versions of TLS, updating the specifications for TLS 1.0, 1.1, and 1.2. In practice, major browser vendors had already begun disabling RC4 by default before the RFC was published, but the formal standard ensured that any compliant implementation would refuse to negotiate it.

TLS 1.3, finalized in 2018, went further and removed the concept of individually negotiable ciphers like RC4 entirely. Its cipher suite list includes only modern AEAD (authenticated encryption with associated data) constructions based on AES-GCM and ChaCha20-Poly1305. There is no mechanism in TLS 1.3 to use RC4 even if both sides wanted to.

Side-Channel Attacks Add Another Dimension

Even when implemented correctly at the protocol level, RC4 can leak secrets through physical side channels. Research has shown that when a chip executes the RC4 algorithm in software, its electromagnetic emissions during the swap-and-lookup operations can be captured and analyzed using machine-learning classifiers to recover information about the internal state.7The Journal of China Universities of Posts and Telecommunications. Electromagnetic side-channel attack based on PSO directed acyclic graph SVM This category of attack does not exploit any mathematical weakness in the cipher itself; instead, it exploits the physical reality that different operations produce subtly different electromagnetic signatures. Side-channel attacks are a concern for any cipher, but RC4’s simple, repetitive structure (the same small loop executed millions of times) can make pattern extraction easier for an attacker with the right equipment.

In practice, electromagnetic side-channel attacks require physical proximity to the target device and specialized measurement equipment, so they are primarily a concern in scenarios involving hardware tokens, smart cards, or embedded controllers where an attacker might have physical access. For ordinary internet traffic, the statistical biases in the keystream are a far more practical threat.

Attempts to Fix RC4

Rather than abandoning RC4 entirely, some researchers have tried to patch its weaknesses by modifying the Key Scheduling Algorithm. One approach replaces the identity permutation (the starting point where S simply contains 0 through 255 in order) with random initial values and also changes how the key influences the permutation process. Testing of this modified KSA showed improvements in randomness and resistance to the known statistical biases, with encryption speed comparable to the original.8Iraqi Journal of Science. A Modified Key Scheduling Algorithm for RC4 Other proposed fixes include discarding the first 256 or even the first 3,072 bytes of keystream output before using any of it for encryption, which sidesteps the early-byte biases at the cost of a slightly slower startup.

None of these variants have gained significant adoption. The cryptographic community has largely concluded that patching RC4 is a losing game. Each fix addresses known biases, but the underlying structure of the PRGA continues to produce output that falls short of what modern analysis demands. Meanwhile, well-analyzed alternatives like ChaCha20 exist that are comparably fast in software, have cleaner security proofs, and were designed from the ground up with modern cryptanalytic techniques in mind. The practical advice from standards bodies and security researchers is to migrate away from RC4, not to try to fix it.

Where RC4 Still Lingers

Despite its formal deprecation, RC4 has not vanished from the wild. Legacy systems are the main culprit. Older industrial control systems, point-of-sale terminals, and embedded devices that were designed a decade or more ago may still use RC4 internally and cannot easily be updated. Some proprietary protocols in enterprise environments also retain RC4 because the cost and risk of replacing the protocol outweigh the perceived threat. In certain countries, government or military systems built on older standards may still rely on RC4-protected communications.

There is also a non-trivial amount of RC4 usage in contexts that are not strictly cryptographic. Some hash-table implementations, random-number generators for simulations, and game engines use RC4 or RC4-like constructions as fast pseudo-random byte generators where cryptographic security is irrelevant. In these applications, the algorithm’s speed and simplicity are the only things that matter, and its statistical imperfections are far below the threshold of concern. This creates an odd afterlife: the algorithm that was once the internet’s most widely deployed cipher now finds more legitimate use as a non-security utility than as actual encryption.

Lessons from RC4’s Rise and Fall

RC4’s trajectory illustrates a pattern that keeps recurring in cryptography. An algorithm is adopted because it is fast, simple, and “good enough” by the standards of its era. It becomes deeply embedded in infrastructure. Researchers begin finding cracks, but the cracks seem theoretical at first and the cost of switching is high. Gradually, the theoretical attacks become practical, and the community enters a painful transition period where the old cipher is known to be weak but has not yet been fully removed. The migration from RC4 in TLS took roughly a decade from the first serious published attacks to the formal prohibition in RFC 7465, and legacy deployments persisted for years beyond that.

One specific lesson is that stream ciphers live or die on the statistical quality of their keystream. Unlike block ciphers, which encrypt fixed-size chunks and can lean on well-studied modes of operation, a stream cipher’s entire security reduces to whether its output looks random. Any detectable pattern, even a subtle one, can be amplified when an attacker collects enough ciphertext. RC4’s biases were small in absolute terms, but the sheer volume of data encrypted under it on the internet gave attackers more than enough material to exploit them. Modern stream ciphers like ChaCha20 were designed with this vulnerability class in mind, and their security arguments rely on much stronger guarantees about output randomness.

RC4 in Education and Research

RC4 remains one of the most commonly taught ciphers in undergraduate cryptography courses, precisely because it is so easy to understand. A student can implement the full algorithm in about 20 lines of code in most programming languages, which makes it an excellent teaching tool for concepts like keystream generation, XOR-based encryption, and the relationship between key scheduling and output quality. It is also a convenient target for teaching cryptanalysis: the biases are well-documented, the attacks have been published in detail, and students can reproduce the early-byte biases with a modest amount of computation.

In research, RC4 continues to serve as a benchmark and reference point. Papers proposing new stream ciphers often compare their constructions against RC4 in terms of speed, memory usage, and resistance to known attack classes. The extensive literature on RC4’s weaknesses also provides a rich body of techniques that cryptanalysts apply to other ciphers. In a sense, RC4’s flaws have been more valuable to the field than its strengths ever were: they taught a generation of researchers what to look for and what to avoid when designing new primitives.