A homograph attack exploits the fact that many characters from different writing systems look identical or nearly identical to the human eye. An attacker swaps one or more characters in a trusted name with a visually similar substitute from another alphabet, creating a fake domain, filename, or username that appears legitimate on screen but actually points somewhere malicious. The trick is effective because the underlying character codes are different even when the visual rendering is indistinguishable, and it works across more surfaces than most people realize.
How Characters Become Weapons
The modern computing world supports well over a hundred thousand characters through the Unicode standard, spanning Latin, Cyrillic, Greek, Armenian, Cherokee, and dozens of other scripts. Many of these scripts evolved independently but converged on similar-looking letterforms. The Cyrillic letter “а” (U+0430) and the Latin letter “a” (U+0061) render as the same glyph in most fonts. The Greek omicron “ο” is a visual twin of the Latin “o.” These lookalike pairs are called homoglyphs, and they exist by the thousands across Unicode’s full range.
An attacker who registers the domain “аpple.com” using a Cyrillic “а” instead of a Latin “a” gets a web address that looks pixel-for-pixel identical to the real apple.com in most browsers and email clients. Behind the scenes, the two addresses resolve to completely different servers. The fake domain can host a convincing login page, serve malware, or intercept credentials. Adversaries use this approach to obfuscate domain names, filenames, and process names so they appear visually identical to their legitimate counterparts.
1ICT Express. Siamese neural network architecture for homoglyph attacks detectionThe deception does not require swapping every character. Replacing a single letter in a long domain is often enough, especially in a context where the reader is skimming quickly. A URL like “micros0ft.com” using a zero for the “o” is a crude version; a URL using a Cyrillic “о” for the Latin “o” is far harder to catch because the substitution is invisible in most typefaces.
Internationalized Domain Names and the Browser Problem
The attack gained broad attention through internationalized domain names, or IDNs. IDNs were introduced so that people around the world could register web addresses in their own scripts, which is a reasonable and necessary feature. A Russian business should be able to have a .рф domain in Cyrillic. A Chinese company should be able to use Chinese characters in its address. The problem is that this same system allows someone to register a domain that mixes scripts in a way designed to deceive.
When you type a non-ASCII domain into a browser, it gets converted behind the scenes to a format called Punycode, which encodes Unicode characters as ASCII. The domain “аpple.com” with a Cyrillic “а” actually resolves as “xn--pple-43d.com” in Punycode. Browsers had to decide whether to show users the pretty Unicode version or the ugly-but-honest Punycode version. For years, many browsers defaulted to showing the Unicode rendering, which made the attack trivially easy.
Major browsers eventually changed their approach. Chrome, Firefox, Safari, and Edge now display the Punycode version whenever a domain mixes characters from multiple scripts or uses characters from scripts associated with homograph abuse. If you visit a domain that combines Latin and Cyrillic letters, your browser will show you the “xn--” Punycode string instead of the human-readable version. This is an effective defense for mixed-script attacks, but it has limits. A domain written entirely in Cyrillic can still look convincing if the target brand happens to be spellable using only Cyrillic lookalikes, and some older or niche browsers may not enforce the same rules.
Beyond Domain Names
Homograph attacks are most commonly discussed in the context of phishing URLs, but the same principle applies wherever computers display text that humans trust at face value. The attack surface extends to filenames, process names, usernames on social platforms, package names in software repositories, and even cryptocurrency wallet addresses.
Filenames are a particularly underappreciated vector. An attacker who sends you an email attachment named “invoice.pdf” might actually be sending “іnvoice.pdf” with a Cyrillic “і” replacing the Latin “i.” On your file system, the two files would coexist as separate entities because their underlying byte sequences differ. Your operating system treats them as distinct files, but your eyes see the same name. Process names in a task manager can be spoofed the same way: a piece of malware running as “svchоst.exe” with a Cyrillic “о” could sit right alongside the legitimate “svchost.exe” and go unnoticed during a casual inspection.
2ICT Express. Siamese neural network architecture for homoglyph attacks detectionIn the software supply chain, package managers like npm, PyPI, and RubyGems identify packages by name. A developer who types “import requеsts” with a Cyrillic “е” instead of “import requests” could pull in a malicious package if an attacker has registered that lookalike name. The code runs the same way, the import statement looks the same on screen, and the error only becomes apparent if someone inspects the raw bytes.
Why Your Eyes Are Not Enough
The core difficulty with homograph attacks is that they exploit the gap between what humans see and what machines process. You were trained from childhood to read by shape recognition, not by analyzing the Unicode code points behind each glyph. When you see “paypal.com” in your browser’s address bar, you pattern-match against a familiar word. You do not and cannot distinguish whether each letter is drawn from the Latin, Cyrillic, or Greek character set.
Font rendering makes this worse. Many popular web fonts are designed to display characters from different scripts with harmonious, consistent styling. A well-designed font makes Cyrillic and Latin letters flow together seamlessly, which is exactly what you want for bilingual documents and exactly what you do not want when trying to spot a homograph attack. Some character pairs are distinguishable in one font and identical in another. The Cyrillic “с” and Latin “c” are indistinguishable in most sans-serif fonts but may show subtle differences in certain serif typefaces. Relying on font-level differences is not a viable defense because you cannot control which font a website or email client uses to render its text.
Context clues sometimes help. If you are visiting a login page and the URL bar shows “xn--” followed by gibberish, that is a strong signal. But many people never look at the URL bar at all, especially on mobile devices where the address is often truncated or hidden behind a tap. And when the attack uses characters from a single script that happens to spell out the target brand, even the Punycode defense falls short.
Automated Detection Approaches
Because human vision is unreliable for this task, researchers have developed automated methods to flag homograph attacks. Early approaches focused on comparing the visual similarity of individual characters using metrics like the Structural Similarity Index, which measures how closely two images resemble each other. A system can render each character in a suspect domain as a small image and compare it against characters in the expected legitimate domain, flagging pairs that score above a similarity threshold.
More recent work has combined character-level visual similarity with statistical patterns in how characters are used together. One research group found that adding features from a language-modeling technique called N-grams on top of the visual similarity features improved classification accuracy by about 1.8 percent and reduced the false-positive rate by roughly 2.2 percent compared to using visual similarity alone.
3arXiv. Improving Homograph Attack ClassificationThose numbers might sound small, but in a system scanning millions of domain lookups per day, a two-percent reduction in false positives means tens of thousands fewer legitimate domains incorrectly flagged as suspicious. The improvement matters at scale. The general trend in this research is toward combining visual features with contextual and linguistic features rather than relying on any single signal.
Neural network approaches have also shown promise. Siamese neural networks, which are designed to learn whether two inputs are the “same” or “different,” can be trained on pairs of legitimate and spoofed domains to learn the subtle byte-level differences that visual inspection misses.
4ICT Express. Siamese neural network architecture for homoglyph attacks detectionHow Homograph Tricks Appear in Phishing Email
Email remains the primary delivery vehicle for homograph attacks. A phishing email that appears to come from “support@аmazon.com” with a Cyrillic “а” is far more convincing than one from a random-looking domain, and many email clients display only the sender’s display name rather than the full address. Even when the full address is visible, the homograph substitution makes it look correct.
Modern spam and phishing detection systems are increasingly aware of character obfuscation as a deliberate evasion tactic. Recent research on adaptive email defense systems has explored using adversarial self-evolution, where one component of the system generates novel evasion tactics, including character obfuscation, while another component learns to detect them.
5arXiv. EvoMail: Self-Evolving Cognitive Agents for Adaptive Spam and Phishing Email DefenseThe arms race here is real. Attackers do not just swap single characters anymore. They combine homoglyph substitutions with other techniques: inserting zero-width Unicode characters that are invisible but change the string’s identity, using right-to-left override characters to reverse the apparent order of text, or mixing in combining diacritical marks that render as invisible on certain platforms. Each layer of obfuscation makes the attack harder for both humans and simple rule-based filters to catch.
Practical Steps You Can Take
Knowing how the attack works gives you a meaningful edge, but awareness alone is not a complete defense. Here are concrete things you can do:
- Type critical URLs yourself. Instead of clicking a link in an email or message to reach your bank, email provider, or shopping account, type the address directly into your browser or use a saved bookmark. This bypasses the homograph entirely because you are entering the characters you intend.
- Watch for Punycode. If your browser’s address bar ever shows a URL starting with “xn--” when you expected a normal domain name, that is a strong warning sign. The site is using internationalized characters, and if you did not expect that, leave.
- Keep browsers updated. Browser vendors continuously refine their IDN display policies. An up-to-date browser is more likely to show you the Punycode version of a suspicious domain than an outdated one.
- Use a password manager. Password managers match credentials to domains at the byte level, not the visual level. If you visit a homograph domain, your password manager will not auto-fill your credentials because the domain does not match what it has stored. This is one of the most reliable passive defenses available.
- Inspect sender addresses in email. Many email clients hide the actual sender address behind a display name. Expand the full header or hover over the sender to see the real address. If it contains unexpected characters or looks subtly off, treat the message with suspicion.
For organizations, technical controls add another layer. DNS-level filtering can block known homograph domains. Email authentication protocols like DMARC, DKIM, and SPF help receiving servers verify whether a message actually came from the domain it claims. These do not specifically target homograph attacks, but they make it harder for an attacker to impersonate a domain that has these protections configured.
The Broader Unicode Security Landscape
Homograph attacks are one piece of a larger category of Unicode-based security problems. The Unicode Consortium itself maintains a technical report, Unicode Technical Report #36, specifically about Unicode security considerations. It catalogs confusable character pairs and provides data tables that software developers can use to build detection systems. The “confusables.txt” file maintained by the Consortium lists thousands of character pairs that are visually similar across scripts.
Other Unicode-based attacks work differently but exploit related trust assumptions. Bidirectional text control characters, for instance, can make a filename that ends in “exe” appear to end in “doc” by reversing the display order of the final characters. Zero-width joiners and non-joiners can create strings that look identical on screen but differ in their byte representation, breaking exact-match security checks. Combining characters can stack invisible diacritical marks onto a letter, producing a visually identical character with a different identity.
These attacks share a common root: the tension between Unicode’s legitimate goal of supporting every human writing system and the security assumption that what you see on screen accurately represents what the computer is processing. That tension is not going away. As Unicode continues to grow and software continues to render text in more sophisticated ways, the gap between visual appearance and underlying data will remain a fertile ground for deception. The defenses are getting better, but so are the attacks, and the fundamental asymmetry, that humans read shapes while computers read code points, is baked into how writing and computing work.
Cryptocurrency and Wallet Address Spoofing
One area where homograph-style attacks cause direct financial harm is cryptocurrency. Wallet addresses are long, seemingly random strings of characters. Most people copy and paste them rather than typing them manually. An attacker who can substitute a lookalike address, whether through clipboard malware or through a homograph trick on a web page displaying the address, can redirect funds to their own wallet with no practical way to reverse the transaction.
The Ethereum Name Service and similar blockchain naming systems were designed partly to solve this problem by letting people use human-readable names instead of raw addresses. But these naming systems introduce their own homograph risks. If you can register “vіtalik.eth” with a Cyrillic “і,” someone sending funds to what they think is a well-known name could end up sending them to an attacker’s wallet instead. The normalization rules these systems use to prevent such registrations are an active area of research and have proven to contain inconsistencies that attackers can exploit.
The stakes in cryptocurrency are unusually high because transactions are irreversible by design. There is no bank to call, no chargeback to file. A successful homograph attack on a wallet address or naming service can result in permanent, unrecoverable loss. This makes the cryptocurrency space a particularly aggressive testing ground for both homograph attacks and the defenses built to stop them.

