What BERT Actually Did: Bidirectional Transformers and Why PBN Content Can’t Just Chase Keywords

What BERT Actually Did Bidirectional Transformers and Why PBN Content Can't Just Chase Keywords

Somewhere around 2019, a particular style of SEO content quietly stopped working, and a lot of people never understood why. The pages that had ranked for years by repeating a target phrase the right number of times, in the right places, at the right density, started slipping. The operators who kept writing that way assumed Google had simply gotten stricter about keyword stuffing. They were missing the real story. Google hadn’t just turned up a dial on an old system — it had switched on a fundamentally different way of reading language, one that made the entire premise of keyword-density optimisation obsolete. That system was called BERT, and understanding what it actually did is the difference between writing PBN content that works in 2026 and writing content that looks like it time-travelled from 2015.

This is a genuinely technical topic, but it isn’t an academic one. The shift BERT represents is the single biggest reason that thin, keyword-stuffed, machine-spun PBN articles are a liability now — and the reason that the content advice we keep repeating, about writing things that read like a real resource, isn’t a vague nicety but a direct response to how Google’s systems read text. We touched on BERT-like understanding in passing in our piece on the decade of algorithm updates that moved Google from keywords to meaning, and enough people asked about it that it deserves its own proper explanation. So here it is, in plain English: what BERT is, what changed when it arrived, and what it means for anyone producing content at network scale.

The World Before BERT: Reading Left To Right

To understand what changed, you have to understand the limitation that came before. Older language models — and the search systems built on them — read text in one direction, like a person running their finger along a line word by word, left to right. Each word was interpreted largely on the basis of the words that came before it. This works passably for simple text, but it falls apart the moment meaning depends on what comes later in the sentence, or on the relationship between words that sit far apart.

The classic example Google itself used to explain the problem involves the query ‘can you get medicine for someone pharmacy.’ A human reads that instantly as: can I pick up a prescription on behalf of another person. The old system fixated on the individual keywords — medicine, pharmacy — and could miss the crucial relational word ‘for someone,’ which is the entire point of the question. Another Google example: ‘2019 brazil traveler to usa need a visa.’ The little word ‘to’ defines the whole meaning — it’s about a Brazilian travelling to the USA, not an American travelling to Brazil. Older systems routinely flattened that distinction, because the small connecting words that carry so much meaning — to, for, of, no — were treated as low-value noise to be largely ignored.

The consequence for content was an entire optimisation culture built around exact-match keywords and density. If the machine was essentially pattern-matching words rather than understanding sentences, then the way to rank was to make sure your target words appeared often enough and prominently enough. Keyword density tools, exact-match anchor obsession, awkwardly repeated phrases jammed into otherwise unnatural sentences — all of it was a rational response to a system that read words rather than meaning. And it worked, for a while.

What BERT Changed: Reading In Both Directions At Once

BERT stands for Bidirectional Encoder Representations from Transformers. Ignore the intimidating name for a moment and focus on one word in it: bidirectional. That’s the whole revolution in a single term.

Instead of reading a sentence left to right, BERT looks at every word in the context of all the other words around it — the ones before and the ones after, simultaneously. The meaning of a word is derived from its entire neighbourhood, in both directions, at the same time. When BERT encounters the word ‘stand’ in a sentence, it doesn’t guess based only on what preceded it; it weighs the words on both sides to work out whether this is ‘stand’ as in stand up, ‘stand’ as in a market stall, or ‘stand’ as in ‘I can’t stand it.’ Context resolves the ambiguity, and context, by definition, lives on both sides of a word.

How It Works, Without The Maths

The underlying machinery is the transformer — a model architecture in which every word in a passage is connected to every other word, and the strength of those connections is calculated dynamically based on how relevant each word is to each other word. This mechanism is called self-attention. In effect, the model is constantly asking, for every word, ‘which other words in this passage matter most for understanding what this one means here?’ The connecting words that older systems threw away — to, for, of, not — turn out to carry enormous weight, because they define relationships between the meaningful words. BERT was the first system Google deployed at search scale that could actually use them.

BERT learned this skill through a clever training trick called masked language modelling. During training, the model is shown billions of sentences with random words hidden, and it has to predict the missing word from the surrounding context on both sides. Do that across a vast enough corpus — BERT was pre-trained on enormous volumes of text including the whole of Wikipedia — and the model develops a genuinely deep statistical sense of how language fits together. It isn’t ‘understanding’ in a human sense, but functionally it captures meaning, relationships, and nuance far better than anything that came before it.

The Scale Of The Shift

When Google announced BERT in October 2019, its own search VP Pandu Nayak called it the biggest leap forward in the past five years, and one of the biggest in the entire history of search. At launch it affected around one in ten English-language queries — already a huge number — and it has since expanded to touch effectively all of them, across many languages. There’s an internal wrinkle worth knowing: according to testimony that emerged during the US antitrust trial, when BERT is used specifically for ranking documents it’s referred to internally as DeepRank, and it took over much of the work previously done by Google’s earlier machine-learning system, RankBrain. The point is that this wasn’t a bolt-on feature. It became part of the core machinery of how Google reads both queries and pages. It belongs to the same long arc Google has been on for a decade — moving, step by step, from matching strings to understanding meaning.

Where BERT Sits In Google’s Lineage

BERT didn’t arrive from nowhere and it didn’t stop with itself. It’s one link in a chain of systems that together moved Google away from keyword matching, and seeing the whole chain makes the direction of travel unmistakable.

  • RankBrain (2015): Google’s first machine-learning ranking system, built to handle the roughly 15% of queries it had never seen before by relating them to known patterns. The first crack in pure keyword matching.
  • Neural Matching (2018): A system for matching queries to documents at the level of concepts rather than words — finding relevant pages even when not a single keyword overlaps, by comparing meaning in a multidimensional space.
  • BERT (2019): Bidirectional language understanding, finally letting Google read the relationships between words in a sentence and grasp the concepts behind them. Used for ranking, it’s DeepRank.
  • MUM (2021): The Multitask Unified Model, reported to be far more powerful than BERT and multimodal — able to work across text, images and languages together. The bridge toward today’s generative systems.

Follow that line forward and you arrive exactly where search is in 2026: AI Overviews and AI Mode, which are only possible because of the language-understanding foundation that BERT and its successors laid. The generative layer everyone is talking about now is built on top of this lineage, not separate from it. Understanding BERT isn’t history for its own sake — it’s understanding the floor that everything current stands on.

Why This Killed Keyword Density For Good

Here’s where the technical story becomes a content story. Once Google could read meaning rather than count words, the entire logic of keyword-density optimisation collapsed — not because Google decided to punish it, but because it simply stopped being the thing the system was looking at.

Think it through. If the machine understands that ‘how to fix a slow laptop,’ ‘why is my computer running so slowly,’ and ‘speed up an old PC’ are all the same underlying intent, then a page doesn’t need to contain the exact phrase a user typed to be the best answer to it. Conversely, a page that repeats an exact phrase twenty times but demonstrates no real grasp of the topic has nothing extra to offer a system that understood the topic in the first place. Keyword repetition went from being the signal to being, at best, irrelevant noise — and at worst, a marker of low-effort content, because genuine writing by a knowledgeable human rarely repeats the same exact phrase unnaturally.

This is why the modern advice is to write about a topic comprehensively rather than to target a keyword repeatedly. A system that reads bidirectionally rewards content that covers a subject with genuine depth and natural language, because that’s what actually demonstrates relevance to the cluster of related questions a topic generates. Density tools became relics. The question stopped being ‘how many times does my keyword appear’ and became ‘does this page genuinely and thoroughly address the thing people are actually asking about.’

What This Means For PBN Content Specifically

For network operators, the BERT shift has consequences that go well beyond ‘write naturally.’ It reshapes what makes a PBN site believable to the systems evaluating it, and it raises the floor on what passable content looks like.

Stuffed, Spun, Or Thin Content Is Now A Tell

Content built for the pre-BERT world — exact-match phrases hammered in at a target density, spun synonyms, thin articles that orbit a keyword without ever saying anything — doesn’t just fail to help. It actively signals low quality to a system specifically built to recognise meaning and coherence. A bidirectional model reading a spun article sees the seams: the unnatural collocations, the sentences that don’t quite cohere, the absence of the relational structure that real explanation has. This is the technical reason behind everything we say about writing PBN content that reads like a genuine resource. ‘Doesn’t scream PBN’ isn’t an aesthetic judgement — it’s a description of content that survives a reading by a system that understands language.

Topical Coherence Is Read, Not Guessed

Because BERT and its successors understand concepts, they can tell whether a site’s content genuinely hangs together as a coherent treatment of a topic, or whether it’s a grab-bag of keyword-targeted pages with no real subject. A former gardening domain rebuilt with articles that actually understand gardening reads as coherent; the same domain stuffed with thin posts about unrelated commercial topics reads as incoherent — and the model can tell the difference far better than the old word-matching systems could. This is why matching a domain to a topic it can genuinely cover matters, a thread that runs through both choosing domains that fit the topic they will cover and the trust and topical signals that actually matter in 2026. The topical fit of a network site is now something Google’s language systems actively read, not a box you tick.

Natural Linking Patterns Read As Natural

The same understanding extends to how content references other things. A page that links out to genuine, relevant authorities in the course of actually discussing a subject looks like real writing; a page whose only outbound link is a single keyword-anchored shot at a commercial target looks like exactly what it is. Bidirectional understanding makes the surrounding context of a link legible — what the page is about, whether the link belongs there, whether the anchor text fits the meaning. This is the deeper reason behind sensible the outbound linking patterns of a real, topically coherent site: in a world where the machine reads context, links have to sit in believable context.

The Practical Takeaway

You don’t need to be able to explain self-attention or masked language modelling to act on what BERT means. The practical instruction is simple, even if the technology behind it isn’t: stop thinking about keywords as ingredients to be measured, and start thinking about topics to be genuinely covered. Write each PBN article the way a knowledgeable person writes about something they actually understand — addressing the real questions around a subject, in natural language, with the connecting words and relationships intact, linking out where a real writer would. That is what a bidirectional system is built to reward, and no amount of density tuning substitutes for it.

The operators still chasing keyword density in 2026 are optimising for a search engine that hasn’t existed for the better part of a decade. The systems reading their content understand language now — imperfectly, statistically, but well enough that meaning beats repetition every time. The content that wins is the content that has something to say and says it clearly. Everything else is writing for a machine that was switched off in 2019.

The day-to-day standards that follow from all of this are laid out across the complete PBN guide for 2026, and if you want to talk through how to lift your network’s content to the level a modern language model rewards, free personalised SEO and PBN advice is available.