Third-party analysis of Google's March 2026 Core Update places the average visibility gain for sites with original data at roughly 22%. The number is the headline. The framework underneath it is what determines whether a site can replicate the gain or merely cite it. Information gain is not a single thing — it operates as three distinct layers, and most "we'll add original research" recovery plans fail because they confuse the layers.
This piece separates the three layers, examines what the 22% figure actually measures, and lists the patterns that won, the patterns that look like gain but aren't, and what the framework means for editorial process between now and Q3 2026.
What "Information Gain" Actually Means in Google's 2026 Framing
Information gain — as the term is used in post-update commentary — collapses three separable layers that need to be priced independently. Conflating them is why generic "publish original research" advice produces poor outcomes.
Layer 1: Original data. Numbers, measurements, observations, primary records that did not exist on the public internet before this article published. A small reader survey with 200 respondents. A scrape of a public dataset cleaned and structured into a comparison. A FOIA disclosure summarized into a table. A field audit. The unit of value is a row of data the reader could not get elsewhere.
Layer 2: Original analysis. Combining existing data in a way no one else has combined it. Cross-referencing two public datasets to surface a pattern. Building a regression on price data already published. Calculating a derived metric from raw numbers nobody has bothered to derive. The data inputs are not new; the synthesis is.
Layer 3: Original framing. A novel way to think about a topic — a category, a distinction, a heuristic — applied to existing knowledge. No new data, no new synthesis, but a structuring claim that reframes what readers do with the same inputs.
The three layers are not interchangeable. Layer 1 is the most expensive to produce and most defensible against replication. Layer 3 is the cheapest but most easily copied once published. Sites that recovered visibility tend to anchor at Layer 1 with Layer 2 and Layer 3 layered on top — not Layer 3 alone re-skinning aggregate content.
The 22% Number — What It Measures and Doesn't
The 22% gain figure originates in third-party visibility tracking that compares pre-update and post-update search visibility for sites tagged as having original data. It is an average across heterogeneous sites and verticals. It is not a guarantee, a floor, or a ceiling. It does not isolate original data as the causal variable — sites with original data also tend to have stronger E-E-A-T signals, named editorial teams, and slower publishing velocity, all of which co-vary with the gain.
What the number does establish is direction. Before the update, sites without original data could rank competitively against sites with original data on many query types. After, the gap widened materially across the sample tracked. The 22% is the magnitude of that gap-widening, averaged.
What the number does not establish: that adding any original data to a previously aggregator-style site will yield 22% gains. The compositional effect — data inserted into a site whose other signals remain aggregator-grade — is much smaller than the average. Sites that gained were not sites that bolted data onto existing pipelines. They were sites whose editorial stance was data-first, with the investment dating to before the update.
Where Original Data Comes From at Reasonable Cost
The objection to original data investment is usually cost. The objection has merit at the highest end — funded, peer-reviewed studies are out of reach for most publishers. But several lower-cost methods produce data Google's framework recognizes as Layer 1, and they are within reach of a single editor with discipline.
Reader surveys at small N. A 150–300 respondent survey of a defined niche audience produces tabulated results no aggregator has. Survey platforms cost under $200 at the volume needed; time investment is one to two weeks for design, distribution, and analysis. Methodology constraint is honest reporting of N, sampling method, and self-selection bias.
Internal product or platform data. Sites operating a tool, calculator, directory, or community accumulate behavioral data. Aggregating it into reportable metrics — most-searched queries, most-saved items, average workflow completion time — produces a report no competitor can publish without owning a comparable platform.
Public datasets cleaned and restructured. Government datasets, regulator filings, and exchange data are public but unusable in raw form. Cleaning and reformatting into a comparison table or derived metric produces something that did not exist in usable form before.
FOIA and public records requests. Topic-specific records requests — agency disclosures, court records, regulatory correspondence — return data no aggregator has summarized. Cost is the time to draft and follow up.
Primary interviews with named sources. A 30-minute call with a domain practitioner, transcribed and quoted, produces editorial content with verifiable provenance. Reach rates are higher than expected when the request is specific.
Structured experiments. Test a vendor or category claim — sample three products, run an A/B, document method and result. The experiment needs to be honest and reproducible, not large.
Manual field audits. Sample N items from a population, examine each on defined criteria, tabulate. Labor-intensive but accessible without funding.
Three Information Gain Patterns That Won After March 2026
Across visibility-gainer cases reviewed in third-party post-update analyses, three patterns recur. They are not the only winning patterns, but they are the most replicable.
Pattern 1: Anchored data plus interpretive framing. A core dataset — survey, audit, scrape — published as a primary artifact, with a sequence of secondary articles that each interpret one slice of the data. The data is the asset; the interpretation articles compound on it. The site doesn't republish the data per article; it cites the same primary source from multiple analytical angles. Visibility gains concentrate in the analytical articles because they each carry a unique synthesis on a defensible source.
Pattern 2: Methodology disclosure as differentiator. Articles that document method — how data was collected, what was excluded, what the limits are — outperform articles that present data without method. The transparency is read as a quality signal both by Google's framework and by readers who have learned to discount unmarked numbers. The cost is one extra section per article; the visibility return is disproportionate.
Pattern 3: Updated data with timestamped revisions. Articles that refresh data on a documented cadence — quarterly, monthly — and surface the revision history outperform static "definitive guide" articles of equivalent content depth. The revision discipline signals an active editorial process rather than a published-and-forgotten asset. It also means the article keeps accumulating freshness signals over time.
Three Patterns That Look Like Gain But Aren't
Several content patterns superficially resemble information gain and produce no measurable benefit, or sometimes negative impact. Recognizing them prevents wasted editorial effort.
Anti-pattern 1: Rephrasing competitors' data. Taking a chart from another publisher, restating the numbers in prose, and citing the source. The information is not gained — it is relocated. The original publisher receives the citation benefit; the rephrasing site receives no Layer 1 credit because no Layer 1 work was done. Common in roundup and "ultimate guide" content.
Anti-pattern 2: Old industry reports cited as fresh insight. Statistics from a 2022 McKinsey or Statista report cited in 2026 content as if current. The data is real but stale; the article inherits no freshness or originality benefit. When the source report itself is paywalled and only cited via secondhand summaries, the citation often fails verification entirely.
Anti-pattern 3: AI-generated "data" presented as research. Tables and figures produced by asking a model to estimate, distribute, or simulate values. The artifact looks like data; it is not data. Verification fails on inspection, and the lexical patterns that produce it cluster with the unedited-AI signals March 2026 actively penalized.
What This Means for Editorial Process
The operational read for sites planning content investment between now and Q3 2026 is that data acquisition has moved upstream of writing. Editorial calendars that lead with topic queues and trail with research now produce worse outcomes than calendars that lead with data assets and derive topic queues from them.
For sites running on AI-assisted production, the productive investment is not better prompts. It is acquiring or collecting Layer 1 data the assistant can analyze, rather than asking the assistant to invent or rephrase. The assistant becomes a synthesis layer over real inputs, which is a defensible workflow. The assistant as primary input source is the workflow that lost ground.
Editorial bandwidth that previously went to formatting and SEO checklists now produces higher returns when reallocated to data work — survey design, FOIA drafting, dataset cleaning, interview scheduling. The skill mix that wins under the post-March framework includes more research operations and less keyword research.
What This Desk Tracks Through Q2-Q3 2026
Three datapoints anchor ongoing tracking on this framework. First, whether the 22% average gain holds, narrows, or widens in subsequent third-party reads as more sites attempt data injection — a narrowing average would signal commoditization of Layer 1; a widening average would signal that the bar is rising. Second, whether Google iterates the framework to weight Layer 2 and Layer 3 differently from Layer 1. Early signals suggest Layer 1 carries the heaviest weight, but that ratio is not fixed and platform behavior between June and September 2026 will reveal direction. Third, the emergence of synthetic-data detection signals. Anti-pattern 3 — AI-generated tables presented as research — is a likely target for the next refinement, and sites currently relying on it should expect tightening.
Honest Limits
The 22% figure is an average from third-party visibility trackers, not a Google-published metric. The variance around the average is substantial; individual sites in the gainer cohort range widely. The three-layer decomposition is editorial — Google's documentation does not name "Layer 1/2/3," and the framing here is this desk's structuring claim, not a verbatim restatement of Search Quality Rater Guidelines or Helpful Content documentation. The patterns and anti-patterns are observed across published case studies and field reading, not statistically inferred from a controlled sample. Sites should treat the framework as a hypothesis to test against their own data, not as a recipe with guaranteed return.
What this desk asserts narrowly: original data investment co-varies with visibility gain in the post-March 2026 environment, and the three-layer decomposition is a useful operational frame for editorial planning. Whether the gain is causal versus co-varying with other E-E-A-T signals is not established by available data. Whether the framework holds beyond Q3 2026 depends on subsequent Google iterations that have not yet shipped.