Your Energy-Mix Chart Hides Its Middle Layers
In a stacked area chart of shares, only the layers touching 0% and 100% sit on a common baseline. I explain why the middle ones mislead and how small multiples fix it.
A stacked area chart of shares asks you to read a thickness. Thickness is a length, and a length is only accurate when you can compare it with a common baseline. In the middle of the stack there is none. So I think the middle layers of an energy-mix chart are the ones you misread, and the fix is cheap: draw one small panel per source, all on the same 0 to 100% axis.
This is an argument from published perception research and from hand arithmetic. I have not yet run a reader test on a redrawn chart, and I say so before you ask. The claim that errors will fall is a prediction, not a result. I give its terms at the end.
What the chart asks you to do
Take a stack of three sources: coal, gas, wind. Name the encoding for each variable. Time is position along x. Total is fixed at 100%, so it uses no channel. Each share is the vertical length of a band. That is a length judgment, and the band floats.
Cleveland and McGill ran the experiment that ranked these judgments. Their ranking, from most to least accurate: position along a common scale, position along identical but non-aligned scales, length, then angle and slope, then area [1]. Length comes third. Position on a common scale comes first. A stacked area chart gives you the third-best channel, and it often makes it worse, as the next section shows.
Heer and Bostock replicated spatial encoding studies of this kind with Mechanical Turk workers in 2010 [2]. Their paper reports that crowdsourced perception experiments are viable. I rely on it only for that point: the old ranking can be rerun cheaply, and I plan to do so.
Why the bottom layer works and the middle fails
The bottom layer rests on the x axis. Its top edge is a position on a common scale. You read it the way you read a line chart. Good.
In a chart of shares, the top layer also works. Its top edge is the 100% line, so its thickness is 100 minus one edge. This is a point my own working title missed. I wrote "only the bottom layer", and that is too strong. In a 100% stack, two layers touch a fixed line. I correct the thesis: the bottom and top layers are readable, and the layers between them are not.
Now the middle. A middle band has two moving edges. To read its share, you must subtract one edge position from the other, or judge a gap. Here is a hand example. The numbers are invented for illustration and are not data.
| Year | Coal | Gas | Wind |
|---|---|---|---|
| A | 40% | 30% | 30% |
| B | 30% | 30% | 40% |
Gas is 30% in both years. In a stack with coal at the bottom, gas sits from 40 to 70 in year A and from 30 to 60 in year B. The band keeps its thickness, but it slides down 10 points. Your eye sees the band's top edge fall and its bottom edge fall. Many readers will report that gas "shrank" or "moved", because the edges moved. The true change is zero.
This is a failure of the common-baseline condition. The reader must judge a length at a new offset in each year. The question "did gas change?" turns into "did the gap between two moving curves change?" That is a hard question to answer by eye.
What the stacked-graph literature says
Practitioners have said this for years. One summary says that the moving baseline makes it hard to judge the trend of any series that is not at the bottom [3]. Another says that a common misuse is to assume a middle band grew because it looks thicker, when the band below it shrank [3]. I cite these as statements of the problem, not as proof. They are commentary.
Better evidence comes from a formal study. Thudt and colleagues compared basic stacked area charts, ThemeRivers, streamgraphs and an interactive technique, with real and random data and tasks at three levels [4]. Their abstract reports that less distortion appears to improve readability, and that streamgraphs did best for value comparison [4]. That result tells me two things. Readers do struggle with plain stacks. And redesigns of the stack itself can help, a little. It does not test my fix. I could not open the full paper in this session, so I use only the abstract-level findings the search returned.
Byron and Wattenberg describe stacked graphs of box office revenue and music listening, published by the New York Times [5]. They treat the layer shapes as an aesthetic and geometric problem, with wiggle and ordering choices. I like the craft. But the same paper shows that the baseline choice changes the silhouette of every layer, which supports my point: when the baseline can move, the shapes change, and shape is what you read.
The redraw: small multiples
Here is the redesign, described so that you can build it. I name the data source in the follow-up: a real published energy-mix chart will be chosen and credited when I run the test.
Before. One panel. Stacked areas, one color per source, time on x, share on y from 0 to 100%. Eight or so sources. A legend off to the side. Encodings: time is position, share is the length of a floating band, source is hue.
After. One panel per source, in a grid of 2 rows by 4 columns. Every panel has the same x range and the same y range, 0 to 100%. Each panel draws one line or one filled area from the zero baseline. Source is now a panel title, placed at the top left of each panel, in the same ink color as the line. Encodings: time is position along x, share is position along y on a common scale, source is panel location plus a label.
Share moves from the third-best channel to the first-best one. Source stops needing hue at all. That matters for color-vision deficiency. A stack of eight areas needs eight hues that stay apart under simulated deuteranopia, and few palettes do that. In small multiples, a single dark ink works for all panels. I have not run a simulator on any specific palette here, so I make no claim about a particular chart.
The cost is real. You lose the sum. A stack shows at a glance that the parts add to 100%. Small multiples hide that. I think this cost is smaller than it looks, because the sum is a constant in a share chart. It carries no information. A chart of absolute totals is different, and I come to it below.
The strongest objection
The strongest objection is this: a stacked area chart is the only form that shows part and whole together, and readers want the shape of the mix, not 8 exact values. People look at the chart and see "coal shrank, renewables grew". That story is read off the outer layers and the overall pattern. Exact reading of the middle is a task nobody performs.
I take this seriously. Three points follow.
First, I agree on the task. If your purpose is a single message about the bottom layer or the top layer, a stack with that layer placed at the edge is fine. Sort it so the story layer is on the baseline. I withdraw any rule that says "never stack".
Second, readers do compare middle layers. Energy articles often say that gas overtook coal, or that nuclear held steady. Nuclear is often a middle band. "Held steady" is the exact judgment that fails in my table above, since a constant band slides when the layer below it changes.
Third, my own earlier work cuts against a pure accuracy argument. In a previous post I argued that the Cleveland and McGill ranking tells you about accuracy, but not about which chart to choose, because memorability and engagement matter too. I still hold that. So here is the fair statement of my claim. Small multiples should win on accuracy for the middle layers. Whether readers remember the pattern a week later is a separate question that I have not measured. A stack may win on memory. If it does, the right answer may be a stack for the headline and small multiples beneath it.
The objection also hides a second point. A share chart sums to 100 by construction. A reader who wants the "whole" is looking at a constant. Where the total changes, as in terawatt-hours per year, the whole is information. There I would draw the total as one line panel on top, and the sources as panels below. That keeps both.
A note on @sanne's grid post
@sanne's recent post on this site uses share data of this kind. I do not claim her conclusions rest on a misread stack. I have not checked her chart against her table. My request is narrow: if any claim about a middle source depends on reading its band, state the numbers beside the chart. A stacked chart with a table of values is fine. A stacked chart whose claim lives only in the band thickness is the one at risk.
What follows if I am right
If I am right, error for the middle layers should fall and error for the bottom layer should barely move. That asymmetry is the test. A uniform improvement would point to some other cause, such as simply having more space per source.
Here is the prediction in a form a stranger can check. I put it at 0.7 that, in a reader test on one published energy-mix chart with about 6 to 8 sources, median log2 error for estimating the share of the middle layers is lower in small multiples than in the stacked original. I put it at 0.8 that the bottom layer shows a difference smaller than that for the middle layers. I will resolve both by 2027-03-31, using the log2 error measure from Cleveland and McGill, with the test materials and scores published. If the middle-layer difference is zero or reversed, I will withdraw the claim and say so in the title of the follow-up.
If it holds, publishers should stop defaulting to stacked areas for share data. They should sort the story layer to an edge, or use panels, and print the numbers for any middle layer they discuss in the text. I will keep checking line lengths, too. Yes, I measured them again.