One long Sample, a case id, an ordered wave column, and the variance of change comes out right
Tutorials
Panel Surveys
Weighting
Estimation
Python
Analyze longitudinal (panel) survey data in Python with svy: stack the waves into one long Sample, estimate levels and net change with the right variance, compute transitions and individual change with lag(), and build longitudinal weights with a nonresponse adjustment per wave.
A panel survey follows the same cases (persons, households, firms) over several waves. In svy a panel is not a new object: it is an ordinary Sample in long format, one row per case and wave, whose Design names two extra columns:
case_id — the column identifying the followed case. It must be non-null and unique within each wave.
wave — the column ordering a case’s rows. Any increasing codes work (1, 2, 3 or 2019, 2021, 2023).
Everything else is the estimation and weighting API you already know. Two things happen automatically once both columns are declared:
The case is the variance PSU when no PSU is declared. A case’s rows are repeated measures of one sampling unit, so they must form one cluster. That is what makes mean("y", by="wave") and its contrasts carry the covariance between waves, so the standard error of a change is the standard error of the individual changes — not the naive “sum of two independent variances”.
Pairing is validated. Duplicated ids within a wave, design columns that vary within a case (a mover must keep the base-wave stratum and PSU), and two waves with no case in common are all refused with a message naming the fix.
What svy does not guess is the estimand: which waves and which weight. You say it with the same where= and use_weight() you use elsewhere.
The bundled panel
svy ships a small synthetic three-wave panel, panel_syn_2026: 1200 cases at wave 1 in 6 strata × 10 PSUs, about 15% attrition per wave, producer longitudinal weights, a binary outcome (emp) and a continuous one (inc). See Datasets for how bundled data works.
╭────────── Design ──────────╮│ Case id case_id ││ Wave wave ││ Stratum stratum ││ PSU psu ││ SSU None ││ Weight w1 ││ With replacement False ││ Prob None ││ Hit None ││ MOS None ││ Population size None ││ Replicate weights None │╰────────────────────────────╯
Every row of a case carries the same w1, stratum and psu: the sampling design is the wave-1 design, and later waves are follow-ups of the same units. The response column resp is "rr" on interviewed rows and "nr" on rows where the case was contacted but not interviewed; cases lost to follow-up simply have no row at later waves.
Producers often deliver one file per wave. svy.combine_samples(kind="panel", case_id=...) stacks them into the long form above and runs the pairing checks. Here we split the bundled panel into per-wave samples to show the round trip:
waves = [ svy.Sample( data.filter(pl.col("wave") == t).drop("wave"), svy.Design(stratum="stratum", psu="psu", wgt="w1"), )for t in (1, 2, 3)]stacked = svy.combine_samples(waves, kind="panel", case_id="case_id")print(stacked.design)
╭──────────── Design ─────────────╮│ Case id case_id ││ Wave wave ││ Stratum (stratum,) ││ PSU (psu,) ││ SSU None ││ Weight combined_wgt ││ With replacement False ││ Prob None ││ Hit None ││ MOS None ││ Population size None ││ Replicate weights None │╰─────────────────────────────────╯
Caller order is time order; the wave column gets codes 1..k labelled "wave 1".."wave k" (an existing wave column present in every input is reused instead). kind="panel" differs from the default kind="cross_sectional" in three ways: the strata are not wave-qualified (the same units recur, so waves are not independent), weights are not averaged over waves (a person is not half a person for appearing twice), and the pairing is validated. The checks:
case_id is present, non-null and unique in every input;
consecutive waves overlap — an empty overlap means two unrelated cross-sections were stacked, and is an error; an overlap under 50% warns;
design columns are constant within a case across waves;
a later wave’s (stratum, PSU) set is a subset of wave 1’s; a PSU with no row at a later wave is a real panel event and warns.
The overlap between consecutive waves is part of the design summary:
drop_nulls=True zero-weights the rows without an interview (inc is null on "nr" rows); it does not drop them, so the design structure is intact. The net change between two waves is a contrast between these estimates, and the covariance between waves is already in the result:
Differences are linear contrasts (exact). The ratio and the percent change are estimated by the delta method on the same design degrees of freedom, matching R’s svycontrast(..., quote((w3 - w1) / w1)). Products, quotients, constants, .log() and .exp() all work inside an estd() expression.
Why the covariance matters
Stack the same three waves as if they were independent cross-sections and the change looks far less precise, because the between-wave covariance that a panel earns is thrown away:
as_cs = svy.combine_samples(waves, kind="cross_sectional", adjust="none")naive = as_cs.estimation.mean("inc", by="wave", drop_nulls=True).contrast(estd(2) - estd(1))print("panel SE :", round(change.estimates[0].se, 2))print("naive SE :", round(naive.estimates[0].se, 2))
panel SE : 20.35
naive SE : 59.83
The panel SE is the standard error of the individual changes, which is what a paired analysis on a wide file would give. The cross-sectional SE treats the two means as independent and is wrong for a panel.
Two estimands, two weights
by="wave" with the base weight w1 estimates each wave’s mean over the cases present at that wave. Attrition is not random, so a comparison across waves mixes real change with the changing composition of respondents. Panel producers ship longitudinal weights for exactly this: lw_12 is nonzero only for cases observed at both waves 1 and 2 and re-inflates them to the wave-1 population; lw_123 does the same through wave 3. They are ordinary columns, selected with use_weight():
╭─────────────────────────────────── Contrast (TAYLOR, df=54) ───────────────────────────────────╮│││contrast est se cv (%) t p_value lci uci││ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ ││ % change 1 → 3, survivors 0.0710 0.0047 6.65 15.0328 6.031e-21 0.0616 0.0805 │││╰────────────────────────────────────────────────────────────────────────────────────────────────╯
The two answers are different questions: with w1, how did the population’s mean income move (entrants and leavers included); with lw_123, how did income change for the people followed through wave 3. svy does not choose for you.
Transitions and individual change with lag()
Anything that pairs a case’s value at one wave with its value at another goes through wrangling.lag(), the one panel primitive. It adds <col>_lag<n>, the value of the column at the case’s wave n steps back — null at the first wave, and null across a skipped wave (a case with rows at waves 1 and 3 gets no lag1 at wave 3, as Stata’s L.y; pass gaps="skip" for the previous observed row instead). Value labels and the variable type carry over. Negative n is a lead.
Here where= picks the wave and drop_nulls=True zero-weights the rows whose lag is null (the first wave) — again a domain, not a subset, so the design df is unchanged. The joint table with its chi-square test is tabulate on the same domain; its where= has R’s subset() semantics, and a null lag outside the domain is not missing data:
ttest(y_pair=), quantile and the transition tables read the weight of the rows in where=. With the base weight w1 that is the cross-sectional wave-2 weight; if you want the survivors’ estimand, select the longitudinal weight first with use_weight("lw_12"). The lag itself does not care — it copies values.
A pooled model across waves
Because the sample is long, a model with the wave as a factor is one glm.fit call, and the case clustering is handled by the design:
fit = panel.glm.fit("inc", x=[Cat("wave"), Cat("age_grp"), "urban"])print(fit)
The wave coefficients are the adjusted net changes against wave 1; fit.contrast(estd("wave_3") - estd("wave_2")) compares two later waves directly.
Building longitudinal weights
When the producer ships no longitudinal weight, or you want to rebuild it with your own cells, the chain is one nonresponse adjustment per wave, scoped with where= to the target wave. On a panel weighting.adjust() does two things it does not do on a cross-section:
a case with earlier rows but none in scope is a nonrespondent — attriters lost to follow-up have no row to filter on, so a plain row filter would never see them and the factor would be wrong;
the factor is applied to the case: every row of the case gets its own weight times the case’s factor, so a case-level base weight stays case-level — one longitudinal weight per case, whichever of its rows you analyze — and a nonrespondent case (interviewed nowhere in scope, or lost) gets 0 on all its rows. A weight column that already varies by wave keeps that variation. Cases outside every cell keep their weight.
respondents_only=True (the default) then drops the nonrespondent rows at the scope waves only; earlier rows stay, since those cases responded then.
The scope must be a set of whole waves for the missing-in-scope rule to apply; a where= mixing other conditions still adjusts, but warns that attriters without a row were not added. Replicate weights, if the design carries them, are adjusted column by column and written back per case in the same way, and each step is recorded on design_history. rake() or calibrate() to base-wave or wave-t controls can follow, exactly as on a cross-section.
TipPropensity classes
For a model-based adjustment, fit glm.fit(..., family="binomial") on the base-wave covariates with a response indicator built from lag(["resp"], n=-1) (the next wave’s status on the wave-1 row), categorize the fitted probabilities into quintiles with wrangling.categorize, and pass the class column as cells=.
Declaring a panel on a long file
A producer’s long file needs no stacking — declare the two columns on the Design:
The same validation runs at construction. If the file is wide (one row per case, inc_2019, inc_2021, …), unpivot it with polars before creating the Sample; a to_long convenience is on the list.