svy on India’s Periodic Labour Force Survey, 2023-24
Labour force
India
Panel data
Variance estimation
Does a rotating panel change the procedure? India’s labour force survey revisits urban households for four quarters. For levels it changes the weight, not the estimator; for change between quarters it halves the standard error; and it moves the level itself.
Periodic Labour Force Survey, PLFS India, rotating panel survey, rotation group bias, change estimation standard error, labour force participation rate, interpenetrating sub-samples, complex survey analysis Python
Summary
TipTL;DR
India’s Periodic Labour Force Survey visits every urban household four times, a quarter apart. A reader asked whether that rotating panel changes the procedure. It does, in three places.
The weight, not the estimator. A quarter’s urban sample is four panels of blocks, each a random draw from the same frame, so a quarterly rate is estimated like any stratified cluster sample. 88 of 90 published cells reproduce at the printed decimal.
The standard error of change. Three blocks in four are back the next quarter. Treat the quarters as independent and the standard error of a quarter-on-quarter change comes out about twice as large as it should: 11 of 27 consecutive-quarter changes are significant at 5 per cent with the overlap accounted for, 2 without it.
The level itself. Labour force participation falls with every visit, from 50.8 per cent at the first to 48.9 at the fourth. That is why the annual urban rate, from first visits only, sits above the average of the four quarterly rates.
The question
After an earlier case study, a reader asked for a PLFS example with a pointed question: it is a rotational panel, so does it make any difference in procedures?
The short answer is that it depends on what you estimate. This page is the long answer, on the 2023-24 unit-level data, checked against both of the survey’s publications for that year.
Background
The Periodic Labour Force Survey is run by India’s National Statistics Office. It has two products. The annual report gives labour force indicators for rural and urban areas, in usual status (a 365-day reference period) and in current weekly status (the last seven days). The quarterly bulletin gives current weekly status for urban areas only, every three months.
The design is set out in the survey’s note on sample design and estimation. First-stage units are villages in rural areas and urban frame survey blocks in towns, drawn with probability proportional to size and with replacement, in two independent sub-samples per stratum. Within a unit, households are split into second-stage strata by how many members completed secondary school, and eight are drawn.
Rural households are visited once. Urban households are visited four times, one quarter apart: a first visit, then three revisits. Each quarter one panel of blocks leaves and a new one enters, so a quarter’s urban sample is always four panels at four different visits:
First visited
2023Q3
2023Q4
2024Q1
2024Q2
2022Q4
visit 4 · 1,434 FSUs
2023Q1
visit 3 · 1,427 FSUs
visit 4 · 1,426 FSUs
2023Q2
visit 2 · 1,408 FSUs
visit 3 · 1,406 FSUs
visit 4 · 1,407 FSUs
2023Q3
visit 1 · 1,441 FSUs
visit 2 · 1,427 FSUs
visit 3 · 1,425 FSUs
visit 4 · 1,425 FSUs
2023Q4
visit 1 · 1,441 FSUs
visit 2 · 1,433 FSUs
visit 3 · 1,432 FSUs
2024Q1
visit 1 · 1,443 FSUs
visit 2 · 1,435 FSUs
2024Q2
visit 1 · 1,443 FSUs
Quarters are calendar quarters: the survey year runs from 2023Q3 (July to September 2023, labelled Q1 in the files) to 2024Q2. Read a column down to see one quarter’s sample. Read a row across to follow one panel through the year. The panels that started before July 2023 are revisits of households first seen in the previous survey year.
The file you actually get
The unit-level release has two person files, and the rotation is split across them:
File
Person records
Sectors
Visits
First visit (PV)
418,159
rural, urban
1
Revisit (PV_R)
504,440
urban
2, 3, 4
The revisit file carries a quarter and a visit number, not a panel. The panel is derived: a record seen in quarter q at visit v belongs to the panel that started in quarter q − v + 1.
build() does this and the rest of the wrangling: it stacks the two files, derives the panel and the keys, and adds the weights and the labour force indicators. Fifteen urban records drawn at random show what comes out, one row per person per visit:
fsu is the serial in the file and fsu_id the block key with the stratum and panel in front. panel runs from −2 to 4: panels 1 to 4 were first visited in the four quarters of 2023-24, and panels 0, −1 and −2 in the last three quarters of 2022-23. cws is the current weekly status code and w_qtr the quarterly weight, both used below.
The panel has to be part of the block key. First-stage unit serial numbers are not block identifiers: 4,256 serials are shared by more than one panel, so keying on the serial alone merges unrelated blocks into one PSU. With the panel in the key, every block sits in one stratum and shows up in consecutive quarters, except 2 that missed a visit. The Apr-Jun 2024 urban sample in the files is then the press note’s, to the unit:
Urban, Apr-Jun 2024
From the files
Press note
First-stage units
5,735
5,735
Households
45,016
45,016
Persons
171,121
171,121
Levels: the rotation changes the weight, not the estimator
Annual estimates use first visits only
The annual report is built from first visits alone, rural and urban. The file’s multiplier is computed for one sub-sample in one quarter, so the weight takes it back to the year and to both sub-samples: divide by 100 (the multiplier is stored with two implied decimals) and by the number of quarters the stratum contributed, and by two again where both sub-samples are present.
base = ( pl.when(pl.col("nss") != pl.col("nsc")) .then(pl.col("mult") /200) .otherwise(pl.col("mult") /100))w_annual = base / pl.col("no_qtr")
The annual sample is the first visits, and the design is the one the survey drew: strata (sector, region, stratum and sub-stratum), blocks as PSUs, drawn with replacement.
Every indicator is a share among persons aged 15 and above, so each is a domain mean of a 0/1 indicator. The unemployment rate is the same mean with the domain narrowed to the labour force.
54 of 54 annual cells match: the three indicators in both usual status and current weekly status, by sex, for rural, urban and all areas. Nothing about the rotation enters here beyond the choice of file.
Quarterly estimates use every visit
The quarterly bulletin uses all four panels of the quarter. The estimation note is explicit about how they combine: a panel’s estimate is the mean of its two sub-samples, and the quarter’s estimate is the mean of its panels. Each panel is a full-scale estimate of the urban population, so the four are averaged, not added. In weights, that is one more division, by the number of panels the stratum has that quarter:
Four panels is the steady state; a stratum that lost a panel to casualties has three. Get the division wrong and the rates barely notice, since a constant factor cancels in a ratio, but every total comes out four times too large.
The quarterly sample is every urban record: all four visits of all four quarters, in one sample. The quarter is a domain, not a separate sample. Estimates for a quarter are the same either way, but keeping the quarters together is what lets svy see the blocks that appear in more than one of them, which is the whole story of the change section below. The PSU count shows it: blocks, not block-quarters.
34 of 36 quarterly cells match, by sex and for all persons. The two that do not come out at 49.848 against 49.9 and 8.948 against 9.0: on the rounding boundary, and within the difference between the two orders the estimation note gives for averaging sub-samples and panels. The published cells cannot tell those orders apart; this page follows the one the note writes for the estimator.
Variance: the official formula is a choice of PSU
The estimation note does not linearise over first-stage units. It uses the two interpenetrating sub-samples: within each stratum, the variance of a ratio is a quarter of the squared difference between the sub-sample estimates of its linearised numerator, summed over strata.
That is exactly svy’s with-replacement variance with two PSUs per stratum. With two PSU totals z₁ and z₂, the estimator’s n/(n − 1) times the sum of squared deviations from their mean is (z₁ − z₂)², and each z is half a sub-sample estimate. So the official variance is a design in which the sub-sample is the PSU:
The by-hand column follows the note to the letter, including its rule that a panel missing one sub-sample borrows the other’s estimate. The last digit of difference comes from how those few panels are weighted.
The block is the other natural PSU. Its standard errors run from 5 per cent below the official ones to 11 per cent above, with about 5,472 degrees of freedom per quarter instead of the 238 that two PSUs per stratum allow.
Both are defensible. This page uses the block from here on, because it is what the design sampled at the first stage and because it is the unit that comes back next quarter.
Change: where the rotation matters
The estimation note measures change over the quarters by “the simple difference” between the two estimates, and gives no variance for it. The variance of a difference is
With independent samples the covariance is zero. Here three panels in four are shared between consecutive quarters, the same households in the same blocks, and a block with high participation this quarter tends to have high participation next quarter. The covariance is large and positive, and it comes off the variance of the change.
In svy there is nothing to add. Keep all four quarters in one sample, let a block keep its PSU across quarters, and take the difference as a contrast. The covariance is estimated from the blocks that appear in both quarters.
╭────────────────────────────── Contrast (TAYLOR, df=9798) ───────────────────────────────╮│││contrast est se cv (%) t p_value lci uci││ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ ││ 2023Q4 minus 2023Q3 0.0059 0.0018 30.04 3.3287 0.0008757 0.0024 0.0093 │││╰─────────────────────────────────────────────────────────────────────────────────────────╯
The same call on a design that gives a block a new PSU every quarter, fsu_qtr_id being the block key with the quarter appended, is what “treat the quarters as independent” looks like. The estimates are identical; only the standard error of the change moves.
╭───────────────────────────── Contrast (TAYLOR, df=22611) ──────────────────────────────╮│││contrast est se cv (%) t p_value lci uci││ ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ ││ 2023Q4 minus 2023Q3 0.0059 0.0034 58.00 1.7243 0.08467 -0.0008 0.0126 │││╰────────────────────────────────────────────────────────────────────────────────────────╯
change
Panel-aware
Quarters as independent
SE
p
SE
p
All: 2023Q3 → 2023Q4
+0.59
0.18
<0.001
0.34
0.085
All: 2023Q4 → 2024Q1
+0.38
0.17
0.027
0.34
0.264
All: 2024Q1 → 2024Q2
-0.18
0.16
0.270
0.34
0.597
All: 2023Q3 → 2024Q2
+0.79
0.28
0.005
0.34
0.020
Women: 2023Q3 → 2023Q4
+1.03
0.24
<0.001
0.46
0.024
Women: 2023Q4 → 2024Q1
+0.59
0.26
0.023
0.46
0.206
Women: 2024Q1 → 2024Q2
-0.37
0.24
0.122
0.46
0.425
Women: 2023Q3 → 2024Q2
+1.25
0.40
0.002
0.45
0.006
Men: 2023Q3 → 2023Q4
+0.28
0.23
0.220
0.44
0.516
Men: 2023Q4 → 2024Q1
+0.26
0.21
0.209
0.44
0.547
Men: 2024Q1 → 2024Q2
+0.32
0.20
0.101
0.44
0.464
Men: 2023Q3 → 2024Q2
+0.87
0.35
0.013
0.44
0.046
Urban LFPR, current weekly status, persons aged 15 and above, percentage points. The last row of each block is three quarters apart, where only one panel in four is shared.
Across the 27 consecutive-quarter changes computed (LFPR, WPR and UR, for all persons, women and men), the panel-aware standard error is 45 to 66 per cent of the independent one, 55 per cent on average. 11 of the changes are significant at 5 per cent once the overlap is counted; 2 are without it. The rise in urban participation from July-September to October-December 2023 is the clearest case: a real movement that a reader of the two bulletins, doing the arithmetic, would call noise.
Three quarters apart, only one panel in four is shared, and the gain shrinks accordingly. The overlap is what earns the precision, and the rotation is what sets the overlap.
The sub-sample design captures the same covariance, since a panel’s blocks keep their sub-sample as they rotate. Its standard errors of change are within a few hundredths of a point of the block design’s. What does not capture it is any analysis that estimates each quarter separately and combines the two standard errors afterwards, whatever the PSU.
The level itself: time in sample
Within a quarter, the four visit groups are four different panels drawn at random from the same frame. Without an effect of being revisited, they would estimate the same thing up to sampling error. They do not:
Visit
1st − 4th (SE)
p
1st
2nd
3rd
4th
LFPR, all
50.8
50.2
49.5
48.9
+1.88 (0.28)
<0.001
LFPR, women
26.1
25.3
24.6
23.7
+2.39 (0.40)
<0.001
LFPR, men
75.0
74.4
73.9
73.6
+1.36 (0.35)
<0.001
WPR, all
47.4
46.8
46.3
45.7
+1.67 (0.28)
<0.001
WPR, women
23.8
23.2
22.4
21.7
+2.14 (0.38)
<0.001
WPR, men
70.5
69.9
69.7
69.3
+1.21 (0.37)
0.001
UR, all
6.7
6.6
6.5
6.5
+0.16 (0.22)
0.466
Urban, current weekly status, persons aged 15 and above, pooled over the four quarters of 2023-24. Per cent.
Participation and employment fall with every visit, for women and for men. Women’s participation drops 2.4 points between the first visit and the fourth. The unemployment rate does not move, because the labour force and employment shrink together: people drift from employed to outside the labour force, not to unemployed. Rotating panels are known for this pattern, often called rotation group bias. One common explanation is that respondents learn the questionnaire: reporting no work in the last seven days skips the day-by-day questions on hours and earnings that follow.
It has a consequence for the published numbers. The annual report’s urban participation rate in current weekly status is built from first visits: 50.8 per cent. The four quarterly bulletins for the same year, same concept, same area, use every visit: 49.3, 49.9, 50.2, 50.1, which average 49.9. Neither is wrong; they are estimated from different visits, and the rotation is the difference.
One caution. The panels in a quarter entered the survey at different times, and the files alone cannot separate a pure effect of being revisited from anything else that differs between panels. What they show is a gradient in the same direction for both sexes, in both indicators, and large relative to its standard error.
What only a panel gives: flows
A cross-section tells you how many people are unemployed each quarter. Only a panel tells you what happened to the people who were. Households keep their identifiers across visits, and so do their members:
This quarter
Next quarter
Employed
Unemployed
Out of labour force
Employed
97.0 (0.09)
1.1 (0.06)
1.9 (0.07)
Unemployed
15.6 (0.62)
79.6 (0.74)
4.9 (0.33)
Out of labour force
1.5 (0.06)
0.3 (0.02)
98.2 (0.06)
Urban adults seen in two consecutive quarters of 2023-24, current weekly status. Row per cent (SE); each row sums to 100.
295,380 of the 303,586 adult records that should have a next visit find one (97 per cent). Each transition is weighted by the earlier quarter’s weight and clustered by block, so a person followed through four quarters counts three times, correlated. The three per cent who could not be followed are not modelled; a production flow estimate would adjust the weight for them, and the time-in-sample drift above is a reminder that a flow into “out of labour force” between visits is partly a change in answers, not in lives.
From 2025, the same question monthly
From January 2025 the survey was redesigned. The rotating panel now covers rural and urban areas, and households are visited in four consecutive months rather than quarters, to produce monthly national estimates. The answer on this page carries over, one step finer: monthly levels need the panel average in the weight, month-on-month change needs the overlap in the variance, and the time-in-sample question applies to both sectors.
Reproducing this page
Everything on this page is computed at render time from the two 2023-24 person files, by the same modules a validation harness runs (python code/validate.py). The press note figures are the only transcribed numbers. After a one-time conversion of the Stata files to parquet, building the 922,599 person-visit records takes about a second, and every estimate on the page a few more.