Release Notes
All notable changes to svy, the Python package for design-based analysis of complex survey data — means, totals, ratios, proportions, regression, weighting, and sample selection — are recorded here. Releases follow Semantic Versioning; the layout follows Keep a Changelog.
Companion packages track their own changes: svy-io (SAS/SPSS/Stata I/O) and svy-rs (internal Rust extension).
Unreleased
Added
- Agent skills ship with svy. Guidance for coding agents, matching the installed version: an entry file (namespace map, a complete example, getting full-precision numbers out, rules, common wrong names) and one file per namespace (io, design, wrangling, estimation, categorical, glm, weighting, sampling). Run
svy-skillsonce per project (orsvy.skills.install()) to link them into.claude/skills/; upgrades then update them. Rerun it when a release adds a skill or after recreating the environment;--copycopies instead. Other packages can ship skills under thesvy.skillsentry-point group, and one run installs them all. - Wrong names say what to use. A name that does not exist on
svy, aSample, aDesign, a sample namespace or a result raisesUnknownNameError(anAttributeError, codeUNKNOWN_NAME) that names the replacement:svy.SurveyDesign→svy.Sample(data, svy.Design(...)),sample.mean→sample.estimation.mean,estimation.tabulate→categorical.tabulate,estimation.proportion→prop,result.est→result.to_polars(), plus close matches.svy.Design(data), unknownDesignarguments (strata=→stratum=,weights=→wgt=,cluster=→psu=) and a pandas frame passed toSampleraiseUsageError(aTypeError) with the fix.help(svy)now shows where to start.
0.32.2 — 2026-10-10
Changed
- Requires polars 2.0 or later (was 1.36.1). The lock and CI test 2.x only; polars 2.0 needs Python 3.10+, within svy’s 3.11+.
Fixed
- Wide frames are fast.
Sample()cost grew with the square of the column count: 6.3 s at 20k rows × 4000 columns, now 38 ms. Every wrangling step is also 2–5× faster on wide frames, since a derived sample no longer deep-copies each variable’s metadata. - Wrangling works on a Sample built with
catalog=. Every step raisedTypeError: cannot pickle '_thread.RLock' object; the derived sample now shares the catalog, asclone()does. - Row order no longer depends on polars internals. Sorts on two or more columns kept tied rows in an arbitrary order. Tied rows now keep their input order in
wrangling.order_by, in selectionorder_by, inshow_dataand inload(order_by=). A seeded draw withorder_byon two or more columns can select different units than before; a single sort column is unaffected. Selection results are joined back in the sample’s row order.Estimate.stratais sorted, as documented, rather than in a different order on each call.describe(top_k=)breaks count ties by level, so the same levels are kept on every run. Sample()collects a LazyFrame. Without a design the frame stayed lazy, and every later step resolved its schema again, with a polarsPerformanceWarningeach time. A design already collected it.- A seeded
load(order_type="random")gives the same rows on every polars version. It shuffled by polars’hash(), which changed in polars 2.0, so the same seed already loaded different rows there; it now uses NumPy’s frozen legacy stream. Seeds wider than 32 bits and negative seeds are used in full. INVALID_TYPEerrors name the type passed.apply_labels, column-name lists,rakemargins and allocationpop_sizereportedgot: strwhatever was passed.- Estimation is faster, and the slowdown since 0.26 is gone. At 1M rows, a stratified clustered mean or total takes 10 ms (was 24–27 ms),
by=62 ms (was 110 ms) andwhere=24 ms (was 47 ms); calibrated designs,prop,median,corrandcovare 20–50% faster. The domain-singleton check and the full-design df behindby=ran eagerly on every call. GLM fits skip a log per row in the separation check.
0.32.1 — 2026-10-08
Added
sample.wrangling.casttakes dtype names:cast("age", "Int64")orcast({"age": "Int64", "region": "String"}), so a cast needs noimport polarsand can be written in a config file. The names are polars’ own (Int8…UInt64,Float32,Float64,String,Boolean,Date,Datetime,Categorical), listed bysvy.wrangling.DtypeName; an unknown name raisesCAST_UNKNOWN_DTYPEwith the closest match. Polars dtypes still work and are needed for parameterised types (a time-zonedDatetime,Enum,Decimal).
0.32.0 — 2026-10-08
Added
Row and column percentages:
tabulate(share_of="row" | "col"). Each cell is a share of its row (or column) instead of the table total, the way offices publish crosstabs. Row shares areestimation.prop(colvar, by=rowvar)(column shares the reverse): same estimates, SEs and intervals, for Taylor and replication, withwhere=; a level absent from a row shows as 0. The Rao-Scott test is unchanged. The title says(row shares)/(column shares)andTable.share_ofrecords it. Two-way tables only, and not withunits="count"orcount_total. Saved tables carry it: schemasvy-result/0.8addsTableData.share_of("total"when read from an older payload).GLM categorical terms read by value label. On a labelled column, coefficients print
area_RURALand categorical marginsRURAL - URBANA; levels stay in code order.Cat(ref=)takes a value label as well as a code (a label shared by two levels raisesCAT_REF_AMBIGUOUS; shared labels in names get their code appended).GLMCoef.termkeeps the engineered column name andto_polars()addsterm_label.glm.fit(use_labels=False)prints codes. Saved fits keep the codes only.Two-way tables print their Rao-Scott tests.
print(tab)on a crosstab ends withRao-Scott F(3.30, 49.48) = 11.67, p < 0.001(second-order) andRao-Scott chi2(3) = 35.21, p < 0.001(first-order); one-way tables are unchanged.ChiSquare,FDistandTDistprint in the same form on their own.ttest(use_labels=)andranktest(use_labels=)show labelled groups andby=levels by their value labels, astabulatedoes:Groups: area = [URBANA vs RURAL], the level rows and── reg = Nortesection headers. None (the default) uses the labels when the variable has any; False shows the codes. The result keeps the codes (GroupLevels.labelsholds the labels), andto_polars("estimates")adds a<group>_labelcolumn (group_level_labelwithtidy=False). Saved results keep the codes only.Choice enums carry a label and a description.
SingletonMethod,DistFamily,LinkFunction,QuantileMethodandRankScoreMethodmembers have.labeland.description, anddescribe()lists(value, label, description)for every member, so applications show svy’s own wording instead of copying it. The enums are unchanged otherwise (stillStrEnum).sample.weighting.trim()annotatesupper,lowerandby(a threshold or a number; one or several columns).Contrasts can be saved.
svy.serialize.to_json(est.contrast(...))writes aContrastData(kind"contrast"): each contrast’sest,se,cv,lci,uci,t,p_valueanddf, withmethod,alpha,dfand the source estimate’sparamandy, whichContrast.paramandContrast.ynow also give;serialize.to_polars()returns the live table. The contrasts’ joint covariance is not saved: take contrasts on the live estimate, which has it.sample.internal_columnslists the columns ofsample.datathat svy made, not the user: the weight-adjustment record’s snapshots (__svy_cells_*,__svy_aux_*), which the design needs to reproduce the adjustment’s variance. Keep them when saving data the design must work with, leave them out when showing it;describe()already does.sample.describe(by=...)describes each group of one or several columns (crossed), a null being a group of its own: every item carriesbyandby_level,to_polars()has a column perbyvariable, and printed names show the group (hh (reg=North)). It is where counts and sums per group come from: a categorical column’slevels, ornandsumof a numeric one.top_k=Nonelists every level instead of the first 10. Saved describe items carryby/by_level, andDescribeResultData.top_kisNonewhen every level was listed (schemasvy-result/0.7).Results record their interval method:
Estimate.ci_method."logit","beta","korn-graubard"or"wilson"for proportions and factor means,"fisher"or"wald"for correlations,"woodruff"for Taylor quantiles, and"wald"(est ± t·se) for means, totals, ratios, covariances and replicate quantiles, so a methods appendix can say which interval was used. Saved results carry it, and GLM fits now also keeptheta/theta_se(negative binomial) and theoffsetcolumn: schemasvy-result/0.7addsEstimateData.ci_method,GLMStatsData.theta,GLMStatsData.theta_seandGLMFitData.offset,Nonewhen read from an older payload.Expr.to_polars()returns the polars expression a svy expression wraps, to filter or derive on a plain polars frame before aSampleexists. It replaces importingto_polars_exprfromsvy.core.expr.Dates in expressions.
svy.date(year, month, day)builds a date from three columns, expressions or integers;Expr.str_to_date(fmt)parses text (svy.col("dob").to_str().pad_left(8, "0").str_to_date("%d%m%Y")reads8082025as 8 August 2025). Both raise on values that are not dates;strict=Falsegives null instead, for codes such as 98.Expr.days_between(other),months_between(other)andyears_between(other)count from the date tootherin days, whole months and whole years, as Stata’sdatediff()does (31 January to 29 February is 0 months; a 29 February birthday completes its year on 1 March).SampleSize().allocate(n, pop_size=...)splits an overall sample size across strata, as a sample-size planning goal besideestimate_propandestimate_mean:method="proportional"(n_h ∝ N_h, or N_h **power;power=0.5is square-root allocation),"neyman"(∝ N_h · S_h, withsigma=per stratum, e.g. from a previous survey) or"equal"(n / H).pop_sizemaps each stratum to its unit count or to any size total the split should follow (households per stratum for a PPS design); its keys are any stratum values, tuples for several stratum columns, independent of a design. Largest-remainder rounding makes the strata sum to n;min_nfloors non-empty strata andcap_at_populationcaps n_h at N_h.SampleSize.ngives{stratum: n_h}, whichsample.sampling.srs(n=...)and thepps_*methods take as is;SampleSize.allocationlists the rows. Withoutn,allocatesplits the overall n of the goal theSampleSizeholds, rounded up:svy.SampleSize().estimate_mean(sigma=12, moe=1.5).allocate(pop_size=N_h);target.from_goal/from_nrecord it. A stratified goal (already one n per stratum) or a two-group comparison cannot be split this way.Design.case_idon several columns.svy.Design(case_id=["cluster", "hh", "line"])identifies a record by the columns together, as a CSPro export does, likestratumandpsutake a list. Everything that uses the case id takes the columns together: uniqueness (duplicates are reported as tuples), nulls, the panel checks, the case as the variance PSU on a panel,combine_samples(kind="panel", case_id=[...]),wrangling.lag, the paneladjust, renames,to_code()and the saved design (DesignData.case_idis a list). Removing one of the columns removes the case id. Results equal those with the same id built as one column.sample.check()checks the data against its design and returns one report with a section per declared part,Nonefor a part the design does not declare:weights(null, non-finite, negative and zero counts; min, max, max/min, sum, mean, Kish design effect and effective sample size over the positive weights),case_id(unique, within wave on a panel; null ids and examples of repeated ones; several id columns together),nesting(PSU codes found in more than one stratum: svy treats a PSU as the (stratum, PSU) pair, as R’snest=TRUEdoes, so this is reported, not refused) andsingletons(strata with one PSU and whether the singleton rule handles them). It never raises on the data’s content;limit=caps the examples.rake’s margin check, the join duplicate check, the panel case-id check anddeff_wshare its code.sample.meta.var_labelsandsample.meta.value_labelsgive every variable’s label and code-to-label mapping, catalog schemes resolved:meta.var_labels["RIAGENDR"]is"Gender",meta.value_labels["RIAGENDR"]is{1: "Male", 2: "Female"}. Both are read-only: an edit raisesLabelsReadOnly(also aTypeError) naming the setter, since a copy would silently not change the metadata;dict(...)gives an editable copy.
Changed
ranktestdefaults to the Kruskal-Wallis score (Wilcoxon for two groups), R’ssvyranktestdefault. Withoutscore=orscore_fn=it raisedINVALID_CHOICE, althoughscoredefaulted to None.SampleSize().compare_propsandestimate_meantake onlymethod="wald".compare_propslisted"miettinen-nurminen","newcombe"and"farrington-manning", andestimate_meanlisted"fleiss", but none were implemented and they raisedNotImplementedError. They now raiseMethodError(INVALID_CHOICE) with a hint, ascompare_meansdoes;var_mode=still sets the Wald variance for two proportions. Thecompare_*docstrings now state that the two groups are independent samples: correlation between panel rounds is not accounted for.SampleSize.targetrecords every argument of the goal, forestimate_prop,estimate_mean,compare_propsandcompare_meansas forallocate: the formula (method,var_mode),alpha,power,two_sides,delta,pop_size,deffandresp_rate, scalars or per-stratum dicts as given (it was None except afterallocate). Stratum keys are kept as given, in the given order:Size.stratumand the keys ofSampleSize.nare the input’s values (tuples for several stratum columns) instead of their text, songoes tosrs(n=...)as is; tables print a tuple asN, u, and a stratum keyed0no longer prints asoverall.read_statareturns Stata integer variables as Int64.byte,intandlongvariables were Float64, andwrite_statawrote integer columns asdouble. Integer columns now keep their type both ways, so a code1matches the value label keyed1. Needs the svy-io release with integer storage (see its changelog).
Removed
BREAKING:
combine_samples(adjust=..., wave_labels=..., kind="cs").adjust=Noneaveraged the weights andadjust="none"kept them; the choice is nowaverage_wgts=(True divides by k, False keeps each wave’s weight, unset picks True for cross-sections and False for a panel). Wave labels are the keys of a mapping, so a label cannot drift from its sample:svy.combine_samples({"2015-16": s1, "2017-18": s2}); a list still works and labels the waves “wave 1”..”wave k”. The"cs"alias ofkind="cross_sectional"is gone.adjust=andwave_labels=raisePARAM_RENAMEDnaming the replacement.BREAKING:
meta.resolve_labels(),meta.resolve_all(),sample.resolve_labels()andsample.labels. Usemeta.var_labelsandmeta.value_labels.meta.set_label()/set_labels()are nowset_var_label()/set_var_labels(), the namesSamplealready used.ResolvedLabelsis no longer exported fromsvy.metadata. Each old name raisesLabelAPIRemoved(also anAttributeError) naming its replacement.BREAKING:
sample.sampling.allocate(),sample.sampling.group_sizes()andsvy.selection.allocate(). Allocation is planning, not a sampling action: usesvy.SampleSize().allocate(n, pop_size=...), withpop_sizethe strata’s counts (e.g. fromsample.estimation.total(...)or your frame) or size totals, and pass.nto the selectors. The"rate"method is gone (a fixed fraction is n = f · N, proportional allocation), as is"size"(pass the size totals aspop_size).
0.31.0 — 2026-09-30
Added
Replication variance for
tabulate,ttestandranktest:method="replication". They took only a Taylor variance, so a replicate design still needed its strata, PSUs and a singleton rule. Withmethod="replication",tabulatere-estimates cell proportions and totals with each replicate weight and computes the Rao-Scott F and chi-square from their replicate covariance.ttestre-estimates the mean, or the two group means and their difference.ranktestre-totals the full-sample influence values with each replicate, as R’ssvyranktestdoes. The df is the replicate df (n_reps - 1unless the design records one), minus 1 for the t-tests and minus k - 1 for the k-sample rank test; it does not shrink on a domain.where=,by=, paired t-tests andscore_fnwork as with Taylor. The singleton rule, FPC and calibration sweep do not apply, and no domain-singleton findings are reported.method=Noneis Taylor, as in estimation: replication is never picked implicitly. Checked against R survey 4.5 (svymean,svytotal,svychisq,svyttest,svyranktestonsvrepdesign) for BRR, Fay-BRR, JK1, JKn, bootstrap and SDR.read_spssandread_stataread zip archives, likeread_sas: the first.sav,.zsavor.pormember, or the first.dtamember, is read, labels included. Needs svy-io with zip support for SPSS and Stata (see its changelog).read_xpt,read_xpt_with_labels,create_from_xptandwrite_xpt, aliases of the SAS functions:read_sasreads SAS Transport, andwrite_saswrites it. With svy-io’s content-based dispatch,read_sasreads transport files under any name (e.g..ssp) and refuses SAS CPORT files with a hint on converting them.catalog_path=labels transport files too.Strata with one PSU inside a domain are detected. A stratum can hold several PSUs in the sample but only one with rows in an estimation domain (
where=crossed with aby=level). Two settings of the rule cover it.domains=picks the variance:"standard"(default, R’s default) uses the usual domain variance;"apply"(centerandscaleonly) handles such a stratum as a singleton, R’soptions(survey.adjust.domain.lonely = TRUE): undercenterits PSU totals (the domain’s and the zero totals of the others) are centred at the grand mean, underscaleit is left out and counted with the singletons.on_domain_singletons=says what svy reports:"ignore"(default, also without a rule) records aDOMAIN_SINGLETON_PSUfinding on the result’s newfindings(Estimate,Table, t-test and rank-test results,GLMFit, whoseto_dict()includes it) and insample.warningsat INFO;"warn"also prints one note under the table (note: 577 domain × stratum pairs with one PSU in the domain (VARSTR=2002 in sex=1; ...); standard domain variance used);"error"raisesDOMAIN_SINGLETON. The finding is never a Python warning.sample.domain_singletons(by=..., where=...)lists the pairs. This applies to every Taylor estimate, table, t-test, rank test and GLM, and matches R. A domain row is one insidewhere=and theby=level with the analysis variables present, whatever its weight, as R’ssubset()keeps zero-weight rows (the degrees of freedom andnkeep counting nonzero weights, as R’s do). A calibrated design has none, as in R.non every estimate row.ParamEst.n,TtestEst.nand a table’sCellEst.ngive the records behind the row: those in its domain (where=,by=level, t-test group, or the table) with a nonzero weight and the estimate’s variables present. It is the count proportion CIs already use, for Taylor and replication alike; a proportion’s category rows carry the domain’s count, each group of a two-group t-test its own, and every cell of a table the table’s count. A GLM fit’sstats.nalready followed the same rule.to_polars()has anncolumn (the printed table does not). Saved results carry it: schemasvy-result/0.5addsParamEstData.n,TtestEstData.nandCellEstData.n,Nonewhen read from an older payload. Withdeff="wr",n / deffis the effective sample size.sample.to_code(data=None)writes the Python that rebuilds the sample:import svy,data = svy.read_parquet(path)whendatapoints to a.parquetfile (the pair ofsvy.write_parquet(sample, path)), andsample = svy.Sample(data, svy.Design(...))with every setting that changes the estimates: replicate weights, the weight-adjustment record and singleton handling. Running it on the same data gives the same design and estimates. The data is the caller’s; metadata (labels) is not included.svy.WgtAdjustmentis exported so the script can name it.A design can be saved and restored.
svy.serialize.serialize(design)/to_json(design)give aDesignData(schemasvy-design/0.1) holding every field: strata, PSUs, weights,PopSize, replicate weights with all their settings and the weight they go with, and the weight-adjustment record.svy.serialize.to_design(data)returns the liveDesign, equal to the original. Estimates onSample(data, to_design(...))equal those on the live sample, SEs included, after poststratification, raking, calibration, standardization and trimming, Taylor and replication alike. The design history is not part of it.The singleton rule is declared on the design:
svy.Singleton.svy.Design(..., singleton=svy.Singleton("center")), orsingleton="center"for short, orsample.update_design(singleton=...), like R’slonely.psuand Stata’ssingleunit().methodis"center","scale","skip","self_representing"(the formercertainty(): R’s and Stata’s “certainty” is svy’sskip),"collapse"or"pool"; the other options are keyword-only and belong to one method:collapse‘susing("smallest","largest","next","previous", a mapping{singleton stratum: target}in the stratum columns’ values, or a callable),within,order_by,descending,rstate, andpool’sname; an option for another method raisesINVALID_SINGLETON_RULE. The rule is intent: svy applies it to the singletons the data has now, when the sample is built and after every filter, recode or design edit, and is idle without singletons.sample.singletonslists the singleton strata of the current data and how the rule handled each (handled:center,collapse -> Center,pool -> __pooled__, …), andsample.n_singletonscounts them; the rule is read back fromsample.design.singleton. An explicit mapping that does not fit the data raises at the next Taylor analysis (SINGLETON_UNMAPPED,SINGLETON_TARGET_MISSING,SINGLETON_TARGET_SINGLETON); its entries for strata that are no longer singletons are ignored with an INFO finding, and a strategy whose mapping changes recordsSINGLETON_COLLAPSE_CHANGED.within=takes any column constant within a stratum, missing values ignored (it silently did nothing for a non-stratum column); an explicit mapping across it raisesSINGLETON_TARGET_OUTSIDE_WITHIN.order_by=is checked the same way: a column taking several values within a stratum raisesSINGLETON_ORDER_BY_INVALID(it used to order the strata by whichever value came first). The rule is indesign_history, saved bysvy.serialize(SingletonData, schemasvy-design/0.1, defaults left out; a callableusingor a Generatorrstatecannot be saved) and written byto_code();reprand the printed design show it (Singleton('collapse', within='region')). It stays through renames (within/order_byfollow them), stratum/PSU/SSU edits, weighting anduse_weight;remove_columnsprotectswithin/order_bycolumns, and the rule goes when the design loses its stratum.combine_samplescarries a rule every input declares (different rules, or an explicit mapping on cross-sections, raise) andadd_stagecarries stage 1’s. Replicate weights are built from the rule’s strata and PSUs undercollapse,poolandself_representing; underskipthe singletons contribute nothing (recorded);create_jk_wgtsandcreate_bs_wgtsraiseSINGLETON_REPLICATESundercenter/scaleor without a rule, where they used to drop the singletons silently.filter_records(check_singletons=True)reports singletons the rule handles at INFO.svy.SingletonMethodnames the methods;svy.utils.deprecatedmarks deprecated API.Strata named by their own values. A collapse mapping (
svy.Singleton("collapse", using={49: 48})), ausingcallable’s return value,candidates_forandcomparetake the stratum columns’ values (a tuple for tuple strata:{(True, 4): (False, 1)}). svy’s key strings ("49","true__by__4") still work; a string that names two strata raises.design.columns(data_columns=None)lists every column a design needs from the data, de-duplicated and in order: design fields,pop_size, replicate weights, the units the replicates were built from, then the weight-adjustment record’snew_wgt,prev_wgt,cellsandaux.data_columnsresolves auto-detected replicate padding.specified_fieldsis unchanged.wrangling.rename_rep_wgts({"w": "final_w", "ps_wgt": "ps_final"})renames replicate-weight sets by prefix: the design’s and those of earlier designs indesign_history. Each column keeps its number and zero-padding (w01→final_w01), only the set’s own columns are renamed, and every design using the set, current or earlier, follows through therename_columnspath. An unknown prefix (the error lists the sets the sample carries), an empty new prefix, a new name that is already a column, and two sets renamed onto the same names are refused before anything changes; a prefix mapped to itself is a no-op.RepWgts.wgt, the full-sample weight the replicates go with.Designfills it from its ownwgt;create_*_wgtsand weighting with replicates pair them with the weight they produced.Threshold.quantile(p)andThreshold.absolute(v). Trimming bounds say what they are: thepquantile of the positive weights, or the value itself. Both compose with the statistic kinds (Threshold.quantile(0.99) + 2 * Threshold("sd")), and reprs are runnable (Threshold.quantile(0.99)).quantile(99)is refused with a pointer to 0.99.Thresholdis now the class andCapits alias, so existing code runs unchanged.svy.serialize.to_polars(data, *, row_index=None, **options)gives a serialized result’s table, the same frame its live result’sto_polars()returns, from the payload alone (for example afterfrom_json). It covers every result kind that serializes (estimates, estimate lists, t-tests, tables, chi-square tests, GLM fits, GLM predictions and describe results), with the same options (tidy,use_labels,component,exponentiate). A serialized estimate stores the labels of the variables and levels it holds, so its label columns come back without the sample’s metadata.row_index="row"adds a first column holding each table row’s position in the payload’s rows; estimate tables are sorted for display, so this is how a table row is matched back to its payload row. The live and payload tables are built by the same function per result kind.ChiSquare.to_polars()andDescribeResult.to_polars(). A one-row frame ofdf,valueandp_value, and one row per described column with its scalar fields (nested frequency tables and percentile lists are left out).GLMFit.alpha, the level of the coefficient intervals, fromglm.fit(alpha=). The printed interval headers follow it ([0.05 0.95]atalpha=0.1; they always read[0.025 0.975]), as does the p-value highlight.GLMFitDatacarries it; a payload without it decodes as 0.05.GLMFit.where_clause, thewhere=domain of the fit, formatted asEstimate.where_clauseis and printed under the modeled variable.GLMFitDatacarries it (defaultNone).
Changed
Breaking:
httpxis optional; a base install makes no network calls. The online dataset catalog moves to the newremoteextra (pip install "svy[remote]", also insvy[all]), andimport svyno longer importshttpx. Without it,datasets.load(source="remote"),catalog(source="remote")anddescribe(source="remote")raiseREMOTE_UNAVAILABLE, whose hint names the extra; under the defaultsource="auto"they return the bundled subsets without a warning (a dataset with no bundled subset, orforce_download=True, raises).source="auto"now reads a full copy already in the local cache (~/.svy/datasets, orSVYLAB_CACHE_DIR) before trying the network, the most recently downloaded version first, so a dataset downloaded once, or a copied cache folder, loads offline;force_download=Trueandsource="remote"still go to the catalog.great-tablesis dropped fromreportandall: nothing used it.Breaking:
ranktesttakes the rank score asscore=;method=is the variance method.ranktest(y, group=..., method="kruskal-wallis")becomesranktest(y, group=..., score="kruskal-wallis"), somethod=means"taylor"or"replication"as intabulate,ttest,glm.fitand estimation. Passing a rank score asmethod=raisesINVALID_CHOICEwith a hint namingscore=.Breaking:
glm.fituses Taylor linearization unlessmethod="replication"is passed. It used replicate weights whenever the design had them, unlike estimation and the categorical tests, where replication is never implicit.glm.fit(..., method="replication")gives the previous replicate standard errors. With the Taylor default on a replicate design, the singleton rule applies andTAYLOR_WITHOUT_DESIGNwarns when the design has no stratum or psu.Breaking:
PROP_CI_BOUNDARYis a note under the table, not a warning. When a proportion is 0 or 1 and the interval method has none there (logit, beta, wilson),prop()andmean(as_factor=True)no longer raise aSvyUserWarning; the estimate prints one line under its table,note: CI undefined at p = 0 or 1 for 19 rows (ci_method='logit'), on every such result, a repeat on an unchanged sample included. The NaN bounds are in the table, so the note is shown by default. The finding is kept likeDOMAIN_SINGLETON_PSU: it keeps its detail and hint, is on the estimate’sfindings(WARNING) and insample.warnings(INFO), and itsextraholdsn_rowsand every cell (by,level,p). A list of estimates prints one note, split by variable (for 5 rows (a: 3, b: 2; ci_method='logit')). Code that relied on the warning, e.g. running with warnings as errors, should checkestimate.findings.to_polars()is unchanged.Breaking: t-tests, rank tests, tables and GLMs refuse unhandled singleton strata.
categorical.ttest,categorical.ranktest,categorical.tabulateandglm.fit(Taylor variance) raiseSINGLETON_ERROR, as estimation always has, when the design has strata with one PSU and no rule for them. They used to leave those strata out of the variance silently. Declare a rule,sample.update_design(singleton=svy.Singleton(...));svy.Singleton("skip")reproduces the old numbers. A GLM on replicate weights is unaffected.Breaking:
__svy_columns are svy’s; everything insample.datais your data or your design’s. svy’s bookkeeping is renamed and kept out of sight: the row indexsvy_row_indexis now__svy_row_index__, the design keysstratum_svy_internal_cols_concatenated(andpsu_,ssu_) are now__svy_stratum_key__,__svy_psu_key__,__svy_ssu_key__, and the singleton variance columns keep their__svy_var_*__names. None of them is shown bysample.data,show_data,dtypes,metaor the printed column count, or written bysvy.write_parquetand the other writers; svy rebuilds them on construction and on every change.sample.dataand the saved file hold the user’s columns plusdesign.columns(), the weight-adjustment record’s__svy_cells_*/__svy_aux_*snapshots included. Thesvy_*selection outputs (svy_prob_selection,svy_sample_weight, …) are unchanged. A column named like the bookkeeping, inSample(...),set_data,update_data,clone(data=),mutate, the value writers (into=),castorwith_row_index, raisesRESERVED_COLUMNnaming it; drop it from a frame taken fromsample._data, or from a file written by an earlier version (its__svy_var_*__columns). Such a file’ssvy_row_indexand*_svy_internal_cols_concatenatedcolumns are now ordinary columns; drop them too.The
PROP_CI_BOUNDARYwarning names thebycolumns (pov=0 in quint=poorest,q=poorest, zone=yfor several) instead of svy’s key column, and lists the cells in domain order.Breaking: the templates are consistent.
controls_margins_templateandcontrol_aux_templatereturn the columns’ own values as keys (an Int column gave{'1': nan}) and take one missing-value option,na="error" | "level" | "drop"(default"error") withna_label;cat_na=andby_na=(also onbuild_aux_matrix) are refused with a pointer tona=.controls_margins_templateno longer includes nulls as a level by default.Breaking: weighting parameters are named for their values.
rake(ll_bound=, up_bound=)becomerake(bounds=(lo, hi)): bounds on the factor g = new/old weight,Noneon a side for open,Nonefor none. They are checked after raking, not enforced, as before (BOUNDS_EXCEEDED, now withexpected={"bounds": [lo, hi]}). A badbounds(not a 2-tuple, a non-number side, NaN or infinity, lo > hi) raisesINVALID_TYPEorINVALID_RANGE.calibrate(bounded=)andcalibrate_matrix(bounded=)becomebounds=, which raisesNOT_SUPPORTED(aNotImplementedError) when a side is set.calibrate_matrix(control=)becomescontrols=.strict=onpoststratify,rake,calibrateandcalibrate_matrixbecomeson_nonconvergence="error" | "warn" | "ignore"(default"error";strict=Trueis"error",strict=Falseis"warn"), coveringMAX_ITER_REACHED,CONVERGENCE_FAILED,CALIBRATION_NOT_METand the trim cycles. The old names are refused withPARAM_RENAMED, whose hint names the new parameter and shows the caller’s values in the new form (bounds=(0.5, None)). Every public weighting method now documents every parameter.Breaking:
standardize,adjustandtrimraise on non-convergence by default. They takeon_nonconvergence="error" | "warn" | "ignore"(default"error") like the other iterative methods: forstandardize’s trim-standardize cycle,adjust’strimming=pass andtrim’s iterations in any domain. They used to keep the result and warn (MAX_ITER_REACHED), which is nowon_nonconvergence="warn". With"error"the call raisesCONVERGENCE_FAILED(fortrim, naming the domains that did not converge) and the sample is left as it was.Breaking:
rake’s trim-rake cycles are capped bytrimming.max_iter, as inpoststratifyandcalibrate;rake(max_iter=)now caps only the raking iterations of each pass. The cycles used to be capped byrake(max_iter=)andtrimming.max_iterwas ignored there. The finding and error nametrimming.max_iter.Breaking: every
on_*parameter means the same."error"raises and leaves the sample as it was;"warn"records the finding and raises it once;"ignore"records it at INFO without raising (it used to record nothing). This applies toon_nonconvergence,wrangling.join(on_unmatched=)(JOIN_UNMATCHED),wrangling.filter_records(on_singletons=)(SINGLETONS_DETECTED) andcombine_samples(on_mixed_design=)(COMBINE_MIXED_DESIGN, which was only logged). Defaults are unchanged.filter_records(on_singletons="error", inplace=True)now leaves the sample unfiltered when it raises, and an unknownon_singletonsvalue is refused withINVALID_CHOICE.Breaking: weighting errors are data. Each input failure has a specific code (
CONTROLS_KEYS_MISMATCH,MARGINS_DISAGREE,RESP_STATUS_UNKNOWN, and others undersvy.errors.WeightingError, aMethodError),expected/gothold structured values (the cells present,{"missing": [...], "extra": [...]},{margin: total},{status: count}) instead of text, and hints use the caller’s names and values.SvyError.to_dict()keeps them structured and JSON-safe. Rake’sINVALID_CONTROL_TOTALS,ZERO_CONTROL_TOTALSandINVALID_SHARESbecomeCONTROLS_*codes.Breaking: weighting targets match the column’s values, with a text fallback. Every keyed input (
poststratify/normalizecontrols and shares, eachrakemargin,standardizeshares,calibratecontrols and domains,adjustresp_mapping) matches a key to the column’s own value first, then to its text form ("1"for 1,"true"/"True"for booleans,"1.0"/"1"for 1.0, ISO strings for dates), so controls stored as JSON work unchanged. Tuple keys match part by part. A key that could mean two levels, or two keys naming one level, is refused. One rule for extra keys everywhere: a key naming no level in scope is an error unless its target is 0 (calibrate used to drop them silently; rake failed in the kernel). A value listed under tworesp_mappingcodes is refused instead of the last one winning.Breaking: one rule for findings. A finding about the data or a sample is recorded in
sample.warningsand raised once as asvy.SvyUserWarning(aUserWarning) at the caller’s line, with the message"[CODE] title: detail".Sample.warnapplies it for every finding. An identical finding (same code and detail) is recorded and raised once per sample state: the state is the sample’s data version, which changes on every data or design rebind (set_design,update_design, wrangling), so the same finding after such a change is a new event and is raised again, while repeating it on an unchanged sample is not. A fork has its own state and does not raise its parent’s findings; the store’s per-code cap (100) counts across states. A finding is raised when it tells the caller something they did not ask for, and only recorded (INFO) when it confirms a consequence of an explicit choice: trim audits,WEIGHT_SUM_CHANGEDfromtrim(redistribute=False)(it was a WARNING), andSELECTION_N_EXCEEDS_GROUPunderwr=True.force=Trueremovals stay raised, since they also report what else was dropped. Findings that svy raised as a bareUserWarning, or only logged, now follow the rule with a code: combining samples (COMBINE_*,PANEL_UNITS_LOST,PANEL_SMALL_OVERLAP,WAVE_LABELS_UNORDERED, recorded on the combined sample),add_stage(STAGE_PSUS_UNMATCHED), selection (SELECTION_*), GLM (GLM_ROWS_DROPPED,GLM_TERM_DROPPED,MAX_ITER_REACHED), removed design fields (DESIGN_FIELDS_REMOVED), uncredited weight adjustments (ADJUSTMENT_NOT_CREDITED), replicate weights removed by a weight change (REP_WGTS_RESET) and cleared singleton handling (SINGLETON_RULE_CLEARED). Where there is no sample (a bareDesign.update, allocation, datasets, SAS export, accessor registration) the warning is raised only, with the same category and line. Code that runs with warnings as errors now sees svy’s findings, including ones that were only recorded before (e.g.TAYLOR_WITHOUT_DESIGN); filter them withwarnings.filterwarnings("ignore", category=svy.SvyUserWarning). Messages now start with the code.In weighting: rows left unadjusted because a cells,
byor trimming-domain value is null are a finding,CELLS_NULL_UNADJUSTED, with the count per column ingot({"zone": 3}) and the number of rows inextra, frompoststratify,normalize,standardize,adjust,trimand calibrate’strimming.by; the rows still keep their previous weight, and nulls outsidewhere=are not a finding. Non-convergence kept underon_nonconvergence="warn"isMAX_ITER_REACHED(rake, and the trim cycles ofpoststratify,standardize,rakeandcalibrate); rake no longer prints “Warning: Raking did not converge”.standardize’s partial domains areDOMAIN_LEVELS_PARTIALand paneladjust’s skipped missing-in-scope rulePANEL_SCOPE_NOT_WAVES. Onlydisplay_iter=Trueprints.on_nonconvergence="error"still raises.Breaking:
adjust(respondents_only=True)drops only the rows the adjustment handled. A row in no adjustment class, because its cell is null or it is outsidewhere=, was not adjusted, so it now keeps its weight and stays in the sample whatever its response status; before, a nonrespondent, unknown or ineligible among them was dropped and its weight left the sample. This matches thewhere=documentation and paneladjust. Null-cell rows are reported byCELLS_NULL_UNADJUSTED, withextra["nonrespondents_kept"]; rows outsidewhere=are the caller’s choice and not a finding.Breaking: domain and category levels keep the column’s type.
ParamEst.by_levelandy_levelheld the kernel’s text:by="zone"gave("3",), a booleanbygave("false",), andmean("zone", as_factor=True)gave"1.0". They now hold the column’s own values ((3,),(False,),1) for every estimator, withwhere=, with severalbyvariables (a tuple of values), and under Taylor and replication alike; so doEstimate.domains,keys(), the t-test’sby_levelandgroup_level, the rank tests’by_levelandgroup_levels, and JSON payloads. A proportion of an integer-valued float column now reports1.0, not1, and string codes such as"01"are no longer read as integers. Test headers print numbers unquoted ([1 vs 2]).contrast()still accepts the old string keys ({"3": 1, "1": -1},"false","1.0"). A level of a date, datetime, time or duration column goes into JSON as its ISO string with its type recorded under a top-level"temporal"field, andfrom_jsongives the value back. Printed tables are unchanged, except thatas_factorlevels print as the column’s values (1, not1.0).Breaking: a bare number is always an absolute trimming bound.
trim(upper=0.9),TrimConfig(upper=0.9)andtrimming=read a number in (0, 1] as a quantile and a larger one as a cap, so an absolute cap of 0.9 (common after normalizing) silently became the 90th percentile. A number now always caps at that value; writeThreshold.quantile(0.9)for the 90th percentile.Estimate.to_polars()andEstimateList.to_polars()are a data view, no longer the printed table. Levels keep their codes under the variable’s name, and each variable with value labels gets a<var>_labelcolumn next to it (by_labelandy_level_labelwithtidy=False). Rows sort by code.EstimateList.to_polars()leads with aycolumn when its members’ variables differ, and anxcolumn when their denominators do. Labels are on by default for both (use_labels=Falsedrops the label columns); before,Estimate.to_polars()gave codes only andEstimateList.to_polars()put labels in place of the codes and variable labels in place of the names. A category level now keeps the type the estimator gives it, so a proportion’s level column is an integer for an integer variable. Printing is unchanged (to_polars_printable()). A label column whose name is already a variable raisesLABEL_COLUMN_CLASH.Serialization schema
svy-result/0.4.EstimateDatagainsas_factor, which decides whether the table has a level column, andlabels.TDistData.dfis nowint | float, so a GLM coefficient’s design df stays an integer. All additive:0.3payloads still decode.Breaking:
sample.singletonis removed; singletons are handled on the design. The rule is declared once withsvy.Design(..., singleton=...)orsample.update_design(singleton=...)(svy.Singleton, above), so there is one place to specify it.sample.singleton.center()becomessvy.Design(..., singleton="center"),.scale()/.skip()/.pool(name=)/.collapse(using=, within=, ...)the method of that name,.certainty()"self_representing"(R’s and Stata’s “certainty” is svy’s"skip"),.handle(method, ...)svy.Singleton(method, ...), and.combine(mapping)svy.Singleton("collapse", using={singleton: target})for strata orsample.wrangling.recode(psu_column, {new: [old, ...]}, replace=True)for PSUs. Accessingsample.singletonraisesSINGLETON_API_REMOVED(also anAttributeError) with these spellings. Its inspection methods are replaced by facts on the sample, likesample.strataandsample.n_psus:detected()/show()/keys()/exists/count/last_resultbysample.singletonsandsample.n_singletons; the collapse helpers (candidates_for,suggest_mapping,compare,rank_strata,strata_profile,variance_contributions,summary()) andraise_error()are gone (helpers for choosing a rule may return in another form).svy.SingletonResult,svy.SingletonSummaryandsvy.SingletonHandlingare removed;svy.SingletonMethodnames the methods. A rule now follows the data: a filter that changes the singletons applies it to them, instead of leaving variance strata built for the old data.Breaking:
update(wgt=X)keeps the design consistent with X.sample.update_design(wgt=X)keeps the record and replicate weights when X is the current weight; on a weight an earlier design indesign_historyproduced, and whose columns are all still in the data, restores that design’s record and replicates (switching back to the raked weight after a trim restores the raking record); on any other column drops the record silently and drops the replicate weights with a warning, unlessrep_wgts=is passed, which pairs them with X.use_weightfollows the same rule. A baredesign.update(wgt=X)keeps the record only if X is itsnew_wgtand the replicates only if X is theirwgt. An explicitwgt_adjustment=orrep_wgts=always wins. Edits to other fields keep both.Breaking: a
Designwhoserep_wgts.wgtorwgt_adjustment.new_wgtis not itswgtraises.Breaking:
Sample(data, design),set_designandupdate_designraise when the data lacks anydesign.columns()entry, record columns included. A frame without them is not the frame the design describes; estimation used to fall back to fixed-weight standard errors with a warning.Breaking:
ignore_reps=Trueleaves the new weight without replicate weights inadjust,normalize,poststratify,standardize,rake,calibrateandcalibrate_matrix. The unadjusted replicates used to stay on the design next to a weight they do not go with. Their columns stay in the data and the previous design keeps them, soupdate_design(wgt=<previous weight>)brings them back.Breaking: wrangling no longer writes new values into a weight column under its name. The design’s weight, every earlier design’s weight, replicate weights, and the weight-adjustment record’s columns (current and earlier,
__svy_cells_*/__svy_aux_*included) are refused as targets ofmutate,recode(replace=True),fill_null,top_code/bottom_code/bottom_and_top_code(replace=True),categorize(replace=True)andcast, inplace or not, and any other wrangling step that would change their values is refused too. An exact widening cast (Float32 to Float64, an integer to a wider integer or to a float that holds it exactly) is allowed. A weight overwritten in place is a new variable that a later weight’s record or the replicates would still read as the old one. TheWEIGHT_OVERWRITEerror names each column and what reads it, and shows both ways out: write under a new name, orrename_columnsfirst (which the design and its history follow), then write the freed name. Strata, PSUs and other design columns stay writable.set_data,update_dataandclone(data=)apply the same rule to a new frame, compared positionally: a protected column whose values change is refused withWEIGHT_OVERWRITE, and a frame with a different number of rows is refused withDATA_ROWS_CHANGEDwhen the sample has a weight-adjustment record, replicate weights or design history (rows changed outside svy cannot be checked against them; the hint points tofilter_records,joinandcombine_samples). A plain design still accepts a new row count. The checks run before anything is rebound, and a later validation failure restores the sample.combine_samples(kind="panel")compares the waves’ replicate designs without their paired weight and pairs them with the combined weight.Estimate.q_methodisNoneexcept on medians and quantiles, and so isEstimateData.q_method(nowstr | None); every other result carriedLinear. Payloads that sayLinearstill decode.Every estimator refuses an
alphaoutside (0, 1), NaN included, withINVALID_RANGE, and a non-number (a string, a bool,None) withINVALID_TYPE:glm.fit,predict,marginsandcontrast,estimation.mean,total,prop,ratio,median,quantile,corrandcov,categorical.tabulate,ttestandranktest,Estimate.contrast/EstimateList.contrast, andSampleSize.estimate_prop,estimate_mean,compare_propsandcompare_means(each value of a per-stratum mapping). The check runs before any work.alpha=0used to give infinite intervals, and a string orNonefailed deep inside.Printed estimates name the estimated variable. Means, totals, medians and quantiles print a
y: <var>line abovewhere:, and ratios printy / x: <y> / <x>. With labels on, the variable label replaces the name. Proportions andas_factormeans are unchanged, since their level column is already headed by the variable. AnEstimateListputs what its members share in the same lines, replacing the: <var>title suffix, and ayorxcolumn for what varies; those columns followuse_labelstoo. The plain-text printer of anEstimateListnow showswhere:, as the rich one did.
Fixed
glm.fit(method="replication")used the Taylor df when the replicate weights recorded none. WithoutRepWgts.df, the replicate fit kept the design df (#PSUs - #strata - (k - 1)), which also shrank on awhere=domain. It now usesn_reps - 1less k - 1, like estimation and the categorical tests, and does not shrink on a domain. Matches Rsvyglmonsvrepdesign(degf = n_reps - 1).The singleton error on a replicate design did not say that the replicates were not used.
tabulate,ttestandranktestcompute a Taylor variance from the stratum and PSU columns even when the design has replicate weights. Estimation does the same unlessmethod="replication"is passed. Without a singleton rule they raiseSINGLETON_ERROR. On a replicate design the error now says that this variance is Taylor linearization and that the replicate weights are not used. Its hint saysmethod="replication"uses the replicates (estimation,tabulate,ttest,ranktestandglm.fit).add_stageresults did not detect singletons. The combined sample was built without its design keys, so a stratum left with one PSU (after a filter, say) went unnoticed and Taylor estimates left it out silently. It is now checked like a constructed sample, and applies a singleton rule carried from stage 1.A zip holding no readable file was reported as missing.
read_sason a.zipwith no.sas7bdat,.xpt,.xportor.sspmember raisedFILE_NOT_FOUND(“No file at x.zip”) for an archive that exists.read_sas,read_spssandread_stata(and theircreate_from_*) now raiseARCHIVE_MEMBER_NOT_FOUND, whose detail lists the archive’s files. Any other file a reader could not find while the data file exists is now named in the error instead of the data file.ranktest(method=svy.RankScoreMethod.VANDER_WAERDEN)raisedUnknown rank method.RankScoreMethodmembers are now accepted as they are, and"vanderWaerden"is accepted as a string.A printed ratio list with several denominators did not say which row was which.
ratio("a", ["b", "c"])printed rows with noxcolumn; it has one now, asto_polars()already did.corr/covrows did not say which pair they were.to_polars(),to_polars_printable()and the printed tables now haveyandxcolumns after anybycolumns, in the order the pairs were requested; with labels on, the display shows variable labels and the data view addsy_label/x_labelwhen a pair variable has one.svy.serialize.to_polarsmatches.keys()end with the pair (("api00", "api99"),("E", "api00", "api99")) instead of repeating"api00"or the domain, and a contrast key may name the pair either way round. Abyvariable namedyorxraisesPAIR_COLUMN_CLASHfromto_polars(). Contrasting a multi-rowcorr/covresult now says that these results carry no between-row covariance, rather than pointing atprop().Nulls outside the
where=domain raised. Withoutdrop_nulls, estimation required every analysis column to be complete on every row, so a nully,bylabel or domain flag on rows outside the domain (people without events after a full join, skip patterns) forceddrop_nulls=True. As in R’ssubset(), a null in a column read only bywhere=now makes the row out-of-domain, and nulls in analysis columns on out-of-domain rows are ignored. Nulls inside the domain still raise, now saying so, and design columns must still be complete on every row. Out-of-domain rows keep contributing to the design, so estimates, SEs and df match R’ssubset(). Coversmean,total,prop,ratio,quantile,median,corr,cov,ttest,ranktestandglm.fit(drop_nulls=False), Taylor and replication;tabulatealready worked this way.cov/corrwithdrop_nulls=Trueunderstated Taylor SEs. A missing value in the second or later column dropped the row from the design, deleting whole PSUs when all their values were missing. It now zeroes the row’s weight, as foryand R’sna.rm=TRUE.Int8, Int16, UInt8 and UInt16 columns crashed the estimators. A stratum, PSU,
pop_size,by,whereor t-testgroupcolumn of one of these dtypes (a Stata byte read with svy-io, or a.cast(pl.Int8)indicator) made every estimator fail withcannot create series from Int8, Taylor and replication alike. The response was not affected. Needs svy-rs with small-integer dtype support (see its changelog).singleton.scale()on a calibrated design scaled each domain by its own strata. R keeps a calibrated design’s out-of-domain rows, whose scores are nonzero, sonstrat/nokstratcounts every stratum of the sample; svy counted the strata of the domain, andby=/where=SEs afterpoststratify,rake,calibrateorstandardizediffered from R. They now match.The singleton rule reached only estimation.
ttestandranktestnever receivedsingleton.center()orsingleton.scale()(the setting was read from an attribute that does not exist),tabulateignored every rule, including the recoded strata ofcertainty(),collapse()andpool(), andglm.fithad no singleton handling. All of them left singleton strata out of the variance, R’slonely.psu = "remove". They now apply the sample’s rule as estimation does, R’s"adjust"and"average"included, withinwhere=domains too, for every GLM family. t statistics, table SEs and the Rao-Scott F, Gaussian and logistic GLM SEs, and rank tests match R.singleton.center()gave the wrong variance for domain totals. A singleton stratum’s PSU total was centered at the mean of every PSU total in the sample, and a singleton stratum without domain rows still contributed. R, which estimates a domain on the subsetted design, averages over the PSUs of the strata holding domain rows and leaves the other strata out. svy now does the same forwhere=,by=and rows dropped for missing values, in totals, per-category totals, means, ratios, proportions, quantiles and correlations. Means, ratios and proportions rarely moved, since their scores sum to zero within the domain; totals did (by 5-7% in the test data). A calibrated design keeps the whole sample, as in R: its scores are nonzero outside the domain. Covariances between domains are unchanged, as in R’ssvyby(covmat=TRUE).Table levels came back in the kernel’s order and sorted as text.
tabulatecells,to_polars(),rowvals/colvals,crosstab()and printed tables listed10before2, and an Enum’s levels alphabetically. Levels are now ordered on the column’s values: an Enum’s in the Enum’s order, numbers numerically. Cells are listed by row level, then column level.trimming=cycles stopped after one pass.poststratify,standardize,rake,calibrateandcalibrate_matrixchecked the trimming bounds right after trimming, where they always hold, so every cycle ended after its first pass andtrimming.max_iterhad no effect: a cap that needed more than one trim-and-readjust pass failed however many cycles were allowed. The bounds are now checked on the readjusted weights. Calls that converged before give the same weights.An unreachable control in a
trimming=cycle isTRIM_INFEASIBLE, notCONVERGENCE_FAILED/MAX_ITER_REACHEDwith a hint to raisemax_iter. When a cycle fails and a control is outside the totals its units can reach within the bounds (a cell, margin level or calibration column; perby=domain), the error names it with its total, reachable range, unit count and, for a cell or level, the bound its units need on average. It followson_nonconvergence:"warn"and"ignore"keep the last cycle’s weights and recordTRIM_INFEASIBLE. Forrakeonly fixed bounds (a number orThreshold.absolute) are diagnosed, since a relative bound is resolved again each cycle.A GLM fit on separated data reported meaningless standard errors, silently. When the response is perfectly predicted on some rows (quasi-complete separation, e.g. a
Catlevel that is all 0), the estimates along that direction are infinite; IRLS stops once the deviance settles and printed the stopped values with SEs that are rounding (sometimes NaN, with a raw numpy warning and t = 0, p = 1) or small and finite with p < 1e-20, as R’s svyglm does. The fit now finds the rows it drove to the boundary and the coefficients they alone determine, for binomial (every link), Poisson and negative binomial fits,where=included. Those coefficients keep their stopped estimate (predict()is unchanged) but their SE, t, p, CI and covariance row and column are NaN, as are the model Wald tests and AIC; with replicate weights, a coefficient separated in any replicate refit is NaN too. AGLM_SEPARATIONfinding names the coefficients, the rows, the combination still identified (_intercept_ + quint_nat_poor), and for aCat, the levels on which the response never varies. The identified coefficients are unchanged and match R and Stata’ssvy: logit.term_test, contrasts, margins and prediction SEs are NaN exactly where they depend on an unidentified coefficient.trimlost or gained weight total silently when redistribution was impossible. Withredistribute=True, a bound on the wrong side of a domain’s mean positive weight (upperbelow it,lowerabove it) ended with every weight at the bound, the total changed and only the INFO audit recorded (converged=True, soon_nonconvergencenever applied). It now raisesTRIM_INFEASIBLEbefore anything is written, whateveron_nonconvergence, naming each failing domain with its bound, weight range, mean,nand total ingot; the hint suggests a bound relative to the weights (Threshold.quantile(0.99),Threshold("median", 3.5)) orredistribute=False. Also applies totrim’sby=/where=domains andadjust(trimming=).create_brr_wgtsandcreate_jk_wgts(paired=True)crashed with a multi-column PSU when strata had to be paired. The replicates now equal those built from a single column holding the same composite key.wrangling.distinct()withoutcolsnever removed duplicates: svy’s row index took part in the check. Only the user’s columns are compared now.add_stagewith a next-stageSamplewhosepsuis the stage-1 PSU madessuequal topsu, and estimation then crashed. The next stage’spsuonly links rows to their stage-1 PSU; the combined design’sssuis nowNone.clean_namesrenamed the selection outputs (svy_sample_weight,svy_prob_selection,svy_number_of_hits,svy_certaintyand their_stage1/_stage2forms) under camel, pascal, kebab, upper or title case; it reserved names svy never writes. They are kept now, and a user column cleaned into one of them gets the_1suffix instead of the selection output.add_stagewith aSampleas the next stage carried that sample’s design keys into the combined sample, where they described the next stage’s strata rather than the combined design’s. They are left behind; svy builds its keys from the combined design.trimandcalibratecheckedwgt_nameonly after doing the work. A refusedtrim(wgt_name=<existing column>)still raised the findings of the trim it then discarded, andcalibratewith a taken name and unmet controls raisedCALIBRATION_NOT_METinstead ofWGT_NAME_EXISTS. Both now check the name first.Taylor
mean,propandratioreturned no rows when one domain had only zero weights, for everyby=level and silently; awhere=to such a domain returned nothing too. The domain now keeps its row with NaN estimate, SE and CI (df 0), as R’ssvymean(subset())and as svy’s replication path already did; the other domains are unaffected. A ratio domain whose weighted denominator is zero is NaN the same way.totalis unchanged (0, SE 0). Kernel errors inmean/total/prop/ratioare no longer turned into an empty result.calibrateandcalibrate_matrixnever checked that the controls were met, except withweights_only=True: a singular system or inconsistent controls stored weights that missed them. Every path now checksX'wagainst the controls (relative tolerance 1e-4, absolute for a zero control; well-posed calibrations in the test suite miss by at most 7.5e-9).on_nonconvergence="error"(the default) raisesCALIBRATION_NOT_METbefore anything is written, with the targets inexpected, the achieved totals ingot(per domain, only the domains that miss, withby=) and the largest relative miss inextra;"warn"and"ignore"keep the weights and record the finding. A trimming cycle keeps its own check (CONVERGENCE_FAILED/MAX_ITER_REACHED).A failed weighting call could leave an
inplace=Truesample half-changed.poststratify(trimming=..., strict=True)wrote the new weight and design before its trim-cycle check, so on failure the sample was modified although the error said it was not;adjustwith an invalidtrimming=failed after its weight and rows were written. Every weighting call now runs on a private copy and is adopted only on success, so a failure leaves the data, design and history as they were, inplace or not (diagnostics recorded before the failure are still kept withinplace=True). The poststratify and standardize trim cycles now run before anything is written.trim(by=...)with nulls inbycrashed on a text column and silently left the rows untrimmed on a numeric one; those rows are left untrimmed and recorded (CELLS_NULL_UNADJUSTED). The same for calibrate’strimming.by. Apoststratifywhose in-scope rows all have a null cell raisesNO_ROWS_IN_SCOPEwith the null counts.calibrate(where=..., trimming=...)lost the diagnostics recorded while calibrating the scoped rows (skipped trimming domains, non-convergence); they are now on the returned sample.calibrate_matrix(by=)with nulls inbycrashed with aTypeError; it raisesBY_NA.poststratifywhosewhere=matched nothing raised a kernelValueError; it raisesNO_ROWS_IN_SCOPE.String keys failed on non-string columns.
poststratify({"1": ...})on an Int column raised a keys mismatch, andrakeon a Boolean column failed with a barenp.False_.total(y, as_factor=True)ignoredas_factorand returned the total of the numeric column. It now gives one row per level with that level’s estimated count, its SE, a Wald interval, the design effect when asked and the covariance across levels and by-groups, so levels can be contrasted (Rsvytotal(~factor(y)),svyby(..., svytotal)). Taylor and replication.Replication
mean(y, as_factor=True)ignoredas_factorand returned the mean of the numeric column. It now gives the per-level shares, as the Taylor path does.as_factor=Truefailed on a string or categorical column (y was cast to float), and withdrop_nulls=Truea missing value became a level of its own (0.0, estimate 0). A missing value now takes the row out of the domain, matching R’sna.rm = TRUE.prop(y, drop_nulls=True)dropped the rows with a missing y, so a PSU with no observed y left the design and the Taylor SE, covariance and intervals were wrong. A missing y (null, NaN or infinity) now takes the row out of the domain, matching R’ssvymean(~factor(y), na.rm = TRUE). Replication estimates were already right.A two-sample t-test or rank test on a numeric group ordered the groups as text, so groups 2 and 10 came out as
(10, 2)and the difference was reported with its sign flipped relative to R. Numeric groups are now ordered by value.Estimate rows came back in a different order on each run. With
by=, the kernel returned domains in hash order, soEstimate.estimates,keys(),domains,covarianceand saved payloads changed order between runs of the same code. Rows are now sorted by domain level, then by category level, in the orderto_polars()lists them: numbers numerically, strings naturally ("a2"before"a10"), and levels differing only in case by their exact text ("B"before"b"). The covariance is permuted with the rows.to_polars()and printed tables use the same tiebreak, so such levels no longer interleave. An ungrouped proportion of a numeric variable now lists10after9, not after1.Levels of an Enum column came back sorted as text. A
by=column, or the variable of a proportion oras_factormean, cast topl.Enumnow lists its levels in the Enum’s order, inestimates,keys(),domains,covariance,to_polars()and printed tables, also when value labels would sort otherwise. Other columns keep the sorted order.Estimate.level_ordersrecords each Enum’s categories, and saved payloads carry them aslevel_orders, sosvy.serialize.to_polarslists rows the same way. Schemasvy-result/0.6.A t-test or rank test group coded
Falseor0printed an empty Level cell. The cell was built withgroup_level or "".Saved results lost NaN and infinity, and
from_jsonfailed on them. JSON has neither, soto_jsonwrote each asnull, which then refused to decode into afloatfield: a singleton domain’s interval, a CV of an estimate at 0 or a NaN quantile limit made the saved result unreadable, and a NaNdeffcame back as not requested.to_jsonstill writesnull, and now records the exact value under a top-level"nonfinite"field keyed by JSON Pointer ({"/estimates/3/cv": "inf"});from_jsonrestores it. Consumers that ignore the field seenullas before. A payload written without the field reads anullplain float as NaN. Schemasvy-result/0.4.Singleton handling was not part of a saved design. A design saved and restored, or a
Samplerebuilt fromsample.design, lost it and gave different SEs.Singletons were detected only when a sample was built or its data or design replaced. A
filter_records,recodeormutatethat left a stratum with one PSU went unnoticed, and estimation ran without the singleton error a freshly built sample raises. Singletons are now detected again after every change to the rows or the stratum/PSU/SSU columns.Singleton variance strata went stale after data changes. They were built once, when the rule was applied: after
collapse(using={"d": "a"}), recoding the stratabintocleft the SE at its value before the recode. They are now rebuilt from the current data.After
update_designchanged the strata or PSUs, singleton detection read the old ones (singleton detection and the estimation check), because the internal stratum/PSU key columns were kept rather than rebuilt.Design.__eq__and__hash__ignoredwgt_adjustment, so a raked design equalled the same design without its record.Wrangling did not know the weight-adjustment record.
remove_columns,keep_columnsandselectnow refuse to drop record columns (including__svy_cells_*and__svy_aux_*snapshots) withoutforce=True.force=Truealso cleans what depended on the column: the record, the design’swgtand the replicate weights that went with it (their columns stay),PopSizecolumns, and stale internal stratum/PSU columns; one warning lists what was removed.rename_columnsandclean_namescarry renames into the record,PopSize,rep_wgts.wgtand earlier designs indesign_history, and an inplacerename_columnsorclean_namesthat cannot be applied leaves the sample untouched.Replicate weight columns were found by pattern, not by the spec. Estimation took every column matching
^prefix\d+$(any case) as a replicate, so a look-alike such asw2023,w21orW1next tow1..w20failed withREP_WEIGHT_COUNT_MISMATCHor, on one path, entered the variance.rename_columnstreated renaming such a column as a replicate rename, and a look-alikew01next to unpaddedw1..w20switched the detected padding. Replicate columns are now the spec’s ownprefix+ 1..n_reps, with the padding and casing under which all of them are present.A failed
set_designorupdate_designleft the sample half-updated, with the rejected design installed and recorded indesign_history. The sample is now left as it was.median()andquantile()always reportedq_methodasLinear. The values used the requested rule, butEstimate.q_methodand the serializedq_methodfield did not. They now give the rule used, Taylor and replication, and the printed header shows it (MEDIAN (TAYLOR, q_method=higher)).GLM predictions and margins truncated their confidence level:
alpha=0.001printed99% CI. It now prints99.9% CI.GLMFit.to_dict()raisedTypeErroron every fitted model: the coefficients and statistics are numpy scalars, which msgspec does not encode. It now returns plain Python values, as doGLMCoef.to_dict()andGLMStats.to_dict().
0.30.0 — 2026-09-25
Added
wrangling.join(other, on, cols=None, into=None, suffix=None)brings variables from another Sample or frame onto this sample’s records: the household weight onto the person file, frame variables onto the sample, response status onto the selected units. A left join that keeps every record in order and leaves the design alone;other’s weight arrives as an ordinary column.onmaps key names when they differ ({"hh": "hh_id"}).colspicks what comes in (default: every non-key column),intonames any of it here asrecode(into=)does (into={"wgt": "hh_wgt"}), andsuffixis appended to the remaining names that already exist. Nothing on this sample is renamed or overwritten: a name that still clashes is refused, design variables included, as are a key that repeats onother(validate="1:1"checks this side too), key types that do not match (integer widths, text with categorical, and a decimal identifier with an integer one when every value is whole, as SPSS and Stata store IDs, are joined), names svy reserves, and names that would read as this sample’s replicate weights. Variable and value labels carry over from a Sample. Unmatched records get nulls and aJOIN_UNMATCHEDwarning with the count (on_unmatched="ignore"|"warn"|"error", ason_singletons);indicator=adds a column marking the matched records. Inner, semi and anti joins areindicator=followed byfilter_records, which checks singletons.selection.add_stagestays the join for chaining a second stage, which rebuilds the design.
0.29.0 — 2026-09-13
Fixed
Taylor quantile and median intervals were centered at
p. The Woodruff interval inverted the weighted CDF atp ± t·se_p, as R’soldsvyquantile(interval.type="Wald")does. It is now centered at the estimated CDF at the quantile,F(q̂) ± t·se_p, as in R’ssvyquantile. On a step CDF the two differ, so limits and SEs changed whereverF(q̂) ≠ p, most on small or tied data.q_method="higher"matchesqrule="math"and"linear"matchesqrule="hf4". A limit whose probability falls outside [0, 1] is now NaN, and so is the SE, instead of being clamped to the sample minimum or maximum. Replicate-weight quantiles are unchanged.Proportion CIs under
where=used the whole-samplen.beta,korn-graubardandwilsonread the domain sample size in the t-adjustment(t(n−1)/t(df))², in Korn–Graubard’s capmin(n, n_eff*), and as the effective sample size at p = 0 or 1. Underwhere=(andwhere=withby=) that count included every out-of-domain row, so awhere=domain got a narrower interval than the same domain throughby=: a few decimals forbeta, but a Korn–Graubard bound at p = 0 hundreds of times too narrow on a small domain of a large survey.nis now the number of rows in the domain with a nonzero weight, for every design and for Taylor and replication alike. Zero-weight rows represent no population units and are not counted, which differs from R’snrow()when a file carries them; R’ssvyciprop(method="beta")also uses the whole sample for calibrated and PPS designs, where svy does not.singleton.skip()dropped the singleton strata from the point estimate. The rows were filtered out before estimation, somean,total,ratio,prop, quantiles, covariances and regressions were computed on the remaining strata alone; only the SE of a total came out as R’slonely.psu="remove". All rows now stay in the estimator and the one-PSU strata contribute nothing to the variance, matching R for every estimator.tabulatewas unaffected. The docstring, which promised that rows were removed, now describes the variance recipe.singleton.scale()did not match R’slonely.psu="average". SEs for means, ratios, proportions and covariances were computed after dropping the singleton strata’s rows, which changes the estimator, and then multiplied by1 − finstead of1/(1 − f); the two errors cancelled only when the singleton strata held exactly a sharefof the weight. The separate full-sample pass for the point estimate read the raw data frame: an unrelated Int8/Int16/UInt8/UInt16 column raisedComputeError,by=andratio()with an integer denominator raised, andwhere=silently returned the whole-sample estimate. There is now a single pass: all rows stay in the estimator, singleton strata contribute no variance, and every Taylor variance, covariance and design effect is multiplied by1/(1 − f). Totals were already right.singleton.scale()inflated every domain by the whole-design singleton fraction. R’slonely.psu="average"countsnstrat/nokstratover the strata present in the rows being estimated, andsubset()andsvyby()drop rows, so aby=level orwhere=domain that leaves whole strata out gets its own fraction. svy applied the full-design1/(1 − f)to every estimate: a domain holding none of the singleton strata was still inflated, and one holding a larger share of them was under-inflated. The factor is now counted per estimate over the strata with in-domain rows, a domain with no singleton stratum is left alone, and a domain lying entirely in singleton strata reports NaN, as in R. Full-sample estimates and domains that touch every stratum are unchanged.
Changed
logit,betaandwilsonreturn NaN bounds at an estimated proportion of 0 or 1, instead of a zero-width[p, p]that read as an interval known with certainty. APROP_CI_BOUNDARYwarning names the affected cells and points toci_method="korn-graubard", which keeps its one-sided interval there. With no residual degrees of freedom every method,korn-graubardincluded, now returns NaN (as R does);betaandkorn-graubardused to skip the t-adjustment and return an interval. A zero SE at 0 < p < 1 still gives[p, p]. The comparisons use a 1e-12 tolerance, so floating-point noise in an SE or a p-hat no longer decides which case applies.Degrees of freedom count PSUs with a nonzero weight, not a positive one, so a PSU carrying only negative calibrated weights is counted (R’s
degf).
0.28.0 — 2026-09-08
Added
Negative binomial, for counts a Poisson cannot hold.
family="negative_binomial"(or"nb"), withVar(mu) = mu + mu^2/thetaand R’sokLinks—log,identity,sqrt.m = sample.glm.fit(y="visits", x=["age", svy.Cat("region")], family="nb") m.fitted.stats.theta, m.fitted.stats.theta_seLeft to itself, theta is estimated by maximum likelihood alongside the coefficients and the standard errors come from the joint
(coefficients, theta)design-based sandwich —survey::svymleover the negative binomial likelihood, the method in Lumley’s Complex Surveys (Appendix E) and whatsjstats::svyglm.nbimplements.stats.theta_seis the design-based standard error of the dispersion itself.Pass
theta=and it is treated as known: no row for it, and the variance conditions on it, matchingsvyglm(family = MASS::negative.binomial(theta)). These are different numbers, not two routes to one — onapistratthe joint standard errors run from 25% under to 13% over the conditional ones, and on Lumley’s own NHANES example 2% under to 12% over. The orthogonality that makes them agree asymptotically is a property of the model, and a design-based sandwich is the thing that declines to assume it.On a replicate-weight design an estimated theta raises: the spread of the refits would only mean something if every replicate re-estimated it. Pass
theta=there.All of it is in the kernel, like every other family — the dispersion loop and the joint variance included. That needed digamma, trigamma and lgamma written out in
svy-rs, for the same reasonnorm_cdfalready was.cauchitandsqrtlinks. Binomial gainscauchit— the inverse Cauchy CDF, the heavy-tailed alternative to probit, so a few observations far out on the linear predictor cannot dominate the fit — and Poisson gainssqrt, the variance-stabilising link for counts.FAMILY_LINKSis now exactly R’sokLinksfor every family, with no links R has that svy does not. Both refuseexponentiate=: exp(β) is not a ratio on either.sample.glm.fit(y="voted", x=["age", svy.Cat("region")], family="binomial", link="cauchit")read_stata(encoding=)is forwarded to the reader instead of being dropped with a warning. A Stata 13 file holding UTF-8 text (the Nigeria GHS-Panel Wave 5 free-text sections) now reads withencoding="utf-8", orencoding="utf8-lossy"to replace undecodable bytes, an optionread_spssandread_sasaccept as well; the parseIoErrorfor that failure carries the engine’s hint naming the option.Panel surveys, with no new type. A panel is a long
SamplewhoseDesign.case_ididentifies the followed case and whoseDesign.waveorders its rows.svy.combine_samples(kind="panel", case_id=...)stacks the waves and validates the pairing (unique id within each wave, consecutive-wave overlap, design columns constant within a case, later waves’ units a subset of wave 1’s);Design(case_id=..., wave=...)declares the same on a long file. When no PSU is declared the case is the variance PSU, somean(y, by="wave")and its contrasts carry the between-wave covariance without being told to — the change SE equals the wide-frame individual-change SE, not the naive independent-waves one. Producer longitudinal weights stay ordinary columns selected withuse_weight(); identical producer replicate weights are accepted across waves.s = svy.combine_samples([w1, w2, w3], kind="panel", case_id="person_id") s.estimation.mean("inc", by="wave").contrast(svy.estd(3) / svy.estd(1) - 1) # percent changewrangling.lag(cols, n=1)— the one panel primitive: the value of a column at the case’s wavensteps back (a lead if negative), null across a skipped wave as Stata’sL.y(gaps="skip"takes the previous observed row). Transitions aretabulate("y_lag1", "y", where=wave == t)orprop(y, by="y_lag1", where=wave == t, drop_nulls=True); paired change is the existingttest(y, y_pair="y_lag1", where=..., drop_nulls=True).weighting.adjust()on a panel applies the nonresponse factor to the case: every row of the case gets its own weight times the case’s factor (so a case-level base weight stays case-level), a nonrespondent case gets 0 on all its rows, and a case with earlier rows but none in scope is a nonrespondent — attriters need no row to be adjusted for.respondents_onlydrops nonrespondents at the scope waves only. Chain one call per wave withwhere=col("wave") == tto build longitudinal weights.categorical.tabulate(where=)— a subpopulation table with R’ssubset()semantics: rows outside the domain keep their design columns and get weight 0, so the PSU structure and the design df are those of the full design (matchingsvymean(~interaction(...), subset(d, ...)),svytotalandsvychisqto 1e-9). Cells are formed from the domain’s categories only, and a null key outside the domain is not missing data.Delta-method contrasts.
svy.estd()expressions now accept*,/, constants,.log()and.exp(): ratios, percent change and log-ratios are estimated by the delta method on the same design df, matching Rsvycontrastto 1e-11. Linear expressions keep the exactL V Lᵀpath; the{key: coef}dict form stays linear.panel_syn_2026— a bundled synthetic three-wave panel (1200 cases, ~15% attrition per wave, producer longitudinal weights, a binary and a continuous outcome) for the panel tutorial and tests.
Fixed
svy.col(...).explode()will not change behaviour under polars 2.0. polars 1.44 deprecates the default of itsempty_as_null, and 2.0 flips it: an empty list would stop yielding a null row and yield no row at all. A row that disappears on a dependency upgrade is one fewer observation in whatever is estimated downstream, so svy passes the argument explicitly and keeps today’s behaviour.explode(empty_as_null=False)selects the other one.This raises the polars floor to 1.36.1, where
empty_as_nullwas added (1.35.1 does not have it). The alternative was a runtime version check, and svy has none anywhere else.On polars ≥ 1.44 the deprecation warning is also gone; because the test suite escalates
DeprecationWarningto an error,test_explodehad been failing there and passing locally only on the lockfile’s pinned 1.39.3.IRLS started from the wrong place for binomial and Poisson. The starting mean is R’s
family$initialize—(w·y + ½)/(w + 1)for binomial andy + 0.1for Poisson — where svy used(y + ½)/2andmax(y, 1e-10), ignoring the weight. IRLS stops on a relative change in the deviance, so on a flat surface where it starts decides where it stops: cloglog onapistratlanded 3e-5 from R’s coefficients, with the deviances agreeing to 14 digits. svy now follows R’s iterate sequence, ending on the same iteration.cauchit’s inverse link is computed the way R’spcauchyis.arctan(η)/π + ½cancels in the tails — three and a half digits gone at η = −1000 — so outside [−1, 1] the tail comes fromarctan(1/η). The forward link is−1/tan(πμ)rather thantan(π(μ − ½)), which is catastrophic as μ approaches 0 or 1.glm.fit()reported failures as results. An emptywhere=domain, or an all-zero weight column, returned a fit of all-zero coefficients (or, for gamma and poisson, aTypeErrorfrom the response check);xand2 * x“fitted” with SE 0 and an F statistic around 1e30; more parameters than observations came back through a pseudoinverse. Each of these now raisesModelErrornaming what is wrong — the collinear term by name, in the model’s own column order.A GLM that ran out of iterations was reported like a converged one.
fit()now warns, giving the iteration count and the tolerance, and so does dropping rows for an invalid weight, dropping aCatwith fewer than two levels among the fitted rows, and — as aModelErrorrather than a polars “duplicate output name” — listing the same predictor twice.where=fitted the complement domain too and threw it away. The predicate now reaches the kernel as a Boolean mask instead of a"true"/"false"by-column, so only the requested domain is fitted (1e6 rows x 20 covariates: 1.22 s → 0.45 s). The estimates are unchanged.glm.predict()andglm.margins()were wrong on an int, float or bool-codedCat. Both rebuilt a dummy column by slicing the level out of the coefficient name and comparing it to the raw column as a string, which numpy answers all-False without raising: every row got the reference-level prediction (off by up to 0.26 in probability on the test model) and every categorical average marginal effect came back exactly 0.0 with SE 0.0. The fitted coefficients were correct throughout, which is what kept it quiet. The design matrix is now rebuilt from the level values in their fitted dtype, through the same polars expression the fit used — and, as a side effect, in one pass instead of a Python loop over object arrays (predict on 1e6 rows: 0.72 s → 0.12 s; a 10-level categorical AME: 5.9 s → 0.8 s).Prediction data is validated. A missing predictor column raises
ModelErrornaming it (was a raw polarsColumnNotFoundErroror a bareKeyError), and a categorical value the fit never saw — including a null — raises instead of being silently coded as the reference level, inpredict()and inmargins(at=)alike.categorical.tabulate()ignoredDesign.pop_size: every table on a design with a finite-population correction reported the fpc-free standard errors (thettestfacade had the same gap). The fpc is now applied, matching Rsvymean(~interaction(...))andsvychisqon anfpc=design.
Changed
stats.scaleis the Pearson dispersion for every family,Σ wᵢ(yᵢ−μᵢ)²/V(μᵢ) / (n_obs − k)on the fit’s own weight scale — Rglm()’s and Stataglm’s “(1/df) Pearson”. Two deliberate departures. Binomial and Poisson used to report a hard-coded 1.0; the estimate is the overdispersion diagnostic a count model needs, and design-based SEs never use it, so nothing else moves. And Rsummary.svyglmreportssvyvar(resid(pearson))instead, a design-weighted variance of the Pearson residuals — not the textbook estimator, and not what someone comparing againstglm()or Stata expects — so that is not copied.The dispersion’s divisor is
n_obs − k, not the design residual df: the dispersion is a moment estimate, while the design df belongs to the t reference distribution. The same model on the same rows used to report a different scale for every design.combine_samples:kind="cross_sectional" | "cs" | "panel"replacesunits="independent" | "shared"; default wave labels are"wave 1".."wave k"instead of"s1".."sk"; the combined weight columnwgt_nameis always created from each wave’s own weight (divided by k underadjust="average"), so the waves may name their weights differently. A panel keeps the base-wave design and now accepts a later wave that lost PSUs (with a warning) instead of requiring identical design units.Design.describe()showsCase idandWaverows, andPSU None (variance: <case_id>)when the case is the fallback PSU; on a panel the design summary lists the case overlap between consecutive waves.
Removed
Design.row_index.Design.case_idis the record identifier on cross-sections and panels alike (non-null, unique; unique within wave on a panel).Design(row_index=...)now raisesTypeError. The hiddensvy_row_indexcolumn is unchanged plumbing.combine_samples(units=...), replaced bykind=without a shim (pre-stable).
0.27.0 — 2026-08-31
Added
svy.combine_samples()— repeated cross-sections and panel waves. Stacks several samples and analyzes them as one stratified design: data pooling, not estimate pooling. Each independent wave contributes its own strata (wave → stratum → PSU), so Taylor variance treats waves as independent without being told to. Caller order is time order, and an existing wave column (NHANESSDDSRVYR, say) is reused and validated rather than replaced.svy.combine_samples([w1, w2, w3], adjust="average") # period-average populationadjust="average"divides the weights by k, which matters only for totals — means, proportions and ratios are invariant to it.units="shared"is the panel case, where the same units recur across waves.Linear contrasts over estimates, via
svy.estd(). AnyEstimateorGLMFitcan be contrasted against its own estimands using the between-estimate covariance:m = sample.estimation.mean("api00", by="stype") m.contrast(svy.estd("E") - svy.estd("H")) fit.contrast({"H vs E": svy.estd("stype_H") - svy.estd("stype_E")})Accepts a contrast expression, a sparse
{key: coefficient}dict, or several named contrasts at once. Keys are the row identitieskeys()lists, and metadata value labels are accepted wherever unambiguous. Inference is t-based on the full design df (R’sdegfconvention), not the per-row domain-aware df.Quantiles and medians raise rather than returning a number: Woodruff intervals do not arise from a linearized score column, so no between-quantile covariance exists to combine.
Marginal effects for categorical predictors.
margins()now returns a discrete contrast per non-reference level —mean_w[mu(x, var=k) - mu(x, var=ref)]with a delta-method SE — instead of skipping the term. Matches R marginaleffects’ factor contrasts; on apistrat thestypecontrasts agree to 1e-9 and the SEs to 2e-10.This fixes a wrong-shaped answer, not just a missing feature:
margins(variables=["stype"])previously returned an empty list, and the defaultmargins()dropped categorical terms with only alog.warning. A variable with no derivative w.r.t. any fitted coefficient now raises instead of being silently skipped.exponentiate=onGLMFit.to_polars()and.show(). Reports exp(β) with exp of the link-scale interval — an odds ratio forlogit, a rate ratio forlog, a hazard ratio forcloglog— and names the column for the link. It refuses the links where exp(β) is not a ratio (identity,probit,inverse,inverse_squared) rather than printing a meaningless number.The interval is the exponentiated link-scale interval, so it is not symmetric about the ratio, and
std_err/statistic/p-value stay on the link scale where the Wald test is computed — a symmetric standard error around a ratio is the mistake this is meant to prevent. Matchesexp(coef(f))andexp(confint(f))from R.offset=onglm.fit(). A known term on the link scale, coefficient fixed at 1 — what makes a rate model possible:s.glm.fit( y="events", x=["age", svy.Cat("region")], family="poisson", offset="log_exposure" ) # coefficients are log rate ratiosCarried through the whole model: IRLS working response, deviance, the sandwich,
predict()andmargins(). Matches Rsurvey4.5 to 2.5e-15 on coefficients and 5e-10 on SEs. Null deviance is the intercept-only fit carrying the offset, found by a one-parameter IRLS the way R’sglm()refits it — the weighted-mean shortcut is only valid without an offset — and reproduces R’snull.devianceexactly.predict()includes the offset, where R’spredict.svyglm(newdata=)drops it. R’s own two answers disagree: on the same rowsfitted(f)returnsexp(Xb + offset)and tracks the observed counts, whilepredict(f, newdata=)returnsexp(Xb). svy followsfitted(), which is also what Stata’spredictafterglm, exposure()does. Passingnewdatawithout the offset column raises rather than silently predicting a rate.Inverse Gaussian family for
glm.fit().family="inverse_gaussian"(canonical linkinverse_squared, R’s1/mu^2) completes the exponential-family set. The kernel already implemented it andDistFamily.INVERSE_GAUSSIANalready named it — only the Python family map withheld it. Matches Rsurvey4.5 to twelve significant figures on coefficients, SEs and deviance across weighted, clustered, stratified and stratified-clustered designs.DistFamilyis exported.svy.DistFamilyandsvy.core.DistFamilynow resolve, matchingLinkFunction, which was already exported.family=andlink=both accept a plain string or their enum.Probit and cloglog links for
glm.fit().family="binomial"now acceptslink="probit"andlink="cloglog"alongsidelogit, completing the binomial link set. Both are wired through the whole model: coefficients and sandwich SEs,predict()on the response scale, andmargins()— average marginal effects and predictive margins, whose delta-method SEs need the link’s second derivative.s.glm.fit(y="insured", x=["age", svy.Cat("region")], family="binomial", link="probit")Verified against R
survey4.5 across weighted, clustered, stratified and stratified-clustered designs — coefficients and SEs to 1e-7 relative, deviance to 1e-14 — and againstmarginaleffects0.32.0 for AME and predictive margins.probitandcloglogmap onto (0, 1), so pairing either with a non-binomial family now raises instead of fitting something meaningless.Sample.weighting.standardize()— direct standardization. Reweights each domain to a common composition so rates are comparable across domains that differ in it. Age is the usual axis, but any composition variable works.s.weighting.standardize( cells="agecat", # composition axis shares={"(0,19]": 55901, "(19,39]": 77670, ...}, # counts or proportions by=["race", "riagendr"], # domains where=svy.col("hi_chol").is_not_null(), # scope )target(g, c) = share(c) × Ŵ_g, so domain totals are preserved and only the within-domain composition is reshaped.Ŵ_gis computed under the same scope that built the cells, which removes the usual footgun — a domain total estimated over a different set of rows than the one being standardized. Domains missing a level are renormalized over the levels they have, with a warning.It is a third name, not a third machinery: it runs on the same engine as
poststratify.standardize(by=None)andpoststratify(shares=…)are deliberately the same operation, each named for the community that uses it.Verified against R
survey::svystandardize4.5 — weights to a relative tolerance of 1e-12, and the NHANESHI_CHOLby race × sex example to 4.8e-14 across all eight domains.Standardized weights are analysis-specific.
wherebakes one variable’s missingness into the weights andbybakes in the domain structure, so estimating a different variable or a different breakdown on the same standardized sample is silently wrong.Calibration-aware Taylor variance. A weighting method that pins population quantities now records what it pinned, and the Taylor path centres the linearized scores against that structure — R
survey’ssvyrecvarpostStrataloop. Coversmean,total,ratio, proportions, quantiles, tabulation,glm, and one- and two-sample t-tests, grouped and ungrouped, under poststratification, raking, GREG calibration and standardization.Verified against survey 4.5 on apiclus1 to twelve significant figures or better, including the cases where the two implementations are easy to get subtly wrong: raking centres on the unweighted margin mean, GREG fits its residual with the previous weights, and cell means are always whole-sample — a subpopulation filter does not restrict them, because the adjustment was not performed inside the subpopulation.
where=on all seven weighting methods. Scope, not domain: matching rows receive the adjustment, and the rest keep their previous weight, so the new column is complete anddesign.wgtcan repoint to it. This is the same keyword and the same reading as estimation’swhere=, with a different consequence — estimation zero-weights for subpopulation variance.
Changed
Requires
svy-rs >= 0.16.0. The GLM offset and the probit/cloglog links call into kernel entry points 0.15.0 does not have, so the pin moves to>=0.16.0,<0.17.0.BREAKING: a family now admits only the links it has a model for.
glm.fit()validates thefamily/linkpairing against theokLinksset of R’s family constructors, so pairings that were never a model are rejected up front instead of failing deep in the kernel — or, in the case ofbinomial+inverse_squared, converging on a meaningless fit and reporting it as a result. Newly rejected:logit/probit/cloglogoutsidebinomial,inverse/inverse_squaredoutside the families whose canonical link they are. Every previously working model still works; what stops working is the combinations that only appeared to.BREAKING:
by=is nowcells=onadjust,normalizeandpoststratify.cellsnames the groups that each receive one derived adjustment factor;bycontinues to mean “repeat independently per group” everywhere it survives (calibrate,trim, andstandardize). Keeping one word for both would have compromised a term of art.BREAKING:
factors=is replaced byshares=. Inpoststratifythis is close to a rename —factorsalready computedf × grand_total, which is shares, unnormalized. What changes is thatsharesare normalized first, so a vector that does not sum to 1 now pins composition instead of silently rescaling the population. Inrakeit is a genuine removal:factorsthere multiplied each category by its own current weight sum, which has no share reading and which no workflow can supply in advance, since raking is iterative.rakegainsshares=(marginal proportions) in its place.BREAKING: a scalar
controlsnow requirescells=None. The rule across every level/share method is that a scalar names one cell and a dict names many;controlssets the total andsharespreserves it. Previouslypoststratify(300, cells="region")silently ignoredcellsand rescaled to the grand total — verified bit-identical tonormalize(300)with no groups.poststratify(331_000_000)with no cells is unaffected and remains the way to pin a known population size.Renamed and removed parameters raise an error naming the replacement rather than reporting an unknown keyword.
Multi-column
cellsbind by tuple key, and key-mismatch errors now echo the keys as written —('R2', 'D1')rather than the internal'R2_&_D1'encoding.rakenow rejects margins that disagree on a population total. Margins were validated only in isolation, so inconsistent totals — the classic reason IPF fails to converge — simply ran tomax_iterand returned whatever they reached. Withshares=the consistency is structural.
Fixed
The missing-values error named a parameter that does not exist.
assert_no_missingtold users to setdrop_missing=True, which is the internal helper’s argument, not a public one — following the advice raisedTypeError. All four call sites gate ondrop_nulls, which is what the message now says.print()on anEstimateListof proportions.prop()over a sequence of variables built one level column per variable, so a diagonal concat unioned them into a staircase of mostly-empty columns — eight conditions gave eight sparse columns, a ~118-character box, and every number truncated to an ellipsis. The members now stack under one sharedlevelcolumn besidey, which for the eight-condition case brings the table to 71 characters at full precision.Only the stacked case changes. A single
prop()still heads its level column with the variable name,mean()over a list is untouched,quantile()keeps itsprobcolumn, and aby=column stays distinct from the level column — including when a variable is itself namedlevel, where the shared column moves aside.categorical.ttestignored the finite population correction. The t-test built its variance without ever asking for an FPC column, so on a design withpop_sizeit returned the answer for sampling with replacement — standard errors too wide,ttoo small. One- and two-sample tests were both affected; on apiclus1 the statistic was off by a factor of1/sqrt(1 − 15/757). It now matches Rsvyttestto fourteen significant figures.One target-resolution path.
normalizeandpoststratifyderived their control arrays independently, with two different label-to-code conventions. Both now resolve every argument form to absolute per-cell targets in one place, which is also the single form crossing into Rust.
Note on variance after a calibrating adjustment.
poststratify,rake,calibrateandstandardizepin quantities in the population, which removes sampling variability. Both variance paths now account for this: replication re-adjusts each replicate column, and Taylor centres the linearized scores within the pinned cells. Previously only replication did, so Taylor standard errors on a calibrated design answered a slightly different question — on the NHANES standardization example they differed from R by roughly −11% to +14%. Point estimates were never affected.The record is honoured only while it still describes the data. Estimating on a different weight, or dropping a column the adjustment referenced, falls back to treating weights as fixed — and says so, once per data/design rebind.
Removed
- BREAKING:
ModelTypeandFitMethodare no longer exported. Both arrived with the monorepo migration and were never wired up — no consumer insvy, its tests or its docs, and neither enum’s values were matched anywhere.ModelTypewas also misleading: it named four family+link combinations whereglm.fit()takes the two separately and now accepts eleven. UseDistFamilyandLinkFunction. The definitions are commented out incore/enumerations.pyrather than deleted, alongside the other parked enums there.
0.26.0 — 2026-08-28
A breaking change to the weighting API, and a substantial cut to import time.
Changed
BREAKING:
Sample.weightingno longer mutates the sample you call it on. All eleven transforming methods — the fourcreate_*_wgts,adjust,normalize,poststratify,rake,calibrate,calibrate_matrix,trim— now return a newSampleand takeinplace: bool = Falseto ask for the old behaviour. Both branches return aSample, so chaining is unchanged ands = s.weighting.…()keeps working exactly as before.This is not a new convention: all 19
wranglingmethods already takeinplace: bool = Falseand default to copying, andsingleton’s five transforms always return a newSample.weightingwas the only namespace that mutated unconditionally, and the only one with no way to ask for a copy —grep -rn inplace src/svy/hit nothing outsidewrangling/.The reassignment idiom hid it, so it surfaced only when two variants branched off one sample — which is what a method comparison, a sensitivity check, or a docs page does:
boot = base.weighting.create_bs_wgts(n_reps=10, rep_prefix="bs") jack = base.weighting.create_jk_wgts(rep_prefix="jk") # before: boot is jack is base. 33 jk* columns in the frame, and a design # still reading method=Bootstrap prefix='bs' -- so an estimate off # `jack` silently used the bootstrap columns and coefficients.What breaks: code that called a weighting method and then read the receiver without capturing the return. Add
inplace=True, or capture the result. Code already written ass = s.weighting.…()needs no change.The isolation covers diagnostics as well as data: an operation that raises under the default leaves the caller’s warning store untouched, because “this call had no effect on my sample” is only true if it covers everything. The error is still raised and still emitted to the log;
inplace=Trueis how you ask for the warning to land on your sample.import svywas paying for scipy before you asked it anything. Importing the package cost ~0.80 s, and 383 ms of that wasscipy.stats, pulled in along the chainsvy.serialize→svy.categorical→svy.estimation.base. Nothing about that import was needed to reach svy — it was needed only by whichever call eventually wanted a quantile.Six modules imported scipy at module scope:
estimation/base.py,regression/base.py,regression/margins.py,engine/size_and_power/size.py,engine/size_and_power/power.py, andutils/hadamard.py. What made this stubborn is that fixing any one of them alone changes nothing — scipy loads once and the other five ridesys.modules, so the cost simply moves. All six had to go together.The uses turned out to be narrow: estimation, regression and margins want the t and F quantiles, size and power the normal, hadamard a single
scipy.linalgcall. Each import now sits in the function that needs it, matching whatcategorical/base.pyhad already been doing in four places — this finishes a refactor that had stalled partway.import svynow costs ~0.27 s andscipy.statsis off the import path entirely; it loads on the first call that actually needs a distribution. Nothing about the numerical results changes.svy.datasetsnow loads on first use rather than on import. The root package imported it purely so it would be reachable as an attribute, which is what a module-level__getattr__handles. Worth ~30 ms — not the ~180 ms an import profiler credits to it, since most of that is polars and numpy, whichdatasetsmerely imported first and which the rest of svy needs anyway.svy.datasets.load(...),from svy import datasetsandimport svy.datasetsall behave exactly as before.
Fixed
BRR validated its strata in two passes, and sent one of them to the wrong subsystem.
create_brr_wgtschecked< 2PSUs first, raisingINSUFFICIENT_PSU, and only then checked for odd counts. A frame carrying both therefore reported the lone-PSU stratum, and revealed the odd ones on the next run — one class of problem per attempt, on a validation that could have been answered in full the first time.The split also borrowed a question that is not BRR’s. A stratum with one PSU is a singleton, which matters to Taylor linearization and has its own subsystem —
Sample.singleton, withcollapse,combine,pool,certainty— none of which the hint mentioned; it said “Combine small strata or check design specification.” For BRR the count is not a singleton question at all: a stratum of 1 and a stratum of 3 fail for exactly the same reason, a PSU with no partner.BRR now asks one question — is the PSU count a multiple of 2 — and reports every stratum that fails it in a single error, naming each with its count:
❌ BRR needs 2 PSUs per variance stratum [ODD_PSU_COUNT] BRR pairs PSUs into variance strata of exactly 2. Found 3 strata whose PSU count is not a multiple of 2, leaving a PSU with no partner. - got: a=1, b=3, c=5“Multiple of 2” rather than “equal to 2” because
create_brr_wgtspairs: a 4-PSU stratum becomes two variance strata of 2 and is valid.INSUFFICIENT_PSUis now the paired jackknife’s alone, which is the one method that genuinely distinguishes the two — it absorbs an odd count into a triplet but still cannot pair a lone PSU.The odd-PSU BRR error pointed at an API that no longer exists. Its hint read
Use method='jk2' which allows 2-3 PSUs per stratum.method='jk2'wascreate_variance_strata’s parameter — a function removed in 0.25.0 — andmethodis not a parameter ofcreate_brr_wgtseither, so the one line telling you what to do instead named a parameter of a gone function on a function that never had it. “2-3 PSUs per stratum” was equally stale:create_jk_wgts(paired=True)pairs any count ≥ 2 and absorbs an odd one into a triplet.The hint now names
create_jk_wgts(paired=True), says why BRR cannot do the same, and tells anyone who needs BRR specifically to fix every stratum listed rather than one. Its test assertedpytest.raises(Exception)and nothing more, which is how the text drifted unnoticed; it now pins the code, thewhere, and that the remedy names a real API.Regenerating replicate weights left the previous method’s design in place.
create_brr_wgtsandcreate_jk_wgtsrecorded their result withDesign.fill_missing(rep_wgts=…), which by definition only fills a field that is currentlyNone— so once any replicate design existed, the record of what had just been built was silently dropped.create_bs_wgtsandcreate_sdr_wgtsusedDesign.updateand were unaffected, which is the whole of the difference: it was an inconsistency inherited from the monorepo migration, never a decision.Building a second set of replicates therefore wrote the new columns and kept the first method’s metadata — and estimation reads the metadata, not the columns:
s = s.weighting.create_bs_wgts(n_reps=10, rep_prefix="bs") s = s.weighting.create_jk_wgts(rep_prefix="jk") # before: 33 jk* columns in the frame, design.rep_wgts reading # method=Bootstrap prefix='bs' n_reps=10. # mean(method="replication") -> Estimate: MEAN (BOOTSTRAP), off `bs*`.Unlike the mutation defect fixed alongside it, this one fired on the plain
s = s.weighting.…()idiom, in both orders, and switching methods mid-analysis is exactly when someone does it — comparing a jackknife against a bootstrap, or moving to JK2 after seeing the replicate count. Both now useupdate: the design describes the columns that were just written, and a declaration that no longer matches them is replaced rather than preserved.
0.25.0 — 2026-08-26
Added
Replicate weights carry the units they were built from.
stratumandpsuon every variant name the columns the replicates were drawn over — a separate question fromDesign.stratum/Design.psu, which describe the analysis design and are what Taylor linearizes over. Producers collapse strata and suppress PSUs for disclosure, publishing a distinct pair (VARSTRAT/VARUNITand its many spellings) alongside — or instead of — the design variables:JackknifeWgts(prefix="rw", n_reps=8, kind="jkn", stratum="VARSTRAT", psu="VARUNIT")They appear in
reprandprint(design), are validated as columns atSampleconstruction, and are rewritten bywrangling.rename_columnsand protected byremove_columnsexactly like design columns. Generators record whatever they used, so a generated design states its provenance rather than leaving a later reader to assume it matches the Design. The Poisson bootstrap records neither, drawing independent per-unit factors and having no units by construction.They take the same shapes
Design’s do —strfor the single collapsed identifier producers usually ship, or a tuple when the unit is several columns together — and the same spellings mean the same thing on both objects, including("region",)staying a one-tuple rather than being unwrapped. Multi-column units are grouped on directly rather than through the internal concatenated columnDesignbuilds, so nothing implementation-specific reaches anything reading provenance back.create_brr_wgtsandcreate_jk_wgts(paired=True)pair PSUs into variance strata themselves.stratum_namenames the created column (the housewgt_nameconvention),order_bypairs adjacent PSUs in that order — what a systematically-sampled frame wants — andshufflepairs at random.stratum/psuon all four generators build from columns the Design does not name, without mutating it.Poisson bootstrap replicate weights (#131).
sample.weighting.create_bs_wgts(kind="poisson")generates Beaumont–Patak generalized bootstrap weights, which need only a weight column. The defaultkind="rao-wu"is the stratified Rao–Wu–Yue rescaling bootstrap and still requirespsuon the design — the guard is deliberately not shared, since the Poisson bootstrap exists precisely for files that have no PSU. Both kinds use the same1/Rper-replicate coefficient; they differ in how the replicates are drawn, not in how the variance is scaled.Beaumont, J.-F. and Patak, Z. (2012). On the generalized bootstrap for sample surveys with special attention to Poisson sampling. International Statistical Review, 80(1), 127–148.
scale: the coefficients a producer publishes (#7). Replicate weights often ship with a documented per-replicate variance coefficient that is not the method’s default. CETIC publishes the ICT Households microdata with0.0065365334145709at R = 200, which bakes in a finite-population correction. Declaring it is now a parameter rather than an arithmetic exercise:svy.RepWeights(method="bootstrap", prefix="REP", n_reps=200, scale=0.0065365334145709)A scalar is broadcast to every replicate; a sequence must match
n_repsand is checked at construction. It replaces, rather than multiplies, the method default:V = Σ_r scale_r · (θ_r − θ̄)². svy’sscaleis Rsvrepdesign’sscale × rscalesfolded into one field, sosvrepdesign(type = "bootstrap", scale = c, rscales = rep(1, R))isscale=chere, with no conversion factor.examples/ict_households.pydrops thesqrt(scale * R)workaround it used to demonstrate.Jackknife designs say which family they are.
JackknifeWgts.kindtakes R’ssvrepdesign(type=)names —"jk1"(unstratified delete-one-PSU,(R−1)/R),"jkn"(stratified,(n_h−1)/n_hper replicate),"jk2"(paired, one replicate per stratum,1.0)."paired"is accepted as an alias for"jk2": it describes the design, not the replication scheme, since a JKn on a paired design is still JKn.The default is
None— unspecified.Noneand"jk1"produce the same number but are different statements, and svy claims only what it knows or was told. A producer withholding design variables is not evidence that the weights are unstratified.Sample.rep_coefsshows which coefficient was applied to which replicate, keyed by the actual column name:>>> sample.rep_coefs {'jk1': 0.6667, 'jk2': 0.6667, 'jk3': 0.6667, 'jk4': 0.5, ...}coefficients()on the variant remains the machine path — a list ofn_repsfloats in replicate order, the shape the kernel takes — but a bare list makes the assignment something a reader has to trust rather than check, and for an unbalanced JKn the assignment is the answer: the same coefficients against a different order give a different standard error.It lives on
Samplerather than onrep_wgtsbecause only the frame resolves the column names: a design declaringprefix="REP"against a file shippingREP001..REP200knows the replicate count but not the padding, so the struct alone would key this onREP1and be confidently wrong. Empty when the design carries no replicate weights, and it raises whatevercoefficients()raises rather than reporting a number it does not have.rep_wgts.coef_sourcenames the provenance alongside it —"scale"(the user asserted them),"derived"(svy computed them and cannot redo it) or"default"(the method’s standard value). A varying vector also prints its distinct values with counts now,0.667 x3, 0.5 x4, rather than first-and-last — which two numbers land at the ends is an artefact of the producer’s replicate order, while the counts are the design.kindis optional once the units are named.jk1/jkn/jk2is expert vocabulary; naming the columns the replicates were built from is not, and requiring both asks for the same fact twice in the harder of the two languages. Withstratumandpsudeclared, svy reads the scheme off the counts — one replicate per PSU across several strata is JKn, per PSU in a single stratum is JK1, one per stratum is JK2 — and records it, at INFO level so the inference is auditable:JackknifeWgts(prefix="jk", n_reps=8, stratum=("region", "urban"), psu="psu") # -> kind='jkn', rep_coefs=0.5 (derived)This used to refuse, on the grounds that deriving a kind would turn “the producer withheld the design” into “these weights are unstratified”. That was right while the counts came from
Design.stratum, which says nothing about how the replicates were drawn; the units declared on the weights are exactly that statement.A
kindthat the units contradict is now an error rather than a warning, but only whenn_repslands exactly on another scheme’s count — then the label is simply wrong and svy can say which it should be, and warning would ship a coefficient wrong by a known factor. Declaringjk2against 8 replicates and 8 PSUs raises and namesjk1/jkn. A count matching no scheme still only warns: a frame subset to fewer PSUs than the weights were built from is not a mislabelling.JKn coefficients are derived when the declared units allow it.
(n_h−1)/n_his the only standard coefficient that is not closed-form inn_reps, so it used to have to be supplied by hand. Givenkind="jkn"plus thestratumandpsuthe replicate weights name (not the Design’s — see above), svy now works it out atSampleconstruction. Akind="jkn"with no units named, or with nothing derivable from them, warns at construction and raises fromcoefficients(). Balanced designs — every stratum the samen_h— take a fast path where the coefficient is uniform and the replicate-to-stratum mapping never has to be found.Unbalanced strata are not a different method, just JKn with a coefficient that varies by stratum, so the only thing missing is that mapping — and a delete-one-PSU replicate states it by zeroing the PSU it drops. svy now reads it back: one aggregation to PSU level, and the recovered vector is identical to what
create_jk_wgtswould have assigned, to the last bit. Only replicate → stratum is needed, since the coefficient is constant within a stratum.It applies only where the delete-one signature holds — every replicate column zeroing exactly one otherwise-nonzero PSU, no PSU dropped twice. A bootstrap, a BRR or a Fay-style variant fails that and falls back to refusing, rather than being handed a confidently wrong vector.
scaleremains the answer for files that do not zero.A declared kind is also checked rather than merely trusted:
jk1andjknimply one replicate per PSU,jk2one per stratum. Mismatches warn rather than raise, since a legitimately subset frame has fewer PSUs than the weight columns were built from. An unspecified kind on a stratified design warns too — svy has evidence the JK1 global is probably wrong, but no claim to act on.
Fixed
kind="jk1"against stratified units silently returned the unstratified coefficient.jk1andjknimply the same replicate count — one per PSU — and differ only in stratification, which the replicate-count check never looks at. So the one mismatch it structurally could not catch was the expensive one:jk1on units declaring several strata used(R−1)/Rwhere(n_h−1)/n_his correct, overstating standard errors bysqrt(R/n_h)— 32% on a 4-strata × 2-PSU design — with no warning. Naming astratumunit says the replicates were drawn within those strata andjk1says they were not, so it now raises and namesjkn. Replicates genuinely drawn without regard to strata are declared by naming onlypsu.A declared JKn with a psu but no stratum silently produced the JK1 coefficient.
(n_h−1)/n_his a per-stratum quantity; with no stratum named, the PSU count fell back to a single group holding every PSU and the derivation returned(R−1)/R— the JK1 global — under a JKn label. On 4 strata × 2 PSUs that is 0.875 where 0.5 is correct, an 8% error in the standard error, with no warning. It now refuses, the same way an unmetkind="jkn"claim already did.Building BRR or JK2 weights no longer destroys the design strata.
create_variance_strataended by overwritingDesign.stratumwith the collapsed variance strata — its only way of handing the result to the generator that ran next. Every later Taylor estimate then silently linearized over pseudo-strata: on a 6-strata × 4-PSU frame the SE moved from 0.858 to 0.641 and df from 18 to 12, with no warning and the original column still sitting in the frame. Pairing is now internal and its output is recorded on the replicate weights, soDesign.stratumis Taylor’s alone and is left exactly as declared.create_jk_wgts(paired=True)no longer undercounts replicates. Called on strata with more than two PSUs — i.e. without the old pairing pre-step — it produced one replicate per original stratum instead of one per variance stratum, in silence: 6 where 12 were correct, understating the variance. It pairs first now. This changes published standard errors for anyone who called it without the pre-step.create_brr_wgtsno longer requires a stratum. It raisedBRR requires 2 PSUs per stratumon any unpaired design, andstratum=Nonewas rejected outright with a hint to go and build variance strata by hand. Both cases now pair and build.A row-level
order_bycolumn paired PSUs more than once. Uniquing on the whole row left the same PSU present several times, so it was assigned to several variance strata and some ended up holding a single PSU — which the BRR kernel then rejected. First occurrence in the frame now wins per PSU.Renaming or dropping a unit column follows through to the replicate weights.
_rep_wgts_with_renamesonly remappedprefixand returned early when no replicate column matched the rename — the common case, since renaming a variance-stratum column touches no replicate column. The units were also absent from the protected set, so dropping one was neither blocked nor cleaned. They are now rewritten on rename, needforce=to drop, and are cleared rather than left dangling when force-dropped.method=Noneno longer auto-selects replication, and Taylor on a replication-only design says so. The default resolved to replication whenever the design carried replicate weights and had nostratum/psu, and to Taylor otherwise. That made the estimator depend on inputs the estimator never reads — replication consumes the replicate columns andcoefficients(), nothing else — so declaring a single design column silently moved the standard error:mean("y"), identical replicate weights throughout wgt only -> Jackknife se=1.252718 wgt + stratum -> Taylor se=1.171192 wgt + psu -> Taylor se=1.002973 wgt + stratum + psu -> Taylor se=1.063071It was worst for JKn, where
stratumandpsuare exactly what svy needs to derive(n_h−1)/n_h: the one action that makes JKn usable was the action that switched JKn off.Noneis now the unstated Taylor default the signatureLiteral["taylor", "replication"] | Nonealways implied, not a third mode. Replication is opt-in: passmethod="replication".Falling through to Taylor alone would have reinstated the worse half of the original bug, which is why the auto-detection existed. Without
stratumorpsuevery row is its own PSU in one stratum, so linearization is SRS-like — on a 200-row file with 8 replicates,df=199against a true 7, silently. That case now emitsTAYLOR_WITHOUT_DESIGNand proceeds, for the explicitmethod="taylor"spelling too: the hazard is in the number, not in who asked for it.A JKn design that cannot be derived warns at construction.
kind="jkn"with nopsuto count from, or with unbalanced strata, built aSamplein silence and failed only when an estimate was requested — at a call site far from the cause. Both dead ends now emitJACKKNIFE_COEFS_UNAVAILABLEnaming which one it hit. TheMethodErrorfromcoefficients()stays lazy on purpose: such aSampleis still perfectly usable for Taylor, for wrangling and forwhere=, so raising at construction would block work that never touches replication. What was missing was the early signal, not the error.The JKn error names both remedies, not one. It named only
scale=, sending anyone whose file does carrypsuoff to hand-compute coefficients svy would have derived for them. It now offers declaringstratum/psufirst, andscale=for the unbalanced case where that cannot help.create_bs_wgts(kind="poisson")fails cleanly against an oldersvy-rs.rust_create_poisson_bs_wgtswas missing from theImportErrorfallback, so the name was never bound and the guard raisedNameErrorinstead of the intended message.MethodError.not_applicableno longer doubles the sentence-ending period. The template appended.to areasonthat ~24 call sites already ended with one, producing…(got psu=None)...RepWeights.dfis annotatedfloat, matching what it has always stored. It was declaredint | Nonewhile the kernels hand back an f64, so a repr readdf=499.0against anintannotation. Widened rather than coerced:dffeeds a t-quantile, which is defined for fractional df, and Satterthwaite-style effective df is fractional by construction. Anintis still accepted and stored unchanged.A user-supplied replicate coefficient was silently discarded for bootstrap, BRR and SDR (#7). Before #131 the override was applied at one site in the kernel —
rscales.unwrap_or_else(|| replicate_coefficients(method, n_reps, fay_coef))— and was therefore method-agnostic by construction. #131 moved coefficient computation into Python and scattered it across four per-variantcoefficients()methods, which gave every method its own chance to forget. Three of four forgot: the field was accepted, stored, length-checked, and then dropped, so a bootstrap declared with a producer’s published scale silently returned the1/Ranswer.coefficients()is now a template method on the shared base. The override is resolved there, at the single point of use, and the variants implement only their own default — they never see it, so a new variant cannot regress this again. No published standard error changes: #131 landed aftersvy-v0.24.1and was never released.Estimation and regression could scale the same design differently.
GLM._rep_coefsre-derived the coefficients by substring-matching a stringified method tag (if "boot" in m), and honoured the user’s override for every method while the estimation path did not. OneDesigntherefore produced differently scaled standard errors fromsample.regression.glm(...)andsample.estimation.mean(...). It is gone; both callrep_wgts.coefficients().A declared paired jackknife used the wrong coefficient. JK2 has one delete-one replicate per stratum and a coefficient of
1.0; the global(R−1)/Runderstates the variance by exactly that factor. Declaringkind="jk2"now gets1.0. Relatedly,kind="jkn"with nothing to compute the per-stratum coefficients from — noscale, no design variables — refuses rather than substituting the JK1 global, which on a 4-strata × 2-PSU design overstates standard errors bysqrt(7/4), 32%. An unmet claim fails; an absent one (kind=None) still falls back to(R−1)/Rexactly as before.Labelled variables were never
is_categorical(#130).Sample(data=df)infersmtypefrom the polars dtype, andimport_labels_from_svyio_metabrought across labels while leaving that untouched — so every labelled variable in a survey recode stayedNumerical Continuousandis_categoricalwasFalse. Deriving a codebook from it classified an entire DHS-style file as continuous. The importer now revisesmtypeon import: it honours a measurement level the file declares, and otherwise asks whether the value labels cover every observed value. On a Stata census file this takes 16 labelled variables from 0 categorical to 16.Measurement type is inferred from label coverage, not label presence. “Has labels” alone would misread a continuous variable that labels only a sentinel code — a “how many TVs?” count labelling just
0 = noneand4 = 4 or moreis not a five-level factor. A variable becomesNOMINALonly when every observed non-null value carries a label; partial coverage, an all-null column, and a column absent from the frame all leavemtypealone, as does a type the user set by hand.ORDINALis assigned only when the file says so, since coverage cannot distinguish it fromNOMINAL.The polars deprecation filter never worked.
filterwarnings = ["error::DeprecationWarning:polars"]looked like a guard against calling deprecated polars APIs and was inert: the module field is compared against the frame the warning is attributed to, and polars setsstacklevelto point at the calling code, so:polarsnever matches. Verified with a probe — a deprecated call passed cleanly under it. The filter is now unqualified, and the suite passes under it.
Changed
Replicate-weight designs are a tagged union of four types, and
svy.RepWeightsis a function (#131). They were a single struct carrying an estimation-method enum plus every method’s parameters. They are now one type per method —BootstrapWgts,JackknifeWgts,BrrWgts,SdrWgts— each carrying exactly the parameters its algorithm has. A bootstrap cannot hold a Fay coefficient, because there is no field for one.svy.RepWeights(method=..., ...)keeps working and returns the variant for the name, so call sites are unaffected. But it is a factory function now, not a class, which breaks two things it used to support:isinstance(x, svy.RepWeights) # TypeError: isinstance() arg 2 must be a type x: svy.RepWeights # no longer a valid annotationUse
svy.RepWgts, the union of the four variants, for both. Code that knows the method at authoring time can construct the variant directly, which is the typed path:BootstrapWgts(prefix="bsw", n_reps=1000, kind="poisson").rscalesis nowscale(user-supplied) orrep_coefs(svy-derived). The single field held both, which made “the user asserted this” indistinguishable from “svy computed this”. They split on who supplied the values — the one axis that is verifiable, since a user declaring stratified JKn supplies the standard coefficient rather than a custom one.scaleis what you pass;rep_coefsis filled bycreate_jk_wgtsand by the JKn derivation, and is shown as(derived)in the design output.This renames a parameter on two released surfaces:
svy.RepWeights(rscales=...)andDesign.update_rep_weights(rscales=...). Both now takescale=(orrep_coefs=), andscaleadditionally accepts a scalar.A parameter that belongs to another method is refused, even at its neutral value.
RepWeights(method="bootstrap", fay_coef=0.0)was accepted as a stored no-op when every method shared one flat signature. It now raises, naming the variant that owns the parameter. Callers that know the method should not be sending it; nothing insvydoes any more.The
methodlabel is display-only, and now genuinely is. Its docstring said nothing read it to choose a code path while eight sites did — seven inweighting/rebuilt a design by round-tripping the type through its own tag and hand-listing every field that had to survive, which is howkind,scaleandpaddingcould vanish across a poststratification. Copies now go throughmsgspec.structs.replace, which preserves the type and drops nothing.Non-default variance coefficients are visible.
scaleandrep_coefsdid not appear inreprorprint(design)at all. A non-standard coefficient is what a reviewer or replicator most needs to see and is set rarely, so rare-and-silent was the wrong combination. A uniform vector prints as the scalar it is.Design.update_rep_weightsis a third the size, with the same behaviour. It resolved nine sentinel parameters by hand and rebuilt the variant from a method name, because seven internal callers needed exactly that. They now usemsgspec.structs.replace, so it keeps only the two jobsreplacecannot do — changing the method, and creating weights where there are none. Unchanged for callers: the same signature, the same first-time error, and a variant-specific field still carries only within the same method.import_labels_from_svyio_metanow requires the frame the metadata describes as its third argument. Resolving a measurement type depends on how well the value labels cover the observed values, which cannot be judged without the data. It is required rather than optional on purpose: an omitted frame would silently fall back to the very behavior this release fixes. The function is internal (svy.engine.io, not exported from the top-levelsvynamespace) and had a single production call site.
Removed
BREAKING:
Sample.weighting.create_variance_strata. Undocumented — no mention in any changelog, README or guide — never called by svy itself, and reachable in practice only through a hint insidecreate_brr_wgts’s error telling you to call it. It was a mandatory pre-step svy made you perform by hand and then punished by clobberingDesign.stratum. The pairing algorithm survives as a private helper with its edge cases intact (odd PSU counts, tuple strata, ordering, reproducible shuffling); itsorder_by,shuffleandintocontrols moved ontocreate_brr_wgtsandcreate_jk_wgts, where they are discoverable.into=is nowstratum_name=.# before # after s = s.weighting.create_variance_strata( s = s.weighting.create_jk_wgts( method="jk2") paired=True) s = s.weighting.create_jk_wgts(paired=True)svy.core.design.make_rep_weights. A strictly weaker duplicate ofsvy.RepWeights— same job, but only four of the parameters, so it silently droppedscale,rep_coefsandkind. It was never exported fromsvyorsvy.core, so it was reachable only asfrom svy.core.design import make_rep_weights, and had no callers in the package. Usesvy.RepWeights(method=..., prefix=..., n_reps=...), which takes the same three positionally.
0.24.1 — 2026-08-11
Fixed
Value labels survive being written. Reading back what
apply_labelshad just stored raisedAttributeError: 'int' object has no attribute 'code'(#127):s.wrangling.apply_labels(categories={"q1": {1: "Yes", 2: "No"}}).labelsVariableMeta.clonecopied throughmsgspec.structs.replace, which does not run__post_init__before msgspec 0.21 — and the floor was>=0.19.0.__post_init__is the one place that turns the authoring dict{1: "Yes"}into theValueLabelpairs the field is declared to hold, so on a resolved 0.20 the dict was stored verbatim and the failure surfaced at the next read rather than at the write. Sinceapply_labelsreachesclonethroughset_value_labels, the whole documented path for labelling a variable was affected on that msgspec.Cloning now goes through the constructor, which normalizes and re-checks the invariants by definition, on any msgspec. The floor moved to 0.21 as well (see below) — two fixes for one defect, deliberately: the pin closes it for
replaceeverywhere in the codebase, and the constructor closes it here regardless of what the pin later becomes. Applied toVariableMeta,LabelandCategoryScheme— the three structs that normalize in__post_init__.
Changed
msgspec floor raised to 0.21.
msgspec>=0.19.0becamemsgspec>=0.21.0(#127).0.21 is the first release where
structs.replaceruns__post_init__; 0.19 and 0.20 skip it, verified directly against each. Any struct that normalizes or validates in__post_init__is therefore only as sound as the resolved msgspec, and svy has several — the value-label defect above is what that looks like in practice, and it reached users because a lockfile resolving 0.21 hid it from the test suite while a docs environment resolving 0.20 published the traceback.Nothing in svy needs a 0.20 or 0.19 runtime, and 0.21.1 was already what every lockfile here resolved, so this only rules out an installation that was never tested.
0.24.0 — 2026-08-10
Added
corrandcov: design-based correlation and covariance (#124). Both are available onsample.estimationwith Taylor and replication variance,by=,where=, domains anddeff.Neither statistic has a direction, so neither takes
y/x. A call names a set of columns through one symmetriccolsargument, and every requested pair is returned as its own row:sample.estimation.corr(("income", "age")) # one pair sample.estimation.corr(["income", "age", "educ"]) # every unique pair sample.estimation.corr([("income", "age"), ("income", "educ")]) # exactly theseThe three spellings are disambiguated by element type and agree wherever they overlap, so a two-element list and a two-element tuple mean the same thing. A flat list yields off-diagonal pairs only — for covariance as much as correlation — so a variance is requested explicitly as
cov(("a", "a")).kind=selects the coefficient ("pearson"today) whilemethod=remains the variance estimator. Since pandas spells the coefficientmethod=, that mix-up is caught by name and redirected; a recognised but unimplemented coefficient reports that it is not supported yet rather than invalid.Correlation is bounded, so its interval is built on Fisher’s z scale and transformed back — the same move
propmakes with a logit, and for the same reason: a symmetric Wald interval can otherwise report bounds outside [-1, 1].ci_method="wald"opts out. Covariance is unbounded and takes Wald.Validated against R survey 4.5: covariance and its SE against
svyvar, correlation and its SE againstsvycontrast’s own delta method over the moment means — R has no correlation SE of its own, so that agreement is an independent check of the linearization rather than a restatement of it — plus stratified and JK1 replicate fixtures. AddsPopParam.CORRandPopParam.COV.
Changed
BREAKING:
deffnow names the SRS reference instead of taking a boolean.deff=Trueanddeff=Falseare rejected with a structuredMethodErrorthat names the replacement; the argument takes"wor","wr"orNone.before now deff=Truedeff="wor"— without replacement, Kish’s design effect. Same numbers as before.deff=Falseomit the argument, or deff=None— deff="wr"— with replacement, the square of Kish’s “deft”A boolean never said which reference it meant. Accepting it silently would leave two spellings for one thing indefinitely, and ignoring it would be worse:
Literalis not enforced at runtime, so a string-only implementation would readdeff=Trueas “off” and quietly stop reporting a design effect the caller asked for. Rejecting it loudly is the only option that cannot mislead."replace"is accepted as an alias for"wr", since that is R’s spelling.The reference describes the denominator — what the design variance is compared against — and says nothing about how the sample was drawn;
deff="wr"is not a claim of with-replacement sampling. The two differ by exactly the finite-population correction1 - n/N, so they agree closely at small sampling fractions and diverge sharply otherwise: at f = 0.8, an evaluation-study shape, they differ five-fold (5.027 against 1.005, both matching R)."wor"infersNfrom the sum of the weights, so it is meaningful only while those weights remain reciprocals of selection probabilities. Afternormalize, or to a lesser degree raking or calibration, that sum is no longer a population count and the design effect is silently wrong — svy cannot detect this in general, since weights normalized to twice the sample size pass every check and yield a plausible wrong answer."wr"has noNin it and is unaffected. The one provable case now raises rather than returning a column of NaN: when the weights sum to no more than the sample size, the correction is zero or negative, which means either rescaled weights or a census with no sampling variance to compare against.The reference is recorded on the result. It appears in the printed header beside the variance method —
Estimate: MEAN (TAYLOR, deff=wr)— is available asEstimate.deff_ref, and round-trips through serialization as an optionaldeff_reffield.to_polars()is deliberately unchanged: that frame carries no method, param or design either, and a deff column there was already provenance-free.Error codes default per class.
SingletonErrorandSvyWarningsErrorfell back to the baseSvyErrorcode when raised without an explicit one, so two distinct failures could surface under the same identifier. Each now carries its own default (#123).svy.serializeraises svy’s structured errors instead of bare built-ins. The serialize module was the last public surface still raisingTypeError/ValueErrorwhere the rest of the library uses theSvyErrorhierarchy — structured errors with a stablecode,expected/gotcontext, and a hint. The newSerializationError(exported assvy.SerializationError) covers the serialize-specific failures, and the unfitted-model case now reuses the sameModelErrorthatGLM.predict()already raises:call before now code serialize(<unsupported type>)TypeErrorSerializationErrorUNSUPPORTED_RESULT_TYPEserialize(<unfitted GLM>)ValueErrorModelErrorMODEL_NOT_FITTEDfrom_json(<no "kind">)ValueErrorSerializationErrorPAYLOAD_MISSING_KINDfrom_json(<unknown "kind">)ValueErrorSerializationErrorPAYLOAD_UNKNOWN_KINDBreaking for callers that catch the old built-ins:
SvyErrorsubclassesException, notTypeError/ValueError. CatchSerializationError/ModelError(orsvy.SvyErrorto cover every svy failure) instead, or match on thecodefield, which is the stable contract.
0.23.0 — 2026-08-05
Added
sample.estimation.quantile()— design-based quantiles with standard errors (#112). Previously only the median carried a standard error; every other quantile was available as a point estimate throughdescribe(percentiles=), with no variance.quantile()estimates any set of probabilities under Taylor linearization or replicate weights, withby=domains andwhere=filters:sample.estimation.quantile("income") # quartiles, the default sample.estimation.quantile("income", p=0.9) # a single Estimate sample.estimation.quantile("income", p=(0.1, 0.5, 0.9), by="region")Standard errors follow Woodruff (1952), the construction behind R’s
svyquantile: the design-based variance of the estimated proportionP(Y <= q)is taken on the probability scale, and the interval comes from inverting the weighted CDF atp ± t·se_p. The reportedseis the back-solved half-width.pfollows the same rule asy: a scalar returns oneEstimate, a sequence returns one per probability.median()is unchanged and remains thep = 0.5case reported asPopParam.MEDIAN.EstimateList— thelistsubclass now returned wherever a call estimates several things at once (mean(["a", "b"]),quantile(y, p=(...))). Printing a bare list previously showed object reprs ([<svy.estimation.estimate.Estimate object at 0x…>, …]); anEstimateListrenders its members as one table and addsto_polars(). It is alist, so indexing, iteration, unpacking, andisinstance(result, list)are unaffected.Serializing a multi-estimate result also works for the first time —
serialize()dispatches on exact type and previously raisedTypeError: No serializer registered for list. The newestimate_listkind wraps the members, each serialized exactly as a standaloneEstimate.
Fixed
Quantile and median confidence limits now invert the CDF with the rule that located the point estimate. The Woodruff interval endpoints were always interpolated linearly, regardless of
q_method. R’soldsvyquantilehands onemethod/fpair to both its pointapproxfunand its endpointapprox; interpolating linearly while estimating with, say,"higher"pulls both endpoints inward and understates the standard error. The error grows where consecutive order statistics are far apart: on a 2000-row stratified design it was 2.0% at the median and 10.3% atp = 0.99, always in the anti-conservative direction. Estimates, standard errors and both confidence limits now agree with R to machine precision at every probability tested from 0.01 to 0.99, including domains.This changes
median()standard errors and confidence limits. Point estimates are unaffected.The Woodruff linearization is centered on its own weighted mean, matching how R computes the variance (
svymean(U, design)). Only thelinear,middleandnearesttie rules moved — up to 4e-4 relative on the confidence limits — because thehigher/lowerinversion snaps to an order statistic and absorbed the difference.median()and default-q_methodresults are unchanged by this one. Seesvy-rs.
0.22.1 — 2026-08-04
Changed
Taylor estimation uses the cores it is given. On a 10-core machine at 1M rows a single-variable mean used 1.15 cores and an 8-variable batched mean reached 1.8 of a possible 8. None of it was a rayon width problem — three pieces of redundant serial work sat around the fan-out, and removing them is what freed the parallelism:
sample.estimationreturned a newEstimationon every attribute access, so the_data_version-keyed caches it carries — the factorized design arrays and prepared design info — were discarded before they could ever be reused, and everysample.estimation.mean(...)re-derived the whole design. The accessor is now retained perSample. Derived samples are handled by an identity check:_replace_dataforks withcopy.copy, which carries the cached accessor over verbatim, and without the check a fork would answer with anEstimationstill bound to its parent’s data.- The reporting metadata on each
Estimate(unique stratum labels, PSU count) was computed withnp.uniqueover the full-length design arrays per estimate. A batched call produces oneEstimateper variable, so an 8-variable mean did 16 full-length passes — 63% of that call’s wall time at 1M rows, all serial. These are properties of the design, not of the variable, so they are memoised on the design cache and invalidate with it. - The Rust kernels stopped indexing the design twice per estimate and now overlap their independent halves — see
svy-rs.
Measured on a 10-core M1 Max at 1M rows, 50 strata, 2000 PSUs:
case before after speedup cores used mean, 1 variable 70.9 ms 22.4 ms 3.2× 1.15 → 1.60 total, 1 variable 83.8 ms 27.4 ms 3.1× 1.12 → 1.86 mean, 8 batched 163.2 ms 38.6 ms 4.2× 1.78 → 3.47 mean by group 144.4 ms 131.9 ms 1.1× 3.45 → 3.46 Thread scaling from 1 to 10 threads went from 1.11× to 1.54× for a single variable and 1.58× to 3.03× for the batched call. A single estimate overlaps two halves rather than fanning out, so its ceiling is ~2× by construction.
Estimates, standard errors and degrees of freedom are unchanged — bit-for-bit, and identical at 1, 2 and 10 threads.
One trade-off worth knowing: retaining the accessor means its cached design arrays (~31 B/row) stay alive as long as the
Samplerather than being freed between calls. Peak memory is unchanged — those arrays were rebuilt on every call before, so the high-water mark was always there (2433.8 MB before vs 2434.5 MB after, over six calls at 10M rows). What is new is that a long-lived process holding many largeSampleobjects idle now holds their design caches too.
Fixed
Replication variance no longer costs O(B²) in the replicate count B. The replicate kernels themselves were never at fault — they already do a single O(n·B) pass — but two sites in the Python prep layer tested column membership once per replicate weight column:
prepare_datawith[c for c in rep_weight_cols if c in df.columns], andEstimation._ensure_float64withc in data.columns and data[c].dtype != ....df.columnsis a property that rebuilds the entire column-name list across the FFI boundary on every access, so B lookups cost O(B²) Python-string constructions;_ensure_float64also materialised a Series per column. A cProfile run at n=25,000 / B=800 putPyDataFrame.columnsat 57% of total runtime, called ~2·B times per estimate.With total work n·B held constant at 20M cells — an identical 160 MB replicate weight matrix in every case — the sweep spanned 36× across cases doing identical arithmetic. Column names are now snapshotted into a set once, and dtypes read from a single
data.schemasnapshot that answers both existence and dtype:n B before after speedup 200,000 100 0.0067 s 0.0059 s 1.1× 50,000 400 0.0219 s 0.0077 s 2.8× 25,000 800 0.0704 s 0.0116 s 6.1× 12,500 1600 0.2442 s 0.0201 s 12.1× Sweep spread drops from 36.2× to 3.4×. This matters most for bootstrap designs at B=1000+; it is negligible for BRR (32–64) and SDR/ACS (80). Results are numerically inert — mean, total, ratio, prop and mean-by-domain are bit-for-bit identical to the previous build.
0.22.0 — 2026-08-02
svy labels values so results print nicely. That is the whole job. Everything that made a label list shareable — concepts across locales, hierarchies, the semantics of why an answer is absent — moves to svy-spec, where a questionnaire gives it meaning. A catalogue without one is machinery with no job.
If you set labels by hand, read them from a .sav/.dta, or print them on estimates, nothing you do changes. If you used missing-value definitions, see Removed.
0.21.1 was never published. It was numbered and its notes written on 2026-07-27, then further work landed before a tag was cut. Its changes ship here, and are kept below under their own headings so nothing is lost — but no
svy==0.21.1exists to install, which is why there is no section for it.
Added
MetadataStore.update(other, *, overwrite=False)— merge one store into another, field by field.Metadata for a variable arrives from several places that each know a different part of it: measurement types inferred from the data, missing-value codes declared by the analyst, question wording carried by an instrument spec. The only way to combine two stores was
set, which replaces a wholeVariableMeta— so applying a spec silently cleared missing codes, because a questionnaire has no concept of them and its record carriesmissing=None. That loss is invisible until an export drops the declarations.updatemerges per field, which means a source can only ever add what it knows and can never clear what it has no opinion about.overwrite=False(the default) fills only gaps, keeping labels you have already chosen;overwrite=Trueletsotherwin where both are set — for a spec whose question wording should be definitive. A fieldotherhas not set is left alone in either mode, which is the property that makes the merge safe.store.update(other) # fill gaps only store.update(other, overwrite=True) # `other` wins on conflicts
Changed
VariableMeta.value_labels,ResolvedLabels.value_labelsandLabel.categorieshold(code, label)pairs, with a mapping accepted when constructing and a dict view for lookup —.labels,.labels,.label_map. JSON object keys are always strings, so adict[Category, str]wrote{101: "Banjul"}and read it back with a string key, silently, after which every join against an integer-coded column missed.Label.categoriesalso dropped_MissingTypefrom its union: msgspec refuses to decode any union containing a custom type, so the struct raised regardless — and nothing ever set the field to the sentinel, sinceNonealready meant “no value labels”.CategorySchemeholds one entry per code,SchemeEntry(code, label), replacing themapping/missing/missing_kindscollections keyed by code. Three of those wereCategory-keyed dicts or sets that survived only through a hand-written encoder; every field is JSON-native now, so the bespoke encoder is gone —to_bytesis a singlemsgspec.json.encode— and a scheme keeps its code types by any route.
Removed
MissingDef, and every API keyed on it:VariableMeta.missing,.na_as_level,.has_missing,.with_missing;ResolvedLabels.missing_codes,.is_missing,.non_missing_labels;MetadataStore.set_missing,.set_na_as_level; the same two onSample; thehas_missingcolumns insummary()andcoverage().svy.metadatano longer exportsMissingDeforMissingKind(the enum stays insvy.core.enumerations).A 99 labelled “Refusal” is the integer 99 with the label “Refusal”. svy reads it, prints it, and forms no opinion. Absence is a polars null, which needs no metadata.
The evidence for removing rather than slimming: import never populated it. svy-io surfaces
MissingRuleandTaggedNA, and svy’s importer reads variable labels and value labels only — so the field’s only source was a hand-set value or svy-spec’s bridge, and its only consumer was writing it back out. Nothing ever acted on it: a declared code did not change an estimate then, and does not now.# before store.set_missing("age", dont_know=[98], refused=[99]) # after — the code is a value, and a value needs a label store.set_value_labels("age", {98: "Don't know", 99: "Refused"})To declare user-missing in an exported
.sav, pass it at the boundary that has the concept:svy_io.write_sav(..., user_missing=[...]).locale, everywhere.CategoryScheme.locale,SchemeRef.locale,LabellingCatalog(locale=)/.locale/.set_locale, thelocale=argument onpick,list,add_scheme,make_scheme,MetadataStore(default_locale=),MetadataStore.set_scheme, andSample.use_scheme.svy does not translate. All
localedid was choose between two schemes registered under one concept — and two concept names do that with no matching algorithm. A label is a string: write"Femme"and svy prints"Femme". svy holds one set of strings; choosing which set is svy-spec’s job.catalog.add_scheme(concept="sex_en", mapping={1: "Male", 2: "Female"}) catalog.add_scheme(concept="sex_fr", mapping={1: "Homme", 2: "Femme"})CategoryScheme.id— it only ever meantconcept:locale. The catalogue is keyed by concept now, one concept holds one scheme, andpick()is a lookup.get,removeandto_labeltake a concept where they took a scheme id.SchemeEntry.parent,.missing,.is_missing, and the lookups over them:parent_of,children_of,codes_of_kind,kind_of,substantive,missing_codes. A scheme is a code→label map. svy has no cascading selects, and why an answer is absent is a questionnaire fact.CategoryScheme.ordered. Order lives in the codes; “is this ordinal” isVariableMeta.mtype.to_label_by_concept, folded intoto_labelnow that concept is the key.171 lines with no caller anywhere in the monorepo, in svy-spec, or in any test:
is_missing_value,recode_for_analysis,display_text,polars_mask,polars_to_analysis,polars_to_display(none exported),SchemeCatalogView, and the no-op seamsvalidate_scheme_missing,normalize_scheme_missing,missing_codes_by_kind.labels.pyno longer imports polars.svy.questionnaire,MetadataStore.import_from_questionnaire, and theSample(questionnaire=)parameter. Describing an instrument is a different job from analysing the data it produced, and svy had come to own a small piece of it: a flat question model with no notion of rosters, ordered scales, or analysis units. That work now lives in svy-spec, which inverts the dependency — svy no longer needs to know what a questionnaire is.This is a removal without a deprecation cycle, which the version number alone does not convey.
Questionnairewas exported fromsvy.questionnaire, but never from the top-levelsvynamespace, never documented, and never used anywhere in svy beyond the oneSample(questionnaire=)hook — which only forwarded toimport_from_questionnaire. A patch bump reflects a path with no known consumer; if you were importing it, pinsvy==0.21.0and migrate at your convenience.To attach instrument metadata, resolve a spec, project it, and merge it in:
from svy_spec.bridge import to_metadata_store from svy_spec.resolve import resolve sample = svy.Sample(data, design, catalog=catalog) sample.meta.update(to_metadata_store(resolve(spec), catalog=catalog), overwrite=True)Use
updaterather than a loop overset: it merges per field, so a field the spec does not model — missing codes you declared, notes you added — is never cleared by applying it.Pass
overwrite=Truewhen the spec is the authority, which it is here.Sample.__init__runsinfer_from_dataframe, which setsmtypeby guessing from each column’s storage type; under the default fill-only merge that guess wins and the spec’s declared level never lands, so an ordered single-select stays Numerical Discrete instead of becoming Categorical Ordinal. Apply the spec first, then any adjustments of your own.MetadataSource.QUESTIONNAIREstays — it is what the bridge sets, and it remains the right provenance for a field-collected variable.
Fixed
The SPSS and SAS writers could not run at all.
_write_spsscalledsvy_io.write_spssand_write_sascalledsvy_io.write_sas; neither name has ever existed. Both calls carried# type: ignore[attr-defined]— the type checker had said so and been silenced.The real API is
write_sav(df, path, *, var_labels, value_labels, user_missing, ...), taking labels as separate arguments rather than onemetadatadict.SAS is more than a rename: ReadStat writes SAS Transport (XPT) only — there is no
sas7bdatwriter — and XPT carries no variable or value labels._write_sasnow writes XPT and warns that the labels did not travel, pointing atwrite_spssorwrite_stata.format=andencoding=are reported as ignored, and_write_spssloses itsencodingparameter, whichwrite_savdoes not take.It survived because the only test used a stub that defined
write_spssandwrite_sasitself. A stub that invents the interface it stands in for cannot catch a call to a function that is not there.A labelled
Table.crosstab()returnedNonefor every estimate. The frame’s rows were replaced with labels while the skeleton they are joined against stayed as codes.A value label did not apply when the code and the value disagreed in type. SPSS stores value-label keys as strings, so a
.savread back gives{"1": "Yes"}against aFloat64column, andResolvedLabels.displayreturned the bare number.displaynow bridges both directions.display_serieswas never affected.ttest_to_markdown()raisedNameErroron any call — it referenced a_stats_summary_linethat does not exist. The docstring promised a summary line the code never had; both are gone.SingletonResult.confignamed a class that does not exist (_SingletonHandlingConfigforSingletonHandlingConfig). Latent only becausefrom __future__ import annotationsleft it unresolved; it would have broken onget_type_hintsor a typed decode.Two tests shared a name, so the second replaced the first and one never ran. Three further tests had no assertions at all.
ruff checkpasses onpackages/svy/{src,tests}for the first time.
0.21.0 — 2026-07-24
Requires svy-rs 0.12.0, which carries the two variance-estimation fixes below; svy-io 0.2.0 is unchanged. Grouped confidence intervals and domain design effects change in this release — point estimates and standard errors do not.
Fixed
by=groups now use their own degrees of freedom. A by-group is a domain, so its df must be counted on the PSUs and strata that group covers. It was instead given the df of the surrounding analysis — the full design with no filter, or thewhere=mask with one — so the same subpopulation got a different interval depending on whether it was reached throughby=orwhere=. Confidence intervals for grouped means, totals, ratios, proportions and medians were consistently too narrow; the effect is negligible for groups spanning most of the sample and large for small ones (22% on a 10-record domain with 6 df rather than 56).est,se,cvare unaffected. Verified against Rsurvey4.5degf(subset(design, ...)).- Design effects no longer count zero-weight rows. Under
drop_nulls, rows with a missing response are kept and zero-weighted rather than dropped; they were still counted in the domain SRS variance’sn, inflatingdefffor any group containing them (~1–2% on the synthetic fixtures). Onlydeffis affected.
Removed
Estimate.degrees_freedom. Degrees of freedom are a per-row property — a domain or by-group is counted on its own active PSUs and strata, so grouped results legitimately carry a different df per cell. The scalar could not represent that: it wasmin()across rows, so a grouped estimate reported its smallest group, and for a by-group inside a domain that meant a headline df of 0. UseParamEst.df(also adfcolumn into_polars()) for the per-row value, andn_psus - n_stratafor the full-design df, which stays at design level under a domain filter.EstimateData.degrees_freedomleaves the serialized payload;SCHEMA_VERSIONmoves tosvy-result/0.2. Strictly a field removal warrants a major bump under the policy inserialize/DESIGN.md; 0.2 was chosen deliberately because no known consumer binds to the field, and the reasoning is recorded there.
Added
ParamEst.df— the design df backing each row’s t-quantile, carried throughto_polars()and the serialized payload. It is deliberately not shown in the printed table: it is constant for most results, so a column would repeat one number down the page and widen every table.
0.20.1 — 2026-07-23
Patch release on top of 0.20.0; svy-rs (0.11.0) and svy-io (0.2.0) are unchanged.
Fixed
tabulatepercent andcount_totalcells used an un-centered variance. A cell percentage is a ratio of two estimated totals, so its variance needs the centered (Hájek) linearization. Because the internal totals flag was inferred fromsum(weights) != 1, scaling weights to sum to 100 (units="percent") or to a caller-suppliedcount_totalrouted the standard error through the un-centered total path, dropping the numerator/denominator covariance term. Cell SEs were inflated by ap-dependent amount (up to ~12% on high-proportion cells) and the confidence interval fell back to Wald, which could dip below zero.units="proportion",units="percent", andcount_total=Nare now the same estimator scaled by a constant and agree exactly; they matchestimation.propand Rsurvey’ssvymean(~interaction(...)). Bareunits="count"is unchanged and still matches R’ssvytotal, and the Rao-Scott chi-square/F test was never affected.
0.20.0 — 2026-07-23
Builds on svy-rs 0.11.0 and svy-io 0.2.0. This release lands the round 7–8 review: correctness fixes across estimation, regression, weighting, size/power, categorical, and the dataset downloader, several of which shift standard errors closer to R survey 4.5.
Added
RepWeights.rscales— exact stratified-JKn variance.RepWeightsgains an optionalrscalestuple (per-replicate variance coefficients, R’sscale×rscalescombined);create_jk_wgtsfills it from the design’s strata and estimation threads it to the Rust kernels. svy-generated JKn weights now reproduce R’sas.svrepdesign(type="JKn")mean/total SEs andmse=TRUEcentering to 13+ digits (df = degf). Absentrscales, each method keeps its global default, so user-supplied replicate weights behave exactly as before unless the file’s documentedrscalesare provided.
Fixed
drop_nullszeroes weights instead of dropping rows (Rna.rm=TRUE/subset()semantics).prepare_dataphysically removed any row with a missing analysis value before the domain machinery ran; under standard skip patterns (ynull outside the domain) this deleted whole PSUs and strata, understating domain SEs — 15% on the reference dataset — and corrupting df. Missing analysis values now keep their rows with main and replicate weights zeroed. Verified against Rsurvey4.5 to 13+ digits; the R-calibrated ttest and ratio fixtures were regenerated with these semantics (the old expectations matched R only on physically-filtered complete-case data).- Float-typed stratum/PSU columns are accepted. Numeric design codes from CSVs (e.g. MEPS
VARSTR/VARPSU) frequently arrive asFloat64; the factorized-design cache cast them straight toCategorical, which polars forbids for floats, crashing estimation with “conversion from f64 to cat failed”. Non-string, non-integer dtypes now route throughUtf8first (float- and int-coded designs produce identical results). - SSU-level FPC is grouped by
(stratum, PSU), not PSU alone. PSU labels are commonly reused across strata, sobuild_fpc_ssu_columnmerged distinct PSUs — valid designs raisedFPC_NOT_CONSTANT, and matchingM_hivalues pooled SSU counts across strata, understating the two-stage SSU FPC. method=Noneauto-detects as documented — replication when the only variance information is replicate weights (no strata/PSU), Taylor otherwise. PreviouslyNonealways meant Taylor, silently giving replication-only designs an SRS-like variance.- Core polish and API consistency (review round 7). Replicate-weight prefix matching is strict
^prefix\d+$(a loosestartswithabsorbed columns likerepwt_flag; a count/n_repsmismatch is a typedDimensionError);set_data/update_data/set_design/update_designrebuild internal concat columns and re-run singleton detection + design validation instead of leaving stale state;describe()reports weighted std/var/quantiles (aweight convention) and computes categorical proportions over all levels;SingletonHandlingenum values are accepted bysingleton.handle();PopSize(psu=..., ssu=None)is accepted for PSU-only FPC;polars_mask()is null-safe; the design-fields cache is bounded (512 entries); importingsvyno longer replaces the host’ssys.excepthook(Rich tracebacks install only onSVY_RICH=1). Deleted the unused content-basedSample.__hash__and the dead_calculate_fpc. - GLM design gaps (round 8). Family-specific unit deviance (matching R
family$dev.resids) and null deviance at the intercept-only fit; deviance/AIC follow Rsurveyexactly (Lumley–Scott dAIC;bicisNone); replicate-weight designs get true replicate variance instead of silently falling back to Taylor SEs;design.pop_sizefeeds per-stratum FPC into the sandwich;Cat(ref=...)with an absent reference level raises a typed error listing observed levels; Cat levels, the response, and the invalid-weight filter are evaluated on in-domain rows underwhere=, eliminating phantom all-zero dummies; covariate/where-column nulls keep-and-zero-weight (preserving PSUs in stratum centering). Validated against Rsurvey4.5 to ~1e-6 or better. - GLM margins rewritten on the fitted frame with delta-method SEs (round 8).
marginsrecomputed from raw sample data with ad-hoc SE formulas; it now averages over exactly the fitted rows (post null-drop, post weight filter, with the domain column), rebuilds interaction columns from counterfactual data, differentiates the full linear predictor for AME, and uses full delta-method SEsg'V(β)gover the design-based covariance (Statavce(delta)convention). Validated against Rsurvey+marginaleffects: points to ~1e-8, SEs to ~1e-4. - Weighting adjustment/calibration/trimming marshalling (round 8).
adjustraises a typed error on unmatched response statuses (was silently encoding them as respondents and inflating weights) and derivesrespondents_onlyfrom the encoded codes (case-insensitive);adjust(trimming=..., update_design_wgts=False)trims the freshly created adjusted weight instead of the caller’s original;calibrate(bounded=True)raisesNotImplementedErrorinstead of being silently ignored; calibration targets are assembled as ordered per-term lists (fixing a “Design matrix label alignment mismatch” on shared numeric codes); the trim-calibrate cycle runs on arrays before writing (a strict non-convergence failure leaves data/design/replicates untouched) and honorsTrimConfig.by/min_cell_size;build_aux_matrixraises on nulls in a continuous auxiliary instead of filling0.0. - Weighting typed errors and sorted control order for the svy-rs 0.11.0 changes:
create_brr_wgtspre-validatesn_repsagainst the Hadamard order (MethodError.invalid_range); raking-bounds violations surface asMethodErrorat all four kernel call sites;normalize()orders control values by sorted group id, matching the kernel andpoststratify. - Wrangling edge cases (round 8).
categorize()closes the outer bin edge (Rcut(include.lowest=TRUE)) so boundary values no longer vanish from tabulations;remove_columns(force=True)cleansdesign.pop_size; a partial replicate-weightrename_columnsraises instead of corrupting theRepWeightsprefix;mutate()specs see same-call redefinitions (dependents no longer read stale values);clean_names()preserves internal concat columns;filter_records()counts and reports Kleene-null-dropped rows;fill_null(strategy="mean")casts integer columns toFloat64for an exact mean;cast(strict=True)raises on lossy float-to-integer casts. - Size and power formulas (round 8).
compare_meansis implemented (was a no-op stub returningNone); non-inferiority sizing keepsepsilonsigned (the old|eps|collapse under-sized NI designs ~5×); the one-mean two-sided clamp that produced astronomically wrongnis removed; one-sided power followssign(delta); pooled two-proportion variance and the optimal allocation ratio are un-inverted; the adjustment pipeline is reordered ton0 → DEFF → FPC → nonresponseso the FPC caps the deff-inflated size towardpop_size; parameter validation (p/moe/sigma/power/deff/resp_rate) raises typedMethodErrorinstead of silently clipping. tabulatecount CIs use the design-df t instead of the normal critical value; with few PSUs (df = 6) count CIs were ~20% too narrow, now matching Statasvy: tabulateand svy’s ownestimation.total.ranktestwith a customscore_fnhonorsby=(each by-level is its own domain, returning one result per level) and group labels reflect the levels actually tested underwhere=/by=(estimates were always correct; only the reported labels were wrong).
Security
- Dataset downloader hardened against a hostile catalog. Slugs from registry JSON flowed unvalidated into cache paths, glob patterns, and tempfile prefixes (a slug like
../../foowrote outside~/.svy/datasets); slugs are now allowlisted at the registry boundary and defensively inpath_for/clear. Downloads without a catalog hash pin the first-seen sha256 (trust-on-first-use) and enforce it thereafter. Plain-http URLs and https→http redirect downgrades are rejected (localhost exempt for development). New error codesDATASET_INVALID_SLUGandDATASET_INSECURE_URL.
0.19.1 — 2026-07-21
Added
- Bundled offline example datasets.
svy.datasets.load/catalog/describenow take asource=argument —"bundled","remote", or"auto"(default: remote if reachable, else bundled). A small, self-consistent synthetic survey — a sampling frame, its household census, and a two-stage sample drawn from that census (design weights sum to the census) — ships inside the wheel, so the docs and your own experiments run fully offline and reproducibly.SVYLAB_OFFLINE=1forces the bundled path. DatasetCatalogtype and richerDatasetmetadata.catalog()returns aDatasetCatalogthat prints as a compact table and drills into any entry’s full metadata with.get(slug)(also.slugs,.to_polars()).Datasetprints as a branded panel and gained anotesfield documenting how a bundled subset was derived from its remote counterpart.
Changed
- All dataset failures route through the
DatasetErrortaxonomy with actionable messages and codes:DATASET_NOT_BUNDLED(lists the available bundled slugs),DATASET_DOWNLOAD_FAILED, andBUNDLED_UNAVAILABLE, alongside the existing not-found, catalog, and integrity errors.
Fixed
SvyErrorpanels render again. The Rich panel path imported its renderers from a module that had since been renamed, so every error silently fell back to plain text; it now renders the branded panel. The panel also stays aligned in HTML/notebook output — the status marker is a width-1 glyph instead of a two-cell emoji — and the title, body, and metadata are spaced for readability.
0.19.0 — 2026-07-12
Added
- Batched multi-variable estimation.
estimation.mean,total,ratio,prop, andmediannow accept a list of columns and return alist[Estimate](one per variable;ratiopairs numerator/denominator element-wise and broadcasts a scalar side). A single string still returns a singleEstimate. For ungrouped Taylor estimation the list form shares one design build across variables and runs them in parallel — 4–13× faster than a manual loop at 1M rows depending on the estimator.by=, replication,drop_nulls, and the singleton scale double-pass transparently fall back to independent per-variable calls (identical results). - A variable may now appear in both
by=andwhere=.whereis domain estimation (out-of-domain weights zeroed) andbygroups on the original values — the two are orthogonal, so the previous guard forbidding overlap is removed. When awherepredicate excludes an entirebylevel (e.g. a “don’t know” code), that level is correctly absent from the results — matching R’sfilter(...) %>% group_by(...)— while every row still contributes to the shared design and degrees of freedom, so surviving groups’ estimates, standard errors, and df are byte-identical. Covers Taylor and replication, all estimands, and multi-by. - Serialization for result objects. New
svy.serializemodule provides stable, versioned serialization of every result type (estimates, t-tests, chi-square, tables, GLM fits/predictions, describe):serialize(result)returns a kind-tagged struct,to_json/to_dictexport, andfrom_jsonround-trips. Payloads carry aSCHEMA_VERSIONfor forward compatibility. - Single-stage designs and explicit population sizes. The design’s
ssu(second-stage unit) is now optional, so single-stage designs no longer need a placeholder. APopSizetype is exported for specifying finite-population sizes (FPC).
Changed
- Estimation now fails fast on unhandled singleton PSUs instead of silently under-reporting the variance. Taylor estimation (
mean,total,prop,ratio,median) raisesSingletonErrorwhen a design has single-PSU strata and no handling strategy was chosen — matching R’soptions(survey.lonely.psu = "fail"). Pick a strategy explicitly withsample.singleton.skip()/.certainty()/.center()/.scale()/.collapse()/.pool(). Previously such strata were dropped from the variance with no error or warning.
Fixed
- Taylor standard errors are now bit-reproducible. The stratified variance summed each stratum’s PSUs in the iteration order of a randomized hash set, so a repeated estimate on identical data could differ in its last digits run-to-run (far below reporting precision, but not reproducible). PSUs are now summed in a canonical order, so
mean/total/ratio/prop/medianreturn identical standard errors across runs. - Stale design cache could return silently wrong results. Estimation design caches were keyed on the identity of the data frame without holding a reference to it; after an in-place mutation freed and reallocated the frame, identity reuse could make a stale entry look current and serve design arrays for the old data. Caches are now keyed on a monotonic per-
Sampledata version bumped on every rebind, so every mutation, weighting, selection, and fork path invalidates correctly. - Replication-design crashes and related correctness fixes. Clone, column keep/remove/rename, and singleton handling now work on replication designs (previously hit stale replicate-weight API usage and could crash).
Exprnow raisesTypeErroron boolean use (and/or/not/chained comparisons) so a malformedwhere=predicate fails loudly instead of silently filtering wrong, and derived samples deep-copy metadata/warnings/design so they no longer share mutable state with the original.
0.18.2 — 2026-05-20
First release tracked in this changelog. For the history prior to 0.18.2, see the Git tags and GitHub Releases.