Release Notes

Version history and notable changes for the svy Python package.

All notable changes to svy, the Python package for design-based analysis of complex survey data — means, totals, ratios, proportions, regression, weighting, and sample selection — are recorded here. Releases follow Semantic Versioning; the layout follows Keep a Changelog.

Companion packages track their own changes: svy-io (SAS/SPSS/Stata I/O) and svy-rs (internal Rust extension).

Unreleased

Added

  • Agent skills ship with svy. Guidance for coding agents, matching the installed version: an entry file (namespace map, a complete example, getting full-precision numbers out, rules, common wrong names) and one file per namespace (io, design, wrangling, estimation, categorical, glm, weighting, sampling). Run svy-skills once per project (or svy.skills.install()) to link them into .claude/skills/; upgrades then update them. Rerun it when a release adds a skill or after recreating the environment; --copy copies instead. Other packages can ship skills under the svy.skills entry-point group, and one run installs them all.
  • Wrong names say what to use. A name that does not exist on svy, a Sample, a Design, a sample namespace or a result raises UnknownNameError (an AttributeError, code UNKNOWN_NAME) that names the replacement: svy.SurveyDesign → svy.Sample(data, svy.Design(...)), sample.mean → sample.estimation.mean, estimation.tabulate → categorical.tabulate, estimation.proportion → prop, result.est → result.to_polars(), plus close matches. svy.Design(data), unknown Design arguments (strata= → stratum=, weights= → wgt=, cluster= → psu=) and a pandas frame passed to Sample raise UsageError (a TypeError) with the fix. help(svy) now shows where to start.

0.32.2 — 2026-10-10

Changed

  • Requires polars 2.0 or later (was 1.36.1). The lock and CI test 2.x only; polars 2.0 needs Python 3.10+, within svy’s 3.11+.

Fixed

  • Wide frames are fast. Sample() cost grew with the square of the column count: 6.3 s at 20k rows × 4000 columns, now 38 ms. Every wrangling step is also 2–5× faster on wide frames, since a derived sample no longer deep-copies each variable’s metadata.
  • Wrangling works on a Sample built with catalog=. Every step raised TypeError: cannot pickle '_thread.RLock' object; the derived sample now shares the catalog, as clone() does.
  • Row order no longer depends on polars internals. Sorts on two or more columns kept tied rows in an arbitrary order. Tied rows now keep their input order in wrangling.order_by, in selection order_by, in show_data and in load(order_by=). A seeded draw with order_by on two or more columns can select different units than before; a single sort column is unaffected. Selection results are joined back in the sample’s row order. Estimate.strata is sorted, as documented, rather than in a different order on each call. describe(top_k=) breaks count ties by level, so the same levels are kept on every run.
  • Sample() collects a LazyFrame. Without a design the frame stayed lazy, and every later step resolved its schema again, with a polars PerformanceWarning each time. A design already collected it.
  • A seeded load(order_type="random") gives the same rows on every polars version. It shuffled by polars’ hash(), which changed in polars 2.0, so the same seed already loaded different rows there; it now uses NumPy’s frozen legacy stream. Seeds wider than 32 bits and negative seeds are used in full.
  • INVALID_TYPE errors name the type passed. apply_labels, column-name lists, rake margins and allocation pop_size reported got: str whatever was passed.
  • Estimation is faster, and the slowdown since 0.26 is gone. At 1M rows, a stratified clustered mean or total takes 10 ms (was 24–27 ms), by= 62 ms (was 110 ms) and where= 24 ms (was 47 ms); calibrated designs, prop, median, corr and cov are 20–50% faster. The domain-singleton check and the full-design df behind by= ran eagerly on every call. GLM fits skip a log per row in the separation check.

0.32.1 — 2026-10-08

Added

  • sample.wrangling.cast takes dtype names: cast("age", "Int64") or cast({"age": "Int64", "region": "String"}), so a cast needs no import polars and can be written in a config file. The names are polars’ own (Int8…UInt64, Float32, Float64, String, Boolean, Date, Datetime, Categorical), listed by svy.wrangling.DtypeName; an unknown name raises CAST_UNKNOWN_DTYPE with the closest match. Polars dtypes still work and are needed for parameterised types (a time-zoned Datetime, Enum, Decimal).

0.32.0 — 2026-10-08

Added

  • Row and column percentages: tabulate(share_of="row" | "col"). Each cell is a share of its row (or column) instead of the table total, the way offices publish crosstabs. Row shares are estimation.prop(colvar, by=rowvar) (column shares the reverse): same estimates, SEs and intervals, for Taylor and replication, with where=; a level absent from a row shows as 0. The Rao-Scott test is unchanged. The title says (row shares) / (column shares) and Table.share_of records it. Two-way tables only, and not with units="count" or count_total. Saved tables carry it: schema svy-result/0.8 adds TableData.share_of ("total" when read from an older payload).

  • GLM categorical terms read by value label. On a labelled column, coefficients print area_RURAL and categorical margins RURAL - URBANA; levels stay in code order. Cat(ref=) takes a value label as well as a code (a label shared by two levels raises CAT_REF_AMBIGUOUS; shared labels in names get their code appended). GLMCoef.term keeps the engineered column name and to_polars() adds term_label. glm.fit(use_labels=False) prints codes. Saved fits keep the codes only.

  • Two-way tables print their Rao-Scott tests. print(tab) on a crosstab ends with Rao-Scott F(3.30, 49.48) = 11.67, p < 0.001 (second-order) and Rao-Scott chi2(3) = 35.21, p < 0.001 (first-order); one-way tables are unchanged. ChiSquare, FDist and TDist print in the same form on their own.

  • ttest(use_labels=) and ranktest(use_labels=) show labelled groups and by= levels by their value labels, as tabulate does: Groups: area = [URBANA vs RURAL], the level rows and ── reg = Norte section headers. None (the default) uses the labels when the variable has any; False shows the codes. The result keeps the codes (GroupLevels.labels holds the labels), and to_polars("estimates") adds a <group>_label column (group_level_label with tidy=False). Saved results keep the codes only.

  • Choice enums carry a label and a description. SingletonMethod, DistFamily, LinkFunction, QuantileMethod and RankScoreMethod members have .label and .description, and describe() lists (value, label, description) for every member, so applications show svy’s own wording instead of copying it. The enums are unchanged otherwise (still StrEnum).

  • sample.weighting.trim() annotates upper, lower and by (a threshold or a number; one or several columns).

  • Contrasts can be saved. svy.serialize.to_json(est.contrast(...)) writes a ContrastData (kind "contrast"): each contrast’s est, se, cv, lci, uci, t, p_value and df, with method, alpha, df and the source estimate’s param and y, which Contrast.param and Contrast.y now also give; serialize.to_polars() returns the live table. The contrasts’ joint covariance is not saved: take contrasts on the live estimate, which has it.

  • sample.internal_columns lists the columns of sample.data that svy made, not the user: the weight-adjustment record’s snapshots (__svy_cells_*, __svy_aux_*), which the design needs to reproduce the adjustment’s variance. Keep them when saving data the design must work with, leave them out when showing it; describe() already does.

  • sample.describe(by=...) describes each group of one or several columns (crossed), a null being a group of its own: every item carries by and by_level, to_polars() has a column per by variable, and printed names show the group (hh (reg=North)). It is where counts and sums per group come from: a categorical column’s levels, or n and sum of a numeric one. top_k=None lists every level instead of the first 10. Saved describe items carry by / by_level, and DescribeResultData.top_k is None when every level was listed (schema svy-result/0.7).

  • Results record their interval method: Estimate.ci_method. "logit", "beta", "korn-graubard" or "wilson" for proportions and factor means, "fisher" or "wald" for correlations, "woodruff" for Taylor quantiles, and "wald" (est ± t·se) for means, totals, ratios, covariances and replicate quantiles, so a methods appendix can say which interval was used. Saved results carry it, and GLM fits now also keep theta / theta_se (negative binomial) and the offset column: schema svy-result/0.7 adds EstimateData.ci_method, GLMStatsData.theta, GLMStatsData.theta_se and GLMFitData.offset, None when read from an older payload.

  • Expr.to_polars() returns the polars expression a svy expression wraps, to filter or derive on a plain polars frame before a Sample exists. It replaces importing to_polars_expr from svy.core.expr.

  • Dates in expressions. svy.date(year, month, day) builds a date from three columns, expressions or integers; Expr.str_to_date(fmt) parses text (svy.col("dob").to_str().pad_left(8, "0").str_to_date("%d%m%Y") reads 8082025 as 8 August 2025). Both raise on values that are not dates; strict=False gives null instead, for codes such as 98. Expr.days_between(other), months_between(other) and years_between(other) count from the date to other in days, whole months and whole years, as Stata’s datediff() does (31 January to 29 February is 0 months; a 29 February birthday completes its year on 1 March).

  • SampleSize().allocate(n, pop_size=...) splits an overall sample size across strata, as a sample-size planning goal beside estimate_prop and estimate_mean: method="proportional" (n_h ∝ N_h, or N_h ** power; power=0.5 is square-root allocation), "neyman" (∝ N_h · S_h, with sigma= per stratum, e.g. from a previous survey) or "equal" (n / H). pop_size maps each stratum to its unit count or to any size total the split should follow (households per stratum for a PPS design); its keys are any stratum values, tuples for several stratum columns, independent of a design. Largest-remainder rounding makes the strata sum to n; min_n floors non-empty strata and cap_at_population caps n_h at N_h. SampleSize.n gives {stratum: n_h}, which sample.sampling.srs(n=...) and the pps_* methods take as is; SampleSize.allocation lists the rows. Without n, allocate splits the overall n of the goal the SampleSize holds, rounded up: svy.SampleSize().estimate_mean(sigma=12, moe=1.5).allocate(pop_size=N_h); target.from_goal / from_n record it. A stratified goal (already one n per stratum) or a two-group comparison cannot be split this way.

  • Design.case_id on several columns. svy.Design(case_id=["cluster", "hh", "line"]) identifies a record by the columns together, as a CSPro export does, like stratum and psu take a list. Everything that uses the case id takes the columns together: uniqueness (duplicates are reported as tuples), nulls, the panel checks, the case as the variance PSU on a panel, combine_samples(kind="panel", case_id=[...]), wrangling.lag, the panel adjust, renames, to_code() and the saved design (DesignData.case_id is a list). Removing one of the columns removes the case id. Results equal those with the same id built as one column.

  • sample.check() checks the data against its design and returns one report with a section per declared part, None for a part the design does not declare: weights (null, non-finite, negative and zero counts; min, max, max/min, sum, mean, Kish design effect and effective sample size over the positive weights), case_id (unique, within wave on a panel; null ids and examples of repeated ones; several id columns together), nesting (PSU codes found in more than one stratum: svy treats a PSU as the (stratum, PSU) pair, as R’s nest=TRUE does, so this is reported, not refused) and singletons (strata with one PSU and whether the singleton rule handles them). It never raises on the data’s content; limit= caps the examples. rake’s margin check, the join duplicate check, the panel case-id check and deff_w share its code.

  • sample.meta.var_labels and sample.meta.value_labels give every variable’s label and code-to-label mapping, catalog schemes resolved: meta.var_labels["RIAGENDR"] is "Gender", meta.value_labels["RIAGENDR"] is {1: "Male", 2: "Female"}. Both are read-only: an edit raises LabelsReadOnly (also a TypeError) naming the setter, since a copy would silently not change the metadata; dict(...) gives an editable copy.

Changed

  • ranktest defaults to the Kruskal-Wallis score (Wilcoxon for two groups), R’s svyranktest default. Without score= or score_fn= it raised INVALID_CHOICE, although score defaulted to None.

  • SampleSize().compare_props and estimate_mean take only method="wald". compare_props listed "miettinen-nurminen", "newcombe" and "farrington-manning", and estimate_mean listed "fleiss", but none were implemented and they raised NotImplementedError. They now raise MethodError (INVALID_CHOICE) with a hint, as compare_means does; var_mode= still sets the Wald variance for two proportions. The compare_* docstrings now state that the two groups are independent samples: correlation between panel rounds is not accounted for.

  • SampleSize.target records every argument of the goal, for estimate_prop, estimate_mean, compare_props and compare_means as for allocate: the formula (method, var_mode), alpha, power, two_sides, delta, pop_size, deff and resp_rate, scalars or per-stratum dicts as given (it was None except after allocate). Stratum keys are kept as given, in the given order: Size.stratum and the keys of SampleSize.n are the input’s values (tuples for several stratum columns) instead of their text, so n goes to srs(n=...) as is; tables print a tuple as N, u, and a stratum keyed 0 no longer prints as overall.

  • read_stata returns Stata integer variables as Int64. byte, int and long variables were Float64, and write_stata wrote integer columns as double. Integer columns now keep their type both ways, so a code 1 matches the value label keyed 1. Needs the svy-io release with integer storage (see its changelog).

Removed

  • BREAKING: combine_samples(adjust=..., wave_labels=..., kind="cs"). adjust=None averaged the weights and adjust="none" kept them; the choice is now average_wgts= (True divides by k, False keeps each wave’s weight, unset picks True for cross-sections and False for a panel). Wave labels are the keys of a mapping, so a label cannot drift from its sample: svy.combine_samples({"2015-16": s1, "2017-18": s2}); a list still works and labels the waves “wave 1”..”wave k”. The "cs" alias of kind="cross_sectional" is gone. adjust= and wave_labels= raise PARAM_RENAMED naming the replacement.

  • BREAKING: meta.resolve_labels(), meta.resolve_all(), sample.resolve_labels() and sample.labels. Use meta.var_labels and meta.value_labels. meta.set_label() / set_labels() are now set_var_label() / set_var_labels(), the names Sample already used. ResolvedLabels is no longer exported from svy.metadata. Each old name raises LabelAPIRemoved (also an AttributeError) naming its replacement.

  • BREAKING: sample.sampling.allocate(), sample.sampling.group_sizes() and svy.selection.allocate(). Allocation is planning, not a sampling action: use svy.SampleSize().allocate(n, pop_size=...), with pop_size the strata’s counts (e.g. from sample.estimation.total(...) or your frame) or size totals, and pass .n to the selectors. The "rate" method is gone (a fixed fraction is n = f · N, proportional allocation), as is "size" (pass the size totals as pop_size).

0.31.0 — 2026-09-30

Added

  • Replication variance for tabulate, ttest and ranktest: method="replication". They took only a Taylor variance, so a replicate design still needed its strata, PSUs and a singleton rule. With method="replication", tabulate re-estimates cell proportions and totals with each replicate weight and computes the Rao-Scott F and chi-square from their replicate covariance. ttest re-estimates the mean, or the two group means and their difference. ranktest re-totals the full-sample influence values with each replicate, as R’s svyranktest does. The df is the replicate df (n_reps - 1 unless the design records one), minus 1 for the t-tests and minus k - 1 for the k-sample rank test; it does not shrink on a domain. where=, by=, paired t-tests and score_fn work as with Taylor. The singleton rule, FPC and calibration sweep do not apply, and no domain-singleton findings are reported. method=None is Taylor, as in estimation: replication is never picked implicitly. Checked against R survey 4.5 (svymean, svytotal, svychisq, svyttest, svyranktest on svrepdesign) for BRR, Fay-BRR, JK1, JKn, bootstrap and SDR.

  • read_spss and read_stata read zip archives, like read_sas: the first .sav, .zsav or .por member, or the first .dta member, is read, labels included. Needs svy-io with zip support for SPSS and Stata (see its changelog).

  • read_xpt, read_xpt_with_labels, create_from_xpt and write_xpt, aliases of the SAS functions: read_sas reads SAS Transport, and write_sas writes it. With svy-io’s content-based dispatch, read_sas reads transport files under any name (e.g. .ssp) and refuses SAS CPORT files with a hint on converting them. catalog_path= labels transport files too.

  • Strata with one PSU inside a domain are detected. A stratum can hold several PSUs in the sample but only one with rows in an estimation domain (where= crossed with a by= level). Two settings of the rule cover it. domains= picks the variance: "standard" (default, R’s default) uses the usual domain variance; "apply" (center and scale only) handles such a stratum as a singleton, R’s options(survey.adjust.domain.lonely = TRUE): under center its PSU totals (the domain’s and the zero totals of the others) are centred at the grand mean, under scale it is left out and counted with the singletons. on_domain_singletons= says what svy reports: "ignore" (default, also without a rule) records a DOMAIN_SINGLETON_PSU finding on the result’s new findings (Estimate, Table, t-test and rank-test results, GLMFit, whose to_dict() includes it) and in sample.warnings at INFO; "warn" also prints one note under the table (note: 577 domain × stratum pairs with one PSU in the domain (VARSTR=2002 in sex=1; ...); standard domain variance used); "error" raises DOMAIN_SINGLETON. The finding is never a Python warning. sample.domain_singletons(by=..., where=...) lists the pairs. This applies to every Taylor estimate, table, t-test, rank test and GLM, and matches R. A domain row is one inside where= and the by= level with the analysis variables present, whatever its weight, as R’s subset() keeps zero-weight rows (the degrees of freedom and n keep counting nonzero weights, as R’s do). A calibrated design has none, as in R.

  • n on every estimate row. ParamEst.n, TtestEst.n and a table’s CellEst.n give the records behind the row: those in its domain (where=, by= level, t-test group, or the table) with a nonzero weight and the estimate’s variables present. It is the count proportion CIs already use, for Taylor and replication alike; a proportion’s category rows carry the domain’s count, each group of a two-group t-test its own, and every cell of a table the table’s count. A GLM fit’s stats.n already followed the same rule. to_polars() has an n column (the printed table does not). Saved results carry it: schema svy-result/0.5 adds ParamEstData.n, TtestEstData.n and CellEstData.n, None when read from an older payload. With deff="wr", n / deff is the effective sample size.

  • sample.to_code(data=None) writes the Python that rebuilds the sample: import svy, data = svy.read_parquet(path) when data points to a .parquet file (the pair of svy.write_parquet(sample, path)), and sample = svy.Sample(data, svy.Design(...)) with every setting that changes the estimates: replicate weights, the weight-adjustment record and singleton handling. Running it on the same data gives the same design and estimates. The data is the caller’s; metadata (labels) is not included. svy.WgtAdjustment is exported so the script can name it.

  • A design can be saved and restored. svy.serialize.serialize(design) / to_json(design) give a DesignData (schema svy-design/0.1) holding every field: strata, PSUs, weights, PopSize, replicate weights with all their settings and the weight they go with, and the weight-adjustment record. svy.serialize.to_design(data) returns the live Design, equal to the original. Estimates on Sample(data, to_design(...)) equal those on the live sample, SEs included, after poststratification, raking, calibration, standardization and trimming, Taylor and replication alike. The design history is not part of it.

  • The singleton rule is declared on the design: svy.Singleton. svy.Design(..., singleton=svy.Singleton("center")), or singleton="center" for short, or sample.update_design(singleton=...), like R’s lonely.psu and Stata’s singleunit(). method is "center", "scale", "skip", "self_representing" (the former certainty(): R’s and Stata’s “certainty” is svy’s skip), "collapse" or "pool"; the other options are keyword-only and belong to one method: collapse‘s using ("smallest", "largest", "next", "previous", a mapping {singleton stratum: target} in the stratum columns’ values, or a callable), within, order_by, descending, rstate, and pool’s name; an option for another method raises INVALID_SINGLETON_RULE. The rule is intent: svy applies it to the singletons the data has now, when the sample is built and after every filter, recode or design edit, and is idle without singletons. sample.singletons lists the singleton strata of the current data and how the rule handled each (handled: center, collapse -> Center, pool -> __pooled__, …), and sample.n_singletons counts them; the rule is read back from sample.design.singleton. An explicit mapping that does not fit the data raises at the next Taylor analysis (SINGLETON_UNMAPPED, SINGLETON_TARGET_MISSING, SINGLETON_TARGET_SINGLETON); its entries for strata that are no longer singletons are ignored with an INFO finding, and a strategy whose mapping changes records SINGLETON_COLLAPSE_CHANGED. within= takes any column constant within a stratum, missing values ignored (it silently did nothing for a non-stratum column); an explicit mapping across it raises SINGLETON_TARGET_OUTSIDE_WITHIN. order_by= is checked the same way: a column taking several values within a stratum raises SINGLETON_ORDER_BY_INVALID (it used to order the strata by whichever value came first). The rule is in design_history, saved by svy.serialize (SingletonData, schema svy-design/0.1, defaults left out; a callable using or a Generator rstate cannot be saved) and written by to_code(); repr and the printed design show it (Singleton('collapse', within='region')). It stays through renames (within/order_by follow them), stratum/PSU/SSU edits, weighting and use_weight; remove_columns protects within/order_by columns, and the rule goes when the design loses its stratum. combine_samples carries a rule every input declares (different rules, or an explicit mapping on cross-sections, raise) and add_stage carries stage 1’s. Replicate weights are built from the rule’s strata and PSUs under collapse, pool and self_representing; under skip the singletons contribute nothing (recorded); create_jk_wgts and create_bs_wgts raise SINGLETON_REPLICATES under center/scale or without a rule, where they used to drop the singletons silently. filter_records(check_singletons=True) reports singletons the rule handles at INFO. svy.SingletonMethod names the methods; svy.utils.deprecated marks deprecated API.

  • Strata named by their own values. A collapse mapping (svy.Singleton("collapse", using={49: 48})), a using callable’s return value, candidates_for and compare take the stratum columns’ values (a tuple for tuple strata: {(True, 4): (False, 1)}). svy’s key strings ("49", "true__by__4") still work; a string that names two strata raises.

  • design.columns(data_columns=None) lists every column a design needs from the data, de-duplicated and in order: design fields, pop_size, replicate weights, the units the replicates were built from, then the weight-adjustment record’s new_wgt, prev_wgt, cells and aux. data_columns resolves auto-detected replicate padding. specified_fields is unchanged.

  • wrangling.rename_rep_wgts({"w": "final_w", "ps_wgt": "ps_final"}) renames replicate-weight sets by prefix: the design’s and those of earlier designs in design_history. Each column keeps its number and zero-padding (w01 → final_w01), only the set’s own columns are renamed, and every design using the set, current or earlier, follows through the rename_columns path. An unknown prefix (the error lists the sets the sample carries), an empty new prefix, a new name that is already a column, and two sets renamed onto the same names are refused before anything changes; a prefix mapped to itself is a no-op.

  • RepWgts.wgt, the full-sample weight the replicates go with. Design fills it from its own wgt; create_*_wgts and weighting with replicates pair them with the weight they produced.

  • Threshold.quantile(p) and Threshold.absolute(v). Trimming bounds say what they are: the p quantile of the positive weights, or the value itself. Both compose with the statistic kinds (Threshold.quantile(0.99) + 2 * Threshold("sd")), and reprs are runnable (Threshold.quantile(0.99)). quantile(99) is refused with a pointer to 0.99. Threshold is now the class and Cap its alias, so existing code runs unchanged.

  • svy.serialize.to_polars(data, *, row_index=None, **options) gives a serialized result’s table, the same frame its live result’s to_polars() returns, from the payload alone (for example after from_json). It covers every result kind that serializes (estimates, estimate lists, t-tests, tables, chi-square tests, GLM fits, GLM predictions and describe results), with the same options (tidy, use_labels, component, exponentiate). A serialized estimate stores the labels of the variables and levels it holds, so its label columns come back without the sample’s metadata. row_index="row" adds a first column holding each table row’s position in the payload’s rows; estimate tables are sorted for display, so this is how a table row is matched back to its payload row. The live and payload tables are built by the same function per result kind.

  • ChiSquare.to_polars() and DescribeResult.to_polars(). A one-row frame of df, value and p_value, and one row per described column with its scalar fields (nested frequency tables and percentile lists are left out).

  • GLMFit.alpha, the level of the coefficient intervals, from glm.fit(alpha=). The printed interval headers follow it ([0.05 0.95] at alpha=0.1; they always read [0.025 0.975]), as does the p-value highlight. GLMFitData carries it; a payload without it decodes as 0.05.

  • GLMFit.where_clause, the where= domain of the fit, formatted as Estimate.where_clause is and printed under the modeled variable. GLMFitData carries it (default None).

Changed

  • Breaking: httpx is optional; a base install makes no network calls. The online dataset catalog moves to the new remote extra (pip install "svy[remote]", also in svy[all]), and import svy no longer imports httpx. Without it, datasets.load(source="remote"), catalog(source="remote") and describe(source="remote") raise REMOTE_UNAVAILABLE, whose hint names the extra; under the default source="auto" they return the bundled subsets without a warning (a dataset with no bundled subset, or force_download=True, raises). source="auto" now reads a full copy already in the local cache (~/.svy/datasets, or SVYLAB_CACHE_DIR) before trying the network, the most recently downloaded version first, so a dataset downloaded once, or a copied cache folder, loads offline; force_download=True and source="remote" still go to the catalog. great-tables is dropped from report and all: nothing used it.

  • Breaking: ranktest takes the rank score as score=; method= is the variance method. ranktest(y, group=..., method="kruskal-wallis") becomes ranktest(y, group=..., score="kruskal-wallis"), so method= means "taylor" or "replication" as in tabulate, ttest, glm.fit and estimation. Passing a rank score as method= raises INVALID_CHOICE with a hint naming score=.

  • Breaking: glm.fit uses Taylor linearization unless method="replication" is passed. It used replicate weights whenever the design had them, unlike estimation and the categorical tests, where replication is never implicit. glm.fit(..., method="replication") gives the previous replicate standard errors. With the Taylor default on a replicate design, the singleton rule applies and TAYLOR_WITHOUT_DESIGN warns when the design has no stratum or psu.

  • Breaking: PROP_CI_BOUNDARY is a note under the table, not a warning. When a proportion is 0 or 1 and the interval method has none there (logit, beta, wilson), prop() and mean(as_factor=True) no longer raise a SvyUserWarning; the estimate prints one line under its table, note: CI undefined at p = 0 or 1 for 19 rows (ci_method='logit'), on every such result, a repeat on an unchanged sample included. The NaN bounds are in the table, so the note is shown by default. The finding is kept like DOMAIN_SINGLETON_PSU: it keeps its detail and hint, is on the estimate’s findings (WARNING) and in sample.warnings (INFO), and its extra holds n_rows and every cell (by, level, p). A list of estimates prints one note, split by variable (for 5 rows (a: 3, b: 2; ci_method='logit')). Code that relied on the warning, e.g. running with warnings as errors, should check estimate.findings. to_polars() is unchanged.

  • Breaking: t-tests, rank tests, tables and GLMs refuse unhandled singleton strata. categorical.ttest, categorical.ranktest, categorical.tabulate and glm.fit (Taylor variance) raise SINGLETON_ERROR, as estimation always has, when the design has strata with one PSU and no rule for them. They used to leave those strata out of the variance silently. Declare a rule, sample.update_design(singleton=svy.Singleton(...)); svy.Singleton("skip") reproduces the old numbers. A GLM on replicate weights is unaffected.

  • Breaking: __svy_ columns are svy’s; everything in sample.data is your data or your design’s. svy’s bookkeeping is renamed and kept out of sight: the row index svy_row_index is now __svy_row_index__, the design keys stratum_svy_internal_cols_concatenated (and psu_, ssu_) are now __svy_stratum_key__, __svy_psu_key__, __svy_ssu_key__, and the singleton variance columns keep their __svy_var_*__ names. None of them is shown by sample.data, show_data, dtypes, meta or the printed column count, or written by svy.write_parquet and the other writers; svy rebuilds them on construction and on every change. sample.data and the saved file hold the user’s columns plus design.columns(), the weight-adjustment record’s __svy_cells_*/__svy_aux_* snapshots included. The svy_* selection outputs (svy_prob_selection, svy_sample_weight, …) are unchanged. A column named like the bookkeeping, in Sample(...), set_data, update_data, clone(data=), mutate, the value writers (into=), cast or with_row_index, raises RESERVED_COLUMN naming it; drop it from a frame taken from sample._data, or from a file written by an earlier version (its __svy_var_*__ columns). Such a file’s svy_row_index and *_svy_internal_cols_concatenated columns are now ordinary columns; drop them too.

  • The PROP_CI_BOUNDARY warning names the by columns (pov=0 in quint=poorest, q=poorest, zone=y for several) instead of svy’s key column, and lists the cells in domain order.

  • Breaking: the templates are consistent. controls_margins_template and control_aux_template return the columns’ own values as keys (an Int column gave {'1': nan}) and take one missing-value option, na="error" | "level" | "drop" (default "error") with na_label; cat_na= and by_na= (also on build_aux_matrix) are refused with a pointer to na=. controls_margins_template no longer includes nulls as a level by default.

  • Breaking: weighting parameters are named for their values. rake(ll_bound=, up_bound=) become rake(bounds=(lo, hi)): bounds on the factor g = new/old weight, None on a side for open, None for none. They are checked after raking, not enforced, as before (BOUNDS_EXCEEDED, now with expected={"bounds": [lo, hi]}). A bad bounds (not a 2-tuple, a non-number side, NaN or infinity, lo > hi) raises INVALID_TYPE or INVALID_RANGE. calibrate(bounded=) and calibrate_matrix(bounded=) become bounds=, which raises NOT_SUPPORTED (a NotImplementedError) when a side is set. calibrate_matrix(control=) becomes controls=. strict= on poststratify, rake, calibrate and calibrate_matrix becomes on_nonconvergence="error" | "warn" | "ignore" (default "error"; strict=True is "error", strict=False is "warn"), covering MAX_ITER_REACHED, CONVERGENCE_FAILED, CALIBRATION_NOT_MET and the trim cycles. The old names are refused with PARAM_RENAMED, whose hint names the new parameter and shows the caller’s values in the new form (bounds=(0.5, None)). Every public weighting method now documents every parameter.

  • Breaking: standardize, adjust and trim raise on non-convergence by default. They take on_nonconvergence="error" | "warn" | "ignore" (default "error") like the other iterative methods: for standardize’s trim-standardize cycle, adjust’s trimming= pass and trim’s iterations in any domain. They used to keep the result and warn (MAX_ITER_REACHED), which is now on_nonconvergence="warn". With "error" the call raises CONVERGENCE_FAILED (for trim, naming the domains that did not converge) and the sample is left as it was.

  • Breaking: rake’s trim-rake cycles are capped by trimming.max_iter, as in poststratify and calibrate; rake(max_iter=) now caps only the raking iterations of each pass. The cycles used to be capped by rake(max_iter=) and trimming.max_iter was ignored there. The finding and error name trimming.max_iter.

  • Breaking: every on_* parameter means the same. "error" raises and leaves the sample as it was; "warn" records the finding and raises it once; "ignore" records it at INFO without raising (it used to record nothing). This applies to on_nonconvergence, wrangling.join(on_unmatched=) (JOIN_UNMATCHED), wrangling.filter_records(on_singletons=) (SINGLETONS_DETECTED) and combine_samples(on_mixed_design=) (COMBINE_MIXED_DESIGN, which was only logged). Defaults are unchanged. filter_records(on_singletons="error", inplace=True) now leaves the sample unfiltered when it raises, and an unknown on_singletons value is refused with INVALID_CHOICE.

  • Breaking: weighting errors are data. Each input failure has a specific code (CONTROLS_KEYS_MISMATCH, MARGINS_DISAGREE, RESP_STATUS_UNKNOWN, and others under svy.errors.WeightingError, a MethodError), expected/got hold structured values (the cells present, {"missing": [...], "extra": [...]}, {margin: total}, {status: count}) instead of text, and hints use the caller’s names and values. SvyError.to_dict() keeps them structured and JSON-safe. Rake’s INVALID_CONTROL_TOTALS, ZERO_CONTROL_TOTALS and INVALID_SHARES become CONTROLS_* codes.

  • Breaking: weighting targets match the column’s values, with a text fallback. Every keyed input (poststratify/normalize controls and shares, each rake margin, standardize shares, calibrate controls and domains, adjust resp_mapping) matches a key to the column’s own value first, then to its text form ("1" for 1, "true"/"True" for booleans, "1.0"/"1" for 1.0, ISO strings for dates), so controls stored as JSON work unchanged. Tuple keys match part by part. A key that could mean two levels, or two keys naming one level, is refused. One rule for extra keys everywhere: a key naming no level in scope is an error unless its target is 0 (calibrate used to drop them silently; rake failed in the kernel). A value listed under two resp_mapping codes is refused instead of the last one winning.

  • Breaking: one rule for findings. A finding about the data or a sample is recorded in sample.warnings and raised once as a svy.SvyUserWarning (a UserWarning) at the caller’s line, with the message "[CODE] title: detail". Sample.warn applies it for every finding. An identical finding (same code and detail) is recorded and raised once per sample state: the state is the sample’s data version, which changes on every data or design rebind (set_design, update_design, wrangling), so the same finding after such a change is a new event and is raised again, while repeating it on an unchanged sample is not. A fork has its own state and does not raise its parent’s findings; the store’s per-code cap (100) counts across states. A finding is raised when it tells the caller something they did not ask for, and only recorded (INFO) when it confirms a consequence of an explicit choice: trim audits, WEIGHT_SUM_CHANGED from trim(redistribute=False) (it was a WARNING), and SELECTION_N_EXCEEDS_GROUP under wr=True. force=True removals stay raised, since they also report what else was dropped. Findings that svy raised as a bare UserWarning, or only logged, now follow the rule with a code: combining samples (COMBINE_*, PANEL_UNITS_LOST, PANEL_SMALL_OVERLAP, WAVE_LABELS_UNORDERED, recorded on the combined sample), add_stage (STAGE_PSUS_UNMATCHED), selection (SELECTION_*), GLM (GLM_ROWS_DROPPED, GLM_TERM_DROPPED, MAX_ITER_REACHED), removed design fields (DESIGN_FIELDS_REMOVED), uncredited weight adjustments (ADJUSTMENT_NOT_CREDITED), replicate weights removed by a weight change (REP_WGTS_RESET) and cleared singleton handling (SINGLETON_RULE_CLEARED). Where there is no sample (a bare Design.update, allocation, datasets, SAS export, accessor registration) the warning is raised only, with the same category and line. Code that runs with warnings as errors now sees svy’s findings, including ones that were only recorded before (e.g. TAYLOR_WITHOUT_DESIGN); filter them with warnings.filterwarnings("ignore", category=svy.SvyUserWarning). Messages now start with the code.

    In weighting: rows left unadjusted because a cells, by or trimming-domain value is null are a finding, CELLS_NULL_UNADJUSTED, with the count per column in got ({"zone": 3}) and the number of rows in extra, from poststratify, normalize, standardize, adjust, trim and calibrate’s trimming.by; the rows still keep their previous weight, and nulls outside where= are not a finding. Non-convergence kept under on_nonconvergence="warn" is MAX_ITER_REACHED (rake, and the trim cycles of poststratify, standardize, rake and calibrate); rake no longer prints “Warning: Raking did not converge”. standardize’s partial domains are DOMAIN_LEVELS_PARTIAL and panel adjust’s skipped missing-in-scope rule PANEL_SCOPE_NOT_WAVES. Only display_iter=True prints. on_nonconvergence="error" still raises.

  • Breaking: adjust(respondents_only=True) drops only the rows the adjustment handled. A row in no adjustment class, because its cell is null or it is outside where=, was not adjusted, so it now keeps its weight and stays in the sample whatever its response status; before, a nonrespondent, unknown or ineligible among them was dropped and its weight left the sample. This matches the where= documentation and panel adjust. Null-cell rows are reported by CELLS_NULL_UNADJUSTED, with extra["nonrespondents_kept"]; rows outside where= are the caller’s choice and not a finding.

  • Breaking: domain and category levels keep the column’s type. ParamEst.by_level and y_level held the kernel’s text: by="zone" gave ("3",), a boolean by gave ("false",), and mean("zone", as_factor=True) gave "1.0". They now hold the column’s own values ((3,), (False,), 1) for every estimator, with where=, with several by variables (a tuple of values), and under Taylor and replication alike; so do Estimate.domains, keys(), the t-test’s by_level and group_level, the rank tests’ by_level and group_levels, and JSON payloads. A proportion of an integer-valued float column now reports 1.0, not 1, and string codes such as "01" are no longer read as integers. Test headers print numbers unquoted ([1 vs 2]). contrast() still accepts the old string keys ({"3": 1, "1": -1}, "false", "1.0"). A level of a date, datetime, time or duration column goes into JSON as its ISO string with its type recorded under a top-level "temporal" field, and from_json gives the value back. Printed tables are unchanged, except that as_factor levels print as the column’s values (1, not 1.0).

  • Breaking: a bare number is always an absolute trimming bound. trim(upper=0.9), TrimConfig(upper=0.9) and trimming= read a number in (0, 1] as a quantile and a larger one as a cap, so an absolute cap of 0.9 (common after normalizing) silently became the 90th percentile. A number now always caps at that value; write Threshold.quantile(0.9) for the 90th percentile.

  • Estimate.to_polars() and EstimateList.to_polars() are a data view, no longer the printed table. Levels keep their codes under the variable’s name, and each variable with value labels gets a <var>_label column next to it (by_label and y_level_label with tidy=False). Rows sort by code. EstimateList.to_polars() leads with a y column when its members’ variables differ, and an x column when their denominators do. Labels are on by default for both (use_labels=False drops the label columns); before, Estimate.to_polars() gave codes only and EstimateList.to_polars() put labels in place of the codes and variable labels in place of the names. A category level now keeps the type the estimator gives it, so a proportion’s level column is an integer for an integer variable. Printing is unchanged (to_polars_printable()). A label column whose name is already a variable raises LABEL_COLUMN_CLASH.

  • Serialization schema svy-result/0.4. EstimateData gains as_factor, which decides whether the table has a level column, and labels. TDistData.df is now int | float, so a GLM coefficient’s design df stays an integer. All additive: 0.3 payloads still decode.

  • Breaking: sample.singleton is removed; singletons are handled on the design. The rule is declared once with svy.Design(..., singleton=...) or sample.update_design(singleton=...) (svy.Singleton, above), so there is one place to specify it. sample.singleton.center() becomes svy.Design(..., singleton="center"), .scale()/.skip()/.pool(name=)/.collapse(using=, within=, ...) the method of that name, .certainty() "self_representing" (R’s and Stata’s “certainty” is svy’s "skip"), .handle(method, ...) svy.Singleton(method, ...), and .combine(mapping) svy.Singleton("collapse", using={singleton: target}) for strata or sample.wrangling.recode(psu_column, {new: [old, ...]}, replace=True) for PSUs. Accessing sample.singleton raises SINGLETON_API_REMOVED (also an AttributeError) with these spellings. Its inspection methods are replaced by facts on the sample, like sample.strata and sample.n_psus: detected()/show()/keys()/exists/count/last_result by sample.singletons and sample.n_singletons; the collapse helpers (candidates_for, suggest_mapping, compare, rank_strata, strata_profile, variance_contributions, summary()) and raise_error() are gone (helpers for choosing a rule may return in another form). svy.SingletonResult, svy.SingletonSummary and svy.SingletonHandling are removed; svy.SingletonMethod names the methods. A rule now follows the data: a filter that changes the singletons applies it to them, instead of leaving variance strata built for the old data.

  • Breaking: update(wgt=X) keeps the design consistent with X. sample.update_design(wgt=X) keeps the record and replicate weights when X is the current weight; on a weight an earlier design in design_history produced, and whose columns are all still in the data, restores that design’s record and replicates (switching back to the raked weight after a trim restores the raking record); on any other column drops the record silently and drops the replicate weights with a warning, unless rep_wgts= is passed, which pairs them with X. use_weight follows the same rule. A bare design.update(wgt=X) keeps the record only if X is its new_wgt and the replicates only if X is their wgt. An explicit wgt_adjustment= or rep_wgts= always wins. Edits to other fields keep both.

  • Breaking: a Design whose rep_wgts.wgt or wgt_adjustment.new_wgt is not its wgt raises.

  • Breaking: Sample(data, design), set_design and update_design raise when the data lacks any design.columns() entry, record columns included. A frame without them is not the frame the design describes; estimation used to fall back to fixed-weight standard errors with a warning.

  • Breaking: ignore_reps=True leaves the new weight without replicate weights in adjust, normalize, poststratify, standardize, rake, calibrate and calibrate_matrix. The unadjusted replicates used to stay on the design next to a weight they do not go with. Their columns stay in the data and the previous design keeps them, so update_design(wgt=<previous weight>) brings them back.

  • Breaking: wrangling no longer writes new values into a weight column under its name. The design’s weight, every earlier design’s weight, replicate weights, and the weight-adjustment record’s columns (current and earlier, __svy_cells_*/__svy_aux_* included) are refused as targets of mutate, recode(replace=True), fill_null, top_code/bottom_code/bottom_and_top_code(replace=True), categorize(replace=True) and cast, inplace or not, and any other wrangling step that would change their values is refused too. An exact widening cast (Float32 to Float64, an integer to a wider integer or to a float that holds it exactly) is allowed. A weight overwritten in place is a new variable that a later weight’s record or the replicates would still read as the old one. The WEIGHT_OVERWRITE error names each column and what reads it, and shows both ways out: write under a new name, or rename_columns first (which the design and its history follow), then write the freed name. Strata, PSUs and other design columns stay writable.

    set_data, update_data and clone(data=) apply the same rule to a new frame, compared positionally: a protected column whose values change is refused with WEIGHT_OVERWRITE, and a frame with a different number of rows is refused with DATA_ROWS_CHANGED when the sample has a weight-adjustment record, replicate weights or design history (rows changed outside svy cannot be checked against them; the hint points to filter_records, join and combine_samples). A plain design still accepts a new row count. The checks run before anything is rebound, and a later validation failure restores the sample.

  • combine_samples(kind="panel") compares the waves’ replicate designs without their paired weight and pairs them with the combined weight.

  • Estimate.q_method is None except on medians and quantiles, and so is EstimateData.q_method (now str | None); every other result carried Linear. Payloads that say Linear still decode.

  • Every estimator refuses an alpha outside (0, 1), NaN included, with INVALID_RANGE, and a non-number (a string, a bool, None) with INVALID_TYPE: glm.fit, predict, margins and contrast, estimation.mean, total, prop, ratio, median, quantile, corr and cov, categorical.tabulate, ttest and ranktest, Estimate.contrast/EstimateList.contrast, and SampleSize.estimate_prop, estimate_mean, compare_props and compare_means (each value of a per-stratum mapping). The check runs before any work. alpha=0 used to give infinite intervals, and a string or None failed deep inside.

  • Printed estimates name the estimated variable. Means, totals, medians and quantiles print a y: <var> line above where:, and ratios print y / x: <y> / <x>. With labels on, the variable label replaces the name. Proportions and as_factor means are unchanged, since their level column is already headed by the variable. An EstimateList puts what its members share in the same lines, replacing the : <var> title suffix, and a y or x column for what varies; those columns follow use_labels too. The plain-text printer of an EstimateList now shows where:, as the rich one did.

Fixed

  • glm.fit(method="replication") used the Taylor df when the replicate weights recorded none. Without RepWgts.df, the replicate fit kept the design df (#PSUs - #strata - (k - 1)), which also shrank on a where= domain. It now uses n_reps - 1 less k - 1, like estimation and the categorical tests, and does not shrink on a domain. Matches R svyglm on svrepdesign(degf = n_reps - 1).

  • The singleton error on a replicate design did not say that the replicates were not used. tabulate, ttest and ranktest compute a Taylor variance from the stratum and PSU columns even when the design has replicate weights. Estimation does the same unless method="replication" is passed. Without a singleton rule they raise SINGLETON_ERROR. On a replicate design the error now says that this variance is Taylor linearization and that the replicate weights are not used. Its hint says method="replication" uses the replicates (estimation, tabulate, ttest, ranktest and glm.fit).

  • add_stage results did not detect singletons. The combined sample was built without its design keys, so a stratum left with one PSU (after a filter, say) went unnoticed and Taylor estimates left it out silently. It is now checked like a constructed sample, and applies a singleton rule carried from stage 1.

  • A zip holding no readable file was reported as missing. read_sas on a .zip with no .sas7bdat, .xpt, .xport or .ssp member raised FILE_NOT_FOUND (“No file at x.zip”) for an archive that exists. read_sas, read_spss and read_stata (and their create_from_*) now raise ARCHIVE_MEMBER_NOT_FOUND, whose detail lists the archive’s files. Any other file a reader could not find while the data file exists is now named in the error instead of the data file.

  • ranktest(method=svy.RankScoreMethod.VANDER_WAERDEN) raised Unknown rank method. RankScoreMethod members are now accepted as they are, and "vanderWaerden" is accepted as a string.

  • A printed ratio list with several denominators did not say which row was which. ratio("a", ["b", "c"]) printed rows with no x column; it has one now, as to_polars() already did.

  • corr/cov rows did not say which pair they were. to_polars(), to_polars_printable() and the printed tables now have y and x columns after any by columns, in the order the pairs were requested; with labels on, the display shows variable labels and the data view adds y_label/x_label when a pair variable has one. svy.serialize.to_polars matches. keys() end with the pair (("api00", "api99"), ("E", "api00", "api99")) instead of repeating "api00" or the domain, and a contrast key may name the pair either way round. A by variable named y or x raises PAIR_COLUMN_CLASH from to_polars(). Contrasting a multi-row corr/cov result now says that these results carry no between-row covariance, rather than pointing at prop().

  • Nulls outside the where= domain raised. Without drop_nulls, estimation required every analysis column to be complete on every row, so a null y, by label or domain flag on rows outside the domain (people without events after a full join, skip patterns) forced drop_nulls=True. As in R’s subset(), a null in a column read only by where= now makes the row out-of-domain, and nulls in analysis columns on out-of-domain rows are ignored. Nulls inside the domain still raise, now saying so, and design columns must still be complete on every row. Out-of-domain rows keep contributing to the design, so estimates, SEs and df match R’s subset(). Covers mean, total, prop, ratio, quantile, median, corr, cov, ttest, ranktest and glm.fit(drop_nulls=False), Taylor and replication; tabulate already worked this way.

  • cov/corr with drop_nulls=True understated Taylor SEs. A missing value in the second or later column dropped the row from the design, deleting whole PSUs when all their values were missing. It now zeroes the row’s weight, as for y and R’s na.rm=TRUE.

  • Int8, Int16, UInt8 and UInt16 columns crashed the estimators. A stratum, PSU, pop_size, by, where or t-test group column of one of these dtypes (a Stata byte read with svy-io, or a .cast(pl.Int8) indicator) made every estimator fail with cannot create series from Int8, Taylor and replication alike. The response was not affected. Needs svy-rs with small-integer dtype support (see its changelog).

  • singleton.scale() on a calibrated design scaled each domain by its own strata. R keeps a calibrated design’s out-of-domain rows, whose scores are nonzero, so nstrat/nokstrat counts every stratum of the sample; svy counted the strata of the domain, and by=/where= SEs after poststratify, rake, calibrate or standardize differed from R. They now match.

  • The singleton rule reached only estimation. ttest and ranktest never received singleton.center() or singleton.scale() (the setting was read from an attribute that does not exist), tabulate ignored every rule, including the recoded strata of certainty(), collapse() and pool(), and glm.fit had no singleton handling. All of them left singleton strata out of the variance, R’s lonely.psu = "remove". They now apply the sample’s rule as estimation does, R’s "adjust" and "average" included, within where= domains too, for every GLM family. t statistics, table SEs and the Rao-Scott F, Gaussian and logistic GLM SEs, and rank tests match R.

  • singleton.center() gave the wrong variance for domain totals. A singleton stratum’s PSU total was centered at the mean of every PSU total in the sample, and a singleton stratum without domain rows still contributed. R, which estimates a domain on the subsetted design, averages over the PSUs of the strata holding domain rows and leaves the other strata out. svy now does the same for where=, by= and rows dropped for missing values, in totals, per-category totals, means, ratios, proportions, quantiles and correlations. Means, ratios and proportions rarely moved, since their scores sum to zero within the domain; totals did (by 5-7% in the test data). A calibrated design keeps the whole sample, as in R: its scores are nonzero outside the domain. Covariances between domains are unchanged, as in R’s svyby(covmat=TRUE).

  • Table levels came back in the kernel’s order and sorted as text. tabulate cells, to_polars(), rowvals/colvals, crosstab() and printed tables listed 10 before 2, and an Enum’s levels alphabetically. Levels are now ordered on the column’s values: an Enum’s in the Enum’s order, numbers numerically. Cells are listed by row level, then column level.

  • trimming= cycles stopped after one pass. poststratify, standardize, rake, calibrate and calibrate_matrix checked the trimming bounds right after trimming, where they always hold, so every cycle ended after its first pass and trimming.max_iter had no effect: a cap that needed more than one trim-and-readjust pass failed however many cycles were allowed. The bounds are now checked on the readjusted weights. Calls that converged before give the same weights.

  • An unreachable control in a trimming= cycle is TRIM_INFEASIBLE, not CONVERGENCE_FAILED / MAX_ITER_REACHED with a hint to raise max_iter. When a cycle fails and a control is outside the totals its units can reach within the bounds (a cell, margin level or calibration column; per by= domain), the error names it with its total, reachable range, unit count and, for a cell or level, the bound its units need on average. It follows on_nonconvergence: "warn" and "ignore" keep the last cycle’s weights and record TRIM_INFEASIBLE. For rake only fixed bounds (a number or Threshold.absolute) are diagnosed, since a relative bound is resolved again each cycle.

  • A GLM fit on separated data reported meaningless standard errors, silently. When the response is perfectly predicted on some rows (quasi-complete separation, e.g. a Cat level that is all 0), the estimates along that direction are infinite; IRLS stops once the deviance settles and printed the stopped values with SEs that are rounding (sometimes NaN, with a raw numpy warning and t = 0, p = 1) or small and finite with p < 1e-20, as R’s svyglm does. The fit now finds the rows it drove to the boundary and the coefficients they alone determine, for binomial (every link), Poisson and negative binomial fits, where= included. Those coefficients keep their stopped estimate (predict() is unchanged) but their SE, t, p, CI and covariance row and column are NaN, as are the model Wald tests and AIC; with replicate weights, a coefficient separated in any replicate refit is NaN too. A GLM_SEPARATION finding names the coefficients, the rows, the combination still identified (_intercept_ + quint_nat_poor), and for a Cat, the levels on which the response never varies. The identified coefficients are unchanged and match R and Stata’s svy: logit. term_test, contrasts, margins and prediction SEs are NaN exactly where they depend on an unidentified coefficient.

  • trim lost or gained weight total silently when redistribution was impossible. With redistribute=True, a bound on the wrong side of a domain’s mean positive weight (upper below it, lower above it) ended with every weight at the bound, the total changed and only the INFO audit recorded (converged=True, so on_nonconvergence never applied). It now raises TRIM_INFEASIBLE before anything is written, whatever on_nonconvergence, naming each failing domain with its bound, weight range, mean, n and total in got; the hint suggests a bound relative to the weights (Threshold.quantile(0.99), Threshold("median", 3.5)) or redistribute=False. Also applies to trim’s by=/where= domains and adjust(trimming=).

  • create_brr_wgts and create_jk_wgts(paired=True) crashed with a multi-column PSU when strata had to be paired. The replicates now equal those built from a single column holding the same composite key.

  • wrangling.distinct() without cols never removed duplicates: svy’s row index took part in the check. Only the user’s columns are compared now.

  • add_stage with a next-stage Sample whose psu is the stage-1 PSU made ssu equal to psu, and estimation then crashed. The next stage’s psu only links rows to their stage-1 PSU; the combined design’s ssu is now None.

  • clean_names renamed the selection outputs (svy_sample_weight, svy_prob_selection, svy_number_of_hits, svy_certainty and their _stage1/_stage2 forms) under camel, pascal, kebab, upper or title case; it reserved names svy never writes. They are kept now, and a user column cleaned into one of them gets the _1 suffix instead of the selection output.

  • add_stage with a Sample as the next stage carried that sample’s design keys into the combined sample, where they described the next stage’s strata rather than the combined design’s. They are left behind; svy builds its keys from the combined design.

  • trim and calibrate checked wgt_name only after doing the work. A refused trim(wgt_name=<existing column>) still raised the findings of the trim it then discarded, and calibrate with a taken name and unmet controls raised CALIBRATION_NOT_MET instead of WGT_NAME_EXISTS. Both now check the name first.

  • Taylor mean, prop and ratio returned no rows when one domain had only zero weights, for every by= level and silently; a where= to such a domain returned nothing too. The domain now keeps its row with NaN estimate, SE and CI (df 0), as R’s svymean(subset()) and as svy’s replication path already did; the other domains are unaffected. A ratio domain whose weighted denominator is zero is NaN the same way. total is unchanged (0, SE 0). Kernel errors in mean/total/prop/ratio are no longer turned into an empty result.

  • calibrate and calibrate_matrix never checked that the controls were met, except with weights_only=True: a singular system or inconsistent controls stored weights that missed them. Every path now checks X'w against the controls (relative tolerance 1e-4, absolute for a zero control; well-posed calibrations in the test suite miss by at most 7.5e-9). on_nonconvergence="error" (the default) raises CALIBRATION_NOT_MET before anything is written, with the targets in expected, the achieved totals in got (per domain, only the domains that miss, with by=) and the largest relative miss in extra; "warn" and "ignore" keep the weights and record the finding. A trimming cycle keeps its own check (CONVERGENCE_FAILED / MAX_ITER_REACHED).

  • A failed weighting call could leave an inplace=True sample half-changed. poststratify(trimming=..., strict=True) wrote the new weight and design before its trim-cycle check, so on failure the sample was modified although the error said it was not; adjust with an invalid trimming= failed after its weight and rows were written. Every weighting call now runs on a private copy and is adopted only on success, so a failure leaves the data, design and history as they were, inplace or not (diagnostics recorded before the failure are still kept with inplace=True). The poststratify and standardize trim cycles now run before anything is written.

  • trim(by=...) with nulls in by crashed on a text column and silently left the rows untrimmed on a numeric one; those rows are left untrimmed and recorded (CELLS_NULL_UNADJUSTED). The same for calibrate’s trimming.by. A poststratify whose in-scope rows all have a null cell raises NO_ROWS_IN_SCOPE with the null counts.

  • calibrate(where=..., trimming=...) lost the diagnostics recorded while calibrating the scoped rows (skipped trimming domains, non-convergence); they are now on the returned sample.

  • calibrate_matrix(by=) with nulls in by crashed with a TypeError; it raises BY_NA. poststratify whose where= matched nothing raised a kernel ValueError; it raises NO_ROWS_IN_SCOPE.

  • String keys failed on non-string columns. poststratify({"1": ...}) on an Int column raised a keys mismatch, and rake on a Boolean column failed with a bare np.False_.

  • total(y, as_factor=True) ignored as_factor and returned the total of the numeric column. It now gives one row per level with that level’s estimated count, its SE, a Wald interval, the design effect when asked and the covariance across levels and by-groups, so levels can be contrasted (R svytotal(~factor(y)), svyby(..., svytotal)). Taylor and replication.

  • Replication mean(y, as_factor=True) ignored as_factor and returned the mean of the numeric column. It now gives the per-level shares, as the Taylor path does.

  • as_factor=True failed on a string or categorical column (y was cast to float), and with drop_nulls=True a missing value became a level of its own (0.0, estimate 0). A missing value now takes the row out of the domain, matching R’s na.rm = TRUE.

  • prop(y, drop_nulls=True) dropped the rows with a missing y, so a PSU with no observed y left the design and the Taylor SE, covariance and intervals were wrong. A missing y (null, NaN or infinity) now takes the row out of the domain, matching R’s svymean(~factor(y), na.rm = TRUE). Replication estimates were already right.

  • A two-sample t-test or rank test on a numeric group ordered the groups as text, so groups 2 and 10 came out as (10, 2) and the difference was reported with its sign flipped relative to R. Numeric groups are now ordered by value.

  • Estimate rows came back in a different order on each run. With by=, the kernel returned domains in hash order, so Estimate.estimates, keys(), domains, covariance and saved payloads changed order between runs of the same code. Rows are now sorted by domain level, then by category level, in the order to_polars() lists them: numbers numerically, strings naturally ("a2" before "a10"), and levels differing only in case by their exact text ("B" before "b"). The covariance is permuted with the rows. to_polars() and printed tables use the same tiebreak, so such levels no longer interleave. An ungrouped proportion of a numeric variable now lists 10 after 9, not after 1.

  • Levels of an Enum column came back sorted as text. A by= column, or the variable of a proportion or as_factor mean, cast to pl.Enum now lists its levels in the Enum’s order, in estimates, keys(), domains, covariance, to_polars() and printed tables, also when value labels would sort otherwise. Other columns keep the sorted order. Estimate.level_orders records each Enum’s categories, and saved payloads carry them as level_orders, so svy.serialize.to_polars lists rows the same way. Schema svy-result/0.6.

  • A t-test or rank test group coded False or 0 printed an empty Level cell. The cell was built with group_level or "".

  • Saved results lost NaN and infinity, and from_json failed on them. JSON has neither, so to_json wrote each as null, which then refused to decode into a float field: a singleton domain’s interval, a CV of an estimate at 0 or a NaN quantile limit made the saved result unreadable, and a NaN deff came back as not requested. to_json still writes null, and now records the exact value under a top-level "nonfinite" field keyed by JSON Pointer ({"/estimates/3/cv": "inf"}); from_json restores it. Consumers that ignore the field see null as before. A payload written without the field reads a null plain float as NaN. Schema svy-result/0.4.

  • Singleton handling was not part of a saved design. A design saved and restored, or a Sample rebuilt from sample.design, lost it and gave different SEs.

  • Singletons were detected only when a sample was built or its data or design replaced. A filter_records, recode or mutate that left a stratum with one PSU went unnoticed, and estimation ran without the singleton error a freshly built sample raises. Singletons are now detected again after every change to the rows or the stratum/PSU/SSU columns.

  • Singleton variance strata went stale after data changes. They were built once, when the rule was applied: after collapse(using={"d": "a"}), recoding the strata b into c left the SE at its value before the recode. They are now rebuilt from the current data.

  • After update_design changed the strata or PSUs, singleton detection read the old ones (singleton detection and the estimation check), because the internal stratum/PSU key columns were kept rather than rebuilt.

  • Design.__eq__ and __hash__ ignored wgt_adjustment, so a raked design equalled the same design without its record.

  • Wrangling did not know the weight-adjustment record. remove_columns, keep_columns and select now refuse to drop record columns (including __svy_cells_* and __svy_aux_* snapshots) without force=True. force=True also cleans what depended on the column: the record, the design’s wgt and the replicate weights that went with it (their columns stay), PopSize columns, and stale internal stratum/PSU columns; one warning lists what was removed. rename_columns and clean_names carry renames into the record, PopSize, rep_wgts.wgt and earlier designs in design_history, and an inplace rename_columns or clean_names that cannot be applied leaves the sample untouched.

  • Replicate weight columns were found by pattern, not by the spec. Estimation took every column matching ^prefix\d+$ (any case) as a replicate, so a look-alike such as w2023, w21 or W1 next to w1..w20 failed with REP_WEIGHT_COUNT_MISMATCH or, on one path, entered the variance. rename_columns treated renaming such a column as a replicate rename, and a look-alike w01 next to unpadded w1..w20 switched the detected padding. Replicate columns are now the spec’s own prefix + 1..n_reps, with the padding and casing under which all of them are present.

  • A failed set_design or update_design left the sample half-updated, with the rejected design installed and recorded in design_history. The sample is now left as it was.

  • median() and quantile() always reported q_method as Linear. The values used the requested rule, but Estimate.q_method and the serialized q_method field did not. They now give the rule used, Taylor and replication, and the printed header shows it (MEDIAN (TAYLOR, q_method=higher)).

  • GLM predictions and margins truncated their confidence level: alpha=0.001 printed 99% CI. It now prints 99.9% CI.

  • GLMFit.to_dict() raised TypeError on every fitted model: the coefficients and statistics are numpy scalars, which msgspec does not encode. It now returns plain Python values, as do GLMCoef.to_dict() and GLMStats.to_dict().

0.30.0 — 2026-09-25

Added

  • wrangling.join(other, on, cols=None, into=None, suffix=None) brings variables from another Sample or frame onto this sample’s records: the household weight onto the person file, frame variables onto the sample, response status onto the selected units. A left join that keeps every record in order and leaves the design alone; other’s weight arrives as an ordinary column. on maps key names when they differ ({"hh": "hh_id"}). cols picks what comes in (default: every non-key column), into names any of it here as recode(into=) does (into={"wgt": "hh_wgt"}), and suffix is appended to the remaining names that already exist. Nothing on this sample is renamed or overwritten: a name that still clashes is refused, design variables included, as are a key that repeats on other (validate="1:1" checks this side too), key types that do not match (integer widths, text with categorical, and a decimal identifier with an integer one when every value is whole, as SPSS and Stata store IDs, are joined), names svy reserves, and names that would read as this sample’s replicate weights. Variable and value labels carry over from a Sample. Unmatched records get nulls and a JOIN_UNMATCHED warning with the count (on_unmatched="ignore"|"warn"|"error", as on_singletons); indicator= adds a column marking the matched records. Inner, semi and anti joins are indicator= followed by filter_records, which checks singletons. selection.add_stage stays the join for chaining a second stage, which rebuilds the design.

0.29.0 — 2026-09-13

Fixed

  • Taylor quantile and median intervals were centered at p. The Woodruff interval inverted the weighted CDF at p ± t·se_p, as R’s oldsvyquantile(interval.type="Wald") does. It is now centered at the estimated CDF at the quantile, F(q̂) ± t·se_p, as in R’s svyquantile. On a step CDF the two differ, so limits and SEs changed wherever F(q̂) ≠ p, most on small or tied data. q_method="higher" matches qrule="math" and "linear" matches qrule="hf4". A limit whose probability falls outside [0, 1] is now NaN, and so is the SE, instead of being clamped to the sample minimum or maximum. Replicate-weight quantiles are unchanged.

  • Proportion CIs under where= used the whole-sample n. beta, korn-graubard and wilson read the domain sample size in the t-adjustment (t(n−1)/t(df))², in Korn–Graubard’s cap min(n, n_eff*), and as the effective sample size at p = 0 or 1. Under where= (and where= with by=) that count included every out-of-domain row, so a where= domain got a narrower interval than the same domain through by=: a few decimals for beta, but a Korn–Graubard bound at p = 0 hundreds of times too narrow on a small domain of a large survey. n is now the number of rows in the domain with a nonzero weight, for every design and for Taylor and replication alike. Zero-weight rows represent no population units and are not counted, which differs from R’s nrow() when a file carries them; R’s svyciprop(method="beta") also uses the whole sample for calibrated and PPS designs, where svy does not.

  • singleton.skip() dropped the singleton strata from the point estimate. The rows were filtered out before estimation, so mean, total, ratio, prop, quantiles, covariances and regressions were computed on the remaining strata alone; only the SE of a total came out as R’s lonely.psu="remove". All rows now stay in the estimator and the one-PSU strata contribute nothing to the variance, matching R for every estimator. tabulate was unaffected. The docstring, which promised that rows were removed, now describes the variance recipe.

  • singleton.scale() did not match R’s lonely.psu="average". SEs for means, ratios, proportions and covariances were computed after dropping the singleton strata’s rows, which changes the estimator, and then multiplied by 1 − f instead of 1/(1 − f); the two errors cancelled only when the singleton strata held exactly a share f of the weight. The separate full-sample pass for the point estimate read the raw data frame: an unrelated Int8/Int16/UInt8/UInt16 column raised ComputeError, by= and ratio() with an integer denominator raised, and where= silently returned the whole-sample estimate. There is now a single pass: all rows stay in the estimator, singleton strata contribute no variance, and every Taylor variance, covariance and design effect is multiplied by 1/(1 − f). Totals were already right.

  • singleton.scale() inflated every domain by the whole-design singleton fraction. R’s lonely.psu="average" counts nstrat/nokstrat over the strata present in the rows being estimated, and subset() and svyby() drop rows, so a by= level or where= domain that leaves whole strata out gets its own fraction. svy applied the full-design 1/(1 − f) to every estimate: a domain holding none of the singleton strata was still inflated, and one holding a larger share of them was under-inflated. The factor is now counted per estimate over the strata with in-domain rows, a domain with no singleton stratum is left alone, and a domain lying entirely in singleton strata reports NaN, as in R. Full-sample estimates and domains that touch every stratum are unchanged.

Changed

  • logit, beta and wilson return NaN bounds at an estimated proportion of 0 or 1, instead of a zero-width [p, p] that read as an interval known with certainty. A PROP_CI_BOUNDARY warning names the affected cells and points to ci_method="korn-graubard", which keeps its one-sided interval there. With no residual degrees of freedom every method, korn-graubard included, now returns NaN (as R does); beta and korn-graubard used to skip the t-adjustment and return an interval. A zero SE at 0 < p < 1 still gives [p, p]. The comparisons use a 1e-12 tolerance, so floating-point noise in an SE or a p-hat no longer decides which case applies.

  • Degrees of freedom count PSUs with a nonzero weight, not a positive one, so a PSU carrying only negative calibrated weights is counted (R’s degf).

0.28.0 — 2026-09-08

Added

  • Negative binomial, for counts a Poisson cannot hold. family="negative_binomial" (or "nb"), with Var(mu) = mu + mu^2/theta and R’s okLinks — log, identity, sqrt.

    m = sample.glm.fit(y="visits", x=["age", svy.Cat("region")], family="nb")
    m.fitted.stats.theta, m.fitted.stats.theta_se

    Left to itself, theta is estimated by maximum likelihood alongside the coefficients and the standard errors come from the joint (coefficients, theta) design-based sandwich — survey::svymle over the negative binomial likelihood, the method in Lumley’s Complex Surveys (Appendix E) and what sjstats::svyglm.nb implements. stats.theta_se is the design-based standard error of the dispersion itself.

    Pass theta= and it is treated as known: no row for it, and the variance conditions on it, matching svyglm(family = MASS::negative.binomial(theta)). These are different numbers, not two routes to one — on apistrat the joint standard errors run from 25% under to 13% over the conditional ones, and on Lumley’s own NHANES example 2% under to 12% over. The orthogonality that makes them agree asymptotically is a property of the model, and a design-based sandwich is the thing that declines to assume it.

    On a replicate-weight design an estimated theta raises: the spread of the refits would only mean something if every replicate re-estimated it. Pass theta= there.

    All of it is in the kernel, like every other family — the dispersion loop and the joint variance included. That needed digamma, trigamma and lgamma written out in svy-rs, for the same reason norm_cdf already was.

  • cauchit and sqrt links. Binomial gains cauchit — the inverse Cauchy CDF, the heavy-tailed alternative to probit, so a few observations far out on the linear predictor cannot dominate the fit — and Poisson gains sqrt, the variance-stabilising link for counts. FAMILY_LINKS is now exactly R’s okLinks for every family, with no links R has that svy does not. Both refuse exponentiate=: exp(β) is not a ratio on either.

    sample.glm.fit(y="voted", x=["age", svy.Cat("region")], family="binomial", link="cauchit")
  • read_stata(encoding=) is forwarded to the reader instead of being dropped with a warning. A Stata 13 file holding UTF-8 text (the Nigeria GHS-Panel Wave 5 free-text sections) now reads with encoding="utf-8", or encoding="utf8-lossy" to replace undecodable bytes, an option read_spss and read_sas accept as well; the parse IoError for that failure carries the engine’s hint naming the option.

  • Panel surveys, with no new type. A panel is a long Sample whose Design.case_id identifies the followed case and whose Design.wave orders its rows. svy.combine_samples(kind="panel", case_id=...) stacks the waves and validates the pairing (unique id within each wave, consecutive-wave overlap, design columns constant within a case, later waves’ units a subset of wave 1’s); Design(case_id=..., wave=...) declares the same on a long file. When no PSU is declared the case is the variance PSU, so mean(y, by="wave") and its contrasts carry the between-wave covariance without being told to — the change SE equals the wide-frame individual-change SE, not the naive independent-waves one. Producer longitudinal weights stay ordinary columns selected with use_weight(); identical producer replicate weights are accepted across waves.

    s = svy.combine_samples([w1, w2, w3], kind="panel", case_id="person_id")
    s.estimation.mean("inc", by="wave").contrast(svy.estd(3) / svy.estd(1) - 1)  # percent change
  • wrangling.lag(cols, n=1) — the one panel primitive: the value of a column at the case’s wave n steps back (a lead if negative), null across a skipped wave as Stata’s L.y (gaps="skip" takes the previous observed row). Transitions are tabulate("y_lag1", "y", where=wave == t) or prop(y, by="y_lag1", where=wave == t, drop_nulls=True); paired change is the existing ttest(y, y_pair="y_lag1", where=..., drop_nulls=True).

  • weighting.adjust() on a panel applies the nonresponse factor to the case: every row of the case gets its own weight times the case’s factor (so a case-level base weight stays case-level), a nonrespondent case gets 0 on all its rows, and a case with earlier rows but none in scope is a nonrespondent — attriters need no row to be adjusted for. respondents_only drops nonrespondents at the scope waves only. Chain one call per wave with where=col("wave") == t to build longitudinal weights.

  • categorical.tabulate(where=) — a subpopulation table with R’s subset() semantics: rows outside the domain keep their design columns and get weight 0, so the PSU structure and the design df are those of the full design (matching svymean(~interaction(...), subset(d, ...)), svytotal and svychisq to 1e-9). Cells are formed from the domain’s categories only, and a null key outside the domain is not missing data.

  • Delta-method contrasts. svy.estd() expressions now accept *, /, constants, .log() and .exp(): ratios, percent change and log-ratios are estimated by the delta method on the same design df, matching R svycontrast to 1e-11. Linear expressions keep the exact L V Lᵀ path; the {key: coef} dict form stays linear.

  • panel_syn_2026 — a bundled synthetic three-wave panel (1200 cases, ~15% attrition per wave, producer longitudinal weights, a binary and a continuous outcome) for the panel tutorial and tests.

Fixed

  • svy.col(...).explode() will not change behaviour under polars 2.0. polars 1.44 deprecates the default of its empty_as_null, and 2.0 flips it: an empty list would stop yielding a null row and yield no row at all. A row that disappears on a dependency upgrade is one fewer observation in whatever is estimated downstream, so svy passes the argument explicitly and keeps today’s behaviour. explode(empty_as_null=False) selects the other one.

    This raises the polars floor to 1.36.1, where empty_as_null was added (1.35.1 does not have it). The alternative was a runtime version check, and svy has none anywhere else.

    On polars ≥ 1.44 the deprecation warning is also gone; because the test suite escalates DeprecationWarning to an error, test_explode had been failing there and passing locally only on the lockfile’s pinned 1.39.3.

  • IRLS started from the wrong place for binomial and Poisson. The starting mean is R’s family$initialize — (w·y + ½)/(w + 1) for binomial and y + 0.1 for Poisson — where svy used (y + ½)/2 and max(y, 1e-10), ignoring the weight. IRLS stops on a relative change in the deviance, so on a flat surface where it starts decides where it stops: cloglog on apistrat landed 3e-5 from R’s coefficients, with the deviances agreeing to 14 digits. svy now follows R’s iterate sequence, ending on the same iteration.

  • cauchit’s inverse link is computed the way R’s pcauchy is. arctan(η)/π + ½ cancels in the tails — three and a half digits gone at η = −1000 — so outside [−1, 1] the tail comes from arctan(1/η). The forward link is −1/tan(πμ) rather than tan(π(μ − ½)), which is catastrophic as μ approaches 0 or 1.

  • glm.fit() reported failures as results. An empty where= domain, or an all-zero weight column, returned a fit of all-zero coefficients (or, for gamma and poisson, a TypeError from the response check); x and 2 * x “fitted” with SE 0 and an F statistic around 1e30; more parameters than observations came back through a pseudoinverse. Each of these now raises ModelError naming what is wrong — the collinear term by name, in the model’s own column order.

  • A GLM that ran out of iterations was reported like a converged one. fit() now warns, giving the iteration count and the tolerance, and so does dropping rows for an invalid weight, dropping a Cat with fewer than two levels among the fitted rows, and — as a ModelError rather than a polars “duplicate output name” — listing the same predictor twice.

  • where= fitted the complement domain too and threw it away. The predicate now reaches the kernel as a Boolean mask instead of a "true"/"false" by-column, so only the requested domain is fitted (1e6 rows x 20 covariates: 1.22 s → 0.45 s). The estimates are unchanged.

  • glm.predict() and glm.margins() were wrong on an int, float or bool-coded Cat. Both rebuilt a dummy column by slicing the level out of the coefficient name and comparing it to the raw column as a string, which numpy answers all-False without raising: every row got the reference-level prediction (off by up to 0.26 in probability on the test model) and every categorical average marginal effect came back exactly 0.0 with SE 0.0. The fitted coefficients were correct throughout, which is what kept it quiet. The design matrix is now rebuilt from the level values in their fitted dtype, through the same polars expression the fit used — and, as a side effect, in one pass instead of a Python loop over object arrays (predict on 1e6 rows: 0.72 s → 0.12 s; a 10-level categorical AME: 5.9 s → 0.8 s).

  • Prediction data is validated. A missing predictor column raises ModelError naming it (was a raw polars ColumnNotFoundError or a bare KeyError), and a categorical value the fit never saw — including a null — raises instead of being silently coded as the reference level, in predict() and in margins(at=) alike.

  • categorical.tabulate() ignored Design.pop_size: every table on a design with a finite-population correction reported the fpc-free standard errors (the ttest facade had the same gap). The fpc is now applied, matching R svymean(~interaction(...)) and svychisq on an fpc= design.

Changed

  • stats.scale is the Pearson dispersion for every family, Σ wᵢ(yᵢ−μᵢ)²/V(μᵢ) / (n_obs − k) on the fit’s own weight scale — R glm()’s and Stata glm’s “(1/df) Pearson”. Two deliberate departures. Binomial and Poisson used to report a hard-coded 1.0; the estimate is the overdispersion diagnostic a count model needs, and design-based SEs never use it, so nothing else moves. And R summary.svyglm reports svyvar(resid(pearson)) instead, a design-weighted variance of the Pearson residuals — not the textbook estimator, and not what someone comparing against glm() or Stata expects — so that is not copied.

  • The dispersion’s divisor is n_obs − k, not the design residual df: the dispersion is a moment estimate, while the design df belongs to the t reference distribution. The same model on the same rows used to report a different scale for every design.

  • combine_samples: kind="cross_sectional" | "cs" | "panel" replaces units="independent" | "shared"; default wave labels are "wave 1".."wave k" instead of "s1".."sk"; the combined weight column wgt_name is always created from each wave’s own weight (divided by k under adjust="average"), so the waves may name their weights differently. A panel keeps the base-wave design and now accepts a later wave that lost PSUs (with a warning) instead of requiring identical design units.

  • Design.describe() shows Case id and Wave rows, and PSU None (variance: <case_id>) when the case is the fallback PSU; on a panel the design summary lists the case overlap between consecutive waves.

Removed

  • Design.row_index. Design.case_id is the record identifier on cross-sections and panels alike (non-null, unique; unique within wave on a panel). Design(row_index=...) now raises TypeError. The hidden svy_row_index column is unchanged plumbing.
  • combine_samples(units=...), replaced by kind= without a shim (pre-stable).

0.27.0 — 2026-08-31

Added

  • svy.combine_samples() — repeated cross-sections and panel waves. Stacks several samples and analyzes them as one stratified design: data pooling, not estimate pooling. Each independent wave contributes its own strata (wave → stratum → PSU), so Taylor variance treats waves as independent without being told to. Caller order is time order, and an existing wave column (NHANES SDDSRVYR, say) is reused and validated rather than replaced.

    svy.combine_samples([w1, w2, w3], adjust="average")  # period-average population

    adjust="average" divides the weights by k, which matters only for totals — means, proportions and ratios are invariant to it. units="shared" is the panel case, where the same units recur across waves.

  • Linear contrasts over estimates, via svy.estd(). Any Estimate or GLMFit can be contrasted against its own estimands using the between-estimate covariance:

    m = sample.estimation.mean("api00", by="stype")
    m.contrast(svy.estd("E") - svy.estd("H"))
    fit.contrast({"H vs E": svy.estd("stype_H") - svy.estd("stype_E")})

    Accepts a contrast expression, a sparse {key: coefficient} dict, or several named contrasts at once. Keys are the row identities keys() lists, and metadata value labels are accepted wherever unambiguous. Inference is t-based on the full design df (R’s degf convention), not the per-row domain-aware df.

    Quantiles and medians raise rather than returning a number: Woodruff intervals do not arise from a linearized score column, so no between-quantile covariance exists to combine.

  • Marginal effects for categorical predictors. margins() now returns a discrete contrast per non-reference level — mean_w[mu(x, var=k) - mu(x, var=ref)] with a delta-method SE — instead of skipping the term. Matches R marginaleffects’ factor contrasts; on apistrat the stype contrasts agree to 1e-9 and the SEs to 2e-10.

    This fixes a wrong-shaped answer, not just a missing feature: margins(variables=["stype"]) previously returned an empty list, and the default margins() dropped categorical terms with only a log.warning. A variable with no derivative w.r.t. any fitted coefficient now raises instead of being silently skipped.

  • exponentiate= on GLMFit.to_polars() and .show(). Reports exp(β) with exp of the link-scale interval — an odds ratio for logit, a rate ratio for log, a hazard ratio for cloglog — and names the column for the link. It refuses the links where exp(β) is not a ratio (identity, probit, inverse, inverse_squared) rather than printing a meaningless number.

    The interval is the exponentiated link-scale interval, so it is not symmetric about the ratio, and std_err/statistic/p-value stay on the link scale where the Wald test is computed — a symmetric standard error around a ratio is the mistake this is meant to prevent. Matches exp(coef(f)) and exp(confint(f)) from R.

  • offset= on glm.fit(). A known term on the link scale, coefficient fixed at 1 — what makes a rate model possible:

    s.glm.fit(
        y="events", x=["age", svy.Cat("region")], family="poisson", offset="log_exposure"
    )  # coefficients are log rate ratios

    Carried through the whole model: IRLS working response, deviance, the sandwich, predict() and margins(). Matches R survey 4.5 to 2.5e-15 on coefficients and 5e-10 on SEs. Null deviance is the intercept-only fit carrying the offset, found by a one-parameter IRLS the way R’s glm() refits it — the weighted-mean shortcut is only valid without an offset — and reproduces R’s null.deviance exactly.

    predict() includes the offset, where R’s predict.svyglm(newdata=) drops it. R’s own two answers disagree: on the same rows fitted(f) returns exp(Xb + offset) and tracks the observed counts, while predict(f, newdata=) returns exp(Xb). svy follows fitted(), which is also what Stata’s predict after glm, exposure() does. Passing newdata without the offset column raises rather than silently predicting a rate.

  • Inverse Gaussian family for glm.fit(). family="inverse_gaussian" (canonical link inverse_squared, R’s 1/mu^2) completes the exponential-family set. The kernel already implemented it and DistFamily.INVERSE_GAUSSIAN already named it — only the Python family map withheld it. Matches R survey 4.5 to twelve significant figures on coefficients, SEs and deviance across weighted, clustered, stratified and stratified-clustered designs.

  • DistFamily is exported. svy.DistFamily and svy.core.DistFamily now resolve, matching LinkFunction, which was already exported. family= and link= both accept a plain string or their enum.

  • Probit and cloglog links for glm.fit(). family="binomial" now accepts link="probit" and link="cloglog" alongside logit, completing the binomial link set. Both are wired through the whole model: coefficients and sandwich SEs, predict() on the response scale, and margins() — average marginal effects and predictive margins, whose delta-method SEs need the link’s second derivative.

    s.glm.fit(y="insured", x=["age", svy.Cat("region")], family="binomial", link="probit")

    Verified against R survey 4.5 across weighted, clustered, stratified and stratified-clustered designs — coefficients and SEs to 1e-7 relative, deviance to 1e-14 — and against marginaleffects 0.32.0 for AME and predictive margins. probit and cloglog map onto (0, 1), so pairing either with a non-binomial family now raises instead of fitting something meaningless.

  • Sample.weighting.standardize() — direct standardization. Reweights each domain to a common composition so rates are comparable across domains that differ in it. Age is the usual axis, but any composition variable works.

    s.weighting.standardize(
        cells="agecat",                          # composition axis
        shares={"(0,19]": 55901, "(19,39]": 77670, ...},   # counts or proportions
        by=["race", "riagendr"],                 # domains
        where=svy.col("hi_chol").is_not_null(),  # scope
    )

    target(g, c) = share(c) × Ŵ_g, so domain totals are preserved and only the within-domain composition is reshaped. Ŵ_g is computed under the same scope that built the cells, which removes the usual footgun — a domain total estimated over a different set of rows than the one being standardized. Domains missing a level are renormalized over the levels they have, with a warning.

    It is a third name, not a third machinery: it runs on the same engine as poststratify. standardize(by=None) and poststratify(shares=…) are deliberately the same operation, each named for the community that uses it.

    Verified against R survey::svystandardize 4.5 — weights to a relative tolerance of 1e-12, and the NHANES HI_CHOL by race × sex example to 4.8e-14 across all eight domains.

    Standardized weights are analysis-specific. where bakes one variable’s missingness into the weights and by bakes in the domain structure, so estimating a different variable or a different breakdown on the same standardized sample is silently wrong.

  • Calibration-aware Taylor variance. A weighting method that pins population quantities now records what it pinned, and the Taylor path centres the linearized scores against that structure — R survey’s svyrecvar postStrata loop. Covers mean, total, ratio, proportions, quantiles, tabulation, glm, and one- and two-sample t-tests, grouped and ungrouped, under poststratification, raking, GREG calibration and standardization.

    Verified against survey 4.5 on apiclus1 to twelve significant figures or better, including the cases where the two implementations are easy to get subtly wrong: raking centres on the unweighted margin mean, GREG fits its residual with the previous weights, and cell means are always whole-sample — a subpopulation filter does not restrict them, because the adjustment was not performed inside the subpopulation.

  • where= on all seven weighting methods. Scope, not domain: matching rows receive the adjustment, and the rest keep their previous weight, so the new column is complete and design.wgt can repoint to it. This is the same keyword and the same reading as estimation’s where=, with a different consequence — estimation zero-weights for subpopulation variance.

Changed

  • Requires svy-rs >= 0.16.0. The GLM offset and the probit/cloglog links call into kernel entry points 0.15.0 does not have, so the pin moves to >=0.16.0,<0.17.0.

  • BREAKING: a family now admits only the links it has a model for. glm.fit() validates the family/link pairing against the okLinks set of R’s family constructors, so pairings that were never a model are rejected up front instead of failing deep in the kernel — or, in the case of binomial + inverse_squared, converging on a meaningless fit and reporting it as a result. Newly rejected: logit/probit/cloglog outside binomial, inverse/inverse_squared outside the families whose canonical link they are. Every previously working model still works; what stops working is the combinations that only appeared to.

  • BREAKING: by= is now cells= on adjust, normalize and poststratify. cells names the groups that each receive one derived adjustment factor; by continues to mean “repeat independently per group” everywhere it survives (calibrate, trim, and standardize). Keeping one word for both would have compromised a term of art.

  • BREAKING: factors= is replaced by shares=. In poststratify this is close to a rename — factors already computed f × grand_total, which is shares, unnormalized. What changes is that shares are normalized first, so a vector that does not sum to 1 now pins composition instead of silently rescaling the population. In rake it is a genuine removal: factors there multiplied each category by its own current weight sum, which has no share reading and which no workflow can supply in advance, since raking is iterative. rake gains shares= (marginal proportions) in its place.

  • BREAKING: a scalar controls now requires cells=None. The rule across every level/share method is that a scalar names one cell and a dict names many; controls sets the total and shares preserves it. Previously poststratify(300, cells="region") silently ignored cells and rescaled to the grand total — verified bit-identical to normalize(300) with no groups. poststratify(331_000_000) with no cells is unaffected and remains the way to pin a known population size.

    Renamed and removed parameters raise an error naming the replacement rather than reporting an unknown keyword.

  • Multi-column cells bind by tuple key, and key-mismatch errors now echo the keys as written — ('R2', 'D1') rather than the internal 'R2_&_D1' encoding.

  • rake now rejects margins that disagree on a population total. Margins were validated only in isolation, so inconsistent totals — the classic reason IPF fails to converge — simply ran to max_iter and returned whatever they reached. With shares= the consistency is structural.

Fixed

  • The missing-values error named a parameter that does not exist. assert_no_missing told users to set drop_missing=True, which is the internal helper’s argument, not a public one — following the advice raised TypeError. All four call sites gate on drop_nulls, which is what the message now says.

  • print() on an EstimateList of proportions. prop() over a sequence of variables built one level column per variable, so a diagonal concat unioned them into a staircase of mostly-empty columns — eight conditions gave eight sparse columns, a ~118-character box, and every number truncated to an ellipsis. The members now stack under one shared level column beside y, which for the eight-condition case brings the table to 71 characters at full precision.

    Only the stacked case changes. A single prop() still heads its level column with the variable name, mean() over a list is untouched, quantile() keeps its prob column, and a by= column stays distinct from the level column — including when a variable is itself named level, where the shared column moves aside.

  • categorical.ttest ignored the finite population correction. The t-test built its variance without ever asking for an FPC column, so on a design with pop_size it returned the answer for sampling with replacement — standard errors too wide, t too small. One- and two-sample tests were both affected; on apiclus1 the statistic was off by a factor of 1/sqrt(1 − 15/757). It now matches R svyttest to fourteen significant figures.

  • One target-resolution path. normalize and poststratify derived their control arrays independently, with two different label-to-code conventions. Both now resolve every argument form to absolute per-cell targets in one place, which is also the single form crossing into Rust.

Note on variance after a calibrating adjustment. poststratify, rake, calibrate and standardize pin quantities in the population, which removes sampling variability. Both variance paths now account for this: replication re-adjusts each replicate column, and Taylor centres the linearized scores within the pinned cells. Previously only replication did, so Taylor standard errors on a calibrated design answered a slightly different question — on the NHANES standardization example they differed from R by roughly −11% to +14%. Point estimates were never affected.

The record is honoured only while it still describes the data. Estimating on a different weight, or dropping a column the adjustment referenced, falls back to treating weights as fixed — and says so, once per data/design rebind.

Removed

  • BREAKING: ModelType and FitMethod are no longer exported. Both arrived with the monorepo migration and were never wired up — no consumer in svy, its tests or its docs, and neither enum’s values were matched anywhere. ModelType was also misleading: it named four family+link combinations where glm.fit() takes the two separately and now accepts eleven. Use DistFamily and LinkFunction. The definitions are commented out in core/enumerations.py rather than deleted, alongside the other parked enums there.

0.26.0 — 2026-08-28

A breaking change to the weighting API, and a substantial cut to import time.

Changed

  • BREAKING: Sample.weighting no longer mutates the sample you call it on. All eleven transforming methods — the four create_*_wgts, adjust, normalize, poststratify, rake, calibrate, calibrate_matrix, trim — now return a new Sample and take inplace: bool = False to ask for the old behaviour. Both branches return a Sample, so chaining is unchanged and s = s.weighting.…() keeps working exactly as before.

    This is not a new convention: all 19 wrangling methods already take inplace: bool = False and default to copying, and singleton’s five transforms always return a new Sample. weighting was the only namespace that mutated unconditionally, and the only one with no way to ask for a copy — grep -rn inplace src/svy/ hit nothing outside wrangling/.

    The reassignment idiom hid it, so it surfaced only when two variants branched off one sample — which is what a method comparison, a sensitivity check, or a docs page does:

    boot = base.weighting.create_bs_wgts(n_reps=10, rep_prefix="bs")
    jack = base.weighting.create_jk_wgts(rep_prefix="jk")
    # before: boot is jack is base. 33 jk* columns in the frame, and a design
    #         still reading method=Bootstrap prefix='bs' -- so an estimate off
    #         `jack` silently used the bootstrap columns and coefficients.

    What breaks: code that called a weighting method and then read the receiver without capturing the return. Add inplace=True, or capture the result. Code already written as s = s.weighting.…() needs no change.

    The isolation covers diagnostics as well as data: an operation that raises under the default leaves the caller’s warning store untouched, because “this call had no effect on my sample” is only true if it covers everything. The error is still raised and still emitted to the log; inplace=True is how you ask for the warning to land on your sample.

  • import svy was paying for scipy before you asked it anything. Importing the package cost ~0.80 s, and 383 ms of that was scipy.stats, pulled in along the chain svy.serialize → svy.categorical → svy.estimation.base. Nothing about that import was needed to reach svy — it was needed only by whichever call eventually wanted a quantile.

    Six modules imported scipy at module scope: estimation/base.py, regression/base.py, regression/margins.py, engine/size_and_power/size.py, engine/size_and_power/power.py, and utils/hadamard.py. What made this stubborn is that fixing any one of them alone changes nothing — scipy loads once and the other five ride sys.modules, so the cost simply moves. All six had to go together.

    The uses turned out to be narrow: estimation, regression and margins want the t and F quantiles, size and power the normal, hadamard a single scipy.linalg call. Each import now sits in the function that needs it, matching what categorical/base.py had already been doing in four places — this finishes a refactor that had stalled partway.

    import svy now costs ~0.27 s and scipy.stats is off the import path entirely; it loads on the first call that actually needs a distribution. Nothing about the numerical results changes.

  • svy.datasets now loads on first use rather than on import. The root package imported it purely so it would be reachable as an attribute, which is what a module-level __getattr__ handles. Worth ~30 ms — not the ~180 ms an import profiler credits to it, since most of that is polars and numpy, which datasets merely imported first and which the rest of svy needs anyway. svy.datasets.load(...), from svy import datasets and import svy.datasets all behave exactly as before.

Fixed

  • BRR validated its strata in two passes, and sent one of them to the wrong subsystem. create_brr_wgts checked < 2 PSUs first, raising INSUFFICIENT_PSU, and only then checked for odd counts. A frame carrying both therefore reported the lone-PSU stratum, and revealed the odd ones on the next run — one class of problem per attempt, on a validation that could have been answered in full the first time.

    The split also borrowed a question that is not BRR’s. A stratum with one PSU is a singleton, which matters to Taylor linearization and has its own subsystem — Sample.singleton, with collapse, combine, pool, certainty — none of which the hint mentioned; it said “Combine small strata or check design specification.” For BRR the count is not a singleton question at all: a stratum of 1 and a stratum of 3 fail for exactly the same reason, a PSU with no partner.

    BRR now asks one question — is the PSU count a multiple of 2 — and reports every stratum that fails it in a single error, naming each with its count:

    ❌ BRR needs 2 PSUs per variance stratum [ODD_PSU_COUNT]
    BRR pairs PSUs into variance strata of exactly 2. Found 3 strata whose PSU
    count is not a multiple of 2, leaving a PSU with no partner.
    - got: a=1, b=3, c=5

    “Multiple of 2” rather than “equal to 2” because create_brr_wgts pairs: a 4-PSU stratum becomes two variance strata of 2 and is valid. INSUFFICIENT_PSU is now the paired jackknife’s alone, which is the one method that genuinely distinguishes the two — it absorbs an odd count into a triplet but still cannot pair a lone PSU.

  • The odd-PSU BRR error pointed at an API that no longer exists. Its hint read Use method='jk2' which allows 2-3 PSUs per stratum. method='jk2' was create_variance_strata’s parameter — a function removed in 0.25.0 — and method is not a parameter of create_brr_wgts either, so the one line telling you what to do instead named a parameter of a gone function on a function that never had it. “2-3 PSUs per stratum” was equally stale: create_jk_wgts(paired=True) pairs any count ≥ 2 and absorbs an odd one into a triplet.

    The hint now names create_jk_wgts(paired=True), says why BRR cannot do the same, and tells anyone who needs BRR specifically to fix every stratum listed rather than one. Its test asserted pytest.raises(Exception) and nothing more, which is how the text drifted unnoticed; it now pins the code, the where, and that the remedy names a real API.

  • Regenerating replicate weights left the previous method’s design in place. create_brr_wgts and create_jk_wgts recorded their result with Design.fill_missing(rep_wgts=…), which by definition only fills a field that is currently None — so once any replicate design existed, the record of what had just been built was silently dropped. create_bs_wgts and create_sdr_wgts used Design.update and were unaffected, which is the whole of the difference: it was an inconsistency inherited from the monorepo migration, never a decision.

    Building a second set of replicates therefore wrote the new columns and kept the first method’s metadata — and estimation reads the metadata, not the columns:

    s = s.weighting.create_bs_wgts(n_reps=10, rep_prefix="bs")
    s = s.weighting.create_jk_wgts(rep_prefix="jk")
    # before: 33 jk* columns in the frame, design.rep_wgts reading
    #         method=Bootstrap prefix='bs' n_reps=10.
    #         mean(method="replication") -> Estimate: MEAN (BOOTSTRAP), off `bs*`.

    Unlike the mutation defect fixed alongside it, this one fired on the plain s = s.weighting.…() idiom, in both orders, and switching methods mid-analysis is exactly when someone does it — comparing a jackknife against a bootstrap, or moving to JK2 after seeing the replicate count. Both now use update: the design describes the columns that were just written, and a declaration that no longer matches them is replaced rather than preserved.

0.25.0 — 2026-08-26

Added

  • Replicate weights carry the units they were built from. stratum and psu on every variant name the columns the replicates were drawn over — a separate question from Design.stratum/Design.psu, which describe the analysis design and are what Taylor linearizes over. Producers collapse strata and suppress PSUs for disclosure, publishing a distinct pair (VARSTRAT/VARUNIT and its many spellings) alongside — or instead of — the design variables:

    JackknifeWgts(prefix="rw", n_reps=8, kind="jkn", stratum="VARSTRAT", psu="VARUNIT")

    They appear in repr and print(design), are validated as columns at Sample construction, and are rewritten by wrangling.rename_columns and protected by remove_columns exactly like design columns. Generators record whatever they used, so a generated design states its provenance rather than leaving a later reader to assume it matches the Design. The Poisson bootstrap records neither, drawing independent per-unit factors and having no units by construction.

    They take the same shapes Design’s do — str for the single collapsed identifier producers usually ship, or a tuple when the unit is several columns together — and the same spellings mean the same thing on both objects, including ("region",) staying a one-tuple rather than being unwrapped. Multi-column units are grouped on directly rather than through the internal concatenated column Design builds, so nothing implementation-specific reaches anything reading provenance back.

  • create_brr_wgts and create_jk_wgts(paired=True) pair PSUs into variance strata themselves. stratum_name names the created column (the house wgt_name convention), order_by pairs adjacent PSUs in that order — what a systematically-sampled frame wants — and shuffle pairs at random. stratum/psu on all four generators build from columns the Design does not name, without mutating it.

  • Poisson bootstrap replicate weights (#131). sample.weighting.create_bs_wgts(kind="poisson") generates Beaumont–Patak generalized bootstrap weights, which need only a weight column. The default kind="rao-wu" is the stratified Rao–Wu–Yue rescaling bootstrap and still requires psu on the design — the guard is deliberately not shared, since the Poisson bootstrap exists precisely for files that have no PSU. Both kinds use the same 1/R per-replicate coefficient; they differ in how the replicates are drawn, not in how the variance is scaled.

    Beaumont, J.-F. and Patak, Z. (2012). On the generalized bootstrap for sample surveys with special attention to Poisson sampling. International Statistical Review, 80(1), 127–148.

  • scale: the coefficients a producer publishes (#7). Replicate weights often ship with a documented per-replicate variance coefficient that is not the method’s default. CETIC publishes the ICT Households microdata with 0.0065365334145709 at R = 200, which bakes in a finite-population correction. Declaring it is now a parameter rather than an arithmetic exercise:

    svy.RepWeights(method="bootstrap", prefix="REP", n_reps=200, scale=0.0065365334145709)

    A scalar is broadcast to every replicate; a sequence must match n_reps and is checked at construction. It replaces, rather than multiplies, the method default: V = Σ_r scale_r · (θ_r − θ̄)². svy’s scale is R svrepdesign’s scale × rscales folded into one field, so svrepdesign(type = "bootstrap", scale = c, rscales = rep(1, R)) is scale=c here, with no conversion factor. examples/ict_households.py drops the sqrt(scale * R) workaround it used to demonstrate.

  • Jackknife designs say which family they are. JackknifeWgts.kind takes R’s svrepdesign(type=) names — "jk1" (unstratified delete-one-PSU, (R−1)/R), "jkn" (stratified, (n_h−1)/n_h per replicate), "jk2" (paired, one replicate per stratum, 1.0). "paired" is accepted as an alias for "jk2": it describes the design, not the replication scheme, since a JKn on a paired design is still JKn.

    The default is None — unspecified. None and "jk1" produce the same number but are different statements, and svy claims only what it knows or was told. A producer withholding design variables is not evidence that the weights are unstratified.

  • Sample.rep_coefs shows which coefficient was applied to which replicate, keyed by the actual column name:

    >>> sample.rep_coefs
    {'jk1': 0.6667, 'jk2': 0.6667, 'jk3': 0.6667, 'jk4': 0.5, ...}

    coefficients() on the variant remains the machine path — a list of n_reps floats in replicate order, the shape the kernel takes — but a bare list makes the assignment something a reader has to trust rather than check, and for an unbalanced JKn the assignment is the answer: the same coefficients against a different order give a different standard error.

    It lives on Sample rather than on rep_wgts because only the frame resolves the column names: a design declaring prefix="REP" against a file shipping REP001..REP200 knows the replicate count but not the padding, so the struct alone would key this on REP1 and be confidently wrong. Empty when the design carries no replicate weights, and it raises whatever coefficients() raises rather than reporting a number it does not have.

    rep_wgts.coef_source names the provenance alongside it — "scale" (the user asserted them), "derived" (svy computed them and cannot redo it) or "default" (the method’s standard value). A varying vector also prints its distinct values with counts now, 0.667 x3, 0.5 x4, rather than first-and-last — which two numbers land at the ends is an artefact of the producer’s replicate order, while the counts are the design.

  • kind is optional once the units are named. jk1/jkn/jk2 is expert vocabulary; naming the columns the replicates were built from is not, and requiring both asks for the same fact twice in the harder of the two languages. With stratum and psu declared, svy reads the scheme off the counts — one replicate per PSU across several strata is JKn, per PSU in a single stratum is JK1, one per stratum is JK2 — and records it, at INFO level so the inference is auditable:

    JackknifeWgts(prefix="jk", n_reps=8, stratum=("region", "urban"), psu="psu")
    # -> kind='jkn', rep_coefs=0.5 (derived)

    This used to refuse, on the grounds that deriving a kind would turn “the producer withheld the design” into “these weights are unstratified”. That was right while the counts came from Design.stratum, which says nothing about how the replicates were drawn; the units declared on the weights are exactly that statement.

    A kind that the units contradict is now an error rather than a warning, but only when n_reps lands exactly on another scheme’s count — then the label is simply wrong and svy can say which it should be, and warning would ship a coefficient wrong by a known factor. Declaring jk2 against 8 replicates and 8 PSUs raises and names jk1/jkn. A count matching no scheme still only warns: a frame subset to fewer PSUs than the weights were built from is not a mislabelling.

  • JKn coefficients are derived when the declared units allow it. (n_h−1)/n_h is the only standard coefficient that is not closed-form in n_reps, so it used to have to be supplied by hand. Given kind="jkn" plus the stratum and psu the replicate weights name (not the Design’s — see above), svy now works it out at Sample construction. A kind="jkn" with no units named, or with nothing derivable from them, warns at construction and raises from coefficients(). Balanced designs — every stratum the same n_h — take a fast path where the coefficient is uniform and the replicate-to-stratum mapping never has to be found.

    Unbalanced strata are not a different method, just JKn with a coefficient that varies by stratum, so the only thing missing is that mapping — and a delete-one-PSU replicate states it by zeroing the PSU it drops. svy now reads it back: one aggregation to PSU level, and the recovered vector is identical to what create_jk_wgts would have assigned, to the last bit. Only replicate → stratum is needed, since the coefficient is constant within a stratum.

    It applies only where the delete-one signature holds — every replicate column zeroing exactly one otherwise-nonzero PSU, no PSU dropped twice. A bootstrap, a BRR or a Fay-style variant fails that and falls back to refusing, rather than being handed a confidently wrong vector. scale remains the answer for files that do not zero.

    A declared kind is also checked rather than merely trusted: jk1 and jkn imply one replicate per PSU, jk2 one per stratum. Mismatches warn rather than raise, since a legitimately subset frame has fewer PSUs than the weight columns were built from. An unspecified kind on a stratified design warns too — svy has evidence the JK1 global is probably wrong, but no claim to act on.

Fixed

  • kind="jk1" against stratified units silently returned the unstratified coefficient. jk1 and jkn imply the same replicate count — one per PSU — and differ only in stratification, which the replicate-count check never looks at. So the one mismatch it structurally could not catch was the expensive one: jk1 on units declaring several strata used (R−1)/R where (n_h−1)/n_h is correct, overstating standard errors by sqrt(R/n_h) — 32% on a 4-strata × 2-PSU design — with no warning. Naming a stratum unit says the replicates were drawn within those strata and jk1 says they were not, so it now raises and names jkn. Replicates genuinely drawn without regard to strata are declared by naming only psu.

  • A declared JKn with a psu but no stratum silently produced the JK1 coefficient. (n_h−1)/n_h is a per-stratum quantity; with no stratum named, the PSU count fell back to a single group holding every PSU and the derivation returned (R−1)/R — the JK1 global — under a JKn label. On 4 strata × 2 PSUs that is 0.875 where 0.5 is correct, an 8% error in the standard error, with no warning. It now refuses, the same way an unmet kind="jkn" claim already did.

  • Building BRR or JK2 weights no longer destroys the design strata. create_variance_strata ended by overwriting Design.stratum with the collapsed variance strata — its only way of handing the result to the generator that ran next. Every later Taylor estimate then silently linearized over pseudo-strata: on a 6-strata × 4-PSU frame the SE moved from 0.858 to 0.641 and df from 18 to 12, with no warning and the original column still sitting in the frame. Pairing is now internal and its output is recorded on the replicate weights, so Design.stratum is Taylor’s alone and is left exactly as declared.

  • create_jk_wgts(paired=True) no longer undercounts replicates. Called on strata with more than two PSUs — i.e. without the old pairing pre-step — it produced one replicate per original stratum instead of one per variance stratum, in silence: 6 where 12 were correct, understating the variance. It pairs first now. This changes published standard errors for anyone who called it without the pre-step.

  • create_brr_wgts no longer requires a stratum. It raised BRR requires 2 PSUs per stratum on any unpaired design, and stratum=None was rejected outright with a hint to go and build variance strata by hand. Both cases now pair and build.

  • A row-level order_by column paired PSUs more than once. Uniquing on the whole row left the same PSU present several times, so it was assigned to several variance strata and some ended up holding a single PSU — which the BRR kernel then rejected. First occurrence in the frame now wins per PSU.

  • Renaming or dropping a unit column follows through to the replicate weights. _rep_wgts_with_renames only remapped prefix and returned early when no replicate column matched the rename — the common case, since renaming a variance-stratum column touches no replicate column. The units were also absent from the protected set, so dropping one was neither blocked nor cleaned. They are now rewritten on rename, need force= to drop, and are cleared rather than left dangling when force-dropped.

  • method=None no longer auto-selects replication, and Taylor on a replication-only design says so. The default resolved to replication whenever the design carried replicate weights and had no stratum/psu, and to Taylor otherwise. That made the estimator depend on inputs the estimator never reads — replication consumes the replicate columns and coefficients(), nothing else — so declaring a single design column silently moved the standard error:

    mean("y"), identical replicate weights throughout
      wgt only             -> Jackknife  se=1.252718
      wgt + stratum        -> Taylor     se=1.171192
      wgt + psu            -> Taylor     se=1.002973
      wgt + stratum + psu  -> Taylor     se=1.063071

    It was worst for JKn, where stratum and psu are exactly what svy needs to derive (n_h−1)/n_h: the one action that makes JKn usable was the action that switched JKn off. None is now the unstated Taylor default the signature Literal["taylor", "replication"] | None always implied, not a third mode. Replication is opt-in: pass method="replication".

    Falling through to Taylor alone would have reinstated the worse half of the original bug, which is why the auto-detection existed. Without stratum or psu every row is its own PSU in one stratum, so linearization is SRS-like — on a 200-row file with 8 replicates, df=199 against a true 7, silently. That case now emits TAYLOR_WITHOUT_DESIGN and proceeds, for the explicit method="taylor" spelling too: the hazard is in the number, not in who asked for it.

  • A JKn design that cannot be derived warns at construction. kind="jkn" with no psu to count from, or with unbalanced strata, built a Sample in silence and failed only when an estimate was requested — at a call site far from the cause. Both dead ends now emit JACKKNIFE_COEFS_UNAVAILABLE naming which one it hit. The MethodError from coefficients() stays lazy on purpose: such a Sample is still perfectly usable for Taylor, for wrangling and for where=, so raising at construction would block work that never touches replication. What was missing was the early signal, not the error.

  • The JKn error names both remedies, not one. It named only scale=, sending anyone whose file does carry psu off to hand-compute coefficients svy would have derived for them. It now offers declaring stratum/psu first, and scale= for the unbalanced case where that cannot help.

  • create_bs_wgts(kind="poisson") fails cleanly against an older svy-rs. rust_create_poisson_bs_wgts was missing from the ImportError fallback, so the name was never bound and the guard raised NameError instead of the intended message.

  • MethodError.not_applicable no longer doubles the sentence-ending period. The template appended . to a reason that ~24 call sites already ended with one, producing …(got psu=None)...

  • RepWeights.df is annotated float, matching what it has always stored. It was declared int | None while the kernels hand back an f64, so a repr read df=499.0 against an int annotation. Widened rather than coerced: df feeds a t-quantile, which is defined for fractional df, and Satterthwaite-style effective df is fractional by construction. An int is still accepted and stored unchanged.

  • A user-supplied replicate coefficient was silently discarded for bootstrap, BRR and SDR (#7). Before #131 the override was applied at one site in the kernel — rscales.unwrap_or_else(|| replicate_coefficients(method, n_reps, fay_coef)) — and was therefore method-agnostic by construction. #131 moved coefficient computation into Python and scattered it across four per-variant coefficients() methods, which gave every method its own chance to forget. Three of four forgot: the field was accepted, stored, length-checked, and then dropped, so a bootstrap declared with a producer’s published scale silently returned the 1/R answer.

    coefficients() is now a template method on the shared base. The override is resolved there, at the single point of use, and the variants implement only their own default — they never see it, so a new variant cannot regress this again. No published standard error changes: #131 landed after svy-v0.24.1 and was never released.

  • Estimation and regression could scale the same design differently. GLM._rep_coefs re-derived the coefficients by substring-matching a stringified method tag (if "boot" in m), and honoured the user’s override for every method while the estimation path did not. One Design therefore produced differently scaled standard errors from sample.regression.glm(...) and sample.estimation.mean(...). It is gone; both call rep_wgts.coefficients().

  • A declared paired jackknife used the wrong coefficient. JK2 has one delete-one replicate per stratum and a coefficient of 1.0; the global (R−1)/R understates the variance by exactly that factor. Declaring kind="jk2" now gets 1.0. Relatedly, kind="jkn" with nothing to compute the per-stratum coefficients from — no scale, no design variables — refuses rather than substituting the JK1 global, which on a 4-strata × 2-PSU design overstates standard errors by sqrt(7/4), 32%. An unmet claim fails; an absent one (kind=None) still falls back to (R−1)/R exactly as before.

  • Labelled variables were never is_categorical (#130). Sample(data=df) infers mtype from the polars dtype, and import_labels_from_svyio_meta brought across labels while leaving that untouched — so every labelled variable in a survey recode stayed Numerical Continuous and is_categorical was False. Deriving a codebook from it classified an entire DHS-style file as continuous. The importer now revises mtype on import: it honours a measurement level the file declares, and otherwise asks whether the value labels cover every observed value. On a Stata census file this takes 16 labelled variables from 0 categorical to 16.

  • Measurement type is inferred from label coverage, not label presence. “Has labels” alone would misread a continuous variable that labels only a sentinel code — a “how many TVs?” count labelling just 0 = none and 4 = 4 or more is not a five-level factor. A variable becomes NOMINAL only when every observed non-null value carries a label; partial coverage, an all-null column, and a column absent from the frame all leave mtype alone, as does a type the user set by hand. ORDINAL is assigned only when the file says so, since coverage cannot distinguish it from NOMINAL.

  • The polars deprecation filter never worked. filterwarnings = ["error::DeprecationWarning:polars"] looked like a guard against calling deprecated polars APIs and was inert: the module field is compared against the frame the warning is attributed to, and polars sets stacklevel to point at the calling code, so :polars never matches. Verified with a probe — a deprecated call passed cleanly under it. The filter is now unqualified, and the suite passes under it.

Changed

  • Replicate-weight designs are a tagged union of four types, and svy.RepWeights is a function (#131). They were a single struct carrying an estimation-method enum plus every method’s parameters. They are now one type per method — BootstrapWgts, JackknifeWgts, BrrWgts, SdrWgts — each carrying exactly the parameters its algorithm has. A bootstrap cannot hold a Fay coefficient, because there is no field for one.

    svy.RepWeights(method=..., ...) keeps working and returns the variant for the name, so call sites are unaffected. But it is a factory function now, not a class, which breaks two things it used to support:

    isinstance(x, svy.RepWeights)  # TypeError: isinstance() arg 2 must be a type
    x: svy.RepWeights  # no longer a valid annotation

    Use svy.RepWgts, the union of the four variants, for both. Code that knows the method at authoring time can construct the variant directly, which is the typed path: BootstrapWgts(prefix="bsw", n_reps=1000, kind="poisson").

  • rscales is now scale (user-supplied) or rep_coefs (svy-derived). The single field held both, which made “the user asserted this” indistinguishable from “svy computed this”. They split on who supplied the values — the one axis that is verifiable, since a user declaring stratified JKn supplies the standard coefficient rather than a custom one. scale is what you pass; rep_coefs is filled by create_jk_wgts and by the JKn derivation, and is shown as (derived) in the design output.

    This renames a parameter on two released surfaces: svy.RepWeights(rscales=...) and Design.update_rep_weights(rscales=...). Both now take scale= (or rep_coefs=), and scale additionally accepts a scalar.

  • A parameter that belongs to another method is refused, even at its neutral value. RepWeights(method="bootstrap", fay_coef=0.0) was accepted as a stored no-op when every method shared one flat signature. It now raises, naming the variant that owns the parameter. Callers that know the method should not be sending it; nothing in svy does any more.

  • The method label is display-only, and now genuinely is. Its docstring said nothing read it to choose a code path while eight sites did — seven in weighting/ rebuilt a design by round-tripping the type through its own tag and hand-listing every field that had to survive, which is how kind, scale and padding could vanish across a poststratification. Copies now go through msgspec.structs.replace, which preserves the type and drops nothing.

  • Non-default variance coefficients are visible. scale and rep_coefs did not appear in repr or print(design) at all. A non-standard coefficient is what a reviewer or replicator most needs to see and is set rarely, so rare-and-silent was the wrong combination. A uniform vector prints as the scalar it is.

  • Design.update_rep_weights is a third the size, with the same behaviour. It resolved nine sentinel parameters by hand and rebuilt the variant from a method name, because seven internal callers needed exactly that. They now use msgspec.structs.replace, so it keeps only the two jobs replace cannot do — changing the method, and creating weights where there are none. Unchanged for callers: the same signature, the same first-time error, and a variant-specific field still carries only within the same method.

  • import_labels_from_svyio_meta now requires the frame the metadata describes as its third argument. Resolving a measurement type depends on how well the value labels cover the observed values, which cannot be judged without the data. It is required rather than optional on purpose: an omitted frame would silently fall back to the very behavior this release fixes. The function is internal (svy.engine.io, not exported from the top-level svy namespace) and had a single production call site.

Removed

  • BREAKING: Sample.weighting.create_variance_strata. Undocumented — no mention in any changelog, README or guide — never called by svy itself, and reachable in practice only through a hint inside create_brr_wgts’s error telling you to call it. It was a mandatory pre-step svy made you perform by hand and then punished by clobbering Design.stratum. The pairing algorithm survives as a private helper with its edge cases intact (odd PSU counts, tuple strata, ordering, reproducible shuffling); its order_by, shuffle and into controls moved onto create_brr_wgts and create_jk_wgts, where they are discoverable. into= is now stratum_name=.

    # before                                    # after
    s = s.weighting.create_variance_strata(     s = s.weighting.create_jk_wgts(
        method="jk2")                               paired=True)
    s = s.weighting.create_jk_wgts(paired=True)
  • svy.core.design.make_rep_weights. A strictly weaker duplicate of svy.RepWeights — same job, but only four of the parameters, so it silently dropped scale, rep_coefs and kind. It was never exported from svy or svy.core, so it was reachable only as from svy.core.design import make_rep_weights, and had no callers in the package. Use svy.RepWeights(method=..., prefix=..., n_reps=...), which takes the same three positionally.

0.24.1 — 2026-08-11

Fixed

  • Value labels survive being written. Reading back what apply_labels had just stored raised AttributeError: 'int' object has no attribute 'code' (#127):

    s.wrangling.apply_labels(categories={"q1": {1: "Yes", 2: "No"}}).labels

    VariableMeta.clone copied through msgspec.structs.replace, which does not run __post_init__ before msgspec 0.21 — and the floor was >=0.19.0. __post_init__ is the one place that turns the authoring dict {1: "Yes"} into the ValueLabel pairs the field is declared to hold, so on a resolved 0.20 the dict was stored verbatim and the failure surfaced at the next read rather than at the write. Since apply_labels reaches clone through set_value_labels, the whole documented path for labelling a variable was affected on that msgspec.

    Cloning now goes through the constructor, which normalizes and re-checks the invariants by definition, on any msgspec. The floor moved to 0.21 as well (see below) — two fixes for one defect, deliberately: the pin closes it for replace everywhere in the codebase, and the constructor closes it here regardless of what the pin later becomes. Applied to VariableMeta, Label and CategoryScheme — the three structs that normalize in __post_init__.

Changed

  • msgspec floor raised to 0.21. msgspec>=0.19.0 became msgspec>=0.21.0 (#127).

    0.21 is the first release where structs.replace runs __post_init__; 0.19 and 0.20 skip it, verified directly against each. Any struct that normalizes or validates in __post_init__ is therefore only as sound as the resolved msgspec, and svy has several — the value-label defect above is what that looks like in practice, and it reached users because a lockfile resolving 0.21 hid it from the test suite while a docs environment resolving 0.20 published the traceback.

    Nothing in svy needs a 0.20 or 0.19 runtime, and 0.21.1 was already what every lockfile here resolved, so this only rules out an installation that was never tested.

0.24.0 — 2026-08-10

Added

  • corr and cov: design-based correlation and covariance (#124). Both are available on sample.estimation with Taylor and replication variance, by=, where=, domains and deff.

    Neither statistic has a direction, so neither takes y/x. A call names a set of columns through one symmetric cols argument, and every requested pair is returned as its own row:

    sample.estimation.corr(("income", "age"))  # one pair
    sample.estimation.corr(["income", "age", "educ"])  # every unique pair
    sample.estimation.corr([("income", "age"), ("income", "educ")])  # exactly these

    The three spellings are disambiguated by element type and agree wherever they overlap, so a two-element list and a two-element tuple mean the same thing. A flat list yields off-diagonal pairs only — for covariance as much as correlation — so a variance is requested explicitly as cov(("a", "a")).

    kind= selects the coefficient ("pearson" today) while method= remains the variance estimator. Since pandas spells the coefficient method=, that mix-up is caught by name and redirected; a recognised but unimplemented coefficient reports that it is not supported yet rather than invalid.

    Correlation is bounded, so its interval is built on Fisher’s z scale and transformed back — the same move prop makes with a logit, and for the same reason: a symmetric Wald interval can otherwise report bounds outside [-1, 1]. ci_method="wald" opts out. Covariance is unbounded and takes Wald.

    Validated against R survey 4.5: covariance and its SE against svyvar, correlation and its SE against svycontrast’s own delta method over the moment means — R has no correlation SE of its own, so that agreement is an independent check of the linearization rather than a restatement of it — plus stratified and JK1 replicate fixtures. Adds PopParam.CORR and PopParam.COV.

Changed

  • BREAKING: deff now names the SRS reference instead of taking a boolean. deff=True and deff=False are rejected with a structured MethodError that names the replacement; the argument takes "wor", "wr" or None.

    before now
    deff=True deff="wor" — without replacement, Kish’s design effect. Same numbers as before.
    deff=False omit the argument, or deff=None
    — deff="wr" — with replacement, the square of Kish’s “deft”

    A boolean never said which reference it meant. Accepting it silently would leave two spellings for one thing indefinitely, and ignoring it would be worse: Literal is not enforced at runtime, so a string-only implementation would read deff=True as “off” and quietly stop reporting a design effect the caller asked for. Rejecting it loudly is the only option that cannot mislead. "replace" is accepted as an alias for "wr", since that is R’s spelling.

    The reference describes the denominator — what the design variance is compared against — and says nothing about how the sample was drawn; deff="wr" is not a claim of with-replacement sampling. The two differ by exactly the finite-population correction 1 - n/N, so they agree closely at small sampling fractions and diverge sharply otherwise: at f = 0.8, an evaluation-study shape, they differ five-fold (5.027 against 1.005, both matching R).

    "wor" infers N from the sum of the weights, so it is meaningful only while those weights remain reciprocals of selection probabilities. After normalize, or to a lesser degree raking or calibration, that sum is no longer a population count and the design effect is silently wrong — svy cannot detect this in general, since weights normalized to twice the sample size pass every check and yield a plausible wrong answer. "wr" has no N in it and is unaffected. The one provable case now raises rather than returning a column of NaN: when the weights sum to no more than the sample size, the correction is zero or negative, which means either rescaled weights or a census with no sampling variance to compare against.

    The reference is recorded on the result. It appears in the printed header beside the variance method — Estimate: MEAN (TAYLOR, deff=wr) — is available as Estimate.deff_ref, and round-trips through serialization as an optional deff_ref field. to_polars() is deliberately unchanged: that frame carries no method, param or design either, and a deff column there was already provenance-free.

  • Error codes default per class. SingletonError and SvyWarningsError fell back to the base SvyError code when raised without an explicit one, so two distinct failures could surface under the same identifier. Each now carries its own default (#123).

  • svy.serialize raises svy’s structured errors instead of bare built-ins. The serialize module was the last public surface still raising TypeError/ValueError where the rest of the library uses the SvyError hierarchy — structured errors with a stable code, expected/got context, and a hint. The new SerializationError (exported as svy.SerializationError) covers the serialize-specific failures, and the unfitted-model case now reuses the same ModelError that GLM.predict() already raises:

    call before now code
    serialize(<unsupported type>) TypeError SerializationError UNSUPPORTED_RESULT_TYPE
    serialize(<unfitted GLM>) ValueError ModelError MODEL_NOT_FITTED
    from_json(<no "kind">) ValueError SerializationError PAYLOAD_MISSING_KIND
    from_json(<unknown "kind">) ValueError SerializationError PAYLOAD_UNKNOWN_KIND

    Breaking for callers that catch the old built-ins: SvyError subclasses Exception, not TypeError/ValueError. Catch SerializationError/ModelError (or svy.SvyError to cover every svy failure) instead, or match on the code field, which is the stable contract.

0.23.0 — 2026-08-05

Added

  • sample.estimation.quantile() — design-based quantiles with standard errors (#112). Previously only the median carried a standard error; every other quantile was available as a point estimate through describe(percentiles=), with no variance. quantile() estimates any set of probabilities under Taylor linearization or replicate weights, with by= domains and where= filters:

    sample.estimation.quantile("income")  # quartiles, the default
    sample.estimation.quantile("income", p=0.9)  # a single Estimate
    sample.estimation.quantile("income", p=(0.1, 0.5, 0.9), by="region")

    Standard errors follow Woodruff (1952), the construction behind R’s svyquantile: the design-based variance of the estimated proportion P(Y <= q) is taken on the probability scale, and the interval comes from inverting the weighted CDF at p ± t·se_p. The reported se is the back-solved half-width.

    p follows the same rule as y: a scalar returns one Estimate, a sequence returns one per probability. median() is unchanged and remains the p = 0.5 case reported as PopParam.MEDIAN.

  • EstimateList — the list subclass now returned wherever a call estimates several things at once (mean(["a", "b"]), quantile(y, p=(...))). Printing a bare list previously showed object reprs ([<svy.estimation.estimate.Estimate object at 0x…>, …]); an EstimateList renders its members as one table and adds to_polars(). It is a list, so indexing, iteration, unpacking, and isinstance(result, list) are unaffected.

    Serializing a multi-estimate result also works for the first time — serialize() dispatches on exact type and previously raised TypeError: No serializer registered for list. The new estimate_list kind wraps the members, each serialized exactly as a standalone Estimate.

Fixed

  • Quantile and median confidence limits now invert the CDF with the rule that located the point estimate. The Woodruff interval endpoints were always interpolated linearly, regardless of q_method. R’s oldsvyquantile hands one method/f pair to both its point approxfun and its endpoint approx; interpolating linearly while estimating with, say, "higher" pulls both endpoints inward and understates the standard error. The error grows where consecutive order statistics are far apart: on a 2000-row stratified design it was 2.0% at the median and 10.3% at p = 0.99, always in the anti-conservative direction. Estimates, standard errors and both confidence limits now agree with R to machine precision at every probability tested from 0.01 to 0.99, including domains.

    This changes median() standard errors and confidence limits. Point estimates are unaffected.

  • The Woodruff linearization is centered on its own weighted mean, matching how R computes the variance (svymean(U, design)). Only the linear, middle and nearest tie rules moved — up to 4e-4 relative on the confidence limits — because the higher/lower inversion snaps to an order statistic and absorbed the difference. median() and default-q_method results are unchanged by this one. See svy-rs.

0.22.1 — 2026-08-04

Changed

  • Taylor estimation uses the cores it is given. On a 10-core machine at 1M rows a single-variable mean used 1.15 cores and an 8-variable batched mean reached 1.8 of a possible 8. None of it was a rayon width problem — three pieces of redundant serial work sat around the fan-out, and removing them is what freed the parallelism:

    • sample.estimation returned a new Estimation on every attribute access, so the _data_version-keyed caches it carries — the factorized design arrays and prepared design info — were discarded before they could ever be reused, and every sample.estimation.mean(...) re-derived the whole design. The accessor is now retained per Sample. Derived samples are handled by an identity check: _replace_data forks with copy.copy, which carries the cached accessor over verbatim, and without the check a fork would answer with an Estimation still bound to its parent’s data.
    • The reporting metadata on each Estimate (unique stratum labels, PSU count) was computed with np.unique over the full-length design arrays per estimate. A batched call produces one Estimate per variable, so an 8-variable mean did 16 full-length passes — 63% of that call’s wall time at 1M rows, all serial. These are properties of the design, not of the variable, so they are memoised on the design cache and invalidate with it.
    • The Rust kernels stopped indexing the design twice per estimate and now overlap their independent halves — see svy-rs.

    Measured on a 10-core M1 Max at 1M rows, 50 strata, 2000 PSUs:

    case before after speedup cores used
    mean, 1 variable 70.9 ms 22.4 ms 3.2× 1.15 → 1.60
    total, 1 variable 83.8 ms 27.4 ms 3.1× 1.12 → 1.86
    mean, 8 batched 163.2 ms 38.6 ms 4.2× 1.78 → 3.47
    mean by group 144.4 ms 131.9 ms 1.1× 3.45 → 3.46

    Thread scaling from 1 to 10 threads went from 1.11× to 1.54× for a single variable and 1.58× to 3.03× for the batched call. A single estimate overlaps two halves rather than fanning out, so its ceiling is ~2× by construction.

    Estimates, standard errors and degrees of freedom are unchanged — bit-for-bit, and identical at 1, 2 and 10 threads.

    One trade-off worth knowing: retaining the accessor means its cached design arrays (~31 B/row) stay alive as long as the Sample rather than being freed between calls. Peak memory is unchanged — those arrays were rebuilt on every call before, so the high-water mark was always there (2433.8 MB before vs 2434.5 MB after, over six calls at 10M rows). What is new is that a long-lived process holding many large Sample objects idle now holds their design caches too.

Fixed

  • Replication variance no longer costs O(B²) in the replicate count B. The replicate kernels themselves were never at fault — they already do a single O(n·B) pass — but two sites in the Python prep layer tested column membership once per replicate weight column: prepare_data with [c for c in rep_weight_cols if c in df.columns], and Estimation._ensure_float64 with c in data.columns and data[c].dtype != .... df.columns is a property that rebuilds the entire column-name list across the FFI boundary on every access, so B lookups cost O(B²) Python-string constructions; _ensure_float64 also materialised a Series per column. A cProfile run at n=25,000 / B=800 put PyDataFrame.columns at 57% of total runtime, called ~2·B times per estimate.

    With total work n·B held constant at 20M cells — an identical 160 MB replicate weight matrix in every case — the sweep spanned 36× across cases doing identical arithmetic. Column names are now snapshotted into a set once, and dtypes read from a single data.schema snapshot that answers both existence and dtype:

    n B before after speedup
    200,000 100 0.0067 s 0.0059 s 1.1×
    50,000 400 0.0219 s 0.0077 s 2.8×
    25,000 800 0.0704 s 0.0116 s 6.1×
    12,500 1600 0.2442 s 0.0201 s 12.1×

    Sweep spread drops from 36.2× to 3.4×. This matters most for bootstrap designs at B=1000+; it is negligible for BRR (32–64) and SDR/ACS (80). Results are numerically inert — mean, total, ratio, prop and mean-by-domain are bit-for-bit identical to the previous build.

0.22.0 — 2026-08-02

svy labels values so results print nicely. That is the whole job. Everything that made a label list shareable — concepts across locales, hierarchies, the semantics of why an answer is absent — moves to svy-spec, where a questionnaire gives it meaning. A catalogue without one is machinery with no job.

If you set labels by hand, read them from a .sav/.dta, or print them on estimates, nothing you do changes. If you used missing-value definitions, see Removed.

0.21.1 was never published. It was numbered and its notes written on 2026-07-27, then further work landed before a tag was cut. Its changes ship here, and are kept below under their own headings so nothing is lost — but no svy==0.21.1 exists to install, which is why there is no section for it.

Added

  • MetadataStore.update(other, *, overwrite=False) — merge one store into another, field by field.

    Metadata for a variable arrives from several places that each know a different part of it: measurement types inferred from the data, missing-value codes declared by the analyst, question wording carried by an instrument spec. The only way to combine two stores was set, which replaces a whole VariableMeta — so applying a spec silently cleared missing codes, because a questionnaire has no concept of them and its record carries missing=None. That loss is invisible until an export drops the declarations.

    update merges per field, which means a source can only ever add what it knows and can never clear what it has no opinion about. overwrite=False (the default) fills only gaps, keeping labels you have already chosen; overwrite=True lets other win where both are set — for a spec whose question wording should be definitive. A field other has not set is left alone in either mode, which is the property that makes the merge safe.

    store.update(other)  # fill gaps only
    store.update(other, overwrite=True)  # `other` wins on conflicts

Changed

  • VariableMeta.value_labels, ResolvedLabels.value_labels and Label.categories hold (code, label) pairs, with a mapping accepted when constructing and a dict view for lookup — .labels, .labels, .label_map. JSON object keys are always strings, so a dict[Category, str] wrote {101: "Banjul"} and read it back with a string key, silently, after which every join against an integer-coded column missed.

    Label.categories also dropped _MissingType from its union: msgspec refuses to decode any union containing a custom type, so the struct raised regardless — and nothing ever set the field to the sentinel, since None already meant “no value labels”.

  • CategoryScheme holds one entry per code, SchemeEntry(code, label), replacing the mapping / missing / missing_kinds collections keyed by code. Three of those were Category-keyed dicts or sets that survived only through a hand-written encoder; every field is JSON-native now, so the bespoke encoder is gone — to_bytes is a single msgspec.json.encode — and a scheme keeps its code types by any route.

Removed

  • MissingDef, and every API keyed on it: VariableMeta.missing, .na_as_level, .has_missing, .with_missing; ResolvedLabels.missing_codes, .is_missing, .non_missing_labels; MetadataStore.set_missing, .set_na_as_level; the same two on Sample; the has_missing columns in summary() and coverage(). svy.metadata no longer exports MissingDef or MissingKind (the enum stays in svy.core.enumerations).

    A 99 labelled “Refusal” is the integer 99 with the label “Refusal”. svy reads it, prints it, and forms no opinion. Absence is a polars null, which needs no metadata.

    The evidence for removing rather than slimming: import never populated it. svy-io surfaces MissingRule and TaggedNA, and svy’s importer reads variable labels and value labels only — so the field’s only source was a hand-set value or svy-spec’s bridge, and its only consumer was writing it back out. Nothing ever acted on it: a declared code did not change an estimate then, and does not now.

    # before
    store.set_missing("age", dont_know=[98], refused=[99])
    # after — the code is a value, and a value needs a label
    store.set_value_labels("age", {98: "Don't know", 99: "Refused"})

    To declare user-missing in an exported .sav, pass it at the boundary that has the concept: svy_io.write_sav(..., user_missing=[...]).

  • locale, everywhere. CategoryScheme.locale, SchemeRef.locale, LabellingCatalog(locale=)/.locale/.set_locale, the locale= argument on pick, list, add_scheme, make_scheme, MetadataStore(default_locale=), MetadataStore.set_scheme, and Sample.use_scheme.

    svy does not translate. All locale did was choose between two schemes registered under one concept — and two concept names do that with no matching algorithm. A label is a string: write "Femme" and svy prints "Femme". svy holds one set of strings; choosing which set is svy-spec’s job.

    catalog.add_scheme(concept="sex_en", mapping={1: "Male", 2: "Female"})
    catalog.add_scheme(concept="sex_fr", mapping={1: "Homme", 2: "Femme"})
  • CategoryScheme.id — it only ever meant concept:locale. The catalogue is keyed by concept now, one concept holds one scheme, and pick() is a lookup. get, remove and to_label take a concept where they took a scheme id.

  • SchemeEntry.parent, .missing, .is_missing, and the lookups over them: parent_of, children_of, codes_of_kind, kind_of, substantive, missing_codes. A scheme is a code→label map. svy has no cascading selects, and why an answer is absent is a questionnaire fact.

  • CategoryScheme.ordered. Order lives in the codes; “is this ordinal” is VariableMeta.mtype.

  • to_label_by_concept, folded into to_label now that concept is the key.

  • 171 lines with no caller anywhere in the monorepo, in svy-spec, or in any test: is_missing_value, recode_for_analysis, display_text, polars_mask, polars_to_analysis, polars_to_display (none exported), SchemeCatalogView, and the no-op seams validate_scheme_missing, normalize_scheme_missing, missing_codes_by_kind. labels.py no longer imports polars.

  • svy.questionnaire, MetadataStore.import_from_questionnaire, and the Sample(questionnaire=) parameter. Describing an instrument is a different job from analysing the data it produced, and svy had come to own a small piece of it: a flat question model with no notion of rosters, ordered scales, or analysis units. That work now lives in svy-spec, which inverts the dependency — svy no longer needs to know what a questionnaire is.

    This is a removal without a deprecation cycle, which the version number alone does not convey. Questionnaire was exported from svy.questionnaire, but never from the top-level svy namespace, never documented, and never used anywhere in svy beyond the one Sample(questionnaire=) hook — which only forwarded to import_from_questionnaire. A patch bump reflects a path with no known consumer; if you were importing it, pin svy==0.21.0 and migrate at your convenience.

    To attach instrument metadata, resolve a spec, project it, and merge it in:

    from svy_spec.bridge import to_metadata_store
    from svy_spec.resolve import resolve
    
    sample = svy.Sample(data, design, catalog=catalog)
    sample.meta.update(to_metadata_store(resolve(spec), catalog=catalog), overwrite=True)

    Use update rather than a loop over set: it merges per field, so a field the spec does not model — missing codes you declared, notes you added — is never cleared by applying it.

    Pass overwrite=True when the spec is the authority, which it is here. Sample.__init__ runs infer_from_dataframe, which sets mtype by guessing from each column’s storage type; under the default fill-only merge that guess wins and the spec’s declared level never lands, so an ordered single-select stays Numerical Discrete instead of becoming Categorical Ordinal. Apply the spec first, then any adjustments of your own.

    MetadataSource.QUESTIONNAIRE stays — it is what the bridge sets, and it remains the right provenance for a field-collected variable.

Fixed

  • The SPSS and SAS writers could not run at all. _write_spss called svy_io.write_spss and _write_sas called svy_io.write_sas; neither name has ever existed. Both calls carried # type: ignore[attr-defined] — the type checker had said so and been silenced.

    The real API is write_sav(df, path, *, var_labels, value_labels, user_missing, ...), taking labels as separate arguments rather than one metadata dict.

    SAS is more than a rename: ReadStat writes SAS Transport (XPT) only — there is no sas7bdat writer — and XPT carries no variable or value labels. _write_sas now writes XPT and warns that the labels did not travel, pointing at write_spss or write_stata. format= and encoding= are reported as ignored, and _write_spss loses its encoding parameter, which write_sav does not take.

    It survived because the only test used a stub that defined write_spss and write_sas itself. A stub that invents the interface it stands in for cannot catch a call to a function that is not there.

  • A labelled Table.crosstab() returned None for every estimate. The frame’s rows were replaced with labels while the skeleton they are joined against stayed as codes.

  • A value label did not apply when the code and the value disagreed in type. SPSS stores value-label keys as strings, so a .sav read back gives {"1": "Yes"} against a Float64 column, and ResolvedLabels.display returned the bare number. display now bridges both directions. display_series was never affected.

  • ttest_to_markdown() raised NameError on any call — it referenced a _stats_summary_line that does not exist. The docstring promised a summary line the code never had; both are gone.

  • SingletonResult.config named a class that does not exist (_SingletonHandlingConfig for SingletonHandlingConfig). Latent only because from __future__ import annotations left it unresolved; it would have broken on get_type_hints or a typed decode.

  • Two tests shared a name, so the second replaced the first and one never ran. Three further tests had no assertions at all.

  • ruff check passes on packages/svy/{src,tests} for the first time.

0.21.0 — 2026-07-24

Requires svy-rs 0.12.0, which carries the two variance-estimation fixes below; svy-io 0.2.0 is unchanged. Grouped confidence intervals and domain design effects change in this release — point estimates and standard errors do not.

Fixed

  • by= groups now use their own degrees of freedom. A by-group is a domain, so its df must be counted on the PSUs and strata that group covers. It was instead given the df of the surrounding analysis — the full design with no filter, or the where= mask with one — so the same subpopulation got a different interval depending on whether it was reached through by= or where=. Confidence intervals for grouped means, totals, ratios, proportions and medians were consistently too narrow; the effect is negligible for groups spanning most of the sample and large for small ones (22% on a 10-record domain with 6 df rather than 56). est, se, cv are unaffected. Verified against R survey 4.5 degf(subset(design, ...)).
  • Design effects no longer count zero-weight rows. Under drop_nulls, rows with a missing response are kept and zero-weighted rather than dropped; they were still counted in the domain SRS variance’s n, inflating deff for any group containing them (~1–2% on the synthetic fixtures). Only deff is affected.

Removed

  • Estimate.degrees_freedom. Degrees of freedom are a per-row property — a domain or by-group is counted on its own active PSUs and strata, so grouped results legitimately carry a different df per cell. The scalar could not represent that: it was min() across rows, so a grouped estimate reported its smallest group, and for a by-group inside a domain that meant a headline df of 0. Use ParamEst.df (also a df column in to_polars()) for the per-row value, and n_psus - n_strata for the full-design df, which stays at design level under a domain filter.
  • EstimateData.degrees_freedom leaves the serialized payload; SCHEMA_VERSION moves to svy-result/0.2. Strictly a field removal warrants a major bump under the policy in serialize/DESIGN.md; 0.2 was chosen deliberately because no known consumer binds to the field, and the reasoning is recorded there.

Added

  • ParamEst.df — the design df backing each row’s t-quantile, carried through to_polars() and the serialized payload. It is deliberately not shown in the printed table: it is constant for most results, so a column would repeat one number down the page and widen every table.

0.20.1 — 2026-07-23

Patch release on top of 0.20.0; svy-rs (0.11.0) and svy-io (0.2.0) are unchanged.

Fixed

  • tabulate percent and count_total cells used an un-centered variance. A cell percentage is a ratio of two estimated totals, so its variance needs the centered (Hájek) linearization. Because the internal totals flag was inferred from sum(weights) != 1, scaling weights to sum to 100 (units="percent") or to a caller-supplied count_total routed the standard error through the un-centered total path, dropping the numerator/denominator covariance term. Cell SEs were inflated by a p-dependent amount (up to ~12% on high-proportion cells) and the confidence interval fell back to Wald, which could dip below zero. units="proportion", units="percent", and count_total=N are now the same estimator scaled by a constant and agree exactly; they match estimation.prop and R survey’s svymean(~interaction(...)). Bare units="count" is unchanged and still matches R’s svytotal, and the Rao-Scott chi-square/F test was never affected.

0.20.0 — 2026-07-23

Builds on svy-rs 0.11.0 and svy-io 0.2.0. This release lands the round 7–8 review: correctness fixes across estimation, regression, weighting, size/power, categorical, and the dataset downloader, several of which shift standard errors closer to R survey 4.5.

Added

  • RepWeights.rscales — exact stratified-JKn variance. RepWeights gains an optional rscales tuple (per-replicate variance coefficients, R’s scale×rscales combined); create_jk_wgts fills it from the design’s strata and estimation threads it to the Rust kernels. svy-generated JKn weights now reproduce R’s as.svrepdesign(type="JKn") mean/total SEs and mse=TRUE centering to 13+ digits (df = degf). Absent rscales, each method keeps its global default, so user-supplied replicate weights behave exactly as before unless the file’s documented rscales are provided.

Fixed

  • drop_nulls zeroes weights instead of dropping rows (R na.rm=TRUE / subset() semantics). prepare_data physically removed any row with a missing analysis value before the domain machinery ran; under standard skip patterns (y null outside the domain) this deleted whole PSUs and strata, understating domain SEs — 15% on the reference dataset — and corrupting df. Missing analysis values now keep their rows with main and replicate weights zeroed. Verified against R survey 4.5 to 13+ digits; the R-calibrated ttest and ratio fixtures were regenerated with these semantics (the old expectations matched R only on physically-filtered complete-case data).
  • Float-typed stratum/PSU columns are accepted. Numeric design codes from CSVs (e.g. MEPS VARSTR/VARPSU) frequently arrive as Float64; the factorized-design cache cast them straight to Categorical, which polars forbids for floats, crashing estimation with “conversion from f64 to cat failed”. Non-string, non-integer dtypes now route through Utf8 first (float- and int-coded designs produce identical results).
  • SSU-level FPC is grouped by (stratum, PSU), not PSU alone. PSU labels are commonly reused across strata, so build_fpc_ssu_column merged distinct PSUs — valid designs raised FPC_NOT_CONSTANT, and matching M_hi values pooled SSU counts across strata, understating the two-stage SSU FPC.
  • method=None auto-detects as documented — replication when the only variance information is replicate weights (no strata/PSU), Taylor otherwise. Previously None always meant Taylor, silently giving replication-only designs an SRS-like variance.
  • Core polish and API consistency (review round 7). Replicate-weight prefix matching is strict ^prefix\d+$ (a loose startswith absorbed columns like repwt_flag; a count/n_reps mismatch is a typed DimensionError); set_data/update_data/set_design/update_design rebuild internal concat columns and re-run singleton detection + design validation instead of leaving stale state; describe() reports weighted std/var/quantiles (aweight convention) and computes categorical proportions over all levels; SingletonHandling enum values are accepted by singleton.handle(); PopSize(psu=..., ssu=None) is accepted for PSU-only FPC; polars_mask() is null-safe; the design-fields cache is bounded (512 entries); importing svy no longer replaces the host’s sys.excepthook (Rich tracebacks install only on SVY_RICH=1). Deleted the unused content-based Sample.__hash__ and the dead _calculate_fpc.
  • GLM design gaps (round 8). Family-specific unit deviance (matching R family$dev.resids) and null deviance at the intercept-only fit; deviance/AIC follow R survey exactly (Lumley–Scott dAIC; bic is None); replicate-weight designs get true replicate variance instead of silently falling back to Taylor SEs; design.pop_size feeds per-stratum FPC into the sandwich; Cat(ref=...) with an absent reference level raises a typed error listing observed levels; Cat levels, the response, and the invalid-weight filter are evaluated on in-domain rows under where=, eliminating phantom all-zero dummies; covariate/where-column nulls keep-and-zero-weight (preserving PSUs in stratum centering). Validated against R survey 4.5 to ~1e-6 or better.
  • GLM margins rewritten on the fitted frame with delta-method SEs (round 8). margins recomputed from raw sample data with ad-hoc SE formulas; it now averages over exactly the fitted rows (post null-drop, post weight filter, with the domain column), rebuilds interaction columns from counterfactual data, differentiates the full linear predictor for AME, and uses full delta-method SEs g'V(β)g over the design-based covariance (Stata vce(delta) convention). Validated against R survey + marginaleffects: points to ~1e-8, SEs to ~1e-4.
  • Weighting adjustment/calibration/trimming marshalling (round 8). adjust raises a typed error on unmatched response statuses (was silently encoding them as respondents and inflating weights) and derives respondents_only from the encoded codes (case-insensitive); adjust(trimming=..., update_design_wgts=False) trims the freshly created adjusted weight instead of the caller’s original; calibrate(bounded=True) raises NotImplementedError instead of being silently ignored; calibration targets are assembled as ordered per-term lists (fixing a “Design matrix label alignment mismatch” on shared numeric codes); the trim-calibrate cycle runs on arrays before writing (a strict non-convergence failure leaves data/design/replicates untouched) and honors TrimConfig.by/min_cell_size; build_aux_matrix raises on nulls in a continuous auxiliary instead of filling 0.0.
  • Weighting typed errors and sorted control order for the svy-rs 0.11.0 changes: create_brr_wgts pre-validates n_reps against the Hadamard order (MethodError.invalid_range); raking-bounds violations surface as MethodError at all four kernel call sites; normalize() orders control values by sorted group id, matching the kernel and poststratify.
  • Wrangling edge cases (round 8). categorize() closes the outer bin edge (R cut(include.lowest=TRUE)) so boundary values no longer vanish from tabulations; remove_columns(force=True) cleans design.pop_size; a partial replicate-weight rename_columns raises instead of corrupting the RepWeights prefix; mutate() specs see same-call redefinitions (dependents no longer read stale values); clean_names() preserves internal concat columns; filter_records() counts and reports Kleene-null-dropped rows; fill_null(strategy="mean") casts integer columns to Float64 for an exact mean; cast(strict=True) raises on lossy float-to-integer casts.
  • Size and power formulas (round 8). compare_means is implemented (was a no-op stub returning None); non-inferiority sizing keeps epsilon signed (the old |eps| collapse under-sized NI designs ~5×); the one-mean two-sided clamp that produced astronomically wrong n is removed; one-sided power follows sign(delta); pooled two-proportion variance and the optimal allocation ratio are un-inverted; the adjustment pipeline is reordered to n0 → DEFF → FPC → nonresponse so the FPC caps the deff-inflated size toward pop_size; parameter validation (p/moe/sigma/power/deff/resp_rate) raises typed MethodError instead of silently clipping.
  • tabulate count CIs use the design-df t instead of the normal critical value; with few PSUs (df = 6) count CIs were ~20% too narrow, now matching Stata svy: tabulate and svy’s own estimation.total.
  • ranktest with a custom score_fn honors by= (each by-level is its own domain, returning one result per level) and group labels reflect the levels actually tested under where=/by= (estimates were always correct; only the reported labels were wrong).

Security

  • Dataset downloader hardened against a hostile catalog. Slugs from registry JSON flowed unvalidated into cache paths, glob patterns, and tempfile prefixes (a slug like ../../foo wrote outside ~/.svy/datasets); slugs are now allowlisted at the registry boundary and defensively in path_for/clear. Downloads without a catalog hash pin the first-seen sha256 (trust-on-first-use) and enforce it thereafter. Plain-http URLs and https→http redirect downgrades are rejected (localhost exempt for development). New error codes DATASET_INVALID_SLUG and DATASET_INSECURE_URL.

0.19.1 — 2026-07-21

Added

  • Bundled offline example datasets. svy.datasets.load / catalog / describe now take a source= argument — "bundled", "remote", or "auto" (default: remote if reachable, else bundled). A small, self-consistent synthetic survey — a sampling frame, its household census, and a two-stage sample drawn from that census (design weights sum to the census) — ships inside the wheel, so the docs and your own experiments run fully offline and reproducibly. SVYLAB_OFFLINE=1 forces the bundled path.
  • DatasetCatalog type and richer Dataset metadata. catalog() returns a DatasetCatalog that prints as a compact table and drills into any entry’s full metadata with .get(slug) (also .slugs, .to_polars()). Dataset prints as a branded panel and gained a notes field documenting how a bundled subset was derived from its remote counterpart.

Changed

  • All dataset failures route through the DatasetError taxonomy with actionable messages and codes: DATASET_NOT_BUNDLED (lists the available bundled slugs), DATASET_DOWNLOAD_FAILED, and BUNDLED_UNAVAILABLE, alongside the existing not-found, catalog, and integrity errors.

Fixed

  • SvyError panels render again. The Rich panel path imported its renderers from a module that had since been renamed, so every error silently fell back to plain text; it now renders the branded panel. The panel also stays aligned in HTML/notebook output — the status marker is a width-1 glyph instead of a two-cell emoji — and the title, body, and metadata are spaced for readability.

0.19.0 — 2026-07-12

Added

  • Batched multi-variable estimation. estimation.mean, total, ratio, prop, and median now accept a list of columns and return a list[Estimate] (one per variable; ratio pairs numerator/denominator element-wise and broadcasts a scalar side). A single string still returns a single Estimate. For ungrouped Taylor estimation the list form shares one design build across variables and runs them in parallel — 4–13× faster than a manual loop at 1M rows depending on the estimator. by=, replication, drop_nulls, and the singleton scale double-pass transparently fall back to independent per-variable calls (identical results).
  • A variable may now appear in both by= and where=. where is domain estimation (out-of-domain weights zeroed) and by groups on the original values — the two are orthogonal, so the previous guard forbidding overlap is removed. When a where predicate excludes an entire by level (e.g. a “don’t know” code), that level is correctly absent from the results — matching R’s filter(...) %>% group_by(...) — while every row still contributes to the shared design and degrees of freedom, so surviving groups’ estimates, standard errors, and df are byte-identical. Covers Taylor and replication, all estimands, and multi-by.
  • Serialization for result objects. New svy.serialize module provides stable, versioned serialization of every result type (estimates, t-tests, chi-square, tables, GLM fits/predictions, describe): serialize(result) returns a kind-tagged struct, to_json / to_dict export, and from_json round-trips. Payloads carry a SCHEMA_VERSION for forward compatibility.
  • Single-stage designs and explicit population sizes. The design’s ssu (second-stage unit) is now optional, so single-stage designs no longer need a placeholder. A PopSize type is exported for specifying finite-population sizes (FPC).

Changed

  • Estimation now fails fast on unhandled singleton PSUs instead of silently under-reporting the variance. Taylor estimation (mean, total, prop, ratio, median) raises SingletonError when a design has single-PSU strata and no handling strategy was chosen — matching R’s options(survey.lonely.psu = "fail"). Pick a strategy explicitly with sample.singleton.skip() / .certainty() / .center() / .scale() / .collapse() / .pool(). Previously such strata were dropped from the variance with no error or warning.

Fixed

  • Taylor standard errors are now bit-reproducible. The stratified variance summed each stratum’s PSUs in the iteration order of a randomized hash set, so a repeated estimate on identical data could differ in its last digits run-to-run (far below reporting precision, but not reproducible). PSUs are now summed in a canonical order, so mean/total/ratio/prop/median return identical standard errors across runs.
  • Stale design cache could return silently wrong results. Estimation design caches were keyed on the identity of the data frame without holding a reference to it; after an in-place mutation freed and reallocated the frame, identity reuse could make a stale entry look current and serve design arrays for the old data. Caches are now keyed on a monotonic per-Sample data version bumped on every rebind, so every mutation, weighting, selection, and fork path invalidates correctly.
  • Replication-design crashes and related correctness fixes. Clone, column keep/remove/rename, and singleton handling now work on replication designs (previously hit stale replicate-weight API usage and could crash). Expr now raises TypeError on boolean use (and/or/not/chained comparisons) so a malformed where= predicate fails loudly instead of silently filtering wrong, and derived samples deep-copy metadata/warnings/design so they no longer share mutable state with the original.

0.18.2 — 2026-05-20

First release tracked in this changelog. For the history prior to 0.18.2, see the Git tags and GitHub Releases.