Release Notes
All notable changes to svy-io, high-speed reading and writing of survey files (SAS, SPSS, Stata) as Polars frames via ReadStat, are recorded here. Releases follow Semantic Versioning; the layout follows Keep a Changelog. Part of the svy project.
Unreleased
0.6.0 — 2026-10-08
Changed
- BREAKING:
write_dtawrites integer columns as Stata integers. Every integer column was cast to Float64 and stored asdouble. Each one (Int8 to Int64, UInt8 to UInt64) is now stored as the smallest Stata type that holds its non-missing values:byte(−127 to 100),int(−32,767 to 32,740) orlong(−2,147,483,647 to 2,147,483,620), anddoubleonly beyond that. Nulls are written as.. Value labels attach to the integer variable, and Stata’sdescribeshows the expected storage. A column holdingTaggedNAvalues is still written asdouble. - BREAKING:
read_dtareturns integer storage as Int64. Statabyte,intandlongvariables were read as Float64, so a code read back as1.0no longer matched the value label keyed"1". They are now Int64 columns. They are not Int8/Int16/Int32 because arithmetic on those polars types overflows without an error..and.ato.zare null, and tagged missings are reported and hydrated as for doubles.floatanddoublestorage still read as Float64.meta["vars"][i]["kind"]gives the storage type:"int8","int16","int32","float","double"or"string"; it was always"double"for numerics. SAS and SPSS files hold only doubles and strings, so their readers are unchanged.read_stata_arrowreturns int64 columns the same way.
Fixed
- The readers work with polars 2.0. Every SAS, SPSS and Stata read failed with
read_ipc() got an unexpected keyword argument 'memory_map', an argument polars 2.0 removed.
0.5.0 — 2026-09-30
Added
- SPSS and Stata readers accept zip archives.
read_savreads the first.sav(else.zsav) member,read_porthe first.por,read_spssthe first.sav,.zsavor.porand dispatches on it, andread_dta/read_stata/read_stata_arrowthe first.dta. As withread_sas, an archive with no matching member raisesFileNotFoundErrorlisting its files, several matches warn (at the caller’s line) and use the first, and the extracted file is removed after the parse. read_sasrecognises SAS Transport (XPT) by content. It dispatched toread_xptonly for.xpt/.xportnames, so transport files published as.sspor.datfailed withrc=5from the sas7bdat parser. The first bytes now decide, for paths, file objects and zip members alike; a zip may hold.xpt,.xportor.sspmembers (a.sas7bdatis still preferred).- SAS CPORT files are refused with a hint.
read_sasandread_xptname the format and point to PROC CIMPORT instead of a generic parse failure. read_xpt(rows_skip=, cols_skip=, encoding=). The native reader already supported row and column selection;encoding(including"utf8-lossy") is now passed to ReadStat as for the other readers.read_xpt(catalog_path=, catalog_encoding=). A.sas7bcatformat catalog labels XPT variables by format name, as it does.sas7bdatones.read_sasforwards the catalog when it hands a file toread_xpt, including a catalog found in a zip.read_sas_arrowreads XPT. It sent every file to the sas7bdat parser, so transport files failed; it now recognises them by content likeread_sasand refuses CPORT with the same hint. The table is returned raw, withoutread_xpt’s temporal coercion.
Fixed
read_sasno longer dropsrows_skipandcols_skipwhen it hands a file toread_xpt.- Catalog labels reach variables whose format has a width. ReadStat names the label set after the format with its width (
WORKSHOP5), so a variable formattedWORKSHOP5.never matched the catalog’sWORKSHOPand silently lost its labels, in.sas7bdatand XPT alike. The width is now ignored when matching;fmtkeeps it.
0.4.0 — 2026-09-08
Added
read_dta(encoding=). ReadStat decodes Stata 13 and older files (format ≤ 117) as Windows-1252 and Stata 14+ as UTF-8; a byte sequence invalid in that encoding failed the whole read with an opaquerc=17. This is what a Stata 13 export holding UTF-8 free text does — the Nigeria GHS-Panel Wave 5 “other, specify” columns carry emoji whose bytes are undefined in CP1252. The option takes an iconv name ("utf-8","latin1", …) and is forwarded toreadstat_set_file_character_encoding, asread_savandread_sasalready do.read_stata_arrowtakes the same option, andread_porgains it too.- Stata 13 and older files are checked for UTF-8 before falling back to Windows-1252. Formats ≤ 117 declare no encoding and ReadStat assumes a code page, which accepts almost any byte, so a UTF-8 file usually decoded silently into mojibake and only failed when it hit one of the five bytes CP1252 leaves undefined (as emoji do).
read_dtanow validates such a file as strict UTF-8 first — data, labels and notes — and reads it as Windows-1252 only if that fails; legacy accented text is essentially never valid UTF-8, so genuinely old files are unaffected. The result is reported inmeta["encoding"]("utf-8","windows-1252", or the explicit value), and an explicitencoding=skips detection. A file that is neither costs a second pass before the rc=17 error. SPSS and SAS headers declare their encoding, so their readers are unchanged. encoding="utf8-lossy"on every reader. The special value (polars’ spelling) decodes as UTF-8 and replaces undecodable bytes with U+FFFD, reporting it inmeta["had_invalid_utf8"]. It matters most for SPSS and SAS: those headers declare an encoding, so ReadStat always transcodes and one stray byte failed the whole read with no way through short of lying about the encoding.
Changed
- Parse errors carry ReadStat’s message.
Failed to parse SAV: rc=17is now… Unable to convert string to the requested encoding (invalid byte sequence) (rc=17)across all readers, and that case appends a hint namingencoding=. had_invalid_utf8covers labels and names too. Only data strings set it before; variable labels, value labels, notes and the file label were replaced silently.
Fixed
cargo testin the native crate builds again. pyo3’sextension-modulefeature was on unconditionally, so the test binary linked without libpython and failed on undefined symbols; it is now a crate feature that maturin enables frompyproject.toml, the same arrangementsvy-rsuses. The SAS reader’s unit tests also still calledparse_sas_implwith the argument list from before the encoding parameters were added.
0.3.0 — 2026-08-26
Added
The public surface is now pinned by a test. Every name in
__all__must resolve, be callable, appear once, and match whatfrom svy_io import *actually yields; each documented alias must be the same object as its target; and every module must be imported by something. This guards a failure that has already happened:svycalledsvy_io.write_spssandsvy_io.write_sas, neither of which has ever existed, and both calls shipped with a# type: ignore[attr-defined]silencing the type checker. Nothing failed until someone ran the writer. Verified with probes — adding a phantom name to__all__fails three of these tests, and adding a module nothing imports fails another.read_spssdispatch andget_user_missing_for_columnare tested. Both were public and referenced by no test.read_spssis not a reader but a dispatcher that picks one by file extension, so the routing is the whole function:.savnow provably produces exactly whatread_savdoes, arguments are forwarded rather than dropped, the match is case-insensitive, and an unrecognized extension raises instead of guessing from content.Readers surface the declared measurement level (#130). Each entry in
meta["vars"]now carriesmeasure—"nominal","ordinal", or"scale"— read from ReadStat’sreadstat_variable_get_measure. A format that carries no such attribute (Stata) or a variable whose writer never set one reportsNonerather than"scale", so a caller can tell “declared continuous” from “never declared”. Worth knowing before relying on it: SPSS defaults numeric variables to"scale"whether or not anyone meant it, so only"nominal"and"ordinal"are positive declarations.
Removed
BREAKING:
VarMeta,ValueLabels,MissingRuleandSvyMetadataare no longer exported. These dataclasses were public and nothing in the package ever constructed one. They also disagreed with what the readers return —SvyMetadata.value_labelswas declaredDict[str, ValueLabels]where a reader hands back alistof plain dicts, andVarMetalacked fields the native layer emits. Anyone importing them to type code againstread_savwas being misled by them. Nothing insvyor in this package referenced them.utils.py. Six functions, 36 statements, 0% coverage — because__init__did not import it and nothing else in the package or the tests did either. Not undertested: unreachable.The benchmark suite. Three of its five tests were the same benchmark:
test_bench_stata_types_13,_14and_15all read one file with no arguments, and reported 29.5 / 29.2 / 29.4 ms — one number, printed three times, under names promising three Stata formats. Half the file was commented out, somake bench-spssandmake bench-sasselected zero tests and exited green. Nothing compared any number to a baseline. It cost ~6s on every test run and an 876 KB data file to say nothing actionable.svy’s harness (bench_kernel.py,check_regression.py, trackedbaselines/) is the shape to copy if svy-io wants perf tracking later.
Fixed
ordered=Truedid nothing. The parameter is public onread_dta,read_sas,as_factor,as_factor_exprandapply_value_labels, and it had no effect on either path. The lazy one passedordering="physical"topl.Categorical, which polars deprecated in 1.32.0 and now ignores — and this package already requires polars ≥ 1.34, so it was inert for every supported version. The eager one never read the argument; it was marked “reserved”. ACategoricalsorts its categories alphabetically, so an education scale came backHigher < None < Primary < Secondary.ordered=Truenow builds apl.Enum, which keeps the order it is given, and the scale sortsNone < Primary < Secondary < Higher. Categories are ordered by the numeric value of the code: readers return value labels keyed by the code’s string form, so the mapping iterates"1", "10", "2"and using that order directly would sequence an 11-category scale wrongly. Non-numeric codes keep their string order.ordered=Truenow refuses what it cannot order.levels="default"andlevels="both"fall back to the raw value for anything unlabelled, so their categories depend on data the label set cannot describe; combining them withordered=Trueraises rather than silently dropping the unlabelled values.ordered=Truewithout value labels raises for the same reason — the code order is what defines the order.Deprecation warnings fail the test run. The suite carried a polars
DeprecationWarningin its summary for releases while the deprecated argument silently did nothing. Note for anyone tempted to narrow the filter back toerror::DeprecationWarning:polars: that form never matches. The module field is compared against the frame the warning is attributed to, and polars setsstacklevelto point at the calling code, so the qualified filter silently does nothing.write_dtasilently dropped value labels (#129). The native writer acceptedvalue_labels_jsonand never read it, so a.dtawas written with its variable labels intact and its value labels gone — no error, and the loss only visible on read-back. The Stata writer now emits a label set per labelled column and points the variable at it, verified againstpandas.io.statafor formats 113 through 119. The set is named after the column (Stata’s ownlabel values v106 v106convention) rather than the SAV writer’s{col}_labels, because dta 113–117 allow only 33 bytes for the name and ReadStat truncates a longer one without complaint.write_dtavalue-label validation. Codes are now rejected when they fall outside[-2147483647, 2147483620](Stata stores a code as int32 and reserves the top of that range for.through.z, so a wider code was truncated silently), when they name a column absent from the frame, and when they target a string column (Stata has no string value labels, unlike SPSS).booland whole-floatcodes are canonicalized to plain integers, so{True: "yes"}writes code1instead of failing at the JSON boundary.
0.2.0 — 2026-07-23
Changed
- BREAKING: unified
user_missingmetadata schema. Three producers emitted three shapes —{var, discrete, ranges}from the native layer,{col, values, range}fromread_sav(user_na=True), whilezap_missinglooked forna_values/na_range— so zapping the missing metadata returned by a realread_savsilently did nothing (the zap tests passed only on hand-crafted dicts). All readers now emit one haven-compatible schema,{col, na_values, na_range}per column, vianormalize_user_missing()(tolerating the legacy shapes). Code that readsuser_missingmetadata must use the new keys.
Fixed
- Native-layer hardening. Encoding parameters (and the SAS
catalog_encoding) now flow intoreadstat_set_file_character_encoding(iconv) instead of being accepted and silently ignored, so legacy code-page files no longer arrive riddled with U+FFFD; invalid encoding names error, and metadata gainshad_invalid_utf8so silent lossy decoding is detectable.n_rowsis counted once per row independent of kept columns (it reported 0 when the first column was skipped); the nativen_max=0off-by-one is fixed andset_row_limitis applied as a defense-in-depth guard for untrusted files.write_xptandwrite_savnow derive and validate string widths from the data (width was hardcoded to 200 / silently capped, leaving truncated files) before writing any bytes. Hand-declared ReadStat externs were replaced with the bindgen bindings so signature drift is caught at build time. - Data-handling edge cases.
n_max=0now opens and validates the file and returns the full schema with zero rows (haven behavior) instead of a schemaless empty frame;write_savencodes categorical columns against their own observed categories with per-column codes under a globalpl.StringCache(was value labels for every cached string in cache order);as_factor_expr(levels="labels")maps unlabelled values to null;LabelledSPSS.to_int()keeps missing valuesNoneinstead of0;_stata_file_formatrejects nonexistent format codes; and_adjust_temporalsnarrows a bareexceptto polars/value errors with a warning. - Column-name collisions.
_normalize_namesdisambiguates two source names that normalize to the same column with numeric suffixes (was a duplicate-rename error) and renamesuser_missingcolumns alongside.
Security
- Private per-call temp dir for zip extraction.
read_sasextracted archive members to predictable paths in the shared system temp dir and never cleaned them up — cross-run collisions and symlink-planting exposure on multi-user machines, plus an unbounded temp leak from thedelete=Falsespool of file-like inputs. All temp artifacts are now scoped to anExitStackclosed right after the native parse; spooled inputs are unlinked.
Build
- Bump
pyo3to 0.29 andbytesto 1.12.1 in the native extension.
0.1.1 — 2026-07-12
Fixed
- SAS datetime values are now decoded correctly. Datetime formats were matched by a
"date"prefix test, soDATETIMEcolumns were sent through the days-since-1960 date path and lost their time-of-day. Datetime formats are now checked first and decoded against a trueDatetimeepoch, preserving the time component. Also fixes anAttributeErroron variables with a null format on the defaultread_savpath. - Numeric ID columns stay numeric. The magnitude/name-based temporal inference heuristics are now gated behind
infer_temporal_formats(opt-in), so numeric identifier columns are no longer coerced to dates/times. as_factor(levels="both")on numeric coded columns no longer mis-handles literal separators; values are stringified before theCategoricalcast.write_savno longer mutates the caller’svalue_labels.- All file-like reader inputs work again (a missing
tempfileimport crashed them). - FFI hardening. Builder pre-allocation is clamped to 65,536 rows so a crafted header can no longer drive a multi-GB eager allocation and abort the process; ReadStat callbacks are wrapped in panic guards that raise Python exceptions instead of aborting; the XPT writer now emits all record batches (rows after the first batch were silently dropped) and handles StringView/dictionary columns.
Packaging
- Source builds and Intel macOS now work. The sdist previously shipped without ReadStat’s C sources (an
includeglob matched zero files), so every source build failed — and Intel Macs always hit that path because nox86_64-apple-darwinwheel was published. The sdist now bundles all ReadStat sources, and prebuilt Intel macOS wheels are published.
0.1.0 — 2026-05
First release tracked in this changelog. For earlier history, see the Git tags.