svy-sae: Small Area Estimation in Python
small area estimation, SAE, Pythonic, statistics, survey sampling
API Stability
svy-sae provides a stable core API for small area estimation, including unit-level and area-level models, benchmarking, and design-consistent extensions.
While the surrounding ecosystem (additional features, performance backends, diagnostics, and documentation) continues to evolve, the primary modeling and estimation interfaces are expected to remain fairly stable.
📧 Feedback welcome: info@svylab.com
🐛 Report issues: GitHub Issues
Overview
svy-sae is the Python package for scalable Small Area Estimation (SAE). It brings rigorous area-level and unit-level modeling workflows to the Python scientific ecosystem, filling the gap for official statistics and applied research.
The package emphasizes practical estimation, reproducibility, and high performance, leveraging the svy ecosystem for design consistency and JAX for computational speed.
Key Capabilities
- Area-Level Models: Fay–Herriot (REML, ML) with Prasad–Rao and parametric bootstrap MSE.
- Unit-Level Models: Battese–Harter–Fuller EBLUP (REML, ML) with analytical and bootstrap MSE.
- Nonlinear Indicators: Molina–Rao EBP with Box–Cox, log, or no transformation — poverty headcount, poverty gap, severity, Gini, quintile share ratio, and custom indicator functions.
- Survey Integration: Direct estimates and design-based variances flow in from svy, which carries the survey design.
- Performance at Scale: JAX-compiled kernels make census-scale work practical — a full EBP bootstrap (10M-person census,
n_mc=200, 2,000 replicates) runs in hours on a laptop, a workload that takes days in single-threaded implementations. - Validation: Estimates are checked against R’s
saepackage on its own benchmark datasets; bootstrap MSE agrees within Monte Carlo error.
Documentation
- Getting Started — Installation and first steps.
- Tutorials — Step-by-step guides for real-world SAE problems.