Converting Shapefiles to GeoPackage with GeoPandas
Load the shapefile into a GeoDataFrame with gpd.read_file(), then write it to a GeoPackage with gdf.to_file(driver="GPKG") — for production workloads, always pass engine="pyogrio" to get vectorized C-level I/O and full GDAL 3.x compliance.
This page is part of the GeoPandas & GeoPackage Integration guide, which covers the full read/write pipeline from driver selection through transaction control.
Why This Matters
Field GIS teams routinely receive data as Shapefiles — a format that splits a single dataset across at least four companion files (.shp, .shx, .dbf, .prj). A single missing sidecar corrupts the entire dataset silently. For offline-first mobile deployments and automated ETL pipelines, this fragility is unacceptable. Converting to GeoPackage format collapses those four or more files into one self-contained SQLite database, adds an automatic R-tree spatial index, and raises the field-name and file-size limits that routinely break Shapefile-based workflows.
Conversion is not only a change of container. A Shapefile carries constraints that a GeoPackage does not, and the conversion is the moment those constraints stop applying — which means it is also the moment truncated field names and coerced types become permanent unless they are repaired here.
Prerequisites
- Python 3.9 or higher
geopandas >= 1.0(switchespyogrioin as the default engine)pyogrio >= 0.7.2(GDAL 3.4+ bindings)shapely >= 2.0- GDAL 3.4+ with GeoPackage driver registered (
GPKGinpyogrio.list_drivers()) - All four Shapefile components present in the same directory (
.shp,.shx,.dbf,.prj)
Verify driver availability before running batch jobs:
import pyogrio
drivers = pyogrio.list_drivers()
assert "GPKG" in drivers, "GeoPackage GDAL driver not registered"
Primary Method
The conversion is a two-call operation: read_file then to_file. The function below adds encoding normalization, SQLite-safe layer naming, and structured logging — the minimum production harness for any batch ETL or serverless spatial pipeline.
import geopandas as gpd
import logging
from pathlib import Path
logging.basicConfig(level=logging.INFO, format="%(levelname)s: %(message)s")
def convert_shp_to_geopackage(
input_shp: str | Path,
output_gpkg: str | Path,
layer_name: str | None = None,
engine: str = "pyogrio",
encoding: str = "utf-8",
) -> None:
"""
Convert a Shapefile to a GeoPackage layer.
Args:
input_shp: Path to the .shp file (companion files must be co-located).
output_gpkg: Destination .gpkg path (created or appended to).
layer_name: SQLite layer identifier; defaults to the stem of input_shp.
engine: GDAL I/O backend — "pyogrio" (default) or "fiona" (legacy).
encoding: Attribute encoding of the .dbf — usually "utf-8" or "latin1".
"""
input_path = Path(input_shp)
output_path = Path(output_gpkg)
if not input_path.exists():
raise FileNotFoundError(f"Shapefile not found: {input_path}")
# 1. Load with explicit engine and .dbf encoding
gdf = gpd.read_file(input_path, engine=engine, encoding=encoding)
# 2. Validate CRS — a missing .prj produces an undefined CRS that will
# corrupt the gpkg_spatial_ref_sys metadata table downstream.
if gdf.crs is None:
raise ValueError(
f"No CRS detected for {input_path.name}. "
"Assign one with gdf.set_crs('EPSG:4326', inplace=True) before export."
)
# 3. Sanitize layer name to SQLite identifier rules
if layer_name is None:
layer_name = input_path.stem
layer_name = layer_name.replace(" ", "_").replace("-", "_").lower()
# 4. Write to GeoPackage; GDAL auto-builds the R-tree spatial index
gdf.to_file(
output_path,
driver="GPKG",
layer=layer_name,
engine=engine,
index=False, # omit the DataFrame row index as a column
)
logging.info(
"Converted %s → %s:%s (%d features)",
input_path.name, output_path.name, layer_name, len(gdf),
)
Step-by-step Walkthrough
1. Confirm all sidecar files are present.
The .shp path you pass to read_file must have its companions in the same directory:
from pathlib import Path
shp = Path("data/field_survey.shp")
for ext in [".shx", ".dbf", ".prj"]:
if not shp.with_suffix(ext).exists():
raise FileNotFoundError(f"Missing companion file: {shp.with_suffix(ext)}")
2. Read with pyogrio and inspect geometry.
pyogrio reads directly into NumPy/Arrow-backed arrays, bypassing Python object overhead. Inspect the result before writing:
gdf = gpd.read_file(shp, engine="pyogrio")
print(gdf.crs) # e.g. EPSG:4326
print(gdf.geom_type.value_counts()) # verify geometry types
print(gdf.shape) # (n_features, n_columns)
3. Assign or reproject the coordinate reference system if needed.
The GeoPackage Specification Deep Dive requires that every feature table entry in gpkg_spatial_ref_sys resolves to a valid EPSG or OGC code. Reproject before export, not after:
# If CRS is already defined but you need WGS 84:
if gdf.crs.to_epsg() != 4326:
gdf = gdf.to_crs("EPSG:4326")
# If .prj was missing (CRS is None):
gdf = gdf.set_crs("EPSG:4326") # assert, do not reproject
4. Write to GeoPackage.
Pass mode="a" to append a layer to an existing file, or omit it (default "w") to create or overwrite the target:
gdf.to_file(
"data/survey_archive.gpkg",
driver="GPKG",
layer="field_survey_2024",
engine="pyogrio",
index=False,
)
5. Scale with batch conversion.
For a directory of Shapefiles, wrap the single-file function in a loop. Use logging.error rather than raising on individual failures so one bad file does not abort the entire run:
def batch_convert_directory(shp_dir: str, gpkg_dir: str) -> None:
out = Path(gpkg_dir)
out.mkdir(parents=True, exist_ok=True)
for shp in Path(shp_dir).glob("*.shp"):
try:
convert_shp_to_geopackage(shp, out / f"{shp.stem}.gpkg")
except Exception as exc:
logging.error("Failed %s: %s", shp.name, exc)
6. Handle large files with Arrow batch reads.
GeoPandas loads the full dataset into RAM. For multi-gigabyte Shapefiles, use pyogrio.open_arrow() to stream record batches, or page through the data:
import pyogrio
# Page through 50,000 features at a time
for offset in range(0, total_features, 50_000):
batch = gpd.read_file(
shp,
engine="pyogrio",
skip_features=offset,
max_features=50_000,
)
mode = "w" if offset == 0 else "a"
batch.to_file(output_gpkg, driver="GPKG", layer=layer_name, mode=mode, index=False)
Verification
After conversion, confirm that GDAL registered the layer and that the R-tree spatial index is present. The quickest check is ogrinfo:
ogrinfo -al -so data/survey_archive.gpkg field_survey_2024
Expected output includes Feature Count, Extent, and Layer SRS WKT. Absence of any means the write failed silently.
From Python, verify the gpkg_contents metadata table and feature count:
import sqlite3
with sqlite3.connect("data/survey_archive.gpkg") as con:
# Verify the layer appears in the GeoPackage contents registry
row = con.execute(
"SELECT table_name, data_type, srs_id FROM gpkg_contents WHERE table_name = ?",
("field_survey_2024",),
).fetchone()
assert row is not None, "Layer missing from gpkg_contents"
print("Layer:", row[0], "| Type:", row[1], "| SRS:", row[2])
# Confirm feature count matches source
count = con.execute("SELECT COUNT(*) FROM field_survey_2024").fetchone()[0]
print(f"Features written: {count}")
# Confirm R-tree index table exists
rtree = con.execute(
"SELECT name FROM sqlite_master WHERE type='table' AND name LIKE 'rtree_%'",
).fetchall()
print("Spatial index tables:", rtree)
Alternative Approaches
CLI via ogr2ogr (no Python dependency).
When you need a quick one-off conversion or are working in a shell script, GDAL’s ogr2ogr is faster to invoke than a Python script:
ogr2ogr \
-f GPKG data/survey_archive.gpkg \
data/field_survey.shp \
-nln field_survey_2024 \
-t_srs EPSG:4326 \
-progress
The -nln flag sets the layer name; -t_srs reprojects on the fly. This approach integrates cleanly with subprocess.run() in Python orchestrators.
Multi-layer GeoPackage (consolidation).
GeoPackage supports multiple feature tables in a single file, which Shapefiles cannot. Consolidate a directory of related Shapefiles into one .gpkg by appending with mode="a":
output_gpkg = Path("data/all_layers.gpkg")
for shp in Path("data/shapefiles").glob("*.shp"):
gdf = gpd.read_file(shp, engine="pyogrio")
mode = "w" if not output_gpkg.exists() else "a"
gdf.to_file(output_gpkg, driver="GPKG", layer=shp.stem, mode=mode, index=False)
This is particularly useful when preparing offline bundles for managing large spatial datasets in memory, since a single open connection serves all layers.
Troubleshooting
DriverError: Could not open datasource
Cause: One or more companion sidecar files (.shx, .dbf) is missing or has a different capitalisation from the .shp.
Fix: Run ls -1 data/field_survey.* to confirm all four files are present. Use ogrinfo data/field_survey.shp to check integrity before calling read_file.
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9
Cause: The .dbf attribute table was written by a legacy Windows application using cp1252 or latin1 encoding. GeoPandas defaults to utf-8.
Fix: Pass the correct encoding explicitly:
gdf = gpd.read_file(shp, engine="pyogrio", encoding="latin1")
If you are unsure of the encoding, try chardet:
import chardet, pathlib
raw = pathlib.Path(shp).with_suffix(".dbf").read_bytes()
print(chardet.detect(raw[:50_000]))
CRSWarning: CRS not set for some of the concatenated objects or srs_id = -1 in gpkg_spatial_ref_sys
Cause: The .prj file is missing, so the CRS is None. GeoPandas writes -1 as the SRS identifier, which violates the OGC GeoPackage Specification and breaks downstream spatial joins.
Fix: Assign the authoritative CRS before export. Use set_crs (assign without reprojection) if you know the correct EPSG code, or to_crs to reproject:
gdf = gdf.set_crs("EPSG:32632") # e.g. UTM Zone 32N
sqlite3.OperationalError: table field_survey already exists
Cause: The target .gpkg already contains a layer with the same name and mode="w" was not passed.
Fix: Either choose a unique layer name, pass mode="a" to append features to the existing table, or explicitly overwrite the entire database with mode="w".
Frequently Asked Questions
Should I convert with GeoPandas or with ogr2ogr?
If the conversion is purely a format change, ogr2ogr is the better tool: it streams, it needs no Python process, and its layer creation options cover indexing and geometry typing directly. GeoPandas earns its place when the conversion also transforms — renaming truncated fields, parsing dates that a Shapefile stored as text, joining an attribute table, filtering rows. Choosing the frame-based path for a straight conversion pays memory for nothing.
How do I recover field names that the Shapefile truncated?
Only from outside the Shapefile, because the truncation happened when it was written and the original names are simply not in the file. If the source came with a data dictionary, apply it as a rename map during conversion; if it did not, the ten-character names are what you have. Recording the rename map alongside the output is worth doing either way — the next conversion from the same source will need it too.
What encoding should I assume for the attribute table?
Assume nothing and state it. Shapefile attribute encoding is recorded, if at all, in an optional .cpg sidecar, and files without one are commonly Latin-1 or a Windows code page rather than UTF-8. Passing encoding= explicitly to the read turns a silent mojibake problem into a correct result or a loud error, and either is better than discovering the issue after the data is published.
Related
- GeoPandas & GeoPackage Integration — parent guide covering driver selection, transaction control, and spatial index management
- GeoPackage Specification Deep Dive — mandatory
gpkg_contentsandgpkg_spatial_ref_systable schemas that govern every conversion - Managing Spatial Reference Systems in SQLite — how to resolve and assign EPSG codes before writing GeoPackage layers
- Managing Large Spatial Datasets in Memory — chunked and Arrow-backed read patterns for multi-gigabyte Shapefiles
- Python Integration & Database Workflows — top-level guide to the full spatial Python stack