Convert CSV to JSON using only Python standard library modules.
The Logic Behind the Conversion
The key tool is csv.DictReader. Instead of returning a list of strings for each row, it treats the first line of the CSV as keys for a dictionary. This eliminates the need to manually track column indices or map headers to values.
Step-by-Step Implementation
How to extract CSV data with DictReader
- Extracting the Data
We open the file with newline="" as the Python documentation recommends to avoid line-ending problems across operating systems.
import csv
def read_csv(csv_path):
with open(csv_path, newline="", encoding="utf-8") as f:
# DictReader automatically maps the header row to dictionary keys
reader = csv.DictReader(f)
return list(reader)
Wrapping the reader in list() loads the whole dataset into memory. This works for small to medium files; for a 2GB CSV, you would iterate through the reader and write to the JSON file line by line.
How to serialize rows to JSON format
- Serializing to JSON
The json.dump method converts a Python list of dictionaries into a valid JSON array.
import json
def write_json(rows, json_path):
with open(json_path, "w", encoding="utf-8") as f:
# indent=2 ensures the output isn't one giant, unreadable line
json.dump(rows, f, indent=2)
How to link the conversion pipeline
- The Final Pipeline
Linking the two functions lets us count the records processed, a useful sanity check for data pipelines.
import csv
import json
def read_csv(csv_path):
with open(csv_path, newline="", encoding="utf-8") as f:
reader = csv.DictReader(f)
return list(reader)
def write_json(rows, json_path):
with open(json_path, "w", encoding="utf-8") as f:
json.dump(rows, f, indent=2)
def csv_to_json(csv_path, json_path):
rows = read_csv(csv_path)
write_json(rows, json_path)
return len(rows)
if __name__ == "__main__":
# Example usage
try:
count = csv_to_json("contacts.csv", "contacts.json")
print(f"Successfully converted {count} rows to contacts.json")
except FileNotFoundError:
print("Error: The source CSV file was not found.")
Real-World Verification
If you have a contacts.csv file with this content:
name,email,city
Alice,[email protected],Austin
Bob,[email protected],Denver
The resulting contacts.json will look like this:
[
{
"name": "Alice",
"email": "[email protected]",
"city": "Austin"
},
{
"name": "Bob",
"email": "[email protected]",
"city": "Denver"
}
]
Performance and Limitations
Why avoid heavy libraries for conversion
For a deeper look at when to prefer this over a heavy library:
- Memory footprint: this approach uses far less RAM than Pandas because it does not create a DataFrame object.
- Speed: for files under 50MB the difference is negligible.
- Edge cases: Python's csv module handles quoted fields such as "New York, NY" automatically, so commas inside data will not break columns.
If you are building a lightweight AI workflow or an LLM agent that must preprocess data before feeding it into a prompt, keeping dependencies at zero speeds up deployment.
All Replies (2)
Want a live back-and-forth? Join the global AI chat room — login to talk.
I'm curious if this handles nested structures or just flat arrays—like nested dictionaries or arrays within rows. The csv.DictReader approach you're using automatically maps the header row to dictionary keys, which means it will naturally handle nested data if the CSV itself contains nested structures encoded as comma-separated values (e.g., "{"key": "value"}" in a cell). For true multi-level data, you'd need to preprocess the CSV or use a different format like JSON.
UTF-8 often fails with weird symbols, but the
csvmodule’sDictReadercan help—it automatically maps headers to keys while handling encoding issues gracefully when you open the file withnewline=""and specifyencoding="utf-8". For stubborn files, tryencoding="utf-8-sig"orencoding="cp1252"as fallbacks. TheDictReaderapproach also avoids manual column tracking, making it cleaner for complex CSVs.