第 9 章 文件与数据格式
9.1 读写文本文件
open(path, mode, encoding) 是文件操作入口,永远搭配 with 使用(自动关闭):
with open("notes.txt", "w", encoding="utf-8") as f:
f.write("第一行\n")
f.writelines(["第二行\n", "第三行\n"])
with open("notes.txt", encoding="utf-8") as f:
content = f.read()
with open("notes.txt", encoding="utf-8") as f:
for line in f:
print(line.rstrip())
编码警告:Windows 下不写 encoding 会用系统默认编码(可能是 GBK),跨平台读写中文文件务必显式 encoding="utf-8"。
常用模式:
| 模式 | 含义 |
|---|
"r" | 只读(默认) |
"w" | 写入(清空原内容) |
"a" | 追加 |
"r+" | 读写 |
"rb" / "wb" | 二进制读写 |
pathlib 写法(推荐)
from pathlib import Path
p = Path("notes.txt")
p.write_text("hello\n", encoding="utf-8")
p.read_text(encoding="utf-8")
p.read_bytes()
9.2 大文件与缓冲
逐行迭代是处理大文件的标准姿势,不要 read() 整个 10GB 文件;
需要"读一段处理一段"用 f.read(size) 或 f.readline();
f.tell() 报告当前位置,f.seek(0) 回到开头(文本模式下 seek 只支持 0)。
with open("huge.log", encoding="utf-8") as f:
errors = [line for line in f if "ERROR" in line]
9.3 JSON:最通用的数据交换格式
import json
data = {"name": "Alice", "age": 25, "tags": ["admin", "dev"]}
s = json.dumps(data, ensure_ascii=False, indent=2)
obj = json.loads(s)
with open("data.json", "w", encoding="utf-8") as f:
json.dump(data, f, ensure_ascii=False, indent=2)
with open("data.json", encoding="utf-8") as f:
data2 = json.load(f)
类型对应表:dict↔object、list↔array、str↔string、int/float↔number、True/False↔true/false、None↔null。元组会被转成数组(往返后变 list)。
自定义类型默认不能序列化,需要 default= 指定转换函数;解析失败抛 json.JSONDecodeError(是 ValueError 的子类)。
9.4 CSV:表格数据
import csv
with open("users.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerow(["name", "age"])
writer.writerow(["Alice", 25])
writer.writerows([["Bob", 30], ["Carol", 28]])
with open("users.csv", newline="", encoding="utf-8") as f:
for row in csv.reader(f):
print(row)
with open("users.csv", newline="", encoding="utf-8") as f:
for row in csv.DictReader(f):
print(row["name"], row["age"])
注意 newline="" 是 csv 模块的官方要求(防止 Windows 下出现空行)。中文 Excel 常见坑:用 utf-8-sig 编码写,Excel 才能正确识别 BOM。真正的电子表格(.xlsx)用第三方库 openpyxl 或 pandas。
9.5 其他实用格式
import configparser
cfg = configparser.ConfigParser()
cfg.read("app.ini", encoding="utf-8")
cfg["db"]["host"]
import pickle
with open("obj.pkl", "wb") as f:
pickle.dump(my_object, f)
with open("obj.pkl", "rb") as f:
obj = pickle.load(f)
import sqlite3
conn = sqlite3.connect("app.db")
conn.execute("CREATE TABLE IF NOT EXISTS users (name TEXT, age INT)")
conn.execute("INSERT INTO users VALUES (?, ?)", ("Alice", 25))
conn.commit()
rows = conn.execute("SELECT * FROM users WHERE age > ?", (18,)).fetchall()
conn.close()
9.6 os / pathlib:文件系统操作
统一推荐 pathlib:
from pathlib import Path
import shutil
Path("a/b/c").mkdir(parents=True, exist_ok=True)
p = Path("old.txt")
p.rename("new.txt")
p.unlink(missing_ok=True)
shutil.copy("a.txt", "b.txt")
shutil.copytree("src", "dst")
shutil.rmtree("tmp_dir")
Path(".").glob("*.py")
Path(".").rglob("*.py")
[x for x in Path(".").iterdir() if x.is_file()]
p.suffix, p.stem, p.parent
p.exists(), p.is_file(), p.is_dir()
9.7 临时文件与环境信息
import tempfile
with tempfile.TemporaryDirectory() as tmpdir:
work = tempfile.Path(tmpdir) if False else None
import os
os.getenv("HOME")
os.environ["MY_FLAG"] = "1"
9.8 本章小结
文件操作永远 with open(...);编码显式 utf-8;
大文件逐行迭代;小文件用 pathlib 的 read_text/write_text;
JSON 记住 dumps/loads(字符串)与 dump/load(文件)两对;
CSV 要 newline="";pickle 不碰不可信来源;
路径操作全面转向 pathlib。
下一章:迭代器、生成器与函数式工具。