本节摘要:本节合并讲三块工程化能力。其一是错误码与异常处理——400/401/403/404/413/429/500/529 各自的常见成因与是否可重试,以及如何用 SDK 的类型化异常类(
BadRequestError/RateLimitError/APIConnectionError等)按「最具体到最一般」的顺序捕获,而非字符串匹配错误消息。其二是批处理(batches)——把大批 Messages 请求异步提交,享 50% 折扣,最多 10 万请求/256MB,适合离线分类、批量摘要等不要求实时的场景。其三是文件 API(files)——上传一次大文件(最大 500MB),在多次请求里用file_id引用,免去重复上传。读完本节,你能正确处理 API 错误,并知道何时用批处理省一半成本、何时用文件 API 复用大文件。
内容来源:Anthropic 官方 Claude API 文档
shared/error-codes.md、python/claude-api/batches.md、python/claude-api/files-api.md(随 Claude Code 分发,从泄露素材库提取),汉化并套用体系化模板。
阅读完本节,你应当能够:
file_id 引用,避免重复上传。Claude API 用标准 HTTP 状态码表达错误。八个常见码及其特性:
| 码 | 错误类型 | 可重试 | 常见成因 |
|---|---|---|---|
| 400 | invalid_request_error |
否 | 请求格式或参数错误 |
| 401 | authentication_error |
否 | API key 无效或缺失 |
| 403 | permission_error |
否 | key 缺少权限 |
| 404 | not_found_error |
否 | 模型 ID 或端点错误 |
| 413 | request_too_large |
否 | 请求体超大小限制 |
| 429 | rate_limit_error |
是 | 请求过多(RPM/TPM/TPD 超限) |
| 500 | api_error |
是 | Anthropic 服务端问题 |
| 529 | overloaded_error |
是 | API 临时过载 |
关键判断:4xx(除 429)通常是请求本身的问题,重试无用,要改请求;429 与 5xx 是临时问题,重试有效。Anthropic 的 SDK 默认自动重试 429 与 5xx,默认指数退避、最多 2 次(max_retries=2)。
💡 限流时看响应头:429 时检查
retry-after(等多少秒)、x-ratelimit-limit-*(你的额度)、x-ratelimit-remaining-*(剩余额度),据此调节流速。
许多 400 是参数问题,而非格式问题:
temperature/top_p/top_k(已移除)、或 budget_tokens(改 adaptive thinking)、或 thinking: {type: "disabled"}(Fable 5 不接受)——都是 400,详见第 2 章 05 节。budget_tokens 必须 < max_tokens,否则 400。user;user/assistant 必须交替(连续同角色会被合并,但首条 assistant 会 400)。claude-sonnet-4.6(点号)而非 claude-sonnet-4-6(连字符)——404。⚠️ 401 的隐蔽成因:同时设了
ANTHROPIC_API_KEY与ANTHROPIC_AUTH_TOKEN时,SDK 会同时发两个头,API 拒绝。OAuth token 要走Authorization: Bearer,不能放进x-api-key。
每个 HTTP 状态码在 SDK 里对应一个类型化异常类。始终用这些类,而非检查错误消息字符串——字符串会变,类名稳定。
各语言异常类名(以 Python/TypeScript 为主):
| HTTP | Python / TypeScript | Go(单错误+分支) |
|---|---|---|
| 400 | BadRequestError |
StatusCode == 400 |
| 401 | AuthenticationError |
== 401 |
| 403 | PermissionDeniedError |
== 403 |
| 404 | NotFoundError |
== 404 |
| 429 | RateLimitError |
== 429 |
| ≥500 | InternalServerError |
default |
| 网络 | APIConnectionError |
(传输错误,非 *anthropic.Error) |
捕获顺序至关重要——从最具体的子类到基类,并为不同处理方式(可重试 vs 不可重试)分设子句:
import anthropic try: response = client.messages.create(...) except anthropic.NotFoundError as e: # 404 —— 模型 ID 错 ... except anthropic.RateLimitError as e: # 429 —— 退避重试 retry_after = int(e.response.headers.get("retry-after", "60")) except anthropic.APIStatusError as e: # 其它非 2xx print(e.status_code, e.message) except anthropic.APIConnectionError as e: # 网络失败(响应前) ...
⚠️ 顺序细节:TypeScript 里
APIConnectionError是APIError的子类,要先判APIConnectionError再判APIError;Python 里两者是平级,顺序不那么敏感。Go 只返回一个*anthropic.Error,用errors.As解包后按StatusCode分支。
.type 字段所有 APIStatusError 子类都有 .type 属性(Python/TS/Go/Java/Ruby/PHP 都有),返回 API 错误类型字符串(如 "rate_limit_error"、"overloaded_error"、"billing_error")。当 HTTP 码不够细时(比如 403 既可能是 permission_error 也可能是 billing_error),用 .type 做更细的分类:
except anthropic.APIStatusError as e: if e.type == "rate_limit_error": ... # 限流 elif e.type == "overloaded_error": ... # 过载
批处理 API(POST /v1/messages/batches)把一批 Messages 请求异步处理,价格只有标准价的 50%。适合不要求实时的离线任务:批量分类、大批摘要、数据集标注、评测打分。
关键事实:
典型流程是「创建 → 轮询 → 取结果」:
import anthropic, time from anthropic.types.message_create_params import MessageCreateParamsNonStreaming from anthropic.types.messages.batch_create_params import Request client = anthropic.Anthropic() # 1. 创建批次:每个请求带一个 custom_id 用于结果匹配 batch = client.messages.batches.create( requests=[ Request( custom_id="request-1", params=MessageCreateParamsNonStreaming( model="claude-haiku-4-5", max_tokens=50, messages=[{"role": "user", "content": "Classify as pos/neg/neu: great!"}], ), ), # ... 更多请求 ], ) # 2. 轮询直到 ended while True: batch = client.messages.batches.retrieve(batch.id) if batch.processing_status == "ended": break time.sleep(60) # 3. 取结果,按类型分流 for result in client.messages.batches.results(batch.id): match result.result.type: case "succeeded": msg = result.result.message print(f"[{result.custom_id}] done") case "errored": if result.result.error.type == "invalid_request": print(f"[{result.custom_id}] 请求错,改后重试") else: print(f"[{result.custom_id}] 服务端错,可重试") case "canceled": print(f"[{result.custom_id}] 已取消") case "expired": print(f"[{result.custom_id}] 过期,重新提交")
💡 批处理 + 缓存叠加:批处理里如果多个请求共享一段大 system prompt(如同一份分析文档),照样可以用
cache_control缓存——50% 折扣与缓存命中可以叠加,成本进一步下降。每个请求的custom_id用于把结果与输入对应,务必唯一。
⚠️ 批处理的限制:
max_tokens: 0的预热请求不能放进批次;流式也不适用(批处理本质异步)。需要实时响应的场景请用普通messages.create+ 流式。
文件 API(client.beta.files.*)让你上传一次大文件(最大 500 MB,每组织 100 GB 存储),在多次 Messages 请求里用 file_id 引用,避免每次重新上传同一份大文档。文件操作(上传/列表/删除)免费,内容用作消息输入时按输入 token 计费。
import anthropic client = anthropic.Anthropic() # 上传(beta) uploaded = client.beta.files.upload( file=("report.pdf", open("report.pdf", "rb"), "application/pdf"), ) print(f"File ID: {uploaded.id}") # 在多次请求里用 file_id 引用,免重复上传 response = client.beta.messages.create( model="claude-opus-4-8", max_tokens=16000, messages=[{ "role": "user", "content": [ {"type": "text", "text": "Summarize the key findings."}, { "type": "document", "source": {"type": "file", "file_id": uploaded.id}, "title": "Q4 Report", "citations": {"enabled": True}, }, ], }], betas=["files-api-2025-04-14"], )
适用场景:同一份大文档(PDF、长文本、图片)要在多个请求里反复用——比如对同一份报告问多个不同问题、同一张图做多种识别。上传一次拿 file_id,后续请求只传 ID,省上传带宽与时间。
⚠️ 平台限制:文件 API 是 beta,需
betas=["files-api-2025-04-14"](SDK 自动加头);在 Amazon Bedrock 与 Google Vertex AI 上不可用。文件持久存在直到删除,注意清理避免占满 100 GB 配额。
| 场景 | 选择 | 理由 |
|---|---|---|
| 实时聊天、agent 交互 | 普通 messages.create + 流式 |
要求低延迟 |
| 离线批量分类/摘要/评测 | 批处理(batches) | 享 50% 折扣,容忍小时级延迟 |
| 同一大文档反复提问 | 文件 API(files) | 上传一次,多次引用 file_id |
| 长重复前缀降本 | 提示缓存(prompt caching) | 命中降到 0.1 倍 |
| 高并发扇出 | 异步客户端 + 流式 + 缓存预热 | 先发一个等首 token 再扇出 |
四者(普通/批处理/文件/缓存)不是互斥的:批处理里可叠加缓存,文件 API 上传的文档内容也可被缓存。组合使用是规模化降本的关键。
BadRequestError/NotFoundError/RateLimitError/APIConnectionError 等类,按「最具体子类到基类」顺序捕获;用 .type 字段做更细分类(如区分 permission_error 与 billing_error)。temperature/budget_tokens、thinking: {type:"disabled"}(Fable 5)、模型 ID 拼错、首条非 user;401 注意别同时设两个 key 环境变量。custom_id)→ 轮询 ended → 取 succeeded/errored/canceled/expired;可与缓存叠加。file_id 引用;文件操作免费、内容按输入 token 计费;beta,需 files-api-2025-04-14,Bedrock/Vertex 不可用。第 2 章到此结束——你已经掌握了 Claude API 的六大核心能力(工具调用、流式、缓存、计数、模型、错误与批处理文件)。第 3 章我们深入两大主力 SDK(Python 与 TypeScript)的工程细节,再合并对照其余六种语言。