AI资讯大全AIPPXP.CN搜索 ↗
工具介绍 · 免费好用

FreeLLMAPI: 34 free big model manufacturers twisted into one OpenAI interface, about 7.4 billion tokens per month freeFreeLLMAPI:把 34 家免费大模型厂商拧成一个 OpenAI 接口,每月约 74 亿 token 免费用

Every AI vendor is giving away free credits - millions of tokens a month, thousands of requests a day. A single look is a toy, but stacked is the available inference capacity of about 7.4 billion tokens per month, covering 474 model families and 635 free endpoints. The problem is that manually stacking is too painful: 34 SDKs, 34 throttling rules, 34 places that will fail. FreeLLMAPI collapses all of this into a local OpenAI-compatible endpoint: a/v1 at the same time says OpenAI, Anthropic, Gemini, Ollama four dialects, automatically picks the fastest model, cuts down the one after being throttled, encrypts the key locally, and helps you stay within the free quota of each house according to the key statistical usage.

每个 AI 厂商都在送免费额度——几百万 token 一个月、几千次请求一天。单看一家是玩具,但叠起来是每月约 74 亿 token 的可用推理能力,覆盖 474 个模型家族、635 个免费端点。问题是手动叠太痛苦:34 套 SDK、34 套限流规则、34 个会失败的地方。FreeLLMAPI 把这一切收成一个本地 OpenAI 兼容端点:一个 /v1 同时说 OpenAI、Anthropic、Gemini、Ollama 四种方言,自动挑最快的模型,被限流就切下一家,密钥加密存在本地,还按 key 统计用量帮你守在每家免费额度内。

2026-10-10 更新 · 免费
FreeLLMAPI:把 34 家免费大模型厂商拧成一个 OpenAI 接口,每月约 74 亿 token 免费用

What Problems Does It Solve?它解决什么问题

The pain point of the free quota is "scattered": Google, Groq, Cerebras, Mistral, OpenRouter, Cloudflare, Cohere, Z.ai (Smart Spectrum), NVIDIA, HuggingFace, but to connect at the same time in your code, the workload is much larger than the paid API.

And the free ecosystem is changing every week: manufacturers launch new models, offline old models, and quietly change quotas. The router of FreeLLMAPI will pull the signed model directory by itself, and it can keep up with changes without git pull, eliminating long-term maintenance.

After the aggregation, there is an additional benefit: the single house will automatically retry the next house (with cooling and key rotation), which is equivalent to turning "34 unstable small faucets" into a stable water pipe.

The most practical meaning for individual developers is that it used to take several accounts and several adaptation layers to get the computing power, but now it only needs to point to one localhost endpoint.

中文

免费额度的痛点是「分散」:Google、Groq、Cerebras、Mistral、OpenRouter、Cloudflare、Cohere、Z.ai(智谱)、NVIDIA、HuggingFace 各给一份,但要在你的代码里同时对接,工作量比接一家付费 API 大得多。

而且免费生态每周都在变:厂商上线新模型、下线旧模型、悄悄改配额。FreeLLMAPI 的 router 会自己拉取签名过的模型目录,无需 git pull 就能跟上变化,省去了长期维护。

聚合之后还有额外收益:单家限流了会自动重试下一家(带冷却与密钥轮换),等于把「34 个不稳定的小水龙头」变成一个稳定的水管。

对个人开发者最实际的意义是:以前需要开好几个账号、写好几个适配层才能拿到的算力,现在只需要指向一个 localhost 端点。

它解决什么问题

Which vendors and clients are supported支持哪些厂商与客户端

Built-in free vendors include Google, Groq, Cerebras, OpenCode Zen, Mistral, OpenRouter, Cloudflare, Cohere, Z.ai (Smart Spectrum), NVIDIA, HuggingFace, and ModelScope (Qwen3/DeepSeek V4/GLM-5) that requires Alibaba Cloud binding, and more than 22 free vendors not listed in the main table.

A custom vendor is also supported: point the chat/embedding/image/audio model to any OpenAI-compatible endpoint, including llama.cpp, LM Studio, vLLM, or local Ollama.

The client side covers mainstream coding agents: Claude Code, Codex CLI, Gemini CLI, Cursor, Cline, Roo Code, OpenCode, Aider, and Continue, Goose, Qwen Code, Kilo Code, Crush, Zed, JetBrains AI, OpenClaw, Hermes Agent, etc., as long as they support OpenAI compatible, Anthropic, Gemini or Ollama protocols.

Most agents can complete the configuration with one command - it will read your current directory, backup the existing configuration and then merge it into it, leaving your original settings unchanged.

中文

内置免费厂商包括 Google、Groq、Cerebras、OpenCode Zen、Mistral、OpenRouter、Cloudflare、Cohere、Z.ai(智谱)、NVIDIA、HuggingFace,以及需要阿里云绑定才能用的 ModelScope(Qwen3 / DeepSeek V4 / GLM-5),另有 22 家以上未在主表列出的免费厂商。

还支持一个自定义厂商:把 chat / embedding / image / audio 模型指向任意 OpenAI 兼容端点,包括 llama.cpp、LM Studio、vLLM 或本地 Ollama。

客户端侧覆盖主流编码智能体:Claude Code、Codex CLI、Gemini CLI、Cursor、Cline、Roo Code、OpenCode、Aider,以及 Continue、Goose、Qwen Code、Kilo Code、Crush、Zed、JetBrains AI、OpenClaw、Hermes Agent 等,只要支持 OpenAI 兼容、Anthropic、Gemini 或 Ollama 协议就能接。

大多数智能体可以一条命令完成配置——它会读取你当前的目录、备份已有配置再合并进去,不动你原来的设置。

Core Ability核心能力

Full coverage of OpenAI-style interfaces: chat/completions, responses (required by Codex CLI), completions (editor gray completion), images/generations, videos/generations, audio/speech, audio/transcriptions, embeddings, models, all support streaming and non-streaming.

Anthropic Messages API:/v1/messages speaks directly to Anthropic's online format, so Claude Code and the official Anthropic SDK can run directly on your heap of free credits.

Native Gemini and Ollama interface: The Gemini CLI can use/v1beta, and the optional Ollama simulation layer can match local model clients such as Zed and JetBrains.

Fusion multi-model synthesis: Request a virtual fusion model, the router will send the question to a set of complementary free models in parallel, and then a referee model will synthesize an answer.

Six routing strategies: Sort the model by real-time speed/capability/stability, and automatically downgrade to the next one with cooling and key rotation at 429 or 5xx.

The tool call and structured output are transparently passed back and forth in OpenAI format (the tool call in plain text format will be recovered as a real tool_calls), and the response_format, seed, logprobs and other sampling parameters are transparently passed one by one according to the manufacturer's ability.

中文

全覆盖的 OpenAI 风格接口:chat/completions、responses(Codex CLI 需要)、completions(编辑器灰字补全)、images/generations、videos/generations、audio/speech、audio/transcriptions、embeddings、models,全部支持流式与非流式。

Anthropic Messages API:/v1/messages 直接说 Anthropic 的线上格式,所以 Claude Code 和官方 Anthropic SDK 可以直接跑在你这堆免费额度上。

原生 Gemini 与 Ollama 接口:Gemini CLI 可以用 /v1beta,可选的 Ollama 模拟层能对上 Zed、JetBrains 等本地模型客户端。

Fusion 多模型合成:请求一个虚拟的 fusion 模型,router 会把问题并行发给一组互补的免费模型,再由一个裁判模型综合出一份答案。

六种路由策略:按实时速度 / 能力 / 稳定性给模型打分排序,429 或 5xx 时带冷却与密钥轮换自动降级到下一家。

工具调用与结构化输出按 OpenAI 格式往返透传(纯文本形式的工具调用会被救回成真正的 tool_calls),response_format、seed、logprobs 等采样参数按厂商能力逐个透传。

Download and Installation Instructions下载与安装说明

The following is the complete source code package of the server-side repository. It is recommended to use Docker self-hosting. The repository comes with compose configuration and installation documents.

After starting the local service (default port 3001), first fill in the API Key of each free vendor on the Keys page, and the router will automatically pull the available model directory.

Then configure your client with a unified local address and key, for example: npx freellmapi setup-claude --url http://localhost: 3001 --api-key unified key, which automatically registers the configuration of Claude Code.

In addition to the source code, the official macOS/Windows desktop, Docker image, Google Play and App Store apps are also available; the free version gets a monthly model snapshot (a model enters the free version 30 days after being added to the real-time directory), and the paid version ($19/year) is synchronized on the same day.

Note: The difference between the free version and the paid version is only in the update timeliness of the model directory, and the routing and aggregation capabilities are consistent; there is also a disclaimer on the terms of use of each manufacturer in the document, which is recommended to be read first and then used on a large scale.

中文

下面提供的是服务端仓库的完整源码包,推荐用 Docker 自托管,仓库里带了 compose 配置与安装文档。

启动后是本地服务(默认端口 3001),先在 Keys 页面把各家免费厂商的 API Key 填进去,router 会自动拉取可用模型目录。

接着用统一的本地地址与密钥配置你的客户端,例如:npx freellmapi setup-claude --url http://localhost:3001 --api-key 统一密钥,它会自动登记 Claude Code 的配置。

除了源码,官方还提供 macOS / Windows 桌面端、Docker 镜像、Google Play 与 App Store 应用;免费版拿到的是月度模型快照(某个模型加入实时目录 30 天后才进入免费版),付费版(19 美元 / 年)当天同步。

注意:免费版与付费版的差异只在模型目录的更新时效,路由与聚合能力一致;文档里另有关于各厂商使用条款的免责声明,建议先读再大规模使用。

Use reminder使用提醒

This site organizes resources from GitHub open source repository tashfeenahmed/freellmapi (mit license), for learning and technical exchanges only. When aggregating free credits, please abide by the terms of service and rate limits of each manufacturer, and do not use the free credits for commercial resale. Although the keys are encrypted and stored locally, it is still recommended to use a separate, revocable Key at any time. In this article, "about 7.4 billion tokens per month" is the self-explanatory caliber of the project README, and the actual available amount changes with the manufacturer's policy, subject to the latest official instructions.

中文

本站整理的资源来自 GitHub 开源仓库 tashfeenahmed/freellmapi(MIT 许可),仅供学习与技术交流。聚合免费额度时请遵守各家厂商的服务条款与速率限制,不要把免费额度用于商业转售;密钥虽然加密存储在本地,但仍建议使用独立的、可随时吊销的 Key。文中「每月约 74 亿 token」为项目 README 的自述口径,实际可用量随厂商政策变动,以官方最新说明为准。

资源下载 · Download

⬇ Download · 点击下载:FreeLLMAPI 源码包(main 分支)(约 6.1 MB)

来自 GitHub 开源仓库 tashfeenahmed/freellmapi(MIT 许可),含服务端路由、桌面端、Docker 部署配置与完整英文 / 中文文档

0阅读0 条评论

阅读与点赞数据保存在你的浏览器本地,欢迎留下你的想法。

评论 文明发言,让讨论更有价值

正能量公益广告今日正能量学一点,用一点;今天种下的种子,会长成明天的能力。去免费下载专区 →广告