SillyTavern Custom API Setup: Claude, GPT & Gemini Without Credit Card
SillyTavern is a powerful AI chat interface that supports multiple LLM backends. However, setting it up with official APIs often requires international credit cards and can be blocked in certain regions. This guide shows you how to configure SillyTavern with a custom API relay that supports Claude, GPT, and Gemini through a single endpoint—with no credit card required.
Why Use a Custom API for SillyTavern?
Official API access from Anthropic, OpenAI, and Google presents several challenges:
- Credit card barriers: All three providers require international credit cards for billing
- Regional restrictions: Direct API access may be blocked or throttled in mainland China
- Complex billing: Managing separate accounts and billing for each provider
- No local payment: WeChat Pay and Alipay aren't accepted by official providers
A unified API relay solves these issues by providing one endpoint for all models, accepting local payment methods, and ensuring stable connectivity.
Prerequisites
Before you start, make sure you have:
- SillyTavern installed (download from GitHub)
- Node.js 18+ installed on your system
- An API key from a relay service (see setup section below)
Step-by-Step Setup Guide
1. Launch SillyTavern
Navigate to your SillyTavern directory and start the server:
cd SillyTavern
node server.js
Open your browser and go to http://localhost:8000
2. Configure API Connection
Click the API Connections icon (plug icon) in the top menu. Select Chat Completion as the API type, then choose your provider:
For Claude (Anthropic)
| Field | Value |
|---|---|
| API Type | Chat Completion (OpenAI-compatible) |
| API URL | https://api.aisafa.xyz/v1 |
| API Key | Your relay API key |
| Model | claude-sonnet-5 or claude-opus-48 |
For GPT Models
| Field | Value |
|---|---|
| API URL | https://api.aisafa.xyz/v1 |
| Model | gpt-56, gpt-4o, or gpt-55 |
For Gemini
| Field | Value |
|---|---|
| API URL | https://api.aisafa.xyz/v1 |
| Model | gemini-pro, gemini-ultra, or gemini-omni |
3. Test the Connection
Click Test Connection to verify your setup. If successful, you'll see a green checkmark and available models listed.
4. Advanced Settings (Optional)
Under Advanced Settings, you can configure:
- Temperature: 0.7-1.0 for creative chat, 0.3-0.5 for focused responses
- Max Response Length: 2048-4096 tokens recommended
- Streaming: Enable for real-time response generation
- Prompt Caching: If supported, reduces cost for repeated context
Obtaining an API Key Without Credit Card
If you don't have an API key yet, relay services like Safa API offer a straightforward solution. Here's what makes them practical:
- No credit card required: Sign up with email and pay via Alipay or WeChat Pay
- Unified endpoint: Access Claude, GPT, and Gemini through
https://api.aisafa.xyz/v1 - Competitive pricing: Often 10-30% cheaper than official rates with volume discounts
- Prompt caching support: Claude's prompt caching works out of the box, reducing costs for long conversations
- Mainland China friendly: Stable connectivity without VPN, optimized routes
To get started, visit aisafa.xyz/register, create an account, top up your balance with local payment, and generate an API key from the dashboard. Copy the key into SillyTavern's API Key field.
Troubleshooting Common Issues
Connection Failed Error
If you see "Connection failed" when testing:
- Verify your API key is correct (no extra spaces)
- Check that the API URL ends with
/v1 - Ensure your relay service account has sufficient balance
- Try disabling any VPN or proxy temporarily
Model Not Found
If a specific model isn't available:
- Verify the model name matches exactly (case-sensitive)
- Check your relay service's model list—not all relays support every model
- Try a different model in the same family (e.g.,
claude-sonnet-5instead ofclaude-opus-48)
Slow Response Times
For laggy responses:
- Enable streaming mode in advanced settings
- Reduce Max Response Length to 2048 tokens
- Check your internet connection stability
- Consider switching to a relay with better routing to your region
Cost Optimization Tips
Keep your SillyTavern API costs under control:
- Use prompt caching: For character cards and long system prompts, caching can cut costs by 90%
- Choose the right model: Use
claude-sonnet-5for general chat, reserveopusfor complex tasks - Limit context length: Set a reasonable conversation history limit (20-30 messages)
- Monitor usage: Check your relay dashboard regularly to track spending
Security Considerations
When using custom APIs with SillyTavern:
- Never share your API key: Treat it like a password
- Use HTTPS only: Ensure the relay endpoint uses SSL/TLS encryption
- Review privacy policy: Understand how your chat data is handled by the relay
- Regenerate keys periodically: Create new API keys every few months
Reputable relay services don't log conversation content and only store metadata for billing purposes. Always read the privacy policy before committing sensitive conversations.
Frequently Asked Questions
Can I use the same API key for Claude, GPT, and Gemini?
Yes, with a unified API relay like Safa API, one key works for all supported models. Just change the model name in SillyTavern's settings to switch between providers.
Is there a free tier or trial available?
Most relay services offer new user credits (typically $1-5 equivalent) to test the service. This is usually enough for 50-200 messages depending on the model.
What happens if I run out of balance mid-conversation?
You'll receive an error message indicating insufficient balance. Simply top up your account and resume—your conversation history in SillyTavern remains intact.
官方直连 · 一个接口接入 Claude / GPT / Gemini · 7×24 稳定
免费注册试用 →