Author: Rafi Kimhi
Date: April 2026
Test Environment: AWS Bedrock + Bedrock Mantle (LiteLLM Gateway)
This analysis compares five AI models integrated into the re:Invent 2025 Knowledge Base Assistant, evaluating performance characteristics based on real-world usage data from production testing. Updated May 2026 to include our custom fine-tuned Llama 3.2 3B model.
| Model | Requests | Avg Latency | Avg Tokens | Avg Cost/Request | Cost per 1K Tokens |
|---|---|---|---|---|---|
| 🦙 Llama 3.2 3B Fine-tuned | 3 | 2,847 ms | 1,456 | $0.0008* | $0.55* |
| Claude Opus 4.6 | 6 | 10,471 ms | 1,948 | $0.0598 | $30.69 |
| GPT-OSS 20B | 3 | 2,592 ms | 1,969 | $0.0007 | $0.35 |
| Qwen3 32B | 3 | 1,313 ms | 1,582 | $0.0006 | $0.41 |
| DeepSeek V3.1 | 3 | 1,726 ms | 1,477 | $0.0010 | $0.69 |
* Llama fine-tuned cost based on SageMaker endpoint inference time (~$0.00028/sec). Excludes endpoint idle time (~$1.01/hr).
| Model | Monthly Cost |
|---|---|
| Claude Opus 4.6 | $1,794 |
| DeepSeek V3.1 | $30 |
| 🦙 Llama 3.2 Fine-tuned* | $24 + endpoint |
| GPT-OSS 20B | $21 |
| Qwen3 32B | $18 |
SageMaker Endpoint (Custom)
Best For:
Avoid For:
LoRA fine-tuned on SageMaker
Native Bedrock Converse API
Best For:
Avoid For:
via Bedrock Mantle
Best For:
Avoid For:
via Bedrock Mantle
Best For:
Avoid For:
via Bedrock Mantle
Best For:
Avoid For:
| Dimension | Claude Opus 4.6 | External Models |
|---|---|---|
| Reasoning Depth | Deep, multi-step | Adequate for RAG |
| Source Citation | Comprehensive | Variable |
| Response Structure | Well-organized | Good |
| Hallucination Risk | Lower | Slightly higher |
| Cost Efficiency | Poor | Excellent |
| Model | Avg Input | Avg Output | Output Ratio |
|---|---|---|---|
| Claude Opus 4.6 | 1,441 | 507 | 35% |
| GPT-OSS 20B | 1,249 | 720 | 58% |
| Qwen3 32B | 1,302 | 280 | 21% |
| DeepSeek V3.1 | 1,205 | 272 | 23% |
Simple queries (what, when, who) → Qwen3 32B
Domain Q&A (re:Invent specific) → 🦙 Llama 3.2 Fine-tuned
Moderate queries (how, explain) → GPT-OSS 20B
Complex queries (analyze, compare) → Claude Opus 4.6
At 1,000 daily requests with 80/15/5 split:
Key Insight: Fine-tuning focused on teaching response style (formatting, tone, structure), not factual knowledge. Factual content comes from RAG via Bedrock Knowledge Base at runtime.
User: rkimhi+sup1@amazon.com | Test Period: April 2026 | Total Requests: 15
| Model | Latency (ms) | Tokens | Cost ($) |
|---|---|---|---|
| Claude Opus 4.6 | 10,146 | 2,424 | 0.0671 |
| Claude Opus 4.6 | 9,698 | 1,887 | 0.0598 |
| Claude Opus 4.6 | 10,709 | 1,747 | 0.0549 |
| Claude Opus 4.6 | 11,851 | 1,983 | 0.0604 |
| Claude Opus 4.6 | 9,841 | 1,908 | 0.0614 |
| Claude Opus 4.6 | 10,582 | 1,740 | 0.0544 |
| GPT-OSS 20B | 2,436 | 1,952 | 0.0006 |
| GPT-OSS 20B | 2,066 | 1,597 | 0.0005 |
| GPT-OSS 20B | 3,273 | 2,357 | 0.0009 |
| Qwen3 32B | 1,425 | 1,659 | 0.0006 |
| Qwen3 32B | 727 | 1,373 | 0.0005 |
| Qwen3 32B | 1,786 | 1,714 | 0.0008 |
| DeepSeek V3.1 | 1,923 | 1,523 | 0.0011 |
| DeepSeek V3.1 | 1,889 | 1,594 | 0.0011 |
| DeepSeek V3.1 | 1,367 | 1,314 | 0.0008 |
Generated from production metrics - April 2026 | Updated May 2026 with fine-tuned model