What it really takes to self-host Qwen3.8-Flash-Next: official BF16 and FP8 sizes, validated GPU topologies, a clearly labeled single-workstation path, license limits, and API break-even math.
Open Models · 2026-09-07
What it takes to self-host Tencent Hunyuan Hy4 preview: official BF16 and FP8 weight sizes, B200/B300/H200 configurations, 200-person and public-API tiers, Apache-2.0 duties, and API break-even math.
Open Models · 2026-09-07
How to self-host GLM-5.3-Flash: its two official FP8 and BF16 weight repositories, a 350 GB RAM single-GPU path, validated 8-GPU serving tiers, MIT terms, best-fit scenarios, and API break-even math.
Open Models · 2026-09-07
Route long-running coding agents by context, modality, tools, and acceptance checks — then use task state, checkpoints, and enforceable limits to keep the loop recoverable.
Coding Agent Routing · 2026-07-24
A practical decision framework for choosing a Chinese LLM or a direct frontier-provider route by product dependency, task constraints, and accepted-output cost—not a generic leaderboard.
LLM Procurement and Routing · 2026-07-24
A practical guide to choosing current ChinaAPI video models for story, movement, reference control, drafts, and edits—without pretending one model wins every job.
Video Model Routing · 2026-07-20
Choose a current ChinaAPI image model by task shape, reference assets, control requirements, and the cost of getting a usable image — not by a generic quality leaderboard.
Image Model Routing · 2026-07-20
Route current ChinaAPI reasoning models by task shape, context window, modality, and usable-answer cost — not a generic quality leaderboard.
Chinese LLM Routing · 2026-07-20
What Chinese LLM, image, and video APIs cost in July 2026 — the sub-$0.20 tier, the 1M-context club, per-second vs per-generation video math, and what to pick by workload.
Pricing · 2026-07-06