Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Roleplay

Best AI Models for Roleplay (RP) and Creative Writing

Model rankings updated August 2026 based on real usage data.

Discover the top AI models for roleplay (RP), character chat and creative writing, ranked by real usage data on OpenRouter. These LLMs excel at maintaining consistent personas, rich dialogue and immersive storytelling across long-context sessions.

Whether you're using Janitor AI, SillyTavern or another frontend, or building your own character chatbot or interactive fiction engine, OpenRouter gives you access to the best roleplay models through a single API.

Browse All ModelsCompare Models

LLM Leaderboard for Roleplay Models

1.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
1.05T
26.3%
2.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
390B
9.8%
3.
Favicon for deepseek
Deepseek V3.2
by deepseek
250B
6.3%
4.
Favicon for google
Gemini 2.5 Flash Lite
by google
233B
5.9%
5.
Favicon for xiaomi
Mimo V2.5
by xiaomi
188B
4.7%
6.
Favicon for stealth
OX Alpha
by stealth
166B
4.2%
7.
Favicon for xiaomi
Mimo V2.5 Pro
by xiaomi
154B
3.9%
8.
Favicon for google
Gemini 3 Flash Preview
by google
150B
3.8%
9.
Favicon for deepseek
Deepseek V4 Pro
by deepseek
148B
3.7%
10.
Favicon for unknown
Others
1.25T
31.5%

Top Roleplay Models on OpenRouter

Based on top weekly usage data from millions of users accessing AI models for roleplay through OpenRouter.

Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0731

13.9T tokens

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows. This is the GA release of DeepSeek V4 Flash.

by deepseek1.31M context$0.03/M input tokens$0.10/M output tokens

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Vision Models
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Speech-to-Text
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections
Favicon for xiaomi

Xiaomi: MiMo-V2.5

12.5T tokens

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

by xiaomi1.05M context$0.119/M input tokens$0.238/M output tokens15% off
Favicon for tencent

Tencent: Hy3

7.74T tokens

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems. With a 256K context window, Hy3 targets long-horizon tasks, including improved coreference resolution, multi-turn constraint tracking, and stable tool-calling that generalizes across agent scaffoldings.

Tencent positions it as a reliable, cost-effective option across coding, document processing, financial analysis, game development, and frontend design, with a strong emphasis on grounded, anti-hallucination behavior that answers when grounded and flags when evidence is missing rather than fabricating.

by tencent262K context$0.0825/M input tokens$0.33/M output tokens
Favicon for openai

OpenAI: GPT-5.6 Luna

6.44T tokens

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier.

by openai1.05M context$0.20/M input tokens$1.20/M output tokens
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash 0423

6.11T tokens

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

by deepseek1.05M context$0.0868/M input tokens$0.1736/M output tokens38% off
Favicon for google

Google: Gemini 3.7 Flash

4.1T tokens

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

by google1.05M context$0.75/M input tokens$3.75/M output tokens50% off
Favicon for z-ai

Z.ai: GLM 5.2

3.53T tokens

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

by z-ai1.05M context$0.4186/M input tokens$1.316/M output tokens70% off
Favicon for z-ai

Z.ai: GLM 5.3 Flash

3.1T tokens

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

by z-ai1.31M context$0.075/M input tokens$0.25/M output tokens50% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Pro 0423

2.06T tokens

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical

by deepseek1.05M context$0.7416/M input tokens$1.483/M output tokens57% off
Favicon for moonshotai

MoonshotAI: Kimi K3

1.69T tokens

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. Its architecture uses KDA and Attention Residuals for computational efficiency.

by moonshotai1.05M context$2.55/M input tokens$12.75/M output tokens