← All Models

qwen3.8-flash-next

Public

Qwen3.8-Flash-Next is an experimental preview of the architecture that will underpin Qwen4, built around a fundamental rethinking of how the core components of modern large language models interact at scale.

103 Downloads

Capabilities

Vision Input
Reasoning

Minimum system memory

96GB

Tags

125B
qwen4exp

README

Qwen3.8 Flash Next

Qwen3.8-Flash-Next is an experimental preview of the architecture that will underpin Qwen4, built around a fundamental rethinking of how the core components of modern large language models interact at scale.

Highlights

  • Hybrid Attention with QSA: Gated DeltaNet is paired with Qwen Sparse Attention (QSA), which operates at the micro-block level rather than selecting individual tokens. This significantly reduces long-context latency for agentic workloads.
  • Efficient Sparse Scaling: the language model has 125B parameters with 6B activated, plus a 51B n-gram embedding and 4B MTP. N-gram embeddings provide a compute-efficient scaling axis that is especially amenable to offloading.
  • Gated Residual: widened residual streams use data-dependent read gates and per-branch scalar write gates for greater expressiveness while preserving training stability.
  • Vision and Agentic Capabilities: native image understanding complements strong performance across coding, professional work, long-horizon agents, and multimodal tool use.
  • Flexible Thinking Control: thinking is enabled by default, with xhigh, medium, and low reasoning-effort levels. Reasoning from previous messages is preserved by default for continuity in multi-turn agentic workflows.
  • Long Context: natively supports up to 262,144 tokens.

Custom Fields

Special features defined by the model author

Reasoning Effort

: select

(default=xhigh)

Controls how much reasoning the model should perform.

Enable Thinking

: boolean

(default=true)

Controls whether the model will think before replying.

Preserve Thinking

: boolean

(default=true)

Preserves reasoning content in all prior assistant turns instead of only the most recent one.

Add Vision IDs

: boolean

(default=false)

Labels images with sequential identifiers in the prompt.

Parameters

Custom configuration options included with this model

Min P Sampling
Disabled
Presence Penalty
Disabled
Repeat Penalty
Disabled
Temperature
1
Top K Sampling
20
Top P Sampling
0.95

Sources

The underlying model files this model uses