CP·CP LEDGER

ImpCC: DeepSeek V4 Flash 0731 (abliterated)

impcc/deepseek-v4-flash-0731
NEWTOOLS

An abliterated (decensored) rebuild of deepseek-ai/DeepSeek-V4-Flash-0731, published by ImpCC and served here on our own hardware. Not affiliated with or endorsed by the base-model authors. 284B total / 13B active parameters, 43 layers, 256 routed experts (top-6) plus 1 shared. Refusal behaviour was removed by a rank-1 orthogonal projection applied to the attention output projections only: 92 of 72,317 tensors differ from the base checkpoint and the remaining 72,225 are byte-identical. There is no quantization step — the weights ship at the base checkpoint's native precision, FP4 (E2M1 with ue8m0 block scales) for the routed experts and FP8 e4m3 elsewhere. Licensed MIT, inherited from the base model. The architecture's native context is 1,048,576 tokens; this endpoint is launched with --max-model-len 131072 and that launch argument, not the card, is what this registration reports. Latency, throughput and uptime are deliberately absent — they are live measurements, and a number frozen into a config file stops being true the moment the next request lands.

MODEL RECORD
Modalities
text
In / Out Price
$0.100 / $0.400
Context
131K
Max Output
Released
Aug 14, 2026
Knowledge Cutoff
PROVIDERS — 1

One provider serves this model, so every request routes to it. When a second one appears, this table becomes the comparison — price, context, quantization and uptime side by side — and the routing mode you pick decides between them.

Provider·Provider ID·Endpoint tag·Context·Max Output·Quant·Input $/MOutput $/M·Cache Read·Latency·Throughput·Uptime 30m·Uptime 1d·ZDR·Discount·Moderated·Impl. cache·Region·Params·
ImpCCimpccimpcc131,072fp4$0.100$0.400$0.037100.00%UnknownNoYES17
one endpoint row — nothing to compare it against yet. The columns above are the full record for ImpCC.

Provider identity, routing tag, per-endpoint pricing, context ceiling, max output, quantization, implicit caching, uptime, endpoint discount, moderation flag, region and zero-data-retention are read from the API.

Latency and throughput are blank here — the spec exposes them only to an authenticated caller, our capture was taken without a key, and nothing on this deployment has reported one yet. Nothing is generated to fill them.

Why an endpoint may be skipped
Provider ignored by accountProvider not allowed by accountProvider blocked by guardrailModel blocked by guardrailProvider not allowed by guardrailZDR violationFree model training violationPaid model publication violation
PRICING

Catalogue pricing next to what individual providers actually post. Every figure on this panel is read from the API. There is no modelled cache-hit rate and no blended “effective price” here, because the API exposes neither of the inputs one would need.

Catalogue vs endpoint input price
One endpoint row
A cheapest/dearest spread needs two prices to sit between. With a single endpoint the catalogue price and the endpoint price are the same number, shown in the matrix beside this.
Full price matrix
Input Price$0.100/M tokens
Output Price$0.400/M tokens
Cache Read$0.037/M tokens
Cache Write$0.130/M tokens
PERFORMANCE

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). The API types both as a percentile block: 0 of 1 endpoint rows here carry a latency reading.

Latency and throughput are blank here — the spec exposes them only to an authenticated caller, our capture was taken without a key, and nothing on this deployment has reported one yet. Nothing is generated to fill them.

This gateway's own round trips
Mean RT3852.1s
Requests695
Tok / request34.3K
Out / request277

Mean request latency, requests served here and tokens served here are measured by this gateway from the requests it served, read back from POST /api/v1/analytics/query. They describe this deployment only — never a model’s usage anywhere else.

EndpointThr p50p75p90p99Lat p50p75p90p99
ImpCC
UPTIME

Percent of requests that succeeded, as reported by the snapshot. These are the three windows the API returns; no other window is shown, and an absent reading renders as an em-dash rather than a zero.

ImpCC impcc5m30m1d100.00%
BENCHMARKS

Artificial Analysis indices, read from the model's own benchmarks block. Present on 0 of 1 models; each index is independently nullable and an absent one is left blank rather than scored. Rank percentile is taken among the models that carry that index.

— · no benchmark block for this model
APPS

Public apps that send the most traffic to this model. Good signal for what real production workloads look like.

Not enough data to display yet.
There is no app-attribution route in the API contract, so no application is named as a consumer of this model. This build does not invent them.
ACTIVITY — THIS GATEWAY

Requests this gateway served for this model, and the tokens they moved. This is our own traffic, not the model’s usage anywhere else.

-7h-5h-4h-2hnow
Prompt951.2KCompletion7.8KReasoning0
Requests695
Tokens23.8M
Spend$2.44
Mean RT3852.1s
Prompt tok23.6M
Completion tok192.7K
Reasoning tok0
Cached tok0

Requests served here, tokens served here, prompt, completion and reasoning token split, spend and mean request latency are measured by this gateway from the requests it served, read back from POST /api/v1/analytics/query. They describe this deployment only — never a model’s usage anywhere else.

Drop-in code to call this model. The API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug.

QUICK START — impcc/deepseek-v4-flash-0731
curl "$CP_BASE_URL/api/v1/chat/completions" \  -H "Authorization: Bearer $CP_API_KEY" \  -H "Content-Type: application/json" \  -d '{    "model": "impcc/deepseek-v4-flash-0731",    "messages": [{ "role": "user", "content": "Hello" }]  }'
FAQ

An abliterated (decensored) rebuild of deepseek-ai/DeepSeek-V4-Flash-0731, published by ImpCC and served here on our own hardware. Not affiliated with or endorsed by the base-model authors. 284B total / 13B active parameters, 43 layers, 256 routed experts (top-6) plus 1 shared. Refusal behaviour was removed by a rank-1 orthogonal projection applied to the attention output projections only: 92 of 72,317 tensors differ from the base checkpoint and the remaining 72,225 are byte-identical. There is no quantization step — the weights ship at the base checkpoint's native precision, FP4 (E2M1 with ue8m0 block scales) for the routed experts and FP8 e4m3 elsewhere. Licensed MIT, inherited from the base model. The architecture's native context is 1,048,576 tokens; this endpoint is launched with --max-model-len 131072 and that launch argument, not the card, is what this registration reports. Latency, throughput and uptime are deliberately absent — they are live measurements, and a number frozen into a config file stops being true the moment the next request lands.

MORE FROM IMPCC
— · this is the only model from ImpCC in the catalogue
CP · CONSENSUS PROTOCOLTerminal Ledger · rendered from this deployment’s own catalogue API