POCKET-Image-Zimage / README.md
SeaWolf-AI's picture
POCKET family: add POCKET-Qwen3.8-Flash-Next (180B on a laptop), cross-link all 8 POCKET repos
d40c018 verified
|
Raw
History Blame Contribute Delete
6.16 kB
metadata
license: apache-2.0
base_model:
  - Tongyi-MAI/Z-Image
pipeline_tag: text-to-image
library_name: diffusers
language:
  - en
  - ko
tags:
  - text-to-image
  - image-generation
  - quantized
  - bitsandbytes
  - nf4
  - 4-bit
  - on-device
  - korean
  - pocket
  - vidraft

πŸ†• POCKET-Qwen3.8-Flash-Next β€” a 180B model running on a laptop with 8 GB VRAM + 32 GB RAM Β· 4.17 tok/s measured.

New VRAM RAM Speed

πŸ†• POCKET-Zimage-CPU β€” photoreal images in 46 s on a CPU only. No GPU, no CUDA, no Python.

New Space RAM

πŸ“š Collections

β–Ά POCKET Models β€” this family (on-device, no GPU) Darwin Family Β· Aether Foundation Β· VKAE Accelerated

πŸ–ΌοΈ POCKET-Image-Zimage β€” 4-bit (NF4) Z-Image for on-device

Pick your build β†’ 35B 26B KR GGUF KR MLX EN GGUF 180B laptop Image NF4 Image CPU

A 4-bit (NF4) quantized build of Z-Image (Apache-2.0), packaged by VIDRAFT for low-VRAM, on-device image generation β€” part of the POCKET line.

  • πŸ“¦ ~6 GB on disk (transformer + text encoder in NF4, VAE in fp16)
  • ⚑ Runs from ~8.6 GB VRAM (β‰ˆ4.5 GB with CPU offload) β€” vs 23.3 GB for bf16
  • 🎯 ~2.7–5Γ— smaller footprint, quality on par with the bf16 base

Usage

import torch
from diffusers import ZImagePipeline   # or ZImageImg2ImgPipeline / ZImageInpaintPipeline

pipe = ZImagePipeline.from_pretrained(
    "FINAL-Bench/POCKET-Image-Zimage", torch_dtype=torch.bfloat16
).to("cuda")
img = pipe("a serene mountain lake at sunrise, photorealistic", num_inference_steps=20).images[0]
img.save("out.png")

Requires bitsandbytes (CUDA). Measured reload + generate peak: 10.9 GB VRAM. For Apple Silicon / CPU, an optimum-quanto int8 build (13.4 GB) is the portable option.

🎨 The full POCKET-Image system

This repo hosts the quantized base model only. The headline character-perfect Korean & multilingual text feature is delivered by the POCKET-Image pipeline, not by these weights alone. Try the full system here:

Base model: Tongyi-MAI/Z-Image (Apache-2.0) Β· Quantization: bitsandbytes NF4 Β· By VIDRAFT.


🧩 The POCKET Family β€” On-device AI by VIDRAFT

Big models, small hardware. No GPU, no cloud.

Models

Demos & tools (Spaces)

πŸ“š Full POCKET collection