BudgetPix: Compute-Adaptive Tokenization for Pixel-Space Image Diffusion

1University of Illinois Urbana-Champaign 2Google
arXiv preprint, 2026
†Work done during an internship at Google.   ‡Joint last authors.

Test-time compute adjustment. BudgetPix dynamically scales inference costs, achieving a strong quality-efficiency tradeoff across JiT, MiniT2I, PixelDiT, 3 popular architectures for pixel-space image diffusion. Percentages show the fraction of tokens used.

Abstract

Most image generation models rely on uniform tokenization, allocating the exact same computational budget to equally-sized image patches. This static paradigm cannot adapt to different resource constraints at inference time, and yields suboptimal quality-cost tradeoff by devoting the same effort to both plain backgrounds and intricate details. We propose BudgetPix, an adaptive tokenization framework that dynamically allocates compute based on visual complexity and spatial layout, enabling flexible computational budgeting at inference time. BudgetPix comprises three key components: (1) an adaptive encoder that maps a fixed-size image to a variable-length token sequence using an entropy-guided quadtree alongside a multi-scale patch embedder; (2) a scale-aware decoder reconstructs fixed-resolution images from multi-scale token sets; and (3) a flexible training and sampling schedule that enables pixel-space denoisers to operate across variable token counts. BudgetPix seamlessly integrates with existing pixel-space diffusion architectures, enabling a single checkpoint to be operated at a wide range of compute budgets. Evaluated on text-to-image generation, BudgetPix matches the fidelity of MiniT2I-L at 512² (GenEval: 0.874 vs. 0.882) and PixelDiT at 1024² (0.725 vs. 0.721) using just 25% of the original compute budget. In class-conditional generation using a MeanFlow backbone, BudgetPix requires merely 60% of the full compute budget to produce images with near-zero quality degradation, observing a marginal 0.8-point increase in FID (3.46 vs. 2.65). Comprehensive assessments by human and VLM judges confirm that BudgetPix establishes a significantly improved quality-efficiency tradeoff over prior budget-adaptive baselines.

BudgetPix Across Budgets

BudgetPix samples with their token layouts. Same prompt, same initial noise; only the token budget changes. Every subject from the paper and its supplementary material.

One-Step Generation

BudgetPix on the one-step Pixel-MeanFlow model, pMF-L/16 at 256² (1 NFE). Same class, same noise; 100% down to 25% of the tokens.

Method Comparison

The base model and the baselines beside BudgetPix at the same prompt, the same noise and the same token budget. ToMe and FeatSim merge nothing at 100% and reproduce the base model.

Method

Overview

Overview. BudgetPix enables compute-adaptive pixel-space diffusion with three components: an encoder, a decoder, and flexible training and inference across token counts. The Encoder and Decoder (purple) are trained from scratch; the Transformer (blue) is initialized from existing baselines.

Non-uniform layout construction

Non-uniform layout construction. Given an image (the original image x during training, or the denoising image x̂ during inference), the objective is to partition it into cells across multiple scales: p × p, 2p × 2p, …, np × np, where n=2j for j ≥ 1. The process begins by dividing the image into np × np cells, which is the coarsest scale. For each cell, we construct a pixel intensity histogram and compute its entropy H. If H < τnp + δ (where τnp is the scale-specific threshold and δ is a global offset), the cell remains a single token. Otherwise, it is subdivided into four cells of the next finer scale (n ← n/2). This recursive subdivision terminates at n = 1, corresponding to the minimum cell size of p × p.

Architecture of the encoder and decoder

BudgetPix Architecture. The encoder receives the noisy input zt and its corresponding spatial layout L, converts the image into a sequence of multi-scale patches. The finest p × p tokens route directly through the fine path, while coarser patches navigate a dedicated coarse path in both the encoder and decoder. In the encoder, each coarse patch undergoes simultaneous downscaling (to p × p) and splitting via a Patch Aggregator (into n2 sub-patches). These streams are processed through MLP layers and a scale mixer before being fused into a single coarse token. In the decoder, each denoised coarse patch is upscaled back to its target resolution (np × np). Concurrently, a Patch Refiner splits the patch into 4 × 4 sub-patches, processes them through an MLP and a lightweight transformer, and fuses them with the upscaled patch to restore local texture. The central diffusion transformer is directly initialized from existing pixel-space models, modified only to operate across variable, dynamically determined token counts rather than a fixed count.

Results

Text-to-image results, MiniT2I-L/16 at 5122

Text-to-image results

(a) Each metric plotted against the wall-clock speed-up over the released model (DPG: DPG-Bench; IR: ImageReward denote metrics). (b) Faithfulness to the method's own full-budget image plotted against the token budget (Sim.: CLIP or DINOv2 feature similarity).

Class-conditional generation on ImageNet, JiT-L and JiT-B

Class-conditional generation on ImageNet

Quality against measured speed-up, JiT-L/32 and MiniT2I-L/16 at 5122

Quality against measured speed-up

BudgetPix on one-step pixel MeanFlow pMF-L/16

One-step generation

BudgetPix provides a better quality-efficiency trade-off compared with the base model and other token merging baselines.

click a plot to enlarge

BibTeX

@article{kara2026budgetpix,
  title={BudgetPix: Compute-Adaptive Tokenization for Pixel-Space Image Diffusion},
  author={Ozgur Kara and Yujia Chen and Daniel Watson and David Forsyth and James Matthew Rehg and Wen-Sheng Chu and Du Tran},
  journal={arXiv preprint},
  year={2026}
}