🎉 75% of content is free forever — Unlock Premium from $10/mo →
CW
💼 Servicesℹ️ About✉️ ContactView Pricing Plansfrom $10

Quantization Techniques Deep Dive

OptimizationQuantization🟢 Free Lesson

Advertisement

Optimization

Quantization Techniques Deep Dive — Reducing Precision Without Sacrificing Quality

Quantization enables running billion-parameter models on consumer hardware. This deep dive covers GPTQ, AWQ, GGUF, and INT4/INT8 methods—theory, implementation, and when to use each approach.

  • Post-Training Quantization — GPTQ and AWQ for minimal quality loss
  • Quantization Formats — INT8, INT4, NF4, and GGUF variants
  • Quantization-Aware Training — QLoRA for fine-tuning quantized models
  • Practical Trade-offs — Quality, speed, and memory comparisons

The art of quantization is knowing what you can throw away without anyone noticing.

Quantization Techniques Deep Dive

Quantization reduces model memory by representing weights with fewer bits. While the concept is simple—map high-precision values to low-precision bins—the implementation details significantly affect quality. This guide covers the major quantization methods and their practical trade-offs.

Quantization Theory

Uniform Quantization

Quantization Error

Group Quantization

GPTQ (Post-Training Quantization)

Theory

Implementation

import torch
import torch.nn as nn
from transformers import AutoModelForCausalLM, AutoTokenizer, GPTQConfig

class GPTQQuantizer:
    """GPTQ quantization implementation."""
    
    def __init__(self, model, tokenizer, calibration_data):
        self.model = model
        self.tokenizer = tokenizer
        self.calibration_data = calibration_data
        
    def quantize_layer(self, layer, n_bits=4, group_size=128):
        """Quantize a single linear layer using GPTQ."""
        W = layer.weight.data  # (out_features, in_features)
        n_rows, n_cols = W.shape
        
        # Compute Hessian approximation using calibration data
        H = self._compute_hessian(layer, self.calibration_data)
        H_inv = torch.linalg.inv(H + 1e-6 * torch.eye(n_cols, device=H.device))
        
        # Quantize column by column
        W_quantized = torch.zeros_like(W)
        
        for col in range(n_cols):
            # Quantize current column
            w_col = W[:, col]
            
            # Compute quantization parameters
            if group_size > 0:
                # Group quantization
                for g in range(0, n_rows, group_size):
                    group = w_col[g:g+group_size]
                    scale = (group.max() - group.min()) / (2**n_bits - 1)
                    zero_point = (-group.min() / scale).round()
                    W_quantized[g:g+group_size, col] = (
                        (group / scale + zero_point).round() * scale - scale * zero_point
                    )
            
            # Compensate for quantization error in remaining columns
            if col < n_cols - 1:
                error = W[:, col] - W_quantized[:, col]
                # Distribute error using Hessian
                W[:, col+1:] -= error.unsqueeze(1) * H_inv[col, col+1:]
        
        return W_quantized
    
    def _compute_hessian(self, layer, data):
        """Compute Hessian approximation from calibration data."""
        H = torch.zeros(layer.in_features, layer.in_features)
        
        for batch in data:
            with torch.no_grad():
                # Forward pass to get input activations
                x = layer._get_input activations(batch)
                H += x.T @ x
        
        H /= len(data)
        return H

GPTQ Configuration

# Using HuggingFace Transformers
from transformers import AutoModelForCausalLM, GPTQConfig

# Configure GPTQ quantization
gptq_config = GPTQConfig(
    bits=4,
    dataset="c4",
    group_size=128,
    desc_act=True,      # Sort by activation magnitude
    damp_percent=0.01,
    true_sequential=True
)

# Load and quantize
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-2-7b-hf",
    quantization_config=gptq_config,
    device_map="auto"
)

# Save quantized model
model.save_pretrained("llama-2-7b-gptq")

AWQ (Activation-Aware Weight Quantization)

Theory

Scale Factor Selection

Implementation

class AWQQuantizer:
    """AWQ quantization with activation-aware scaling."""
    
    def __init__(self, model, calibration_data):
        self.model = model
        self.calibration_data = calibration_data
    
    def find_optimal_scales(self, layer, n_bits=4, group_size=128):
        """Find optimal per-channel scales for AWQ."""
        W = layer.weight.data
        n_out, n_in = W.shape
        
        # Compute activation importance
        importance = self._compute_importance(layer)
        
        # Initialize scales
        scales = torch.ones(n_in, device=W.device)
        
        # Grid search for optimal scales
        best_error = float('inf')
        best_scales = scales.clone()
        
        for alpha in [0.5, 0.75, 1.0, 1.25, 1.5]:
            candidate_scales = importance.pow(alpha)
            candidate_scales = candidate_scales / candidate_scales.mean()
            
            # Quantize with candidate scales
            W_scaled = W * candidate_scales.unsqueeze(0)
            W_quant = self._quantize(W_scaled, n_bits, group_size)
            W_dequant = W_quant / candidate_scales.unsqueeze(0)
            
            # Compute error
            error = (W - W_dequant).pow(2).mean()
            
            if error < best_error:
                best_error = error
                best_scales = candidate_scales
        
        return best_scales
    
    def _compute_importance(self, layer):
        """Compute channel importance from activations."""
        importance = torch.zeros(layer.in_features)
        
        for batch in self.calibration_data:
            with torch.no_grad():
                x = layer._get_input_activations(batch)
                importance += x.abs().mean(dim=0)
        
        return importance / len(self.calibration_data)

GGUF (GGML Unified Format)

Overview

Quantization Types

TypeBitsQualitySpeedUse Case
Q2_K2LowFastestExtreme compression
Q3_K_M3ModerateFastLow memory
Q4_04GoodFastBalanced
Q4_K_M4BetterFastRecommended
Q5_K_M5Very GoodModerateHigh quality
Q6_K6ExcellentSlowerNear-lossless
Q8_08LosslessSlowestMaximum quality

GGUF Conversion

# Converting to GGUF format using llama.cpp
# Step 1: Convert HuggingFace model to GGML
python convert.py /path/to/model --outtype f16

# Step 2: Quantize to desired format
./quantize /path/to/model.ggml /path/to/model-q4_k_m.gguf q4_k_m

# Step 3: Run inference
./main -m /path/to/model-q4_k_m.gguf -p "Hello, world" -n 100

INT4 vs INT8 Quantization

Quality Comparison

Model SizeINT8 PPLINT4 PPLFP16 PPLINT8 DeltaINT4 Delta
1B12.112.811.9+0.2+0.9
7B5.725.915.68+0.04+0.23
13B5.485.625.45+0.03+0.17
70B3.123.213.10+0.02+0.11

Memory and Speed

FormatMemory (7B)Inference SpeedGPU Support
FP1614 GBBaselineAll GPUs
INT87 GB1.0-1.2×A100, RTX 30xx+
INT43.5 GB0.8-1.0×Limited
NF43.5 GB0.7-0.9×CPU only

Quantization-Aware Training (QAT)

QLoRA

import torch
from transformers import (
    AutoModelForCausalLM,
    BitsAndBytesConfig,
    TrainingArguments,
    Trainer
)
from peft import LoraConfig, get_peft_model, prepare_model_for_kbit_training

# NF4 quantization config
bnb_config = BitsAndBytesConfig(
    load_in_4bit=True,
    bnb_4bit_quant_type="nf4",
    bnb_4bit_compute_dtype=torch.float16,
    bnb_4bit_use_double_quant=True,
)

# Load quantized model
model = AutoModelForCausalLM.from_pretrained(
    "meta-llama/Llama-2-7b-hf",
    quantization_config=bnb_config,
    device_map="auto"
)

# Prepare for training
model = prepare_model_for_kbit_training(model)

# Add LoRA adapters
lora_config = LoraConfig(
    r=64,
    lora_alpha=16,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj"],
    lora_dropout=0.05,
    bias="none",
    task_type="CAUSAL_LM"
)

model = get_peft_model(model, lora_config)
model.print_trainable_parameters()
# trainable params: 33,554,432 || all params: 3,776,126,976 || trainable%: 0.8887

Practical Deployment

Choosing the Right Quantization

ScenarioRecommendedWhy
Production servingGPTQ INT4 or AWQFast inference, good quality
CPU deploymentGGUF Q4_K_MOptimized for CPU
Fine-tuningQLoRA NF4Enables training on consumer GPUs
Edge devicesGGUF Q2_K or Q3_KMinimal memory
Maximum qualityINT8 or FP16No quality loss

Benchmarking Results

Architecture Diagram
Model: Llama-2-7B
Hardware: RTX 4090 (24 GB)

| Format  | Memory | Tokens/s | PPL  |
|---------|--------|----------|------|
| FP16    | 14 GB  | 85       | 5.68 |
| GPTQ-4  | 4 GB   | 95       | 5.91 |
| AWQ-4   | 4 GB   | 98       | 5.89 |
| GGUF-Q4 | 4.2 GB | 72*      | 5.93 |

*GGUF runs on CPU, slower than GPU

Practice Exercises

  1. Mathematical: Calculate the theoretical memory savings of quantizing a 13B parameter model from FP16 to INT4. Include the overhead of quantization constants and group scales.

  2. Implementation: Implement a simple INT8 quantizer with per-channel scaling and compare its perplexity degradation against per-tensor quantization.

  3. Analysis: Compare GPTQ and AWQ on a 7B model using the same calibration dataset. Which method produces better quality at 3-bit quantization?

  4. Research: Investigate mixed-precision quantization where different layers use different bit widths. What is the optimal allocation strategy?


What to Learn Next

-> LoRA and PEFT Efficient fine-tuning using low-rank adaptation.

-> Pruning for LLMs Reducing model size through weight pruning.

-> Low-Rank Factorization SVD decomposition and weight sharing techniques.

-> Hardware-Aware LLM Design Optimizing models for GPU memory hierarchy.

-> Model Merging and Fusion Combining multiple fine-tuned models.

-> LLM Inference Optimization Speeding up model inference for production.

Need Expert LLM Help?

Get personalized tutoring, project support, or professional consulting.

Advertisement