pytorch-quantization

Mar 11, 2025

8c6d69d · Mar 11, 2025

Name	Name	Last commit message	Last commit date
parent directory ..
docs	docs	TensorRT 10.9 OSS Release. (#4381 )	Mar 11, 2025
examples	examples	TensorRT 10.9 OSS Release. (#4381 )	Mar 11, 2025
pytorch_quantization	pytorch_quantization	TensorRT 10.9 OSS Release. (#4381 )	Mar 11, 2025
src	src	TensorRT 10.0 Release	Apr 3, 2024
tests	tests	TensorRT 10.9 OSS Release. (#4381 )	Mar 11, 2025
.coveragerc	.coveragerc	TensorRT OSS release v7.2.1	Oct 20, 2020
.gitignore	.gitignore	TensorRT OSS release v7.2.1	Oct 20, 2020
.pylintrc	.pylintrc	TensorRT OSS release v7.2.1	Oct 20, 2020
.style.yapf	.style.yapf	TensorRT OSS release v7.2.1	Oct 20, 2020
CONTRIBUTING.md	CONTRIBUTING.md	10.2 GA release update (#3998 )	Jul 11, 2024
LICENSE	LICENSE	TensorRT OSS release v7.2.1	Oct 20, 2020
MANIFEST.in	MANIFEST.in	TensorRT OSS release v7.2.1	Oct 20, 2020
README.md	README.md	TensorRT 10.9 OSS Release. (#4381 )	Mar 11, 2025
VERSION	VERSION	10.2 GA release update (#3998 )	Jul 11, 2024
requirements.txt	requirements.txt	TensorRT OSS 21.02 release	Feb 5, 2021
setup.cfg	setup.cfg	TensorRT OSS release v7.2.1	Oct 20, 2020
setup.py	setup.py	TensorRT 10.0 Release	Apr 3, 2024

README.md

Note: Pytorch Quantization development has transitioned to the TensorRT Model Optimizer. All developers are encouraged to use the TensorRT Model Optimizer to benefit from the latest advancements on quantization and compression. While the Pytorch Quantization code will remain available, it will no longer receive further development.

Pytorch Quantization

PyTorch-Quantization is a toolkit for training and evaluating PyTorch models with simulated quantization. Quantization can be added to the model automatically, or manually, allowing the model to be tuned for accuracy and performance. Quantization is compatible with NVIDIAs high performance integer kernels which leverage integer Tensor Cores. The quantized model can be exported to ONNX and imported by TensorRT 8.0 and later.

Install

Binaries

pip install pytorch-quantization --extra-index-url https://pypi.ngc.nvidia.com

From Source

git clone https://github.com/NVIDIA/TensorRT.git
cd tools/pytorch-quantization

Install PyTorch and prerequisites

pip install -r requirements.txt
# for CUDA 10.2 users
pip install torch>=1.9.1
# for CUDA 11.1 users
pip install torch>=1.9.1+cu111

Build and install pytorch-quantization

# Python version >= 3.7, GCC version >= 5.4 required
python setup.py install

NGC Container

pytorch-quantization is preinstalled in NVIDIA NGC PyTorch container, e.g. nvcr.io/nvidia/pytorch:22.12-py3

Resources

Pytorch Quantization Toolkit userguide
Quantization Basics whitepaper

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Files

pytorch-quantization

pytorch-quantization

README.md

Pytorch Quantization

Install

Binaries

From Source

NGC Container

Resources

Files

pytorch-quantization

Directory actions

More options

Directory actions

More options

Latest commit

History

pytorch-quantization

Folders and files

parent directory

README.md

Pytorch Quantization

Install

Binaries

From Source

NGC Container

Resources