What huggingface/transformers shipped
Written by FoxPlug from public releases; not affiliated with Hugging Face. An automatic summary of the public release, pull request and commit data of github.com/huggingface/transformers. Hugging Face did not write it and does not use or endorse FoxPlug. Every line links to the public change it describes.
Get a weekly update like this for your product, free
Week of September 21, 2026
What shipped
- Added support for Nemotron Omni architecture [20]. Pull request #46509
- Added Nemotron3Diarization model support [26]. Pull request #49056
- Added MLX recipe for Executorch [10]. Pull request #48910
- Added Trainer.end method [12]. Pull request #48875
- Integrated PEFT with Dtensor-based tensor parallelism [36]. Pull request #48485
- Added image processing tester init [11]. Pull request #48829
- Added support for loading GGUF models via Chat CLI [33]. Pull request #49031
- Fixed legacy Gemma 1
hidden_actconfiguration remapping [2]. Pull request #49084 - Fixed exporters import compatibility with torch versions below 2.9 [1]. Pull request #49124
- Enabled ROCm flash attention 3 routing and fixture generation for gpt-oss [13]. Pull request #46837
Why it matters
This week includes two new model architectures (Nemotron Omni and Nemotron3Diarization), expanded deployment options through MLX recipes and GGUF support in Chat CLI, and fixes for critical issues across attention mechanisms, configuration handling, and compatibility layers. New trainer functionality and improved device handling strengthen the core training pipeline.
Changelog entry
- Support for Nemotron Omni architecture Pull request #46509
- Support for Nemotron3Diarization Pull request #49056
- MLX recipe for Executorch Pull request #48910
- Trainer.end method Pull request #48875
- PEFT integration with Dtensor-based tensor parallelism Pull request #48485
- GGUF model loading in Chat CLI Pull request #49031
- Image processing tester initialization Pull request #48829
- Fix legacy Gemma 1
hidden_actremapping in config post-init Pull request #49084 - Fix exporters import on torch < 2.9 (
is_contiguous_or_false) Pull request #49124 - ROCm flash attention 3 routing and fixture generation for gpt-oss Pull request #46837
- Use
grouped_mmon TPU devices under torch.compile Pull request #49097 - Preserve eos ids declared in config Pull request #49082
- Fix MaskFormerSwin attention mask dtype to follow hidden states Pull request #49032
- Use namespaced dataset ids in docs and PyTorch examples Pull request #49005
- Fix RGB early-return skipping PNG tRNS compositing Pull request #49092
- Fix NaN in Parakeet eager attention with padded batches Pull request #49070
- Fix mask creation not being skipped under torch.compile Pull request #48975
- Add
use_associative_scanconfig flag to Zamba to avoid torch.compile slowdown Pull request #48331
This week: Nemotron Omni & Nemotron3Diarization support, MLX & GGUF integration, Trainer.end method, PEFT tensor parallelism, plus fixes for attention, configs & compatibility [20,26,10,12,36,33]
Transformers now supports Nemotron Omni and Nemotron3Diarization architectures, expands deployment with MLX recipes and GGUF loading in Chat CLI, introduces Trainer.end for better training control, and ships PEFT integration with Dtensor-based tensor parallelism. This week also includes critical fixes for attention mechanisms, configuration handling, and cross-version compatibility to strengthen both inference and training workflows.