Skip to content
← Back to today

transformers v5.19.0: EmbeddingGemma2, MoE router logits, EP changes

Release v5.19.0 adds EmbeddingGemma2 multimodal embeddings, makes MoE models return router logits with output_router_logits=True, updates expert parallelism and cache handling, and changes attention/continuous-batching internals.

What changed

Added EmbeddingGemma2; MoE models now return router logits when output_router_logits=True; expert parallelism defaults to token-dispatch for Qwen3 MoE and Mellum; cache fixes and per-layer cache configs.

Why it matters

MoE users must adapt code to the new router-logit outputs; embedding and retrieval engineers can evaluate EmbeddingGemma2's 768-dimensional multimodal vectors; infra may need to adjust attention and CB code.

Who should care
DevelopersResearchersEnterprises
Worth trying

Yes — Use EmbeddingGemma2 for multimodal 768-dimensional embeddings; MoE and cache fixes affect inference.

Related