Blog
Articles & insights
Software development, artificial intelligence, open source and experience reports.

Qwen3.8-Flash-Next: Open-Weights LLM Revolution & Hybrid Architecture
Discover Qwen3.8-Flash-Next, Alibaba's open-weights hybrid architecture revolutionizing LLM efficiency with QSA, Gated DeltaNet, and a 1M context window.

Qwen3.8-Flash-Next: Architecture, Benchmarks, and LLM Integration Guide
Discover Qwen3.8-Flash-Next, Alibaba's new hybrid LLM architecture combining sparse attention, Gated DeltaNet, and extreme efficiency.
Sep 2, 2026
MiniMax H3: The New Standard in Omni-Modal Video Generation
Discover MiniMax H3, an open-source omni-modal system capable of generating videos with native stereo audio up to 2K resolution.
Aug 29, 2026
Huihui-Qwen3.8-27B Abliterated: Technical Guide and GGUF Analysis
Technical analysis of the Huihui-Qwen3.8-27B-abliterated-GGUF model, an innovative open-source approach to removing LLM censorship through targeted weight modification.
Aug 29, 2026
MiniMax Music 3: Long-Form AI Music Generation Model
Discover MiniMax Music 3, a hybrid AI music generation model capable of producing complete five-minute tracks in 32 kHz stereo.
Aug 29, 2026
Qwen3.8-27B: The New Standard for Open and Agentic LLMs
Discover Qwen3.8-27B, an open 27B-parameter model with advanced agentic and multimodal capabilities, optimized with Unsloth Dynamic V3.0 GGUF.
Aug 29, 2026
Qwen3.8-Flash-Next: Architectural Innovation for AI Agents
The new Qwen3.8-Flash-Next model redefines LLM efficiency with its hybrid architecture, long-context processing capabilities, and state-of-the-art performance.
Aug 29, 2026
MiniMax-H3: Accelerate Your AI Video Generation with PDD LoRAs
Discover how new acceleration LoRAs for MiniMax-H3, based on Parallel Decoding Distillation (PDD), are transforming AI video generation by drastically reducing inference time.
Aug 29, 2026
Qwen3.8-27B OBLITERATED V3: The Complete Guide to Abliteration
Discover Qwen3.8-27B OBLITERATED V3, an open-source LLM surgically modified to eliminate refusals and moralizing, offering total freedom for research and development.
Aug 29, 2026
Ornith-1.5-35B-A3B: The Self-Improving AI for Coding
Discover Ornith-1.5-35B-A3B, the new open-source model redefining self-improvement and agentic coding performance, surpassing models of equivalent size.
Aug 29, 2026
Qwen3.8-27B-Uncensored GGUF: Technical Guide and Local Deployment
Discover Qwen3.8-27B-Uncensored in GGUF format: how ablation, MTP, and advanced quantization enable the local deployment of a high-performance LLM without standard filtering.
Aug 29, 2026
The Rise of Chinese AI Models: A Credible Economic Alternative?
Faced with soaring costs from OpenAI and Anthropic, US companies are turning to Chinese AI models, which are more competitive and performant. An analysis of a new technological era.
Jul 10, 2026
Meta and AI: New Paid Model and Strategic Pivot
Meta is taking a strategic turn by launching a paid version of its new AI models. We analyze the impacts on developers and the broader technological ecosystem.
Jul 10, 2026
OpenAI: The Rollout of GPT-5.6 and Its Technical Implications
OpenAI rolls out GPT-5.6 models, lifting government restrictions. A technical analysis of the implications for developers, architecture, and AI agent security.
Jul 10, 2026
OpenAI: GPT-5.6 and ChatGPT Work, a New Era for Automation
OpenAI launches GPT-5.6 and ChatGPT Work. A technical analysis of these new advancements and their impact on AI integration in DevOps workflows and enterprise software architectures.
Jul 10, 2026
Meta launches Muse: The new AI model for image generation
Meta launches Muse, its first AI model for image generation. We analyze its technical aspects, advertising implications, and its impact on the digital content ecosystem.
Jul 10, 2026
AI and Governance: What are the Stakes for the Tech Industry?
Recent discussions between AI leaders and the US administration raise crucial questions about technological sovereignty and the future of LLMs.
Jul 6, 2026
Gradio-Lite: Run your AI interfaces serverless in the browser
Gradio-Lite allows running Gradio applications directly in the browser using WebAssembly. A revolution for serverless deployment of AI models.
Jul 6, 2026
AI Agent Observability and Evaluation: The smolagents and Phoenix Duo
Discover how to integrate Arize Phoenix with your smolagents to ensure total observability, debug complex workflows, and evaluate your model performance.
Jun 16, 2026
Simplify AI Training with 🤗 Accelerate: The Complete Guide
Discover 🤗 Accelerate, the Hugging Face library that simplifies training your AI models on any hardware (GPU, TPU) without modifying your PyTorch code.
Jun 12, 2026
Habana Gaudi2 vs Nvidia A100: AI Performance Comparison
Discover the performance of Habana Gaudi2 processors against Nvidia A100 GPUs. A technical analysis to optimize the training and inference of your AI models.
Jun 11, 2026
Hugging Face and Pollen Robotics: Towards AI-Driven Robotics
Hugging Face acquires Pollen Robotics to democratize open-source, AI-driven robotics. Discover how this union will transform Embodied AI and LLM-based robot control.
Jun 10, 2026
Convert your models to ONNX with Optimum: The complete guide
Learn how to convert your Transformers models to ONNX format with the 🤗 Optimum library. Optimize your models for production, reduce latency, and simplify your deployments.
Jun 10, 2026
Implementing MCP Servers in Python: Create a Shopping Assistant with Gradio
Discover how to implement an MCP server in Python to build a shopping assistant with Virtual Try-On (VTON) and Gradio, leveraging the power of the Model Context Protocol.
Jun 9, 2026
Simplicity at the Heart of High-Performance Neural Networks
Bigger is not always better. Back to basics: why simplicity and rigor in building neural networks are the true keys to AI performance.
May 29, 2026
The 20 Best AI Chatbots in 2026
Discover the 20 best AI chatbots in 2026 to boost your sales, automate marketing, and improve productivity. A comprehensive comparison by use case.
May 29, 2026
AI and Biology: The Arc Virtual Cell Challenge Explained
The Arc Virtual Cell Challenge uses foundation models to simulate cellular behavior. Discover how AI is redefining research in molecular biology.
May 29, 2026
Xet on Hugging Face: Optimize Your Dataset Versioning
The integration of Xet into the Hugging Face Hub revolutionizes massive dataset versioning. Discover how this solution optimizes storage, speed, and collaboration.
May 29, 2026
Gemini 3.5 Flash: Efficiency at the Heart of AI Agents
Google unveils Gemini 3.5 Flash, an ultra-high-performance and efficient model designed to accelerate the large-scale deployment of autonomous AI agents.
May 29, 2026
Reeboot Fleet: we built the console we were missing to run our Raspberry Pis
Reeboot Fleet, our Raspberry Pi fleet management SaaS, is now in public access. One-command install, Tailscale terminal, targeted deployments — from €1.40 / device / month.
May 27, 2026
Hugging Face TGI Now Supports vLLM and TensorRT-LLM
Hugging Face announces multi-backend support for TGI, now enabling the use of vLLM and TensorRT-LLM for LLM inference in production. Increased flexibility for performance.
May 26, 2026
How to become a prompt engineer?
Becoming a prompt engineer is an achievable goal for curious tech profiles. Discover the key skills, the current market, and how to train yourself to master generative AI.
May 26, 2026
Hugging Face Introduces DOIs for Datasets and Models
Hugging Face now offers DOI (Digital Object Identifier) assignment for datasets and models, facilitating their citation and traceability in scientific research.
May 26, 2026
Personal Copilot: How to train your own coding assistant
Discover how to train your own personalized coding assistant using compact open-source models for increased security and better business relevance.
May 25, 2026
GLM-5.1: 754B parameters — Z.ai's agentic engineering flagship
Z.ai's GLM-5.1 is a 754B MoE model built for agentic engineering. It leads on SWE-Bench Pro (58.4%), scores 95.3% on AIME 2026, and is MIT licensed.
Apr 8, 2026
Gemma 4 31B: Google's multimodal model with 256K context and thinking mode
Google's Gemma 4 31B is a dense 30.7B multimodal model supporting text, images, and video with a 256K context window, native thinking mode, function calling, and 140+ languages — released under Apache
Apr 3, 2026
Chroma Context-1: the 20B agentic search model that edits its own context
Chroma's Context-1 is a 20B MoE agentic search model trained for multi-hop retrieval. It decomposes queries, calls tools in parallel, prunes irrelevant documents mid-search, and runs at up to 10x the
Apr 3, 2026
Cohere Transcribe: a 2B ASR model that tops the English leaderboard
Cohere Labs' Transcribe 03-2026 is a 2B Conformer-based ASR model ranked #1 on the English ASR leaderboard with a 5.42 average WER, supporting 14 languages at 524x real-time speed — faster and more ac
Apr 3, 2026
GLM-5: 744B parameters, 40B active — Z.ai's open-source frontier model
Z.ai's GLM-5 is a 744B MoE model with 40B active parameters, trained on 28.5T tokens. It scores 92.7% on AIME 2026, 77.8% on SWE-bench Verified, and is the best open-source model on HMMT Nov 2025.
Mar 30, 2026Voxtral-4B: Mistral's open-weights TTS model that speaks 9 languages in real time
Mistral just released Voxtral-4B-TTS, an open-weights text-to-speech model with 20 preset voices, 9 languages, and 70 ms latency — built for real-time voice agents and production deployment.
Mar 30, 2026
Qianfan-OCR: Baidu's 4B model that beats Gemini on document parsing
Baidu's Qianfan-OCR is a 4B end-to-end document understanding model that ranks #1 on OmniDocBench v1.5 — beating Gemini 3 Pro and DeepSeek-OCR-v2 on tables, formulas, layout, and key information extra
Mar 30, 2026
Qwen3.5-27B Distilled by Claude 4.6 Opus: A Local Reasoning Powerhouse
Discover how Jackrong distilled Claude 4.6 Opus reasoning into Qwen3.5-27B — a 28B open-source model that thinks for 9+ minutes autonomously, runs on a single GPU, and rivals frontier AI for coding an
Mar 24, 2026
Nemotron Cascade 2: NVIDIA's 30B model that won the math and coding Olympics
NVIDIA's Nemotron Cascade 2 is a 30B MoE model with only 3B activated parameters — and it just won gold medals at the 2025 International Mathematical Olympiad and International Olympiad in Informatics
Mar 24, 2026
NVIDIA Nemotron-3 Super: a 120B MoE model that runs on a single GPU
A deep dive into NVIDIA Nemotron-3 Super, a 120B-parameter Mixture of Experts model with only 12B active parameters, 1 million token context, and configurable reasoning — deployable on a single B200 G
Mar 20, 2026
Mistral Small 4: One Unified Model to Rule Reasoning, Code, and Vision
A deep dive into Mistral Small 4, the new model from Mistral AI that merges advanced reasoning, code generation, and multimodal capabilities into a 119-billion parameter Mixture of Experts architectur
Mar 18, 2026
Protect Yourself from the Enedis Scam
How to identify and avoid fraudulent calls impersonating Enedis.
Feb 21, 2026
The Photo Revolution on iPhone
Google launches a Snapseed camera app for iPhone with professional tools.
Feb 20, 2026Today's Tech News
ROG Strix SCAR 18, VPN and health: what you need to know.
Feb 19, 2026
Protect Your AI Conversations
How malicious extensions steal your ChatGPT history.
Jan 8, 2026Windows Ecosystem Embraces Android
A seamless fusion between your smartphone and your computer.
Jan 8, 2026