NeuReality Blog

New

Reproducing (and Beating) a Published Inference Benchmark: A DeepSeek V4 Pro Debugging Story

I set out to do something that sounds simple: take a published benchmark for DeepSeek V4 Pro on B200, reproduce it, and see if we could beat it.

  • Cost per Token
  • Distributed Inference
  • GPU Scale-Out
  • GPU Utilization
  • Inference Networking
  • Inference Operating System
  • Inference Orchestration
  • LLM Serving
  • Token Factory
Read More
  • Introducing The Hidden Supply Podcast

    AI Infrastructure, Exposed

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    08.07.26

    |Written by Elan Neiger

    Read More
  • AI Coding Agents Are Exposing the Missing Operating Layer in Production AI

    The most visible AI use case is showing what production AI actually costs to run.

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    17.06.26

    |Written by Ronli Vignanski, Elan Neiger

    Read More
  • Why Standard NICs Fail AI Inference — And What an AI NIC Does Differently

    • AI NIC
    • AI-SuperNIC
    • GPU Scale-Out
    • In-Network Compute
    • Inference Networking
    • LLM Serving
    • SuperNIC
    • Ultra Ethernet

    01.06.26

    |Written by Lior Khermosh

    Read More
  • Behind the LLM Math: The Serving Equations That Actually Matter

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    11.05.26

    |Written by Or Zipori

    Read More
  • The Rise of the AI Token Factory — And Why Inference Infrastructure Must Change

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    05.05.26

    |Written by Elan Neiger

    Read More
  • LLM Serving: Why So Hard??

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    29.03.26

    |Written by Or Zipori

    Read More
  • LLM Inference Parallelism: A Salad of Acronyms

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    17.03.26

    |Written by Or Zipori

    Read More
  • What Is an AI SuperNIC, And Why It’s the Missing Piece in Your AI Infrastructure

    • AI NIC
    • AI-SuperNIC
    • GPU Scale-Out
    • In-Network Compute
    • Inference Networking
    • Inference Orchestration
    • SuperNIC
    • Ultra Ethernet

    24.03.26

    |Written by Elan Neiger

    Read More
  • Scaling LLM Inference with llm-d and NeuReality Inference Serving Stack

    In this blog post, we share our perspective on the emerging generative-AI-at-scale framework landscape, and describe our experience with the llm-d framework.

    • Cost per Token
    • Distributed Inference
    • Inference Networking
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving

    25.12.25

    |Written by Irit Fastovskiy, Vadim Eisenberg, Or Zipori

    Read More
  • Scaling Multimodal Pipelines: Efficient Vision Understanding for the AI Era

    12.11.25

    |Written by Irit Fastovskiy

    Read More
  • NeuReality Redefines the AI Head Node with Arm Neoverse V3

    NeuReality Redefines the AI Head Node with Arm Neoverse V3

    25.08.25

    Read More
  • Leveraging AI Inference for Transformative Telecom Solutions

    Leveraging AI Inference for Transformative Telecom Solutions

    23.07.25

    Read More
Loading posts…