NeuReality Blog

New

Reproducing (and Beating) a Published Inference Benchmark: A DeepSeek V4 Pro Debugging Story

I set out to do something that sounds simple: take a published benchmark for DeepSeek V4 Pro on B200, reproduce it, and see if we could beat it.

  • Cost per Token
  • Distributed Inference
  • GPU Scale-Out
  • GPU Utilization
  • Inference Networking
  • Inference Operating System
  • Inference Orchestration
  • LLM Serving
  • Token Factory
Read More
  • Introducing The Hidden Supply Podcast

    AI Infrastructure, Exposed

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    08.07.26

    |Written by Elan Neiger

    Read More
  • AI Coding Agents Are Exposing the Missing Operating Layer in Production AI

    The most visible AI use case is showing what production AI actually costs to run.

    • Cost per Token
    • Distributed Inference
    • GPU Utilization
    • Inference Operating System
    • Inference Orchestration
    • LLM Serving
    • Token Factory

    17.06.26

    |Written by Ronli Vignanski, Elan Neiger

    Read More
  • Why Standard NICs Fail AI Inference — And What an AI NIC Does Differently

    • AI NIC
    • AI-SuperNIC
    • GPU Scale-Out
    • In-Network Compute
    • Inference Networking
    • LLM Serving
    • SuperNIC
    • Ultra Ethernet

    01.06.26

    |Written by Lior Khermosh

    Read More
Loading posts…