Saksham Arora · Chimera

Tag: Inference

8 items with this tag.

  • Oct 01, 2026

    Gemini 4 Argon Costs the Same Per Token and Twice Per Task

    • AI
    • LLMs
    • Evals
    • Inference
  • Sep 28, 2026

    Ember-1 Cut 40% of the Tokens and Kept the Score

    • AI
    • LLMs
    • Inference
    • Evals
  • Sep 19, 2026

    CoreWeave Put 16,635 Tokens Per GPU on a Single Rack

    • Systems
    • Hardware
    • Inference
    • Networking
  • Aug 26, 2026

    OpenAI's First Chip Wins on Latency, Not FLOPs

    • Systems
    • Hardware
    • Inference
    • Latency
  • Aug 20, 2026

    Cerebras CS-4 Shipped a 2 Microsecond Interconnect, and That's the Number That Matters

    • Systems
    • Hardware
    • Inference
    • Latency
  • Jul 13, 2026

    GPT-5.6 Beats Claude by 13 Points on One Benchmark, Loses by 16 on Another

    • LLM
    • Benchmarks
    • Systems
    • Inference
  • Jul 10, 2026

    NVIDIA Killed the Draft Model and Got 4x Throughput For Free

    • LLM
    • Inference
    • Systems
    • Speculative-Decoding
  • Jul 06, 2026

    The 518-Token Bug: When GPT-5.5's Reasoning Gets Cut Off at Fixed Intervals

    • LLM
    • Inference
    • Systems
    • Debugging

  • GitHub
  • LinkedIn
  • X
  • Portfolio