Get API access

LLM benchmarks show improved reasoning and efficiency

20 statements, Jul 26, 2026 14:22 UTC to Sep 23, 2026 17:37 UTC
Analysis
Created
Updated
GOOGLEMICROSOFTNVIDIAMSFT
All topics

Statements (20)

  1. Bullish

    The technology of minimal-norm univariate two-layer ReLU networks with skip connections achieves exact global optimality for binary classification.

    arxiv.org
  2. Bullish

    CompKV framework improves sparse attention speedup for LLM inference by up to 6.85x over full attention.

    arxiv.org
  3. Bullish

    AI model training efficiency improves with hybrid algorithms combining sharded data parallelism and federated learning.

    arxiv.org
  4. Bullish

    CoRA-NAS framework achieves mean Spearman correlations of 0.786 on TransNAS-Bench-101, indicating improved accuracy prediction for neural architecture search.

    arxiv.org
  5. Bullish

    CoRA-NAS framework achieves mean Spearman correlations of 0.715 on NAS-Bench-101, indicating improved accuracy prediction for neural architecture search.

    arxiv.org
  6. Bullish

    CoRA-NAS framework achieves mean Spearman correlations of 0.946 on NAS-Bench-201, indicating improved accuracy prediction for neural architecture search.

    arxiv.org
  7. Bullish

    CoRA-NAS framework achieves mean Spearman correlations of 0.894 on NATS-SSS, indicating improved accuracy prediction for neural architecture search.

    arxiv.org
  8. Bullish

    CoRA-NAS framework achieves mean Spearman correlations of 0.715 on NAS-Bench-201/CIFAR-100, indicating improved accuracy prediction for neural architecture search.

    arxiv.org
  9. Bullish

    DSE-VTG achieves state-of-the-art performance on training-free video temporal grounding benchmarks.

    arxiv.org
  10. Bullish

    Sliding Window Attention with sinks requires no post-training and is extremely fast and low memory.

    arxiv.org
  11. Bullish

    Large Language Models using Sliding Window Attention with sinks achieve 2 to 10 times higher performance than Linear Attention models on long-context reasoning tasks.

    arxiv.org
  12. Bullish

    Switching to Sliding Window Attention instead of post-training Linear Attention models reduces inference memory cost.

    arxiv.org
  13. Bullish

    Reinforcement learning on tasks with verified statistical reward substantially improves reliability of LLM agents in hypothesis testing.

    arxiv.org
  14. Bullish

    Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines including GPT-5.4 and DeepSeekV4-Pro.

    arxiv.org
  15. Bullish

    Skill Entropy RL improves Qwen3-4B-Instruct on Skill^2-Bench from 34.4% to 68.4%.

    arxiv.org
  16. Bullish

    Compass achieves state-of-the-art performance on multimodal crack segmentation with 90% depth modality missing, using only 2.58M parameters.

    arxiv.org
  17. Neutral

    Behavioral Grammar uses a compact 0.88M-parameter causal Transformer for host runtime behavior analysis.

    arxiv.org
  18. Bullish

    Microsoft Corporation may benefit from improved mixed-precision post-training quantization performance for Vision Transformers.

    arxiv.org
  19. Bullish

    ViT-Small/16 and ResNet-50 models achieve the best validation accuracy when using a spectral inner step with a Muon outer step.

    arxiv.org
  20. Bullish

    SpecAHD reduces held-out objective cost by up to 57.7% on four routing problems using LLM backbones.

    arxiv.org