JAN 2025

AI MODEL TESTING LABORATORY

PRIVATE RESEARCH NOTES — DO NOT CIRCULATE

Re-run with Q5_K_S? —JB
Purchase second GPU? —JB
Best performer to date. —JB
Date Subject Apparatus Procedures Data Observations
15 Jan 2025
Llama 3.1 8B Instruct
Meta AI Foundation
NVIDIA RTX 3090
Ollama v0.4.0 · Q4_K_M quant
MMLU-Pro GPQA IFEval
ACCESS VRAM peaked at 14.2GB of 24GB. Throughput 47 tok/s sustained. Six-hour stability trial completed without incident. Quantization noise acceptable for reasoning benchmarks.
12 Jan 2025
Qwen2.5 14B Instruct
Alibaba DAMO Academy
RTX 3090 + 64GB system
llama.cpp b4000 · Q5_K_M
HumanEval MBPP MultiPL-E LiveCodeBench
ACCESS Eight layers offloaded to CPU. Throughput degraded to 23 tok/s. Code generation remarkably capable in C++ and Python. Rust compilation tasks failed on context length, not competence.
08 Jan 2025
Mistral Small 24B
Mistral AI, Paris
Dual RTX 3090 NVLink
vLLM 0.6.5 · BF16 native
MATH GSM8K BBH
ACCESS Native 24B architecture fits 48GB aggregate without quantization. 89 tok/s sustained throughput. Chain-of-thought reasoning markedly superior to any quantized competitor tested.

Supplementary Materials

Roleplaying evaluations, creative writing assessments, character consistency trials, and exploratory prompt engineering exercises. Materials unsuitable for standardized tabulation but valuable for qualitative analysis.

Retrieve from Filing Cabinet →
Primary Investigator
Approved for
internal distribution
Date of Review
— 1 —