_LOCAL_INFERENCE_BENCHMARKING_ARCHAEOLOGY_
| _DATE_ | _MODEL_ID_ | _RESOURCE_ALLOC_ | _TEST_SUITE_ | _OUTPUT_ | _FIELD_NOTES_ |
|---|---|---|---|---|---|
| 2025.01.15_03:47 | Llama-3.1-8B-Instruct meta-llama/Llama-3.1-8B | RTX 3090 24GB Ollama v0.4.0 · Q4_K_M | MMLU-Pro GPQA IFEval | [ACCESS_LOG] → | Peak VRAM 14.2GB. 47 tok/s sustained. Temp 0.7, top_p 0.9. Six hour stress test completed without OOM. Quantization artifacts minimal on reasoning tasks. |
| 2025.01.12_22:15 | Qwen2.5-14B-Instruct Qwen/Qwen2.5-14B-Instruct | RTX 3090 + 64GB RAM llama.cpp b4000 · Q5_K_M | HumanEval MBPP MultiPL-E LiveCodeBench | [ACCESS_LOG] → | Partial CPU offloading required, 23 tok/s with 8 layers on CPU. Coding performance exceptional on C++ and Python. Rust compilation tests failed due to context length limits, not capability. |
| 2025.01.08_17:33 | Mistral-Small-24B-Instruct mistralai/Mistral-Small-Instruct-2409 | Dual RTX 3090 NVLink vLLM 0.6.5 · BF16 | MATH GSM8K BBH | [ACCESS_LOG] → | Native 24B architecture, no quantization. 48GB combined VRAM fully utilized. 89 tok/s sustained. Chain-of-thought reasoning superior to quantized alternatives. NVLink overhead negligible. |
Extended evaluations, roleplaying scenarios, creative writing benchmarks, character card tests, and exploratory prompt engineering sessions that resist standardization.
Initialize_Transfer_SEGMENT_FAULT_RECOVERED_ © 2025 AI_MODEL_TESTING_LAB