AI_MODEL_TESTING_LAB

_LOCAL_INFERENCE_BENCHMARKING_ARCHAEOLOGY_

Uptime
GPU RTX_3090
47% VRAM
3 Tests_Today

Model_Tests_Results // LIVE_FEED

_DATE_ _MODEL_ID_ _RESOURCE_ALLOC_ _TEST_SUITE_ _OUTPUT_ _FIELD_NOTES_
2025.01.15_03:47 Llama-3.1-8B-Instruct meta-llama/Llama-3.1-8B RTX 3090 24GB Ollama v0.4.0 · Q4_K_M MMLU-Pro GPQA IFEval [ACCESS_LOG] → Peak VRAM 14.2GB. 47 tok/s sustained. Temp 0.7, top_p 0.9. Six hour stress test completed without OOM. Quantization artifacts minimal on reasoning tasks.
2025.01.12_22:15 Qwen2.5-14B-Instruct Qwen/Qwen2.5-14B-Instruct RTX 3090 + 64GB RAM llama.cpp b4000 · Q5_K_M HumanEval MBPP MultiPL-E LiveCodeBench [ACCESS_LOG] → Partial CPU offloading required, 23 tok/s with 8 layers on CPU. Coding performance exceptional on C++ and Python. Rust compilation tests failed due to context length limits, not capability.
2025.01.08_17:33 Mistral-Small-24B-Instruct mistralai/Mistral-Small-Instruct-2409 Dual RTX 3090 NVLink vLLM 0.6.5 · BF16 MATH GSM8K BBH [ACCESS_LOG] → Native 24B architecture, no quantization. 48GB combined VRAM fully utilized. 89 tok/s sustained. Chain-of-thought reasoning superior to quantized alternatives. NVLink overhead negligible.

OTHER_TESTS_PORTAL

Extended evaluations, roleplaying scenarios, creative writing benchmarks, character card tests, and exploratory prompt engineering sessions that resist standardization.

Initialize_Transfer