common_init_: fitting params to device memory ...
llama_prepare_model_devices: using device CUDA0 (GPU A) - 8192 MiB free
load_tensors: CUDA0 model buffer size = 0.00 MiB
load_tensors: CUDA_Host model buffer size = 0.00 MiB
llama_context: CUDA_Host output buffer size = 0.98 MiB
llama_kv_cache: CUDA0 KV buffer size = 0.00 MiB
llama_kv_cache: CUDA0 KV buffer size = 0.00 MiB
llama_context: CUDA0 compute buffer size = 71.02 MiB
llama_context: CUDA_Host compute buffer size = 13.02 MiB
common_fit_params: fitting params to free memory took 0.38 seconds
llama_prepare_model_devices: using device CUDA0 (GPU A) - 8192 MiB free
load_tensors: CPU_Mapped model buffer size = 461.43 MiB
load_tensors: CUDA0 model buffer size = 1623.70 MiB
llama_context: CUDA_Host output buffer size = 0.98 MiB
llama_kv_cache: CUDA0 KV buffer size = 104.00 MiB
llama_kv_cache: CUDA0 KV buffer size = 104.00 MiB
llama_context: CUDA0 compute buffer size = 71.02 MiB
llama_context: CUDA_Host compute buffer size = 13.02 MiB
