期望部署glm5.2,但报错如下:
ftllm server /mnt/nvme1n1/vllm/GLM-5.2-FP8 --dtype fp8 --moe_dtype fp8 --device cudapp=4
2026-06-21 08:11:48,515 81876 server.py[line:234] INFO: Namespace(command='server', version=False, model='/mnt/nvme1n1/vllm/GLM-5.2-FP8', path='', threads=-1, low=False, dtype='fp8', moe_dtype='fp8', moe_atype='', atype='auto', kv_cache_dtype='auto', cuda_embedding=False, kv_cache_limit='auto', max_batch=-1, chunked_prefill_size=-1, device='cudapp=4', tp='', moe_device='', moe_device_layers=-1, moe_experts=-1, cache_history='', cache_fast='', enable_thinking='', cuda_shared_expert='true', enable_amx='false', tokens=-1, page_size=-1, prefix_cache='', prefix_cache_snapshot_interval_pages=-1, prefix_cache_snapshot_max_per_request=-1, prefix_cache_snapshot_max_records=-1, gpu_mem_ratio=0.9, cuda_slab=0, mtp=0, custom='', lora='', cache_dir='', dtype_config='', ori='', tool_call_parser='auto', chat_template='', model_name='', host='0.0.0.0', port=8080, api_key='', temperature=None, top_p=None, top_k=None, repeat_penalty=None, think='false', hide_input=False, dev_mode=False)
[device] cudapp expand: cudapp=4 => {'cuda:0': 1, 'cuda:1': 1, 'cuda:2': 1, 'cuda:3': 1}
Load libnuma.so.1
CPU Instruction Info: [AVX2: ON] [AVX512F: ON] [AVX512_VNNI: ON] [AVX512_BF16: ON] [AMX: ON]
Load libfastllm_tools.so
FastLLM Error: Unsupport graph model type glm_moe_dsa
Press any key to exit...
FastLLM Error: Unsupport graph model type glm_moe_dsa^Cinto exit
^Cinto exit
Segmentation fault (core dumped)
期望部署glm5.2,但报错如下:
ftllm server /mnt/nvme1n1/vllm/GLM-5.2-FP8 --dtype fp8 --moe_dtype fp8 --device cudapp=4
2026-06-21 08:11:48,515 81876 server.py[line:234] INFO: Namespace(command='server', version=False, model='/mnt/nvme1n1/vllm/GLM-5.2-FP8', path='', threads=-1, low=False, dtype='fp8', moe_dtype='fp8', moe_atype='', atype='auto', kv_cache_dtype='auto', cuda_embedding=False, kv_cache_limit='auto', max_batch=-1, chunked_prefill_size=-1, device='cudapp=4', tp='', moe_device='', moe_device_layers=-1, moe_experts=-1, cache_history='', cache_fast='', enable_thinking='', cuda_shared_expert='true', enable_amx='false', tokens=-1, page_size=-1, prefix_cache='', prefix_cache_snapshot_interval_pages=-1, prefix_cache_snapshot_max_per_request=-1, prefix_cache_snapshot_max_records=-1, gpu_mem_ratio=0.9, cuda_slab=0, mtp=0, custom='', lora='', cache_dir='', dtype_config='', ori='', tool_call_parser='auto', chat_template='', model_name='', host='0.0.0.0', port=8080, api_key='', temperature=None, top_p=None, top_k=None, repeat_penalty=None, think='false', hide_input=False, dev_mode=False)
[device] cudapp expand: cudapp=4 => {'cuda:0': 1, 'cuda:1': 1, 'cuda:2': 1, 'cuda:3': 1}
Load libnuma.so.1
CPU Instruction Info: [AVX2: ON] [AVX512F: ON] [AVX512_VNNI: ON] [AVX512_BF16: ON] [AMX: ON]
Load libfastllm_tools.so
FastLLM Error: Unsupport graph model type glm_moe_dsa
Press any key to exit...
FastLLM Error: Unsupport graph model type glm_moe_dsa^Cinto exit
^Cinto exit
Segmentation fault (core dumped)