Skip to content

希望部署glm5.2 #697

Description

@WangHHY19931001

期望部署glm5.2,但报错如下:
ftllm server /mnt/nvme1n1/vllm/GLM-5.2-FP8 --dtype fp8 --moe_dtype fp8 --device cudapp=4
2026-06-21 08:11:48,515 81876 server.py[line:234] INFO: Namespace(command='server', version=False, model='/mnt/nvme1n1/vllm/GLM-5.2-FP8', path='', threads=-1, low=False, dtype='fp8', moe_dtype='fp8', moe_atype='', atype='auto', kv_cache_dtype='auto', cuda_embedding=False, kv_cache_limit='auto', max_batch=-1, chunked_prefill_size=-1, device='cudapp=4', tp='', moe_device='', moe_device_layers=-1, moe_experts=-1, cache_history='', cache_fast='', enable_thinking='', cuda_shared_expert='true', enable_amx='false', tokens=-1, page_size=-1, prefix_cache='', prefix_cache_snapshot_interval_pages=-1, prefix_cache_snapshot_max_per_request=-1, prefix_cache_snapshot_max_records=-1, gpu_mem_ratio=0.9, cuda_slab=0, mtp=0, custom='', lora='', cache_dir='', dtype_config='', ori='', tool_call_parser='auto', chat_template='', model_name='', host='0.0.0.0', port=8080, api_key='', temperature=None, top_p=None, top_k=None, repeat_penalty=None, think='false', hide_input=False, dev_mode=False)
[device] cudapp expand: cudapp=4 => {'cuda:0': 1, 'cuda:1': 1, 'cuda:2': 1, 'cuda:3': 1}
Load libnuma.so.1
CPU Instruction Info: [AVX2: ON] [AVX512F: ON] [AVX512_VNNI: ON] [AVX512_BF16: ON] [AMX: ON]
Load libfastllm_tools.so
FastLLM Error: Unsupport graph model type glm_moe_dsa
Press any key to exit...
FastLLM Error: Unsupport graph model type glm_moe_dsa^Cinto exit
^Cinto exit
Segmentation fault (core dumped)

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions