Skip to content

vllm.config.vllm

Classes:

Functions:

OptimizationLevel

Bases: IntEnum

Optimization level enum.

Attributes:

  • O0 –

    O0 : No optimization. no compilation, no cudagraphs, no other

  • O1 –

    O1: Quick optimizations. Dynamo+Inductor compilation and Piecewise

  • O2 –

    O2: Full optimizations. -O1 as well as Full and Piecewise cudagraphs.

  • O3 –

    O3: Currently the same as -O2s.

Source code in vllm/config/vllm.py
class OptimizationLevel(IntEnum):
    """Optimization level enum."""

    O0 = 0
    """O0 : No optimization. no compilation, no cudagraphs, no other
    optimization, just starting up immediately"""
    O1 = 1
    """O1: Quick optimizations. Dynamo+Inductor compilation and Piecewise
    cudagraphs"""
    O2 = 2
    """O2: Full optimizations. -O1 as well as Full and Piecewise cudagraphs."""
    O3 = 3
    """O3: Currently the same as -O2s."""

O0 = 0 class-attribute instance-attribute

O0 : No optimization. no compilation, no cudagraphs, no other optimization, just starting up immediately

O1 = 1 class-attribute instance-attribute

O1: Quick optimizations. Dynamo+Inductor compilation and Piecewise cudagraphs

O2 = 2 class-attribute instance-attribute

O2: Full optimizations. -O1 as well as Full and Piecewise cudagraphs.

O3 = 3 class-attribute instance-attribute

O3: Currently the same as -O2s.

VllmConfig

Dataclass which contains all vllm-related configuration. This simplifies passing around the distinct configurations in the codebase.

Methods:

Attributes:

Source code in vllm/config/vllm.py
 346
 347
 348
 349
 350
 351
 352
 353
 354
 355
 356
 357
 358
 359
 360
 361
 362
 363
 364
 365
 366
 367
 368
 369
 370
 371
 372
 373
 374
 375
 376
 377
 378
 379
 380
 381
 382
 383
 384
 385
 386
 387
 388
 389
 390
 391
 392
 393
 394
 395
 396
 397
 398
 399
 400
 401
 402
 403
 404
 405
 406
 407
 408
 409
 410
 411
 412
 413
 414
 415
 416
 417
 418
 419
 420
 421
 422
 423
 424
 425
 426
 427
 428
 429
 430
 431
 432
 433
 434
 435
 436
 437
 438
 439
 440
 441
 442
 443
 444
 445
 446
 447
 448
 449
 450
 451
 452
 453
 454
 455
 456
 457
 458
 459
 460
 461
 462
 463
 464
 465
 466
 467
 468
 469
 470
 471
 472
 473
 474
 475
 476
 477
 478
 479
 480
 481
 482
 483
 484
 485
 486
 487
 488
 489
 490
 491
 492
 493
 494
 495
 496
 497
 498
 499
 500
 501
 502
 503
 504
 505
 506
 507
 508
 509
 510
 511
 512
 513
 514
 515
 516
 517
 518
 519
 520
 521
 522
 523
 524
 525
 526
 527
 528
 529
 530
 531
 532
 533
 534
 535
 536
 537
 538
 539
 540
 541
 542
 543
 544
 545
 546
 547
 548
 549
 550
 551
 552
 553
 554
 555
 556
 557
 558
 559
 560
 561
 562
 563
 564
 565
 566
 567
 568
 569
 570
 571
 572
 573
 574
 575
 576
 577
 578
 579
 580
 581
 582
 583
 584
 585
 586
 587
 588
 589
 590
 591
 592
 593
 594
 595
 596
 597
 598
 599
 600
 601
 602
 603
 604
 605
 606
 607
 608
 609
 610
 611
 612
 613
 614
 615
 616
 617
 618
 619
 620
 621
 622
 623
 624
 625
 626
 627
 628
 629
 630
 631
 632
 633
 634
 635
 636
 637
 638
 639
 640
 641
 642
 643
 644
 645
 646
 647
 648
 649
 650
 651
 652
 653
 654
 655
 656
 657
 658
 659
 660
 661
 662
 663
 664
 665
 666
 667
 668
 669
 670
 671
 672
 673
 674
 675
 676
 677
 678
 679
 680
 681
 682
 683
 684
 685
 686
 687
 688
 689
 690
 691
 692
 693
 694
 695
 696
 697
 698
 699
 700
 701
 702
 703
 704
 705
 706
 707
 708
 709
 710
 711
 712
 713
 714
 715
 716
 717
 718
 719
 720
 721
 722
 723
 724
 725
 726
 727
 728
 729
 730
 731
 732
 733
 734
 735
 736
 737
 738
 739
 740
 741
 742
 743
 744
 745
 746
 747
 748
 749
 750
 751
 752
 753
 754
 755
 756
 757
 758
 759
 760
 761
 762
 763
 764
 765
 766
 767
 768
 769
 770
 771
 772
 773
 774
 775
 776
 777
 778
 779
 780
 781
 782
 783
 784
 785
 786
 787
 788
 789
 790
 791
 792
 793
 794
 795
 796
 797
 798
 799
 800
 801
 802
 803
 804
 805
 806
 807
 808
 809
 810
 811
 812
 813
 814
 815
 816
 817
 818
 819
 820
 821
 822
 823
 824
 825
 826
 827
 828
 829
 830
 831
 832
 833
 834
 835
 836
 837
 838
 839
 840
 841
 842
 843
 844
 845
 846
 847
 848
 849
 850
 851
 852
 853
 854
 855
 856
 857
 858
 859
 860
 861
 862
 863
 864
 865
 866
 867
 868
 869
 870
 871
 872
 873
 874
 875
 876
 877
 878
 879
 880
 881
 882
 883
 884
 885
 886
 887
 888
 889
 890
 891
 892
 893
 894
 895
 896
 897
 898
 899
 900
 901
 902
 903
 904
 905
 906
 907
 908
 909
 910
 911
 912
 913
 914
 915
 916
 917
 918
 919
 920
 921
 922
 923
 924
 925
 926
 927
 928
 929
 930
 931
 932
 933
 934
 935
 936
 937
 938
 939
 940
 941
 942
 943
 944
 945
 946
 947
 948
 949
 950
 951
 952
 953
 954
 955
 956
 957
 958
 959
 960
 961
 962
 963
 964
 965
 966
 967
 968
 969
 970
 971
 972
 973
 974
 975
 976
 977
 978
 979
 980
 981
 982
 983
 984
 985
 986
 987
 988
 989
 990
 991
 992
 993
 994
 995
 996
 997
 998
 999
1000
1001
1002
1003
1004
1005
1006
1007
1008
1009
1010
1011
1012
1013
1014
1015
1016
1017
1018
1019
1020
1021
1022
1023
1024
1025
1026
1027
1028
1029
1030
1031
1032
1033
1034
1035
1036
1037
1038
1039
1040
1041
1042
1043
1044
1045
1046
1047
1048
1049
1050
1051
1052
1053
1054
1055
1056
1057
1058
1059
1060
1061
1062
1063
1064
1065
1066
1067
1068
1069
1070
1071
1072
1073
1074
1075
1076
1077
1078
1079
1080
1081
1082
1083
1084
1085
1086
1087
1088
1089
1090
1091
1092
1093
1094
1095
1096
1097
1098
1099
1100
1101
1102
1103
1104
1105
1106
1107
1108
1109
1110
1111
1112
1113
1114
1115
1116
1117
1118
1119
1120
1121
1122
1123
1124
1125
1126
1127
1128
1129
1130
1131
1132
1133
1134
1135
1136
1137
1138
1139
1140
1141
1142
1143
1144
1145
1146
1147
1148
1149
1150
1151
1152
1153
1154
1155
1156
1157
1158
1159
1160
1161
1162
1163
1164
1165
1166
1167
1168
1169
1170
1171
1172
1173
1174
1175
1176
1177
1178
1179
1180
1181
1182
1183
1184
1185
1186
1187
1188
1189
1190
1191
1192
1193
1194
1195
1196
1197
1198
1199
1200
1201
1202
1203
1204
1205
1206
1207
1208
1209
1210
1211
1212
1213
1214
1215
1216
1217
1218
1219
1220
1221
1222
1223
1224
1225
1226
1227
1228
1229
1230
1231
1232
1233
1234
1235
1236
1237
1238
1239
1240
1241
1242
1243
1244
1245
1246
1247
1248
1249
1250
1251
1252
1253
1254
1255
1256
1257
1258
1259
1260
1261
1262
1263
1264
1265
1266
1267
1268
1269
1270
1271
1272
1273
1274
1275
1276
1277
1278
1279
1280
1281
1282
1283
1284
1285
1286
1287
1288
1289
1290
1291
1292
1293
1294
1295
1296
1297
1298
1299
1300
1301
1302
1303
1304
1305
1306
1307
1308
1309
1310
1311
1312
1313
1314
1315
1316
1317
1318
1319
1320
1321
1322
1323
1324
1325
1326
1327
1328
1329
1330
1331
1332
1333
1334
1335
1336
1337
1338
1339
1340
1341
1342
1343
1344
1345
1346
1347
1348
1349
1350
1351
1352
1353
1354
1355
1356
1357
1358
1359
1360
1361
1362
1363
1364
1365
1366
1367
1368
1369
1370
1371
1372
1373
1374
1375
1376
1377
1378
1379
1380
1381
1382
1383
1384
1385
1386
1387
1388
1389
1390
1391
1392
1393
1394
1395
1396
1397
1398
1399
1400
1401
1402
1403
1404
1405
1406
1407
1408
1409
1410
1411
1412
1413
1414
1415
1416
1417
1418
1419
1420
1421
1422
1423
1424
1425
1426
1427
1428
1429
1430
1431
1432
1433
1434
1435
1436
1437
1438
1439
1440
1441
1442
1443
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
2308
2309
2310
2311
2312
2313
2314
2315
2316
2317
2318
2319
2320
2321
2322
2323
2324
2325
2326
2327
2328
2329
2330
2331
2332
2333
2334
2335
2336
2337
2338
2339
2340
2341
2342
2343
2344
2345
2346
2347
2348
2349
2350
2351
2352
2353
2354
2355
2356
2357
2358
2359
2360
2361
2362
2363
2364
2365
2366
2367
2368
2369
2370
2371
2372
2373
2374
2375
2376
2377
2378
2379
2380
2381
2382
2383
2384
2385
2386
2387
2388
2389
2390
2391
2392
2393
2394
2395
2396
2397
2398
2399
2400
2401
2402
2403
2404
2405
2406
2407
2408
2409
2410
2411
2412
2413
2414
2415
2416
2417
2418
2419
2420
2421
2422
2423
2424
2425
2426
2427
2428
2429
2430
2431
2432
2433
2434
2435
2436
2437
2438
2439
2440
2441
2442
2443
2444
2445
2446
2447
2448
2449
2450
2451
2452
2453
2454
2455
2456
2457
2458
2459
2460
2461
2462
2463
2464
2465
2466
2467
2468
2469
2470
2471
2472
2473
2474
2475
2476
2477
2478
2479
2480
2481
2482
2483
2484
2485
2486
2487
2488
2489
2490
2491
2492
2493
2494
2495
2496
2497
2498
2499
2500
2501
2502
2503
2504
2505
2506
2507
2508
2509
2510
2511
2512
2513
2514
2515
2516
2517
2518
2519
2520
2521
2522
2523
2524
2525
2526
2527
2528
2529
2530
2531
2532
2533
2534
2535
2536
2537
2538
2539
2540
2541
2542
2543
2544
2545
2546
2547
2548
2549
2550
2551
2552
2553
2554
2555
2556
2557
2558
2559
2560
2561
2562
2563
2564
2565
2566
2567
2568
2569
2570
2571
2572
2573
2574
2575
2576
2577
2578
2579
2580
2581
2582
2583
2584
2585
2586
2587
2588
2589
2590
2591
2592
2593
2594
2595
2596
2597
2598
2599
2600
2601
2602
2603
2604
2605
2606
2607
2608
2609
2610
2611
2612
2613
2614
2615
2616
2617
2618
2619
2620
2621
2622
2623
2624
2625
2626
2627
2628
2629
2630
2631
2632
2633
2634
2635
2636
2637
2638
2639
2640
2641
2642
2643
2644
2645
2646
2647
2648
2649
2650
2651
2652
2653
2654
2655
2656
2657
2658
2659
2660
2661
2662
2663
2664
2665
2666
2667
2668
2669
2670
2671
2672
2673
2674
2675
2676
2677
2678
2679
2680
2681
2682
2683
2684
2685
2686
2687
2688
2689
2690
2691
2692
2693
2694
2695
2696
2697
2698
2699
2700
2701
2702
2703
2704
2705
2706
2707
2708
2709
2710
2711
2712
2713
2714
2715
2716
2717
2718
2719
2720
2721
2722
2723
2724
2725
2726
2727
2728
2729
2730
2731
2732
2733
2734
2735
2736
2737
2738
2739
2740
2741
2742
2743
2744
2745
2746
2747
2748
2749
2750
2751
2752
2753
2754
2755
2756
2757
2758
2759
2760
2761
2762
2763
2764
2765
2766
2767
2768
2769
2770
2771
2772
2773
2774
2775
2776
2777
2778
2779
2780
2781
2782
2783
2784
2785
2786
2787
2788
2789
2790
2791
2792
2793
2794
2795
2796
2797
2798
2799
2800
2801
2802
2803
2804
2805
2806
2807
2808
2809
2810
2811
2812
2813
2814
2815
2816
2817
2818
2819
2820
2821
2822
2823
2824
2825
2826
2827
2828
2829
2830
2831
2832
2833
2834
2835
2836
2837
2838
2839
2840
2841
2842
2843
2844
2845
2846
2847
2848
2849
2850
2851
2852
2853
2854
2855
2856
2857
2858
2859
2860
2861
2862
2863
2864
2865
2866
2867
2868
2869
2870
2871
2872
2873
2874
2875
2876
2877
2878
2879
2880
2881
2882
2883
2884
2885
2886
2887
2888
2889
2890
2891
2892
2893
2894
2895
2896
2897
2898
2899
2900
2901
2902
2903
2904
2905
2906
2907
2908
2909
2910
2911
2912
2913
2914
2915
2916
2917
2918
2919
2920
2921
2922
2923
2924
2925
2926
2927
2928
2929
2930
2931
2932
2933
2934
2935
2936
2937
2938
2939
2940
2941
2942
2943
2944
2945
2946
2947
2948
2949
2950
2951
2952
2953
2954
2955
2956
2957
2958
2959
2960
2961
2962
2963
2964
2965
2966
2967
2968
2969
2970
2971
2972
2973
2974
2975
2976
2977
2978
2979
2980
2981
2982
2983
2984
2985
2986
2987
2988
2989
2990
2991
2992
2993
2994
2995
2996
2997
2998
2999
3000
3001
3002
3003
3004
3005
3006
3007
3008
3009
3010
3011
3012
3013
3014
3015
3016
3017
3018
3019
3020
3021
3022
3023
3024
3025
3026
3027
3028
3029
3030
3031
3032
3033
3034
3035
3036
3037
3038
3039
3040
3041
3042
3043
3044
3045
3046
3047
3048
3049
3050
3051
3052
3053
3054
3055
3056
3057
3058
3059
3060
3061
3062
3063
3064
3065
3066
3067
3068
3069
3070
3071
3072
3073
3074
3075
3076
3077
3078
3079
3080
3081
3082
3083
3084
3085
3086
3087
3088
3089
3090
3091
3092
3093
3094
3095
3096
3097
3098
3099
3100
3101
3102
3103
3104
3105
3106
3107
3108
3109
3110
3111
3112
3113
3114
3115
3116
3117
3118
3119
3120
3121
3122
3123
3124
3125
3126
3127
3128
3129
3130
3131
3132
3133
3134
3135
3136
3137
3138
3139
3140
3141
3142
3143
3144
3145
3146
3147
3148
3149
3150
3151
3152
3153
3154
3155
3156
3157
3158
3159
3160
3161
3162
3163
3164
3165
3166
3167
3168
3169
3170
3171
3172
3173
3174
3175
3176
3177
3178
3179
3180
3181
3182
3183
3184
3185
3186
3187
3188
3189
3190
3191
3192
3193
3194
3195
3196
3197
3198
3199
3200
3201
3202
3203
3204
3205
3206
3207
3208
3209
3210
3211
3212
3213
3214
3215
3216
3217
3218
3219
3220
3221
3222
3223
3224
3225
3226
3227
3228
3229
3230
3231
3232
3233
3234
3235
3236
3237
3238
3239
3240
3241
3242
3243
3244
3245
3246
3247
3248
3249
3250
3251
3252
3253
3254
3255
3256
3257
3258
3259
3260
3261
3262
3263
3264
3265
3266
3267
3268
3269
3270
3271
3272
3273
3274
3275
3276
3277
3278
3279
3280
3281
3282
3283
3284
3285
3286
3287
3288
3289
3290
3291
3292
3293
3294
3295
3296
3297
3298
3299
3300
3301
3302
3303
3304
3305
3306
3307
3308
3309
3310
3311
3312
3313
3314
3315
3316
3317
3318
3319
3320
3321
3322
3323
3324
3325
3326
3327
3328
3329
3330
3331
3332
3333
3334
3335
3336
3337
3338
3339
3340
3341
3342
3343
3344
3345
3346
3347
3348
3349
3350
3351
3352
3353
3354
3355
3356
3357
3358
3359
3360
3361
3362
3363
3364
3365
3366
3367
3368
3369
3370
3371
3372
3373
3374
3375
3376
3377
3378
3379
3380
3381
3382
3383
3384
3385
3386
3387
3388
3389
3390
3391
3392
3393
3394
3395
3396
3397
3398
3399
3400
3401
3402
3403
3404
3405
3406
3407
3408
3409
3410
3411
3412
3413
3414
3415
3416
3417
3418
3419
3420
3421
3422
3423
3424
3425
3426
3427
3428
3429
3430
3431
3432
3433
3434
3435
3436
3437
3438
3439
3440
3441
3442
3443
3444
3445
3446
3447
3448
3449
3450
3451
3452
3453
3454
3455
3456
3457
3458
3459
3460
3461
3462
3463
3464
3465
3466
3467
3468
@config(config=ConfigDict(arbitrary_types_allowed=True))
class VllmConfig:
    """Dataclass which contains all vllm-related configuration. This
    simplifies passing around the distinct configurations in the codebase.
    """

    # TODO: use default_factory once default constructing ModelConfig doesn't
    # try to download a model
    model_config: ModelConfig = None  # type: ignore[assignment]
    """Model configuration."""
    cache_config: CacheConfig = Field(default_factory=CacheConfig)
    """Cache configuration."""
    parallel_config: ParallelConfig = Field(default_factory=ParallelConfig)
    """Parallel configuration."""
    scheduler_config: SchedulerConfig = Field(
        default_factory=SchedulerConfig.default_factory,
    )
    """Scheduler configuration."""
    device_config: DeviceConfig = Field(default_factory=DeviceConfig)
    """Device configuration."""
    load_config: LoadConfig = Field(default_factory=LoadConfig)
    """Load configuration."""
    offload_config: OffloadConfig = Field(default_factory=OffloadConfig)
    """Model weight offloading configuration."""
    attention_config: AttentionConfig = Field(default_factory=AttentionConfig)
    """Attention configuration."""
    aux_output_config: AuxOutputConfig = Field(default_factory=AuxOutputConfig)
    """Execution auxiliary output configuration."""
    engram_config: EngramConfig | None = None
    """N-gram embedding storage and sharding settings."""
    mamba_config: MambaConfig = Field(default_factory=MambaConfig)
    """Mamba configuration."""
    kernel_config: KernelConfig = Field(default_factory=KernelConfig)
    """Kernel configuration."""
    lora_config: LoRAConfig | None = None
    """LoRA configuration."""
    speculative_config: SpeculativeConfig | None = None
    """Speculative decoding configuration."""
    watermark_config: WatermarkConfig | None = None
    """Text watermarking configuration."""
    diffusion_config: DiffusionConfig | None = None
    """Diffusion LLM (dLLM) configuration."""

    structured_outputs_config: StructuredOutputsConfig = Field(
        default_factory=StructuredOutputsConfig
    )
    """Structured outputs configuration."""
    observability_config: ObservabilityConfig = Field(
        default_factory=ObservabilityConfig
    )
    """Observability configuration."""
    logging_config: LoggingConfig = Field(default_factory=LoggingConfig)
    """Logging configuration."""
    quant_config: QuantizationConfig | None = None
    """Quantization configuration."""
    compilation_config: CompilationConfig = Field(default_factory=CompilationConfig)
    """`torch.compile` and cudagraph capture configuration for the model.

    As a shorthand, one can append compilation arguments via
    -cc.parameter=argument such as `-cc.mode=3` (same as `-cc='{"mode":3}'`).

    You can specify the full compilation config like so:
    `{"mode": 3, "cudagraph_capture_sizes": [1, 2, 4, 8]}`
    """
    profiler_config: ProfilerConfig = Field(default_factory=ProfilerConfig)
    """Profiling configuration."""
    kv_transfer_config: KVTransferConfig | None = None
    """The configurations for distributed KV cache transfer."""
    kv_events_config: KVEventsConfig | None = None
    """The configurations for event publishing."""
    ec_transfer_config: ECTransferConfig | None = None
    """The configurations for distributed EC cache transfer."""
    ec_manager_config: EncoderCacheManagerConfig = Field(
        default_factory=EncoderCacheManagerConfig
    )
    """The configurations for custom encoder cache manager."""
    reasoning_config: ReasoningConfig | None = None
    """The configurations for reasoning model."""
    # some opaque config, only used to provide additional information
    # for the hash computation, mainly used for testing, debugging or out of
    # tree config registration.
    additional_config: dict | SupportsHash = Field(default_factory=dict)
    """Additional config for specified platform. Different platforms may
    support different configs. Make sure the configs are valid for the platform
    you are using. Contents must be hashable."""
    instance_id: str = ""
    """The ID of the vLLM instance."""
    optimization_level: OptimizationLevel = OptimizationLevel.O2
    """The optimization level. These levels trade startup time cost for
    performance, with -O0 having the best startup time and -O3 having the best
    performance. -O2 is used by default. See OptimizationLevel for full
    description."""

    performance_mode: PerformanceMode = "balanced"
    """Performance mode for runtime behavior, 'balanced' is the default.
    'interactivity' favors low end-to-end per-request latency at small batch
    sizes (fine-grained CUDA graphs, latency-oriented kernels).
    'throughput' favors aggregate tokens/sec at high concurrency (larger CUDA
    graphs, more aggressive batching, throughput-oriented kernels)."""

    weight_transfer_config: WeightTransferConfig | None = None
    """The configurations for weight transfer during RL training."""

    shutdown_timeout: int = Field(default=0, ge=0)
    """Shutdown grace period for in-flight requests. Shutdown will be delayed for
    up to this amount of time to allow already-running requests to complete. Any
    remaining requests are aborted once the timeout is reached.
    """

    def compute_hash(self, include_version: bool = True) -> str:
        """WARNING: Whenever a new field is added to this config,
        ensure that it is included in the factors list if
        it affects the computation graph.

        Provide a hash that uniquely identifies all the configs
        that affect the structure of the computation
        graph from input ids/embeddings to the final hidden states,
        excluding anything before input ids/embeddings and after
        the final hidden states.

        Args:
            include_version: Include the vLLM version in the hash.

        """
        factors: list[Any] = []

        # summarize vllm config
        vllm_factors: list[Any] = []
        if include_version:
            from vllm import __version__

            vllm_factors.append(__version__)
        if self.model_config:
            vllm_factors.append(self.model_config.compute_hash())
            if (
                self.compilation_config
                and getattr(self.compilation_config, "compile_mm_encoder", False)
                and self.model_config.multimodal_config
            ):
                vllm_factors.append(self.model_config.multimodal_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.cache_config:
            vllm_factors.append(self.cache_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.parallel_config:
            vllm_factors.append(self.parallel_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.scheduler_config:
            vllm_factors.append(self.scheduler_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.device_config:
            vllm_factors.append(self.device_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.load_config:
            vllm_factors.append(self.load_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.offload_config:
            vllm_factors.append(self.offload_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.attention_config:
            vllm_factors.append(self.attention_config.compute_hash())
        else:
            vllm_factors.append("None")
        vllm_factors.append(
            self.engram_config.compute_hash()
            if self.engram_config is not None
            else "None"
        )
        if self.lora_config:
            vllm_factors.append(self.lora_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.speculative_config:
            vllm_factors.append(self.speculative_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.structured_outputs_config:
            vllm_factors.append(self.structured_outputs_config.compute_hash())
        if self.profiler_config:
            vllm_factors.append(self.profiler_config.compute_hash())
        else:
            vllm_factors.append("None")
        vllm_factors.append(self.observability_config.compute_hash())
        if self.quant_config:
            pass  # should be captured by model_config.quantization
        if self.compilation_config:
            vllm_factors.append(self.compilation_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.kernel_config:
            vllm_factors.append(self.kernel_config.compute_hash())
        else:
            vllm_factors.append(None)
        if self.kv_transfer_config:
            vllm_factors.append(self.kv_transfer_config.compute_hash())
        else:
            vllm_factors.append("None")
        if self.ec_transfer_config:
            vllm_factors.append(self.ec_transfer_config.compute_hash())
        else:
            vllm_factors.append("None")
        vllm_factors.append(self.aux_output_config.compute_hash())
        if self.additional_config:
            if isinstance(additional_config := self.additional_config, dict):
                additional_config_hash = safe_hash(
                    json.dumps(additional_config, sort_keys=True).encode(),
                    usedforsecurity=False,
                ).hexdigest()
            else:
                additional_config_hash = additional_config.compute_hash()
            vllm_factors.append(additional_config_hash)
        else:
            vllm_factors.append("None")
        factors.append(vllm_factors)

        hash_str = safe_hash(str(factors).encode(), usedforsecurity=False).hexdigest()[
            :10
        ]
        return hash_str

    @property
    def is_mm_encoder_only(self) -> bool:
        mm_config = (
            self.model_config.multimodal_config
            if self.model_config is not None
            else None
        )
        return bool(mm_config and mm_config.mm_encoder_only)

    @property
    def max_concurrent_batches(self) -> int:
        # PP requires PP-size concurrent batches to fill the pipeline.
        # Async scheduling requires 2 concurrent batches to overlap.
        pp_size = self.parallel_config.pipeline_parallel_size
        if self.scheduler_config.async_scheduling:
            if self.use_v2_model_runner:
                return pp_size + 1
            # V1 Model Runner does not fully support async scheduling with PP.
            if pp_size <= 1:
                return 2
        return pp_size

    @property
    def max_in_flight_tokens(self) -> int:
        # Upper bound on tokens that are scheduled but not yet settled (freed):
        # every concurrent batch may hold up to a full `max_num_batched_tokens`.
        # Recycling-aware KV cache specs (sliding-window, chunked-local) reserve
        # for this because out-of-window blocks are freed on the processed-token
        # basis, so in-flight steps transiently keep their blocks.
        return (
            self.max_concurrent_batches * self.scheduler_config.max_num_batched_tokens
        )

    @property
    def num_speculative_tokens(self) -> int:
        if (
            self.speculative_config is not None
            and self.speculative_config.num_speculative_tokens is not None
        ):
            return self.speculative_config.num_speculative_tokens
        if (
            self.diffusion_config is not None
            and self.diffusion_config.canvas_length is not None
        ):
            return self.diffusion_config.canvas_length
        return 0

    @property
    def num_lookahead_tokens(self) -> int:
        """KV slots to reserve past the tokens the target model is scheduled for.

        The drafter writes KV for positions beyond the target model's query
        range, so every component that reserves blocks must add this margin:
        the scheduler through `allocate_slots`, and the worker warmup, which
        builds its own `SchedulerOutput`s. Consumers must read this property
        rather than re-deriving their own per-method lookahead, so the
        scheduler and warmup cannot drift apart.
        """
        speculative_config = self.speculative_config
        if speculative_config is None:
            return 0
        dspark_fill_in = speculative_config.use_dspark() and not getattr(
            speculative_config.draft_model_config.hf_config, "sample_from_anchor", True
        )
        if speculative_config.use_dflash() or dspark_fill_in:
            # Fill-in drafting uses a bonus query plus one query per draft token.
            # DSpark's anchor-sampling layout does not need the extra slot.
            return self.num_speculative_tokens + 1
        if speculative_config.use_eagle() or speculative_config.uses_draft_model():
            # Anchor-sampling DSpark drafts a block of num_speculative_tokens
            # query tokens in which the anchor itself is the first prediction
            # position (no separate bonus query), so it needs exactly
            # num_speculative_tokens lookahead slots.
            return self.num_speculative_tokens
        return 0

    @property
    def num_prefill_lookahead_tokens(self) -> int:
        """Prefill tokens past the computed range that the drafter reads.

        Mid-prefill the drafter consumes tokens the target model has not been
        scheduled for yet, so every component that has to keep them available
        must apply this margin: the scheduler, which never ends a chunk within
        it and shifts encoder scheduling by it, and the KV cache manager, which
        treats the trailing `this - 1` tokens as re-prefillable rather than
        finalized. Consumers must read this property rather than re-deriving
        their own per-method lookahead, so those components cannot drift apart.
        """
        speculative_config = self.speculative_config
        if speculative_config is None or not speculative_config.use_eagle():
            return 0
        if speculative_config.use_multi_module_mtp():
            # Each MTP module reads one token further ahead than the one before
            # it, so the chain needs num_speculative_tokens of runway at a
            # chunked-prefill boundary.
            return self.num_speculative_tokens
        # Eagle-family drafters read only the immediate next token.
        return 1

    @property
    def uniform_decode_query_len(self) -> int:
        """Query length of every request in a uniform decode batch.

        A decode step submits one query for the newly sampled token plus one
        for each draft token, so the widest uniform decode batch the scheduler
        can build is `max_num_seqs * uniform_decode_query_len` tokens. Anything
        that has to cover a decode batch reads this, so the sizing rule cannot
        drift between the places that apply it.

        This deliberately does not derive from the KV slots a drafter reserves
        past the target's query range, which is a *reservation* contract rather
        than a query-length one. The two do not differ by a constant: DFlash
        reserves `num_speculative_tokens + 1` slots yet still verifies `1 +
        num_speculative_tokens` queries, while EAGLE reserves
        `num_speculative_tokens` and verifies the same `1 + n`. Deriving one
        from the other would under-size EAGLE by a full request width, which is
        the failure this property exists to prevent.
        """
        return 1 + self.num_speculative_tokens

    @property
    def use_cumem_cudagraph_pool(self) -> bool:
        """Whether CUDA graphs go to the cuMem pool that sleep offloads."""
        from vllm.platforms import current_platform

        model_config = self.model_config
        return (
            model_config is not None
            and model_config.sleep_mode_offload_cudagraph
            and model_config.enable_sleep_mode
            and model_config.sleep_mode_backend == "cumem"
            and self.compilation_config.cudagraph_mode != CUDAGraphMode.NONE
            and current_platform.is_cuda()
        )

    @property
    def use_v2_model_runner(self) -> bool:
        if self.attention_config.hisparse_config is not None:
            if envs.VLLM_USE_V2_MODEL_RUNNER is False:
                raise ValueError(
                    "HiSparse requires Model Runner V2; remove "
                    "VLLM_USE_V2_MODEL_RUNNER=0."
                )
            return True

        if getattr(self, "watermark_config", None) is not None:
            if envs.VLLM_USE_V2_MODEL_RUNNER is False:
                logger.info_once(
                    "Watermarking requires Model Runner V2 and overrides "
                    "VLLM_USE_V2_MODEL_RUNNER=0."
                )
            return True

        use_v2_model_runner = envs.VLLM_USE_V2_MODEL_RUNNER
        if use_v2_model_runner is not None:
            return use_v2_model_runner

        from vllm.platforms import current_platform

        model_config = self.model_config
        if model_config is not None and current_platform.is_rocm():
            architectures = getattr(model_config, "architectures", ())
            if any(arch in ROCM_DEFAULT_MRV1_ARCHITECTURES for arch in architectures):
                # This default is a speed preference, not a claim that V1 can
                # serve the config, so it yields where V1 cannot. It yields by
                # falling through to the checks below, not by selecting V2.
                v1_unsupported = self._get_v1_model_runner_unsupported_features()
                if not v1_unsupported:
                    logger.warning_once(
                        "Defaulting to V1 model runner on ROCm for model "
                        "architectures: %s",
                        ", ".join(architectures),
                    )
                    return False
                logger.warning_once(
                    "Skipping the ROCm V1 model runner default for %s: V1 does "
                    "not support %s.",
                    ", ".join(architectures),
                    ", ".join(v1_unsupported),
                )

        if not HAS_TRITON:
            logger.warning_once(
                "Model Runner V2 requires Triton; using the V1 model runner instead."
            )
            return False

        unsupported = self._get_v2_model_runner_unsupported_features()
        if unsupported:
            logger.warning_once(
                "Model Runner V2 does not yet support %s; using the V1 model "
                "runner instead.",
                ", ".join(unsupported),
            )
            return False

        return True

    def _is_dflash_candidate_draft(self) -> bool:
        """Whether the DFlash draft has a candidate head, by the architecture the
        speculator selects on (v1/worker/gpu/spec_decode/__init__.py)."""
        spec = self.speculative_config
        if spec is None or spec.method != "dflash":
            return False
        draft_config = getattr(spec, "draft_model_config", None)
        if draft_config is None:
            return False
        return bool(
            {"DFlash2DraftModel", "LiLiCorrDraftModel"}.intersection(
                draft_config.architectures or []
            )
        )

    def _dflash_needs_multi_kv_group(self) -> bool:
        """Whether a DFlash draft mixes sliding-window and full attention."""
        spec = self.speculative_config
        if spec is None or spec.method != "dflash":
            return False
        draft_config = getattr(spec, "draft_model_config", None)
        if draft_config is None:
            return False
        layer_types = getattr(draft_config.hf_config, "layer_types", None) or []
        num_sliding = sum(lt == "sliding_attention" for lt in layer_types)
        return 0 < num_sliding < len(layer_types)

    def _uses_breakable_cudagraph_by_default(self) -> bool:
        model_config = self.model_config
        if model_config is None:
            return False

        architectures = set(model_config.architectures)
        return bool(architectures & default_breakable_cudagraph_architectures())

    def _uses_breakable_cudagraph_for_batch_invariance(self) -> bool:
        """Avoid freezing runtime-M tile lookup in compiled forward (#54243).
        Breakable graphs look up tuned bf16, unquantized qkv/o/gate_up/down tiles
        at capture; lm_head runs outside compiled forward and does not benefit."""
        from vllm.model_executor.determinism import batch_invariant_configs as bi
        from vllm.platforms import current_platform

        model = self.model_config
        if (
            not envs.VLLM_BATCH_INVARIANT
            or model is None
            or model.enforce_eager
            or model.dtype != torch.bfloat16
            or model.quantization is not None
            or not current_platform.is_cuda()
        ):
            return False
        family = bi._get_tuned_matmul_arch_family(
            current_platform.get_device_capability()
        )
        if family is None or family not in bi._BATCH_INVARIANT_MATMUL_TUNED_CONFIGS:
            return False
        table = bi._BATCH_INVARIANT_MATMUL_TUNED_CONFIGS[family]
        parallel = self.parallel_config
        tp = parallel.tensor_parallel_size
        hidden = model.get_hidden_size()
        head = model.get_head_size()
        heads = model.get_num_attention_heads(parallel)
        kv_heads = model.get_num_kv_heads(parallel)
        shapes = [((heads + 2 * kv_heads) * head, hidden), (hidden, heads * head)]
        intermediate = getattr(model.hf_text_config, "intermediate_size", None)
        # Per-layer sizes (e.g. Gemma3n) are not modeled.
        if isinstance(intermediate, int):
            shapes += [(2 * intermediate // tp, hidden), (hidden, intermediate // tp)]
        return any(shape in table for shape in shapes)

    def _maybe_enable_breakable_cudagraph(self) -> bool:
        if (
            "VLLM_USE_BREAKABLE_CUDAGRAPH" not in os.environ
            and self._uses_breakable_cudagraph_by_default()
        ):
            os.environ["VLLM_USE_BREAKABLE_CUDAGRAPH"] = "1"
            logger.info_once(
                "Auto-enabling VLLM_USE_BREAKABLE_CUDAGRAPH=1. "
                "Set VLLM_USE_BREAKABLE_CUDAGRAPH=0 to opt out."
            )
        elif (
            "VLLM_USE_BREAKABLE_CUDAGRAPH" not in os.environ
            and envs.VLLM_BATCH_INVARIANT
            and self._uses_breakable_cudagraph_for_batch_invariance()
        ):
            os.environ["VLLM_USE_BREAKABLE_CUDAGRAPH"] = "1"
            logger.info_once(
                "VLLM_BATCH_INVARIANT=1: auto-enabling VLLM_USE_BREAKABLE_CUDAGRAPH=1. "
                "Set VLLM_USE_BREAKABLE_CUDAGRAPH=0 to opt out."
            )

        from vllm.compilation.breakable_cudagraph import (
            is_breakable_cudagraph_enabled,
        )

        enabled = is_breakable_cudagraph_enabled()
        if enabled:
            self.compilation_config.mode = CompilationMode.NONE
        return enabled

    @property
    def needs_dp_coordinator(self) -> bool:
        """Determine if the DPCoordinator process is needed.

        The DPCoordinator is needed in two cases:
        1. For MoE models with DP > 1: to handle wave coordination
           (even in external LB mode, since wave coordination runs in the coordinator)
        2. For non-MoE models in internal/hybrid LB mode: to collect and publish
           queue stats for load balancing across DP ranks

        Returns:
            True if DPCoordinator process is needed, False otherwise.

        """
        # For non-MoE models, only need coordinator in internal/hybrid LB mode
        # (for stats collection).
        return self.parallel_config.data_parallel_size > 1 and (
            self.model_config is None
            or self.model_config.is_moe
            or not self.parallel_config.data_parallel_external_lb
        )

    def enable_trace_function_call_for_thread(self) -> None:
        """Set up function tracing for the current thread,
        if enabled via the `VLLM_TRACE_FUNCTION` environment variable.
        """
        if envs.VLLM_TRACE_FUNCTION:
            tmp_dir = tempfile.gettempdir()
            # add username to tmp_dir to avoid permission issues
            tmp_dir = os.path.join(tmp_dir, getpass.getuser())
            filename = (
                f"VLLM_TRACE_FUNCTION_for_process_{os.getpid()}"
                f"_thread_{threading.get_ident()}_at_{datetime.now()}.log"
            ).replace(" ", "_")
            log_path = os.path.join(
                tmp_dir,
                "vllm",
                f"vllm-instance-{self.instance_id}",
                filename,
            )
            os.makedirs(os.path.dirname(log_path), exist_ok=True)
            enable_trace_function_call(log_path)

    @staticmethod
    def _get_quantization_config(
        model_config: ModelConfig, load_config: LoadConfig
    ) -> QuantizationConfig | None:
        """Get the quantization config."""
        from vllm.platforms import current_platform

        if model_config.quantization is not None:
            from vllm.model_executor.model_loader.weight_utils import get_quant_config

            quant_config = get_quant_config(model_config, load_config)
            capability_tuple = current_platform.get_device_capability()

            if capability_tuple is not None:
                capability = capability_tuple.to_int()
                if capability < quant_config.get_min_capability():
                    raise ValueError(
                        f"The quantization method {model_config.quantization} "
                        "is not supported for the current GPU. Minimum "
                        f"capability: {quant_config.get_min_capability()}. "
                        f"Current capability: {capability}."
                    )
            supported_dtypes = quant_config.get_supported_act_dtypes()
            if model_config.dtype not in supported_dtypes:
                raise ValueError(
                    f"{model_config.dtype} is not supported for quantization "
                    f"method {model_config.quantization}. Supported dtypes: "
                    f"{supported_dtypes}"
                )
            quant_config.maybe_update_config(
                model_config.model,
                hf_config=model_config.hf_config,
                revision=model_config.revision,
            )
            return quant_config
        return None

    @staticmethod
    def get_quantization_config(
        model_config: ModelConfig, load_config: LoadConfig
    ) -> QuantizationConfig | None:
        import copy

        # For some reason, the _ version of this modifies the model_config
        # object, so using deepcopy to avoid this problem.
        return VllmConfig._get_quantization_config(
            copy.deepcopy(model_config), load_config
        )

    def with_hf_config(
        self,
        hf_config: PreTrainedConfig,
        architectures: list[str] | None = None,
    ) -> "VllmConfig":
        if architectures is not None:
            hf_config = copy.deepcopy(hf_config)
            hf_config.architectures = architectures
        elif hf_config.architectures is None:
            from transformers.models.auto.modeling_auto import (
                MODEL_FOR_CAUSAL_LM_MAPPING_NAMES,
            )

            if hf_config.model_type in MODEL_FOR_CAUSAL_LM_MAPPING_NAMES:
                hf_config = copy.deepcopy(hf_config)
                hf_config.architectures = [
                    MODEL_FOR_CAUSAL_LM_MAPPING_NAMES[hf_config.model_type]
                ]

        model_config = copy.deepcopy(self.model_config)

        # In Transformers v5, tie_word_embeddings belongs to the config of the class
        # that can see both layers to be tied. For example:
        #
        # SomeVLModel:
        #   self.language_model = SomeLanguageModel(SomeVLTextConfig)
        #   self.vision_model = SomeVisionModel(SomeVLVisionConfig)
        #
        # SomeVLModelForMultimodalLM:
        #   self.model = SomeVLModel(SomeVLConfig)
        #   self.lm_head = nn.Linear()
        #
        # Therefore, tie_word_embeddings is defined in SomeVLConfig and is not present
        # in SomeVLTextConfig*. In vLLM, the lm_head belongs to the language_model, so
        # we must ensure that tie_word_embeddings is set in the language_model's config.
        #
        # *For some models, SomeVLTextConfig may also have a tie_word_embeddings field.
        # This is only the case if SomeVLTextConfig is also used for a text only version
        # of the same model. For example:
        #
        # SomeVLModelForCausalLM:
        #   self.model = SomeLanguageModel(SomeVLTextConfig)
        #   self.lm_head = nn.Linear()
        #
        # Therefore, the presence of tie_word_embeddings in SomeVLTextConfig cannot
        # be used as a signal for whether tie_word_embeddings should be copied from
        # hf_config to the language_model config.
        if model_config.is_multimodal_model and hasattr(
            model_config.hf_config, "tie_word_embeddings"
        ):
            tie_word_embeddings = model_config.hf_config.tie_word_embeddings
            hf_config.get_text_config().tie_word_embeddings = tie_word_embeddings

        model_config.hf_config = hf_config
        model_config.model_arch_config = model_config.get_model_arch_config()
        model_config.is_submodel_config = True

        return replace(self, model_config=model_config)

    def _set_config_default(self, config_obj: Any, key: str, value: Any) -> None:
        """Set config attribute to default if not already set by user.

        Args:
            config_obj: Configuration object to update.
            key: Attribute name.
            value: Default value (static or callable).

        """
        if getattr(config_obj, key) is None:
            # Some config values are known before initialization and are
            # hard coded.
            # Other values depend on the user given configuration, so they are
            # implemented with lambda functions and decided at run time.
            setattr(config_obj, key, value(self) if callable(value) else value)

    def _apply_optimization_level_defaults(self, defaults: dict[str, Any]) -> None:
        """Apply optimization level defaults using self as root.

        Recursively applies values from defaults into nested config objects.
        Only fields present in defaults are overwritten.

        If the user configuration does not specify a value for a default field
        and if the default field is still None after all user selections are
        applied, then default values will be applied to the field. User specified
        fields will not be overridden by the default.

        Args:
            defaults: Dictionary of default values to apply.

        """

        def apply_recursive(config_obj: Any, config_defaults: dict[str, Any]) -> None:
            """Recursively apply defaults to config_obj, using self as root."""
            for key, value in config_defaults.items():
                if not hasattr(config_obj, key):
                    continue

                current = getattr(config_obj, key)
                if isinstance(value, dict) and is_dataclass(current):
                    apply_recursive(current, value)
                else:
                    self._set_config_default(config_obj, key, value)

        apply_recursive(self, defaults)

    def _maybe_override_dynamic_sd_cudagraph_mode(self) -> None:
        speculative_config = self.speculative_config
        if (
            speculative_config is None
            or not speculative_config.uses_dynamic_speculative_decoding()
            or not self.compilation_config.cudagraph_mode.has_full_cudagraphs()
            or self.use_v2_model_runner
        ):
            return

        logger.warning_once(
            "Dynamic speculative decoding changes the target verification "
            "length at runtime. Overriding cudagraph_mode from %s to "
            "PIECEWISE for reliability. Use VLLM_USE_V2_MODEL_RUNNER=1 "
            "if you want to use full CUDA graphs.",
            self.compilation_config.cudagraph_mode.name,
        )
        self.compilation_config.cudagraph_mode = CUDAGraphMode.PIECEWISE

    def _maybe_disable_dynamic_sd_for_data_parallel(self) -> None:
        speculative_config = self.speculative_config
        if (
            speculative_config is None
            or not speculative_config.uses_dynamic_speculative_decoding()
            or self.parallel_config.data_parallel_size <= 1
        ):
            return

        logger.warning_once(
            "Dynamic speculative decoding is not supported with data "
            "parallelism because data-parallel ranks can select different "
            "speculative-token counts, causing DP divergence and deadlocks. "
            "Disabling num_speculative_tokens_per_batch_size and falling back "
            "to static num_speculative_tokens=%d.",
            speculative_config.num_speculative_tokens,
        )
        speculative_config.num_speculative_tokens_per_batch_size = None

    def _normalize_piecewise_cudagraph_mode(
        self, *, breakable_cudagraph_enabled: bool
    ) -> None:
        compilation_config = self.compilation_config
        if (
            compilation_config.cudagraph_mode.requires_piecewise_compilation()
            and compilation_config.mode != CompilationMode.VLLM_COMPILE
            and not breakable_cudagraph_enabled
        ):
            fallback_mode = compilation_config.cudagraph_mode.without_piecewise()
            logger.info_once(
                "Cudagraph mode %s is not compatible with compilation mode %s. "
                "Overriding to %s.",
                compilation_config.cudagraph_mode,
                compilation_config.mode,
                fallback_mode,
            )
            compilation_config.cudagraph_mode = fallback_mode
            if fallback_mode == CUDAGraphMode.NONE:
                compilation_config.max_cudagraph_capture_size = 0
                compilation_config.cudagraph_capture_sizes = []

    def _post_init_kv_transfer_config(self) -> None:
        """Update KVTransferConfig based on top-level configs in VllmConfig.

        Right now, this function reads the offloading settings from
        CacheConfig and configures the KVTransferConfig accordingly.
        """
        # KV offloading is only activated when kv_offloading_size is set.
        if (kv_offloading_size := self.cache_config.kv_offloading_size) is None:
            return

        kv_offloading_backend = self.cache_config.kv_offloading_backend

        # If no KVTransferConfig is provided, create a default one.
        if self.kv_transfer_config is None:
            self.kv_transfer_config = KVTransferConfig()

        if kv_offloading_backend == "native":
            if envs.VLLM_USE_SIMPLE_KV_OFFLOAD:
                config_connector = "SimpleCPUOffloadConnector"
            else:
                config_connector = "OffloadingConnector"
            self.kv_transfer_config.kv_connector = config_connector
            self.kv_transfer_config.kv_connector_extra_config.update(
                {"cpu_bytes_to_use": kv_offloading_size * (1 << 30)}
            )
        elif kv_offloading_backend == "lmcache":
            # Default to LMCache multi-process (MP) mode. The actual KV
            # storage capacity is managed by the standalone LMCache server
            # process, so ``kv_offloading_size`` is not propagated here.
            # ``LMCacheMPConnector`` falls back to ``tcp://localhost:5555``
            # when host/port are not provided via extra_config.
            self.kv_transfer_config.kv_connector = "LMCacheMPConnector"

        # This is the same for all backends
        self.kv_transfer_config.kv_role = "kv_both"

    def _verify_aux_output_compatibility(self) -> None:
        """Reject configurations unsupported by enabled auxiliary outputs."""
        if not self.aux_output_config.enabled:
            return
        from vllm.platforms import current_platform

        # In-tree platforms only wire AuxOutput to MRV2. TPU and out-of-tree
        # platforms bring their own model runners and validate AuxOutput
        # support themselves.
        if not self.use_v2_model_runner and not (
            current_platform.is_tpu() or current_platform.is_out_of_tree()
        ):
            raise ValueError(
                "AuxOutput Connector requires Model Runner V2; set "
                "VLLM_USE_V2_MODEL_RUNNER=1."
            )
        if self.model_config.runner_type != "generate":
            raise ValueError("AuxOutput Connector only supports generate runners.")
        if not self.model_config.is_moe:
            raise ValueError("AuxOutput Connector only supports MoE models.")
        if not self.cache_config.enable_prefix_caching:
            raise ValueError("AuxOutput Connector requires prefix caching.")
        if (
            self.speculative_config is not None
            and self.speculative_config.enable_adaptive_verification
        ):
            raise ValueError(
                "--enable-return-routed-experts is incompatible with "
                "adaptive speculative verification."
            )
        if self.parallel_config.pipeline_parallel_size > 1:
            raise ValueError(
                "--enable-return-routed-experts is incompatible with "
                "pipeline parallelism (PP > 1)."
            )
        if (
            self.parallel_config.decode_context_parallel_size > 1
            or self.parallel_config.prefill_context_parallel_size > 1
        ):
            raise ValueError(
                "--enable-return-routed-experts is incompatible with "
                "context parallelism (DCP/PCP > 1)."
            )

        kv_transfer_config = self.kv_transfer_config
        if kv_transfer_config is not None:
            for connector_name in (
                "NixlConnector",
                "NixlPullConnector",
                "NixlPushConnector",
                "MoRIIOConnector",
                "MooncakeConnector",
            ):
                if kv_transfer_config.has_connector(connector_name):
                    raise ValueError(
                        "--enable-return-routed-experts is incompatible with "
                        f"{connector_name}; PD auxiliary output is not supported."
                    )

    def _verify_kv_transfer_compat(self) -> None:
        """Reject configurations that silently corrupt KV transfers."""
        if (
            self.kv_transfer_config is None
            or self.kv_transfer_config.kv_connector is None
        ):
            return

        # PyTorch's expandable_segments allocator uses CUDA VMM, which can
        # remap a virtual address range to different physical pages over the
        # engine's lifetime. KV connectors that pin KV cache memory (e.g.
        # NixlConnector via ibv_reg_mr, MooncakeConnector) end up with their
        # registrations pointing at stale physical pages after any remap,
        # producing RDMA failures like IBV_WC_REM_ACCESS_ERR /
        # NIXL_ERR_REMOTE_DISCONNECT at the first inter-node KV transfer.
        # We can't enumerate every in-tree and out-of-tree connector that
        # pins memory, so we conservatively reject the combination whenever
        # any KV connector is configured.
        #
        # CuMem allocator is exempt: CuMemAllocator.use_memory_pool toggles
        # expandable_segments off around its pool (see #40812), so the KV
        # cache allocated within that context lands on stable physical pages
        # even when the env var is set.
        if "expandable_segments:True" not in os.environ.get(
            "PYTORCH_CUDA_ALLOC_CONF", ""
        ):
            return
        if self.model_config is not None and (self.model_config.enable_cumem_allocator):
            return

        raise ValueError(
            f"KV connector {self.kv_transfer_config.kv_connector} is "
            "incompatible with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True "
            "unless enable_cumem_allocator is also enabled. PyTorch's CUDA VMM "
            "allocator can remap KV cache virtual addresses to different "
            "physical pages, invalidating any pinned/registered KV memory "
            "(e.g. IB memory regions registered by NIXL or Mooncake). Either "
            "unset expandable_segments:True or enable the cumem allocator "
            "(sleep mode does this automatically and also "
            "routes KV allocations through CuMemAllocator's pool, where "
            "expandable_segments is automatically disabled)."
        )

    def _verify_sampling_replay_config(self) -> None:
        model_config = self.model_config
        if model_config is None or not model_config.return_sampling_mask:
            return
        if not self.use_v2_model_runner:
            raise ValueError("sampling distribution replay requires Model Runner V2")
        speculative_config = self.speculative_config
        if (
            speculative_config is not None
            and speculative_config.enable_adaptive_verification
        ):
            raise ValueError(
                "sampling distribution replay with speculative decoding "
                "requires fixed verification boundaries; disable adaptive "
                "verification"
            )
        if model_config.is_diffusion:
            raise ValueError(
                "sampling distribution replay does not support diffusion models"
            )
        if model_config.logits_processors:
            raise ValueError(
                "sampling distribution replay does not support custom logits processors"
            )
        if model_config.logprobs_mode != "processed_logprobs":
            raise ValueError(
                "sampling distribution replay requires "
                "logprobs_mode='processed_logprobs' so that returned logprobs "
                "are normalized over the same nucleus as the sampling mask"
            )

    def _verify_trace_replay_config(self) -> None:
        model_config = self.model_config
        if model_config is None or not model_config.enable_trace_replay:
            return
        if not self.use_v2_model_runner:
            raise ValueError("trace replay requires Model Runner V2")

    def _check_supports_watermarking(
        self,
        config: "SamplingParams | BeamSearchParams | None" = None,
        *,
        custom_sampler: bool = False,
    ) -> bool:
        watermark_config = getattr(self, "watermark_config", None)
        if watermark_config is None:
            if config is not None and config.watermarking is True:
                logger.warning_once(
                    "Watermarking is enabled for this request, but the engine has no "
                    "watermark configuration. This and subsequent requests will run "
                    "without watermarking.",
                    scope="global",
                )
            return False
        if self.speculative_config is not None:
            speculative_config = self.speculative_config
            if speculative_config.draft_sample_method != "probabilistic":
                raise ValueError(
                    "Speculative decoding with watermarking requires "
                    "draft_sample_method='probabilistic'."
                )
            if speculative_config.rejection_sample_method != "standard":
                raise ValueError(
                    "Speculative decoding with watermarking requires "
                    "rejection_sample_method='standard'."
                )
            if (
                speculative_config.parallel_drafting
                and speculative_config.method != "dspark"
            ):
                raise ValueError(
                    "Parallel speculative drafting is not supported with watermarking."
                )
            if speculative_config.method not in ("dspark", "eagle", "eagle3", "mtp"):
                raise ValueError(
                    "Watermarking supports only autoregressive model-based "
                    "speculative decoding."
                )
            if (
                not watermark_config.allow_target_only_watermarking
                and not watermark_config.supports_speculative_decoding
            ):
                raise ValueError(
                    f"The '{watermark_config.algorithm}' watermarking algorithm "
                    "does not support speculative decoding. Set "
                    "allow_target_only_watermarking=true to leave draft tokens "
                    "unwatermarked."
                )
            if (
                watermark_config.allow_target_only_watermarking
                and not watermark_config.supports_speculative_decoding
            ):
                logger.warning_once(
                    "Target-only watermarking leaves accepted draft tokens "
                    "unwatermarked, weakening detectability in proportion to the "
                    "share of output tokens supplied by accepted drafts.",
                    scope="global",
                )
            if watermark_config.algorithm == "dual_key_gumbel" and (
                watermark_config.alpha != get_field(WatermarkConfig, "alpha").default
            ):
                logger.warning_once(
                    "Speculative decoding selects the watermark key by token role: "
                    "draft tokens use key A, recovery and bonus tokens use key B. "
                    "The configured alpha=%s is not used.",
                    watermark_config.alpha,
                    scope="global",
                )
        if custom_sampler:
            raise ValueError(
                "Model-specific custom samplers are not supported with watermarking."
            )
        if config is None:
            return True
        if config.watermarking is False:
            return False

        from vllm.sampling_params import BeamSearchParams, SamplingParams

        if isinstance(config, BeamSearchParams):
            logger.warning_once(
                "Watermarking is enabled, but beam search cannot be watermarked. "
                "This and subsequent beam search requests will run without "
                "watermarking.",
                scope="global",
            )
            return False
        if not isinstance(config, SamplingParams):
            raise TypeError(f"Unsupported watermarking config: {type(config).__name__}")
        if config.trace_decode_token_ids is not None:
            logger.warning_once(
                "Watermarking is enabled, but trace replay cannot be watermarked. "
                "This and subsequent trace replay requests will run without "
                "watermarking.",
                scope="global",
            )
            return False
        if config.temperature == 0:
            logger.warning_once(
                "Watermarking is enabled, but greedy decoding "
                "(temperature=0) cannot be watermarked. This and subsequent "
                "greedy requests will use ordinary greedy sampling.",
                scope="global",
            )
            return False
        return True

    def _resolve_and_verify_engram_config(self) -> None:
        """Resolve defaults and validate n-gram embedding settings."""
        model_config = self.model_config
        speculative_config = self.speculative_config
        # Draft configs inherit the target's communication groups and settings.
        # Validate the target because the draft may disable n-gram embeddings.
        if (
            speculative_config is not None
            and model_config is speculative_config.draft_model_config
        ):
            model_config = speculative_config.target_model_config
        if (
            model_config is not None
            and model_config.architecture == "DeepseekV41ForCausalLM"
            and getattr(model_config.hf_text_config, "engram_layer_ids", None)
            and self.parallel_config.use_ubatching
        ):
            raise ValueError(
                "DeepSeek V4.1 Engram does not support DBO or microbatching. "
                "Disable --enable-dbo and set --ubatch-size to 0."
            )
        if self.engram_config is None:
            if not model_has_engram_layers(model_config):
                return
            self.engram_config = EngramConfig()
        self.engram_config.verify_model_config(model_config)
        self.engram_config.resolve_dp_shared_memory(self.parallel_config)
        self.engram_config.verify_parallel_config(self.parallel_config)
        logger.info_once("Resolved Engram configuration: %s", str(self.engram_config))

    def __post_init__(self):
        """Verify configs are valid & consistent with each other."""
        # To give each torch profile run a unique instance name.
        self.instance_id = f"{time.time_ns()}"

        if self.model_config is not None and self.model_config.is_submodel_config:
            # with_hf_config() view: the parent config was already validated,
            # and this view's empty architecture list makes the model-dependent checks
            # below unsafe (e.g. use_mla resolves the architecture registry).
            return

        self._resolve_mm_encoder_only()

        if self.is_mm_encoder_only and self.cache_config.enable_prefix_caching:
            # Such an instance publishes encoder embeddings and runs no language
            # model, so it holds no KV cache for prefix caching to reuse and its
            # coordinator would have no group to manage. Disable before
            # `try_verify_and_update_config` so model config hooks (e.g. the
            # hybrid mamba hook setting `mamba_block_size`) already see prefix
            # caching as disabled.
            logger.info(
                "Disabling prefix caching: this instance runs the "
                "multi-modal encoder only."
            )
            self.cache_config.enable_prefix_caching = False

        if self.performance_mode != "balanced":
            logger.info_once("Performance mode set to '%s'.", self.performance_mode)

        self.try_verify_and_update_config()
        self._resolve_and_verify_engram_config()

        self._check_supports_watermarking()
        # Models may have supplied their own DCP defaults above; anything still
        # unset falls back to the stock ones.
        self.parallel_config.set_dcp_defaults()

        if self.model_config is not None:
            self.model_config.verify_with_parallel_config(self.parallel_config)
            self.model_config.verify_dual_chunk_attention_config(self.load_config)

            self.parallel_config.is_moe_model = self.model_config.is_moe

        if (
            self.model_config is not None
            and self.model_config.multimodal_config is not None
            and self.model_config.multimodal_config.language_model_only
            and self.compilation_config.cudagraph_mm_encoder
        ):
            raise ValueError(
                "--language-model-only is incompatible with "
                "cudagraph_mm_encoder=True, since it disables all multimodal "
                "inputs and the multimodal encoder is never run. Please "
                "disable one of them."
            )

        self._verify_sampling_replay_config()
        self._verify_trace_replay_config()

        # A NIXL side is either fully replicated or fully DCP-sharded; MLA only.
        if (
            self.kv_transfer_config is not None
            and self.kv_transfer_config.has_connector("NixlConnector")
        ):
            dcp_size = self.parallel_config.decode_context_parallel_size
            transfer_tp_size = max(
                self.parallel_config.tensor_parallel_size,
                self.parallel_config.prefill_context_parallel_size,
            )
            assert dcp_size in (1, transfer_tp_size), (
                f"decode_context_parallel_size={dcp_size} must be 1 or equal "
                f"to the NIXL transfer parallel size={transfer_tp_size}."
            )
            if self.model_config is not None:
                assert self.model_config.use_mla or dcp_size == 1, (
                    "PD with decode_context_parallel_size > 1 is only "
                    "supported for MLA models."
                )
        if self.lora_config is not None:
            self.lora_config.verify_with_model_config(self.model_config)

        if (
            self.mamba_config.enable_stochastic_rounding
            and self.cache_config.mamba_ssm_cache_dtype != "float16"
        ):
            raise ValueError(
                "Stochastic rounding for Mamba cache requires "
                "the SSM cache to be float16. Please set it explicitly, "
                "by specifying `--mamba-ssm-cache-dtype float16`, or disable "
                "stochastic rounding by not specifying "
                "`--enable-mamba-cache-stochastic-rounding`."
            )

        if self.quant_config is None and self.model_config is not None:
            self.quant_config = VllmConfig._get_quantization_config(
                self.model_config, self.load_config
            )

        # "dummy" reads no weights at all, and the sharded formats read a vLLM
        # state dict, which stores tied word embeddings under the lm_head only.
        # Neither can tell us what the original checkpoint contained.
        if self.model_config is not None and self.load_config.load_format not in (
            "dummy",
            "sharded_state",
            "runai_streamer_sharded",
        ):
            self.model_config.maybe_untie_word_embeddings()

        if (
            self.quant_config is not None
            and self.model_config is not None
            and hasattr(self.quant_config, "use_deep_gemm")
            and self.quant_config.use_deep_gemm is None
        ):
            from vllm.utils.deep_gemm import should_auto_disable_deep_gemm

            model_type = getattr(self.model_config.hf_text_config, "model_type", None)
            if should_auto_disable_deep_gemm(model_type):
                self.quant_config.use_deep_gemm = False
                logger.warning_once(
                    "Auto-disabled DeepGemm for model_type=%s on Blackwell. "
                    "DeepGemm E8M0 scale format causes accuracy degradation "
                    "for this architecture. Falling back to CUTLASS. "
                    "To disable DeepGemm globally, set VLLM_USE_DEEP_GEMM=0.",
                    model_type,
                )

        from vllm.platforms import current_platform
        from vllm.v1.executor.abstract import Executor

        executor_backend = self.parallel_config.distributed_executor_backend
        executor_class = Executor.get_class(self)
        executor_supports_async_sched = executor_class.supports_async_scheduling()
        uses_rocm_deepep_ht_dbo = (
            current_platform.is_rocm()
            and self.parallel_config.enable_dbo
            and self.parallel_config.all2all_backend == "deepep_high_throughput"
        )

        if self.scheduler_config.async_scheduling:
            # Async scheduling explicitly enabled, hard fail any incompatibilities.
            # Currently, async scheduling only support eagle speculative
            # decoding.
            if uses_rocm_deepep_ht_dbo:
                raise ValueError(
                    "Async scheduling is not compatible with ROCm DeepEP "
                    "high-throughput DBO. Please use --no-async-scheduling or "
                    "select a different all2all backend."
                )
            if self.speculative_config is not None:
                if (
                    self.speculative_config.method not in get_args(EagleModelTypes)
                    and self.speculative_config.method not in get_args(NgramGPUTypes)
                    and self.speculative_config.method != "draft_model"
                    and self.speculative_config.method != "dspark"
                    and self.speculative_config.method != "dflash"
                ):
                    raise ValueError(
                        "Currently, async scheduling is only supported "
                        "with EAGLE/MTP/Draft Model/NGram GPU/DSpark/DFlash "
                        "kind of speculative decoding"
                    )
                if self.speculative_config.disable_padded_drafter_batch:
                    raise ValueError(
                        "Async scheduling is not compatible with "
                        "disable_padded_drafter_batch=True."
                    )
            if not executor_supports_async_sched:
                raise ValueError(
                    f"`{executor_backend}` does not support async scheduling yet."
                )
        elif self.scheduler_config.async_scheduling is None:
            # Enable async scheduling unless there is an incompatible option.
            if (
                self.model_config is not None
                and self.model_config.runner_type == "pooling"
            ):
                # The current implementation of asynchronous scheduling negatively
                # impacts performance of pooling models, so we disable by default.
                logger.debug(
                    "Disabling asynchronous scheduling by default for pooling model."
                )
                self.scheduler_config.async_scheduling = False
            elif (
                self.speculative_config is not None
                and self.speculative_config.method not in get_args(EagleModelTypes)
                and self.speculative_config.method not in get_args(NgramGPUTypes)
                and self.speculative_config.method != "draft_model"
                and self.speculative_config.method != "dspark"
                and self.speculative_config.method != "dflash"
            ):
                logger.warning_once(
                    "Async scheduling not supported with %s-based "
                    "speculative decoding and will be disabled.",
                    self.speculative_config.method,
                )
                self.scheduler_config.async_scheduling = False
            elif (
                self.speculative_config is not None
                and self.speculative_config.disable_padded_drafter_batch
            ):
                logger.warning_once(
                    "Async scheduling is not compatible with "
                    "disable_padded_drafter_batch=True and will be disabled.",
                )
                self.scheduler_config.async_scheduling = False
            elif not executor_supports_async_sched:
                logger.warning_once(
                    "Async scheduling will be disabled because it is not supported "
                    "with the `%s` distributed executor backend. ",
                    executor_backend,
                )
                self.scheduler_config.async_scheduling = False
            elif uses_rocm_deepep_ht_dbo:
                logger.warning_once(
                    "Async scheduling is disabled for ROCm DeepEP "
                    "high-throughput DBO because that combination can corrupt "
                    "DP+EP generation accuracy."
                )
                self.scheduler_config.async_scheduling = False
            elif (
                self.parallel_config.pipeline_parallel_size > 1
                and not self.use_v2_model_runner
            ):
                logger.warning_once(
                    "Async scheduling is disabled because the V1 model runner "
                    "does not support it with pipeline parallelism."
                )
                self.scheduler_config.async_scheduling = False
            else:
                self.scheduler_config.async_scheduling = True

        if self.parallel_config.disable_nccl_for_dp_synchronization is None:
            if self.scheduler_config.async_scheduling:
                if self.parallel_config.data_parallel_size > 1 and (
                    self.model_config is None or self.model_config.is_moe
                ):
                    logger.info_once(
                        "Disabling NCCL for DP synchronization "
                        "when using async scheduling.",
                    )
                self.parallel_config.disable_nccl_for_dp_synchronization = True
            else:
                self.parallel_config.disable_nccl_for_dp_synchronization = False

        if (
            self.speculative_config is not None
            and self.scheduler_config.async_scheduling
            and self.model_config is not None
            and not self.model_config.disable_cascade_attn
        ):
            logger.warning_once(
                "Disabling cascade attention (not yet compatible with "
                "async speculative decoding).",
            )
            self.model_config.disable_cascade_attn = True

        if (
            self.observability_config.per_request_spec_decode_metrics != "none"
            and self.speculative_config is None
        ):
            raise ValueError(
                "--per-request-spec-decode-metrics requires speculative decoding "
                "to be enabled (via --speculative-config)."
            )

        if (
            self.model_config is not None
            and self.model_config.multimodal_config is not None
            and self.model_config.multimodal_config.mm_tensor_ipc == "torch_shm"
            and os.environ.get("VLLM_WORKER_MULTIPROC_METHOD") != "spawn"
        ):
            raise ValueError(
                "torch_shm is known to fail without "
                "VLLM_WORKER_MULTIPROC_METHOD set to spawn"
            )

        if (
            self.model_config is not None
            and self.scheduler_config.enable_chunked_prefill
            and self.model_config.dtype == torch.float32
            and current_platform.get_device_capability() == (7, 5)
        ):
            logger.warning_once(
                "Turing devices tensor cores do not support float32 matmul. "
                "To workaround this limitation, vLLM will set 'ieee' input "
                "precision for chunked prefill triton kernels."
            )

        if self.model_config is not None and self.model_config.enforce_eager:
            self.compilation_config.mode = CompilationMode.NONE
            self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE
            if self.parallel_config.enable_fault_tolerance:
                # Keep JIT warmup: in-inference Triton compilation latency
                # spikes can delay peer-fault detection past its deadline.
                logger.warning_once(
                    "Enforce eager set, disabling torch.compile and CUDAGraphs. "
                    "This is equivalent to setting -cc.mode=none "
                    "-cc.cudagraph_mode=none"
                )
            else:
                logger.warning_once(
                    "Enforce eager set, disabling torch.compile, CUDAGraphs, and "
                    "JIT kernel warmup. This is equivalent to setting "
                    "-cc.mode=none -cc.cudagraph_mode=none and "
                    "--kernel_config.enable_jit_warmup=False"
                )
                self.kernel_config.enable_jit_warmup = False

        if os.environ.get("TORCH_COMPILE_DISABLE") == "1":
            logger.warning_once(
                "TORCH_COMPILE_DISABLE is set, disabling torch.compile. "
                "This is equivalent to setting -cc.mode=none"
            )
            self.compilation_config.mode = CompilationMode.NONE

        breakable_cudagraph_enabled = self._maybe_enable_breakable_cudagraph()

        if not breakable_cudagraph_enabled and (
            self.compilation_config.backend == "eager"
            or (
                self.compilation_config.mode is not None
                and self.compilation_config.mode != CompilationMode.VLLM_COMPILE
            )
        ):
            logger.warning_once(
                "Inductor compilation was disabled by user settings, "
                "optimizations settings that are only active during "
                "inductor compilation will be ignored."
            )

        def has_blocked_weights():
            if self.quant_config is not None:
                if hasattr(self.quant_config, "weight_block_size"):
                    return self.quant_config.weight_block_size is not None
                elif hasattr(self.quant_config, "has_blocked_weights"):
                    return self.quant_config.has_blocked_weights()
            return False

        # Enable quant_fp8 CUDA ops (TODO disable in follow up)
        # On H100 the CUDA kernel is faster than
        # native implementation
        # https://github.com/vllm-project/vllm/issues/25094
        if has_blocked_weights():
            custom_ops = self.compilation_config.custom_ops
            if "-quant_fp8" not in custom_ops:
                custom_ops.append("+quant_fp8")

        current_platform.apply_config_platform_defaults(self)

        if self.compilation_config.mode is None:
            if self.optimization_level > OptimizationLevel.O0:
                self.compilation_config.mode = CompilationMode.VLLM_COMPILE
            else:
                self.compilation_config.mode = CompilationMode.NONE

        # By default, enable torch wrapping only when using custom Inductor lowering
        if self.compilation_config.ir_enable_torch_wrap is None:
            self.compilation_config.ir_enable_torch_wrap = (
                self.compilation_config.mode == CompilationMode.VLLM_COMPILE
                and self.compilation_config.backend == "inductor"
            )

        if all(s not in self.compilation_config.custom_ops for s in ("all", "none")):
            if (
                self.compilation_config.backend == "inductor"
                and self.compilation_config.mode != CompilationMode.NONE
            ):
                self.compilation_config.custom_ops.append("none")
            else:
                self.compilation_config.custom_ops.append("all")

        # This populates IR op priorities,
        # must happen after compilation mode and backend are decided,
        # but before fusion defaults are applied as those may depend on op priority.
        self.kernel_config.set_platform_defaults(self)

        default_config = OPTIMIZATION_LEVEL_TO_CONFIG[self.optimization_level]
        self._apply_optimization_level_defaults(default_config)
        if self.kernel_config.enable_flashinfer_autotune is None:
            raise ValueError(
                "KernelConfig.enable_flashinfer_autotune must be set after applying "
                "optimization level defaults."
            )

        self._maybe_disable_dynamic_sd_for_data_parallel()
        self._maybe_override_dynamic_sd_cudagraph_mode()

        if (
            self.attention_config.hisparse_config is None
            and self.kv_transfer_config is not None
            and self.kv_transfer_config.has_connector("HiSparseConnector")
        ):
            self.attention_config.hisparse_config = HiSparseConfig()

        if self.attention_config.hisparse_config is not None:
            if not current_platform.is_cuda():
                raise ValueError("HiSparse currently requires NVIDIA CUDA.")
            if self.parallel_config.pipeline_parallel_size > 1:
                raise ValueError("HiSparse does not support pipeline parallelism.")
            if self.parallel_config.decode_context_parallel_size > 1:
                raise ValueError(
                    "HiSparse does not support decode context parallelism."
                )
            if self.compilation_config.cudagraph_mode == CUDAGraphMode.FULL:
                raise ValueError(
                    "HiSparse does not support cudagraph_mode=FULL; use "
                    "FULL_AND_PIECEWISE (the default), which captures FULL graphs "
                    "for decode batches."
                )
            if not self.scheduler_config.scheduler_reserve_full_isl:
                # Without it, async loads admitted against free host blocks can
                # each wait on host pages the others hold, and waiting requests
                # are never preempted to free them.
                raise ValueError(
                    "HiSparse requires --scheduler-reserve-full-isl; remove "
                    "--no-scheduler-reserve-full-isl."
                )
            if self.model_config is not None and not hasattr(
                self.model_config.hf_config, "index_topk"
            ):
                raise ValueError(
                    "HiSparse is only supported for DSA models with index_topk."
                )
            if self.kv_transfer_config is not None and (
                self.kv_transfer_config.kv_connector
                not in (
                    None,
                    "NixlConnector",
                    "MooncakeStoreConnector",
                    "MultiConnector",
                )
            ):
                logger.warning(
                    "HiSparse host-resident KV is configured with connector "
                    "%s. NixlConnector (GPU-staged host imports) and "
                    "MooncakeStoreConnector (shared-store offload) are the "
                    "validated paths; other connectors are treated as "
                    "debug/fallback paths.",
                    self.kv_transfer_config.kv_connector,
                )

        self._normalize_piecewise_cudagraph_mode(
            breakable_cudagraph_enabled=breakable_cudagraph_enabled
        )

        from vllm.utils.torch_utils import HAS_OPAQUE_TYPE

        if HAS_OPAQUE_TYPE:
            # On torch >= 2.11 the hoisted OpaqueObject approach supersedes
            # fast_moe_cold_start, so force it off.
            self.compilation_config.fast_moe_cold_start = False
        elif self.compilation_config.fast_moe_cold_start is None:
            # resolve default behavior: try to be as safe as possible
            # this config is unsafe if any spec decoding draft model has a MOE.
            # We'll conservatively turn it off if we see spec decoding.
            self.compilation_config.fast_moe_cold_start = (
                self.speculative_config is None
            )

        self._set_max_num_scheduled_tokens()

        if current_platform.support_static_graph_mode():
            # if cudagraph_mode has full cudagraphs, we need to check support
            if model_config := self.model_config:
                if (
                    self.compilation_config.cudagraph_mode.has_full_cudagraphs()
                    and model_config.pooler_config is not None
                ):
                    logger.warning_once(
                        "Pooling models do not support full cudagraphs. "
                        "Overriding cudagraph_mode to PIECEWISE."
                    )
                    self.compilation_config.cudagraph_mode = CUDAGraphMode.PIECEWISE
                elif (
                    model_config.is_encoder_decoder
                    and self.compilation_config.cudagraph_mode
                    not in (CUDAGraphMode.NONE, CUDAGraphMode.FULL_DECODE_ONLY)
                ):
                    logger.info_once(
                        "Encoder-decoder models do not support %s. "
                        "Overriding cudagraph_mode to FULL_DECODE_ONLY.",
                        self.compilation_config.cudagraph_mode.name,
                    )
                    self.compilation_config.cudagraph_mode = (
                        CUDAGraphMode.FULL_DECODE_ONLY
                    )

            # Check if KV connector requires PIECEWISE mode for CUDA graphs
            if (
                self.kv_transfer_config is not None
                and self.kv_transfer_config.is_kv_transfer_instance
                and self.compilation_config.cudagraph_mode.has_full_cudagraphs()
            ):
                # Lazy import to avoid circular dependencies
                from vllm.distributed.kv_transfer.kv_connector.factory import (
                    KVConnectorFactory,
                )

                connector_cls = KVConnectorFactory.get_connector_class(
                    self.kv_transfer_config
                )
                if connector_cls.requires_piecewise_for_cudagraph(
                    self.kv_transfer_config.kv_connector_extra_config
                ):
                    logger.warning_once(
                        "KV connector %s requires PIECEWISE CUDA graph mode "
                        "due to layerwise async operations that cannot be "
                        "captured in CUDA graphs. "
                        "Overriding cudagraph_mode from %s to PIECEWISE.",
                        connector_cls.__name__,
                        self.compilation_config.cudagraph_mode.name,
                    )
                    self.compilation_config.cudagraph_mode = CUDAGraphMode.PIECEWISE

            # disable cudagraph when enforce eager execution
            if self.model_config is not None and self.model_config.enforce_eager:
                logger.info_once("Cudagraph is disabled under eager mode")
                self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE
                # override related settings when enforce eager
                self.compilation_config.max_cudagraph_capture_size = 0
                self.compilation_config.cudagraph_capture_sizes = []
            else:
                self.compilation_config.cudagraph_num_of_warmups = 1

            self._set_cudagraph_sizes()

        else:
            self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE

        if self.cache_config.kv_sharing_fast_prefill:
            if (
                self.speculative_config is not None
                and self.speculative_config.use_eagle()
            ):
                raise ValueError(
                    "Fast prefill optimization for KV sharing is not "
                    "compatible with EAGLE as EAGLE requires correct logits "
                    "for all tokens while fast prefill gives incorrect logits "
                    "for prompt tokens."
                )

            logger.warning_once(
                "--kv-sharing-fast-prefill requires changes on model side for "
                "correctness and to realize prefill savings."
            )

        if (
            self.model_config
            and self.model_config.architecture == "WhisperForConditionalGeneration"
            and os.environ.get("VLLM_WORKER_MULTIPROC_METHOD") != "spawn"
        ):
            logger.warning_once(
                "Whisper is known to have issues with "
                "forked workers. If startup is hanging, "
                "try setting 'VLLM_WORKER_MULTIPROC_METHOD' "
                "to 'spawn'."
            )

        if (
            self.kv_events_config is not None
            and self.kv_events_config.enable_kv_cache_events
            and not self.cache_config.enable_prefix_caching
        ):
            logger.warning_once(
                "KV cache events are on, but prefix caching is not enabled. "
                "Use --enable-prefix-caching to enable."
            )
        if (
            self.kv_events_config is not None
            and self.kv_events_config.publisher != "null"
            and not self.kv_events_config.enable_kv_cache_events
        ):
            logger.warning_once(
                "KV cache events are disabled, "
                "but the scheduler is configured to publish them. "
                "Modify KVEventsConfig.enable_kv_cache_events "
                "to True to enable."
            )
        current_platform.check_and_update_config(self)

        # After the platform hook, which has the last word on async scheduling.
        if (
            self.diffusion_config is not None
            and self.scheduler_config.scheduler_cls is None
        ):
            scheduler_name = (
                "DiffusionAsyncScheduler"
                if self.scheduler_config.async_scheduling
                else "DiffusionScheduler"
            )
            self.scheduler_config.scheduler_cls = (
                f"vllm.v1.core.sched.diffusion_scheduler.{scheduler_name}"
            )

        self._normalize_piecewise_cudagraph_mode(
            breakable_cudagraph_enabled=breakable_cudagraph_enabled
        )

        self._resolve_mm_embedding_inputs()
        self._resolve_mm_processor_device()
        self._resolve_mm_video_decode_device()
        self._validate_mm_processor_device()

        if self.use_v2_model_runner:
            self._disable_cudagraphs_for_v2_stock_torch_compile()
            self._validate_v2_model_runner()
        else:
            self._validate_v1_model_runner()

        self._validate_profiler_config()
        self._validate_batch_sharded_sampling()
        self._validate_adaptive_verification()

        # Re-compute compile ranges after platform-specific config updates
        # (e.g., XPU may lower max_num_batched_tokens when MLA is enabled)
        self._set_compile_ranges()

        if self.parallel_config.all2all_backend == "moonep":
            if (
                self.model_config is not None
                and self.model_config.quantization is not None
            ):
                raise ValueError(
                    "The moonep all2all backend currently supports unquantized "
                    "BF16 models only; got "
                    f"quantization={self.model_config.quantization!r}. Use a "
                    "different --all2all-backend for quantized models."
                )
            if (
                self.model_config is not None
                and self.model_config.dtype != torch.bfloat16
            ):
                raise ValueError(
                    "The moonep all2all backend currently supports BF16 models "
                    f"only; got dtype={self.model_config.dtype}. Use a "
                    "different --all2all-backend or --dtype bfloat16."
                )
            if self.parallel_config.enable_eplb:
                raise ValueError(
                    "The moonep all2all backend does not support EPLB yet: "
                    "EPLB rearranges expert parameters in a layout MoonEP's "
                    "replicated [E+B] weights do not follow. Disable "
                    "--enable-eplb or use a different --all2all-backend."
                )
            if self.parallel_config.expert_placement_strategy != "linear":
                raise ValueError(
                    "The moonep all2all backend requires linear expert "
                    "placement: its load-time all-gather assumes each rank "
                    "holds a contiguous chunk of the global expert range. Got "
                    "--expert-placement-strategy "
                    f"{self.parallel_config.expert_placement_strategy!r}."
                )
            # Enforced here rather than in set_splitting_ops_for_v1 so it
            # holds for every compilation mode, and keyed on use_all2all so
            # PCP/SP-only topologies are covered too: MoonEP dispatch/combine
            # are eager-only and must not be captured.
            if (
                self.parallel_config.use_all2all
                and self.compilation_config.cudagraph_mode != CUDAGraphMode.NONE
            ):
                logger.info(
                    "MoonEP: Disabling CUDA Graphs since the MoonEP "
                    "integration is currently eager-only."
                )
                self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE

        # Do this after all the updates to compilation_config.mode
        effective_dp_size = (
            self.parallel_config.data_parallel_size
            if self.model_config is None or self.model_config.is_moe
            else 1
        )
        self.compilation_config.set_splitting_ops_for_v1(
            all2all_backend=self.parallel_config.all2all_backend,
            data_parallel_size=effective_dp_size,
        )

        # final check of cudagraph mode after all possible updates
        if current_platform.is_cuda_alike():
            if (
                self.compilation_config.cudagraph_mode.has_full_cudagraphs()
                and self.model_config is not None
                and not self.model_config.disable_cascade_attn
                and not self.compilation_config.cudagraph_mode.has_piecewise_cudagraphs()  # noqa: E501
            ):
                logger.warning_once(
                    "No piecewise cudagraph for executing cascade attention. "
                    "Will fall back to eager execution if a batch runs into "
                    "cascade attentions."
                )

            if self.compilation_config.cudagraph_mode.requires_piecewise_compilation():
                assert (
                    self.compilation_config.mode == CompilationMode.VLLM_COMPILE
                    or envs.VLLM_USE_BREAKABLE_CUDAGRAPH
                ), (
                    "Compilation mode should be CompilationMode.VLLM_COMPILE "
                    "when cudagraph_mode piecewise cudagraphs is used, "
                    f"cudagraph_mode={self.compilation_config.cudagraph_mode}"
                )
        if (
            self.model_config
            and envs.VLLM_BATCH_INVARIANT
            and not self.model_config.disable_cascade_attn
        ):
            self.model_config.disable_cascade_attn = True
            logger.warning_once(
                "Disabling cascade attention when VLLM_BATCH_INVARIANT is enabled.",
            )

        if self.parallel_config.use_ubatching:
            a2a_backend = self.parallel_config.all2all_backend
            assert a2a_backend in [
                "deepep_low_latency",
                "deepep_high_throughput",
                "nixl_ep",
            ], (
                "Microbatching currently only supports the deepep_low_latency, "
                "deepep_high_throughput, and nixl_ep all2all backends. "
                f"{a2a_backend} is not supported. To fix use "
                "--all2all-backend=deepep_low_latency, "
                "--all2all-backend=deepep_high_throughput, or "
                "--all2all-backend=nixl_ep and install the matching kernels."
            )

            if not self.model_config.disable_cascade_attn:
                self.model_config.disable_cascade_attn = True
                logger.warning_once("Disabling cascade attention when DBO is enabled.")

        if not self.instance_id:
            self.instance_id = random_uuid()[:5]

        if self.reasoning_config is not None and self.model_config is not None:
            self.reasoning_config.initialize_token_ids(self.model_config)
            if not self.reasoning_config.enabled:
                logger.warning_once(
                    "Auto-initialization of reasoning token IDs failed. "
                    "Please check whether your reasoning parser has implemented "
                    "the `reasoning_start_str` and `reasoning_end_str`."
                )

        # Resolve kv_offloading-derived connector name into kv_transfer_config
        # before the HMA check below, which inspects the connector class.
        self._post_init_kv_transfer_config()
        self._verify_aux_output_compatibility()

        # Hybrid KV cache manager (HMA) runtime rules:
        # - Explicit enable (--no-disable-kv-cache-manager): error if runtime
        #   disables it
        # - No preference: auto-disable for unsupported features or connector configs
        # - Explicit disable (--disable-kv-cache-manager): always respect it
        need_disable_hybrid_kv_cache_manager = False
        # logger should only print warning message for hybrid models. As we
        # can't know whether the model is hybrid or not now, so we don't log
        # warning message here and will log it later.
        if not current_platform.support_hybrid_kv_cache():
            # Hybrid KV cache manager is not supported on non-GPU platforms.
            need_disable_hybrid_kv_cache_manager = True
        if (
            self.model_config is not None
            and self.model_config.attention_chunk_size is not None
        ):
            if (
                self.speculative_config is not None
                and self.speculative_config.use_eagle()
            ):
                # Hybrid KV cache manager is not yet supported with chunked
                # local attention + eagle.
                need_disable_hybrid_kv_cache_manager = True
            elif not envs.VLLM_ALLOW_CHUNKED_LOCAL_ATTN_WITH_HYBRID_KV_CACHE:
                logger.warning(
                    "There is a latency regression when using chunked local"
                    " attention with the hybrid KV cache manager. Disabling"
                    " it, by default. To enable it, set the environment "
                    "VLLM_ALLOW_CHUNKED_LOCAL_ATTN_WITH_HYBRID_KV_CACHE=1."
                )
                # Hybrid KV cache manager is not yet supported with chunked
                # local attention.
                need_disable_hybrid_kv_cache_manager = True

        if self.scheduler_config.disable_hybrid_kv_cache_manager is None:
            # Auto-disable HMA only when the connector config does not support it.
            if self.kv_transfer_config is not None:
                from vllm.distributed.kv_transfer.kv_connector.factory import (
                    KVConnectorFactory,
                )

                if not KVConnectorFactory.supports_hma_config(self.kv_transfer_config):
                    need_disable_hybrid_kv_cache_manager = True
                    logger.warning(
                        "Turning off hybrid kv cache manager because "
                        "`--kv-transfer-config` selects a KV connector that "
                        "does not support it. Impact: hybrid SSM models "
                        "(e.g. Jamba, Bamba) require HMA and will fail at "
                        "startup without it; models with sliding window "
                        "attention will run with reduced performance. "
                        "To add HMA support to a KV connector, subclass "
                        "`SupportsHMA` defined in kv_connector/v1/base.py "
                        "(for MultiConnector, all child connectors must "
                        "support HMA)."
                    )
            self.scheduler_config.disable_hybrid_kv_cache_manager = (
                need_disable_hybrid_kv_cache_manager
            )
        elif (
            self.scheduler_config.disable_hybrid_kv_cache_manager is False
            and need_disable_hybrid_kv_cache_manager
        ):
            raise ValueError(
                "Hybrid KV cache manager was explicitly enabled but is not "
                "supported in this configuration. Consider omitting the "
                "--no-disable-hybrid-kv-cache-manager flag to let vLLM decide"
                " automatically."
            )

        if self.scheduler_config.disable_hybrid_kv_cache_manager is None:
            # Default to enable HMA if not explicitly disabled by user or logic above.
            self.scheduler_config.disable_hybrid_kv_cache_manager = False

        if (
            self.attention_config.hisparse_config is not None
            and self.scheduler_config.disable_hybrid_kv_cache_manager
        ):
            raise ValueError(
                "HiSparse requires the hybrid KV cache manager; remove "
                "--disable-hybrid-kv-cache-manager or use connectors that "
                "support HMA."
            )

        if self.compilation_config.debug_dump_path:
            self.compilation_config.debug_dump_path = (
                self.compilation_config.debug_dump_path.absolute().expanduser()
            )
        if envs.VLLM_DEBUG_DUMP_PATH is not None:
            env_path = Path(envs.VLLM_DEBUG_DUMP_PATH).absolute().expanduser()
            if self.compilation_config.debug_dump_path:
                logger.warning(
                    "Config-specified debug dump path is overridden"
                    " by VLLM_DEBUG_DUMP_PATH to %s",
                    env_path,
                )
            self.compilation_config.debug_dump_path = env_path

        # Enable quant_fp8 CUDA ops (TODO disable in follow up)
        # On H100 the CUDA kernel is faster than
        # native implementation
        # https://github.com/vllm-project/vllm/issues/25094
        if has_blocked_weights():
            custom_ops = self.compilation_config.custom_ops
            if "-quant_fp8" not in custom_ops:
                custom_ops.append("+quant_fp8")

        self._verify_kv_transfer_compat()
        if self.use_cumem_cudagraph_pool:
            # NCCL graph registration pins the offloaded pool; workers inherit this.
            value = os.environ.setdefault("NCCL_GRAPH_REGISTER", "0")
            if value != "0":
                logger.warning(
                    "NCCL_GRAPH_REGISTER=%s pins the CUDA graph pool during sleep.",
                    value,
                )
        # Log the custom passes that are enabled
        self.compilation_config.pass_config.log_enabled_passes()

    def _set_max_num_scheduled_tokens(self):
        """In most cases, the scheduler may schedule a batch with as many tokens as the
        worker is configured to handle.
        """
        if self.speculative_config is not None:
            scheduled_token_delta = (
                self.speculative_config.max_num_new_slots_for_drafting
            )
            max_num_batched_tokens = self.scheduler_config.max_num_batched_tokens
            if self.scheduler_config.max_num_scheduled_tokens is None:
                self.scheduler_config.max_num_scheduled_tokens = max_num_batched_tokens

            if self.scheduler_config.max_num_scheduled_tokens <= 0:
                raise ValueError(
                    "max_num_scheduled_tokens is set to"
                    f" {self.scheduler_config.max_num_scheduled_tokens} based on"
                    " the speculative decoding settings, which does not allow"
                    " any tokens to be scheduled. Increase max_num_batched_tokens"
                    " to accommodate the additional draft token slots, or decrease"
                    " num_speculative_tokens."
                )
            if self.scheduler_config.max_num_scheduled_tokens < 8192:
                logger.warning_once(
                    "max_num_scheduled_tokens is set to"
                    f" {self.scheduler_config.max_num_scheduled_tokens} based on"
                    " the speculative decoding settings. This may lead to suboptimal"
                    " performance. Consider increasing max_num_batched_tokens to"
                    " accommodate the additional draft token slots, or decrease"
                    " num_speculative_tokens.",
                )

            if max_num_batched_tokens <= scheduled_token_delta:
                raise ValueError(
                    "VllmConfig does not have enough slots to schedule a token and"
                    " support the speculative decoding settings."
                    f" Got {max_num_batched_tokens=} and {scheduled_token_delta=}."
                )

    def _set_cudagraph_sizes(self):
        """VLLM defines the default candidate list of batch sizes for CUDA graph
        capture as:

        ```python
        default_max_graph_size = 1024 if is_data_center_blackwell else 512
        decode_query_len = self.uniform_decode_query_len
        max_graph_size = min(
            max_num_seqs * decode_query_len * 2, default_max_graph_size
        )
        # 1, 2, 4, then multiples of 8 up to 256 and then multiples of 16
        # up to max_graph_size
        cudagraph_capture_sizes = [1, 2, 4] + list(range(8, 256, 8)) + list(
            range(256, max_graph_size + 1, 16))

        `max_num_batched_tokens` is also appended to the list if it fits
        within `max_cudagraph_capture_size`, so the max batch size is captured
        even when off-stride. Uniform decode sizes are appended when they fit
        within the platform's default capture ceiling, since they need not land
        on an 8- or 16-token stride.

        In the end, `vllm_config.compilation_config.cudagraph_capture_sizes`
        will be the final sizes to capture cudagraph (in ascending order).

        These sizes are used to capture and reuse CUDA graphs for
        performance-critical paths (e.g., decoding). Capturing enables
        significantly faster kernel dispatch by avoiding Python overhead. The
        list is then filtered based on `max_num_batched_tokens` (e.g., 8192 on
        most GPUs), which controls the total allowed number of tokens in a
        batch. Since each sequence may have a variable number of tokens, the
        maximum usable batch size will depend on actual sequence lengths.

        Example:
            With `max_num_batched_tokens = 8192`, and typical sequences
            averaging ~32 tokens, most practical batch sizes fall below 256.
            However, the system will still allow capture sizes up to the
            platform default if shape and memory permit.

        Note:
            If users explicitly specify cudagraph capture sizes in the
            compilation config, those will override this default logic.
            At runtime:

            - If batch size <= one of the `cudagraph_capture_sizes`, the closest
            padded CUDA graph will be used.
            - If batch size > largest `cudagraph_capture_sizes`, cudagraph will
            not be used.

        """
        if (
            self.model_config is not None
            and not self.model_config.enforce_eager
            and self.compilation_config.cudagraph_mode != CUDAGraphMode.NONE
        ):
            # determine the initial max_cudagraph_capture_size
            max_cudagraph_capture_size = (
                self.compilation_config.max_cudagraph_capture_size
            )
            # Decode sizes to cover, in tokens. Populated only when the default
            # is computed here, so an explicit capture range is left exactly as
            # configured.
            uniform_decode_sizes: list[int] = []
            if max_cudagraph_capture_size is None:
                from vllm.platforms import current_platform

                default_max_graph_size = (
                    1024 if current_platform.is_device_capability_family(100) else 512
                )
                decode_query_len = self.uniform_decode_query_len
                max_num_seqs = self.scheduler_config.max_num_seqs
                max_cudagraph_capture_size = min(
                    max_num_seqs * decode_query_len * 2, default_max_graph_size
                )
                if decode_query_len > 1:
                    # A uniform decode batch is decode_query_len tokens per
                    # request, so the widest one is far outside this ceiling.
                    # Coverage comes from appending the decode sizes rather than
                    # extending the token-strided grid. Extending that grid to
                    # the widest decode size produces 581 sizes at
                    # max_num_seqs=512 and 16 draft tokens, versus 100 with the
                    # request-count grid.
                    #
                    # The grid would not buy decode coverage anyway. Dispatch
                    # requires an exact multiple of decode_query_len, so a
                    # token-strided entry is only usable when it happens to be
                    # one; at query length 17 a captured 560 rounds to 561 and
                    # is rejected. Scaling a request-count grid keeps every
                    # entry usable and the count comparable to the non-
                    # speculative case.
                    def request_counts(max_reqs: int) -> list[int]:
                        # At most the platform default number of requests,
                        # mirroring the one-token-per-request decode ceiling.
                        max_reqs = min(max_reqs, default_max_graph_size)
                        counts = [n for n in (1, 2, 4) if n <= max_reqs]
                        counts += list(range(8, min(max_reqs + 1, 256), 8))
                        counts += list(range(256, max_reqs + 1, 16))
                        return sorted(set(counts + [max_reqs]))

                    # Dynamic speculative decoding picks the draft width from
                    # the batch size, so a decode step is only uniform within a
                    # tier and each tier needs its own sizes. Scaling by the
                    # widest one alone leaves the narrower tiers short: the
                    # manager rounds a capture size up to a multiple of the
                    # tier's query length and drops it once the implied request
                    # count exceeds max_num_seqs, so at query length 3 sizes
                    # built from 17 stop covering at 227 of 256 requests.
                    decode_tiers = [(decode_query_len, max_num_seqs)]
                    speculative_config = self.speculative_config
                    if (
                        speculative_config is not None
                        and speculative_config.uses_dynamic_speculative_decoding()
                    ):
                        from vllm.v1.spec_decode.dynamic.utils import (
                            build_dynamic_sd_schedule_lookup,
                        )

                        schedule = (
                            speculative_config.num_speculative_tokens_per_batch_size
                        )
                        assert schedule is not None
                        # Read the tiers off the dense lookup the scheduler
                        # runs on, so the clamp against num_speculative_tokens
                        # and the carry-forward through gaps and the tail
                        # cannot drift from it. Validation lives elsewhere; an
                        # invalid schedule keeps the single-tier default.
                        try:
                            dense_schedule = build_dynamic_sd_schedule_lookup(
                                schedule,
                                vllm_max_batch_size=max_num_seqs,
                                vllm_num_speculative_tokens=self.num_speculative_tokens,
                            )
                        except ValueError:
                            pass
                        else:
                            # Ascending batch size, so the last write per
                            # query length is the widest batch running at it.
                            widest_batch: dict[int, int] = {}
                            for batch_size, num_spec in enumerate(
                                dense_schedule[1:], start=1
                            ):
                                widest_batch[num_spec + 1] = batch_size
                            decode_tiers = list(widest_batch.items())

                    uniform_decode_sizes = sorted(
                        {
                            n * query_len
                            for query_len, tier_max_reqs in decode_tiers
                            for n in request_counts(tier_max_reqs)
                            if n * query_len <= max_cudagraph_capture_size
                        }
                    )
                elif max_num_seqs <= max_cudagraph_capture_size:
                    uniform_decode_sizes = [max_num_seqs]
            max_num_tokens = self.scheduler_config.max_num_batched_tokens
            max_cudagraph_capture_size = min(max_num_tokens, max_cudagraph_capture_size)

            assert max_cudagraph_capture_size >= 1, (
                "Maximum cudagraph size should be greater than or equal to 1 "
                "when using cuda graph."
            )

            # determine the cudagraph_capture_sizes
            if self.compilation_config.cudagraph_capture_sizes is not None:
                assert len(self.compilation_config.cudagraph_capture_sizes) > 0, (
                    "cudagraph_capture_sizes should contain at least one element "
                    "when using cuda graph."
                )
                # de-duplicate the sizes provided by the config
                dedup_sizes = list(set(self.compilation_config.cudagraph_capture_sizes))
                cudagraph_capture_sizes = [
                    i for i in dedup_sizes if i <= max_num_tokens
                ]
                # sort to make sure the sizes are in ascending order
                cudagraph_capture_sizes.sort()
            else:
                if self.performance_mode == "interactivity":
                    # Fine-grained CUDA graphs at small batch sizes
                    # for minimal padding overhead
                    interactivity_max = min(max_cudagraph_capture_size, 32)
                    cudagraph_capture_sizes = list(range(1, interactivity_max + 1))
                else:
                    cudagraph_capture_sizes = [
                        i for i in [1, 2, 4] if i <= max_cudagraph_capture_size
                    ]
                if max_cudagraph_capture_size >= 8:
                    # Step size 8 for small batch sizes, up to 256(not included)
                    cudagraph_capture_sizes += list(
                        range(8, min(max_cudagraph_capture_size + 1, 256), 8)
                    )
                if max_cudagraph_capture_size >= 256:
                    # Step size 16 for larger batch sizes
                    cudagraph_capture_sizes += list(
                        range(256, max_cudagraph_capture_size + 1, 16)
                    )
                # ensure max_num_tokens is captured if within max capture size
                if (
                    max_num_tokens <= max_cudagraph_capture_size
                    and max_num_tokens not in cudagraph_capture_sizes
                ):
                    cudagraph_capture_sizes.append(max_num_tokens)
                # Preserve the platform's default capture ceiling. Larger
                # uniform decode batches fall back to eager execution unless
                # users explicitly configure wider capture sizes.
                cudagraph_capture_sizes += [
                    size for size in uniform_decode_sizes if size <= max_num_tokens
                ]
                # de-duplicate and sort the sizes
                cudagraph_capture_sizes = sorted(set(cudagraph_capture_sizes))

            # user-specific compilation_config.max_cudagraph_capture_size get
            # truncated to valid_max_size when they are inconsistent.
            valid_max_size = (
                cudagraph_capture_sizes[-1] if cudagraph_capture_sizes else 0
            )
            if (
                self.compilation_config.max_cudagraph_capture_size is not None
                and self.compilation_config.max_cudagraph_capture_size != valid_max_size
            ):
                # raise error only when both two flags are user-specified
                # and they are inconsistent with each other
                if self.compilation_config.cudagraph_capture_sizes is not None:
                    raise ValueError(
                        "customized max_cudagraph_capture_size"
                        f"(={self.compilation_config.max_cudagraph_capture_size}) "
                        "should be consistent with the max value of "
                        f"cudagraph_capture_sizes(={valid_max_size})"
                    )

                logger.warning(
                    "Truncating max_cudagraph_capture_size to %d",
                    valid_max_size,
                )
            # always set the final max_cudagraph_capture_size
            self.compilation_config.max_cudagraph_capture_size = valid_max_size

            if self.compilation_config.cudagraph_capture_sizes is not None and len(
                cudagraph_capture_sizes
            ) < len(self.compilation_config.cudagraph_capture_sizes):
                # If users have specified capture sizes, we only need to
                # compare the lens before and after modification since the modified
                # list is only the subset of the original list.
                logger.warning(
                    (
                        "cudagraph_capture_sizes specified in compilation_config"
                        " %s is overridden by config %s"
                    ),
                    self.compilation_config.cudagraph_capture_sizes,
                    cudagraph_capture_sizes,
                )
            # always write back the final sizes
            self.compilation_config.cudagraph_capture_sizes = cudagraph_capture_sizes

        else:
            # no cudagraph in use
            self.compilation_config.max_cudagraph_capture_size = 0
            self.compilation_config.cudagraph_capture_sizes = []

        # complete the remaining process.
        self.compilation_config.post_init_cudagraph_sizes()

    def _set_compile_ranges(self):
        """Set the compile ranges for the compilation config."""
        compilation_config = self.compilation_config
        computed_compile_ranges_endpoints = []

        # The upper bound of the compile ranges is the max_num_batched_tokens.
        compile_range_end = self.scheduler_config.max_num_batched_tokens
        if compile_range_end is not None:
            computed_compile_ranges_endpoints.append(compile_range_end)

        # Add the compile ranges for flashinfer/aiter.
        if compilation_config.pass_config.fuse_allreduce_rms:
            tp_size = self.parallel_config.tensor_parallel_size
            from vllm._aiter_ops import rocm_aiter_ops

            max_size: int | None = None
            if rocm_aiter_ops.is_custom_all_reduce_enabled():
                from vllm.distributed.device_communicators.aiter_custom_all_reduce import (  # noqa: E501
                    AiterCustomAllreduce,
                )

                max_size = AiterCustomAllreduce.effective_max_size()
            else:
                max_size = compilation_config.pass_config.flashinfer_max_size(tp_size)
            if max_size is not None and self.model_config is not None:
                assert isinstance(self.model_config.dtype, torch.dtype)
                max_token_num = max_size // (
                    self.model_config.get_hidden_size()
                    * self.model_config.dtype.itemsize
                )
                if compile_range_end is not None and max_token_num < compile_range_end:
                    computed_compile_ranges_endpoints.append(max_token_num)
                else:
                    logger.debug(
                        "Max num batched tokens below allreduce-rms fusion threshold, "
                        "allreduce-rms fusion will be enabled for all num_tokens."
                    )

        if compilation_config.pass_config.fuse_rope_kvcache:
            max_token_num = (
                compilation_config.pass_config.rope_kvcache_fusion_max_token_num
            )
            if max_token_num is not None:
                if compile_range_end is not None and max_token_num < compile_range_end:
                    computed_compile_ranges_endpoints.append(max_token_num)
                else:
                    logger.debug(
                        "Max num batched tokens below rope+kvcache fusion threshold, "
                        "rope+kvcache fusion enabled for num_tokens <= %d.",
                        compile_range_end,
                    )

        if compilation_config.pass_config.fuse_qk_norm_rope_kvcache:
            max_token_num = (
                compilation_config.pass_config.rope_kvcache_fusion_max_token_num
            )
            if max_token_num is not None:
                if compile_range_end is not None and max_token_num < compile_range_end:
                    computed_compile_ranges_endpoints.append(max_token_num)
                else:
                    logger.debug(
                        "Max num batched tokens below qk_norm+rope+kvcache "
                        "fusion threshold, fusion enabled for "
                        "num_tokens <= %d.",
                        compile_range_end,
                    )

        if compilation_config.compile_ranges_endpoints is not None:
            for x in compilation_config.compile_ranges_endpoints:
                assert isinstance(x, int)
                assert x > 0, f"Invalid compile range endpoint: {x}"
                if compile_range_end is not None and x < compile_range_end and x > 1:
                    computed_compile_ranges_endpoints.append(x)
        compilation_config.compile_ranges_endpoints = sorted(
            computed_compile_ranges_endpoints
        )

    def try_verify_and_update_config(self):
        if self.model_config is None:
            return

        # Avoid running try_verify_and_update_config multiple times
        if getattr(self.model_config, "config_updated", False):
            return
        self.model_config.config_updated = True

        architecture = self.model_config.architecture
        if architecture is None:
            return

        from vllm.model_executor.models import ModelRegistry
        from vllm.model_executor.models.config import (
            MODELS_CONFIG_MAP,
            HybridAttentionMambaModelConfig,
        )

        cls = MODELS_CONFIG_MAP.get(architecture, None)
        if cls is None:
            # `architecture` may be an HF base-model name (e.g. "Mamba2Model"
            # when `architectures` is omitted); normalize to the resolved arch
            # so per-arch config hooks are not skipped.
            architecture = ModelRegistry._normalize_arch(
                architecture, self.model_config
            )
            cls = MODELS_CONFIG_MAP.get(architecture, None)
        if cls is not None:
            cls.verify_and_update_config(self)

        if self.model_config.is_hybrid:
            HybridAttentionMambaModelConfig.verify_and_update_config(self)

        if self.model_config.convert_type == "classify":
            # Maybe convert ForCausalLM into ForSequenceClassification model.
            from vllm.model_executor.models.adapters import SequenceClassificationConfig

            SequenceClassificationConfig.verify_and_update_config(self)

        if hasattr(self.model_config, "model_weights") and is_runai_obj_uri(
            self.model_config.model_weights
        ):
            if self.load_config.load_format == "auto":
                logger.info(
                    "Detected Run:ai model config. "
                    "Overriding `load_format` to 'runai_streamer'"
                )
                self.load_config.load_format = "runai_streamer"
            elif self.load_config.load_format not in (
                "modelexpress",
                "runai_streamer",
                "runai_streamer_sharded",
            ):
                raise ValueError(
                    f"To load a model from object storage (S3/GCS/Azure), "
                    f"'load_format' must be 'modelexpress', 'runai_streamer' or "
                    f"'runai_streamer_sharded', "
                    f"but got '{self.load_config.load_format}'. "
                    f"Model: {self.model_config.model}"
                )

    def compile_debug_dump_path(self) -> Path | None:
        """Returns a rank-aware path for dumping
        torch.compile debug information.
        """
        if self.compilation_config.debug_dump_path is None:
            return None
        tp_rank = self.parallel_config.rank
        dp_rank = self.parallel_config.data_parallel_index
        append_path = f"rank_{tp_rank}_dp_{dp_rank}"
        path = self.compilation_config.debug_dump_path / append_path
        return path

    def __str__(self):
        return (
            f"model={self.model_config.model!r}, "
            f"speculative_config={self.speculative_config!r}, "
            f"tokenizer={self.model_config.tokenizer!r}, "
            f"skip_tokenizer_init={self.model_config.skip_tokenizer_init}, "
            f"tokenizer_mode={self.model_config.tokenizer_mode}, "
            f"revision={self.model_config.revision}, "
            f"tokenizer_revision={self.model_config.tokenizer_revision}, "
            f"trust_remote_code={self.model_config.trust_remote_code}, "
            f"dtype={self.model_config.dtype}, "
            f"max_seq_len={self.model_config.max_model_len}, "
            f"download_dir={self.load_config.download_dir!r}, "
            f"load_format={self.load_config.load_format}, "
            f"tensor_parallel_size={self.parallel_config.tensor_parallel_size}, "  # noqa
            f"pipeline_parallel_size={self.parallel_config.pipeline_parallel_size}, "  # noqa
            f"data_parallel_size={self.parallel_config.data_parallel_size}, "  # noqa
            f"decode_context_parallel_size={self.parallel_config.decode_context_parallel_size}, "  # noqa
            f"dcp_comm_backend={self.parallel_config.dcp_comm_backend}, "  # noqa
            f"disable_custom_all_reduce={self.parallel_config.disable_custom_all_reduce}, "  # noqa
            f"quantization={self.model_config.quantization}, "
            f"quantization_config={self.model_config.quantization_config}, "  # noqa
            f"enforce_eager={self.model_config.enforce_eager}, "
            f"aux_output_config={self.aux_output_config!r}, "
            f"kv_cache_dtype={self.cache_config.cache_dtype}, "
            f"device_config={self.device_config.device}, "
            f"structured_outputs_config={self.structured_outputs_config!r}, "
            f"observability_config={self.observability_config!r}, "
            f"seed={self.model_config.seed}, "
            f"served_model_name={self.model_config.served_model_name}, "
            f"enable_prefix_caching={self.cache_config.enable_prefix_caching}, "
            f"enable_chunked_prefill={self.scheduler_config.enable_chunked_prefill}, "  # noqa
            f"pooler_config={self.model_config.pooler_config!r}, "
            f"compilation_config={self.compilation_config!r}, "
            f"kernel_config={self.kernel_config!r}"
        )

    def _resolve_mm_embedding_inputs(self) -> None:
        """Accept embedding inputs, tensor optional, on disaggregated consumers.

        An EC consumer loads embeddings from its connector. A KV consumer
        receives the prompt KV produced from those embeddings, so it does not
        need the tensors either. On every other deployment a missing tensor is
        a client error and must keep failing fast in the frontend.
        """
        model_config = self.model_config
        if model_config is None:
            return
        mm_config = model_config.multimodal_config
        if mm_config is None:
            return

        ec_config = self.ec_transfer_config
        kv_config = self.kv_transfer_config
        # Derived, so overwrite unconditionally rather than honouring a value
        # that was set by hand.
        mm_config.allow_missing_mm_embeddings = (
            ec_config is not None and ec_config.is_ec_consumer
        ) or (kv_config is not None and kv_config.is_kv_consumer)
        if not mm_config.allow_missing_mm_embeddings:
            return

        if not mm_config.enable_mm_embeds:
            # Allowing missing tensors still requires enabling embedding inputs
            # for the frontend to accept metadata-only requests.
            mm_config.enable_mm_embeds = True
            logger.info_once(
                "EC/KV consumer: accepting pre-computed-embedding inputs, "
                "which this role is sent by definition."
            )
        logger.info_once(
            "EC/KV consumer: pre-computed-embedding inputs may "
            "omit the embedding tensor."
        )

    def _resolve_mm_encoder_only(self) -> None:
        """Enable encoder-only mode for a dedicated EC producer."""
        ec_config = self.ec_transfer_config
        if ec_config is None or not ec_config.is_encode_only:
            return

        model_config = self.model_config
        mm_config = model_config.multimodal_config if model_config is not None else None
        if mm_config is None:
            raise ValueError(
                "An EC producer-only instance requires a multimodal model."
            )
        mm_config.mm_encoder_only = True

    def _resolve_mm_processor_device(self) -> None:
        """Settle `--mm-processor-device=auto` now that the EC role is known.

        "auto" means "the accelerator, but only where the processor has it to
        itself and its output can be handed over without a copy back to host":
        an encode-only instance whose tensor transport carries device tensors.
        Every other deployment keeps the processor on CPU.

        An explicit device -- from `--mm-processor-device` or straight from
        `mm_processor_kwargs` -- is already folded in by `MultiModalConfig`, so
        it is left alone here and validated by `_validate_mm_processor_device`.
        """
        model_config = self.model_config
        if model_config is None:
            return
        mm_config = model_config.multimodal_config
        if mm_config is None:
            return
        if mm_config.get_mm_processor_device_type() is not None:
            return

        from vllm.platforms import current_platform

        device_type = current_platform.device_type
        if device_type in ("", "cpu"):
            return

        ec_config = self.ec_transfer_config
        # An EC producer that is not also a consumer runs no forward pass and
        # allocates no KV cache, so frontend accelerator work has the device to
        # itself.
        if ec_config is None or not ec_config.is_encode_only:
            return

        if mm_config.mm_tensor_ipc != "torch_shm":
            # Any other transport serializes host bytes, so the output would be
            # copied back, and that copy costs more than running the transform
            # on device saves.
            logger.info_once(
                "EPD encoder instance: keeping the multi-modal processor on CPU "
                "because mm_tensor_ipc=%s cannot carry device tensors. Add "
                "--mm-tensor-ipc=torch_shm to run it on the accelerator.",
                mm_config.mm_tensor_ipc,
            )
            return

        mm_config.mm_processor_kwargs = {
            **(mm_config.mm_processor_kwargs or {}),
            "device": device_type,
        }
        logger.info_once(
            "EPD encoder instance: running the multi-modal processor on %s. "
            "Override with --mm-processor-device=cpu.",
            device_type,
        )

    def _resolve_mm_video_decode_device(self) -> None:
        """Default video decoding to torchcodec GPU backend for EPD encoder-only
        instance if the mm processor runs on CUDA.

        The processor consumes the decoded frames on-device in that case, so
        keeping the frames on the GPU skips the host round-trip through the
        CPU media path. An explicit codec/backend choice in
        `--media-io-kwargs` is left alone, and the default is skipped where
        torchcodec (or its FFmpeg runtime) is unavailable.
        """
        if self.model_config is None or self.model_config.multimodal_config is None:
            return
        mm_config = self.model_config.multimodal_config

        ec_config = self.ec_transfer_config
        # An EC producer that is not also a consumer runs no forward pass and
        # allocates no KV cache, so frontend accelerator work has the device to
        # itself.
        if ec_config is None or not ec_config.is_encode_only:
            return

        from vllm.platforms import current_platform

        device_type = current_platform.device_type
        if (
            device_type != "cuda"
            or mm_config.get_mm_processor_device_type() != device_type
        ):
            return

        # User set video backend or device explicitly
        video_kwargs = mm_config.media_io_kwargs.setdefault("video", {})
        if "backend" in video_kwargs or "device" in video_kwargs:
            return

        from vllm.utils.import_utils import check_torchcodec_available

        try:
            check_torchcodec_available()
        except (ImportError, RuntimeError):
            # torchcodec is not installed, or is installed without a usable
            # FFmpeg runtime (it raises rather than returning False).
            logger.info_once(
                "EPD encoder instance: keeping CPU video decoding because "
                "torchcodec is not available (needs a CUDA build with FFmpeg)."
            )
            return

        video_kwargs["backend"] = "torchcodec"
        video_kwargs["device"] = device_type
        logger.info_once(
            "EPD encoder instance: decoding video with NVDEC (torchcodec device=%s).",
            device_type,
        )

    def _validate_mm_processor_device(self) -> None:
        """Hand the EC config to `MultiModalConfig`, which owns the rule."""
        model_config = self.model_config
        if model_config is None:
            return
        mm_config = model_config.multimodal_config
        if mm_config is None:
            return

        mm_config.validate_mm_processor_device(self.ec_transfer_config)

    def _get_v2_model_runner_unsupported_features(self) -> list[str]:
        """Collect features not yet supported by the V2 model runner."""
        unsupported: list[str] = []
        speculative_config = self.speculative_config

        # V2 does not implement the external_launcher (torchrun) PP-output
        # broadcast that V1 uses to keep all ranks in sync (broadcast_pp_output).
        if (
            self.parallel_config.distributed_executor_backend == "external_launcher"
            and self.parallel_config.pipeline_parallel_size > 1
        ):
            unsupported.append("pipeline parallelism with external_launcher")

        if speculative_config is not None:
            if speculative_config.method in (
                "suffix",
                "medusa",
                "mlp_speculator",
                "custom_class",
            ):
                unsupported.append(f"speculative method '{speculative_config.method}'")

            # V2 EagleSpeculator does not support parallel_drafting (for P-Eagle).
            # DFlash and DSpark use parallel drafting natively in V2 via their
            # own speculators.
            if (
                speculative_config.parallel_drafting
                and speculative_config.method not in ("dflash", "dspark")
            ):
                unsupported.append("parallel drafting for EAGLE speculative decoding")

            # The V2 draft-model speculator has no token mapping between the
            # draft and target vocabularies (TLI is TBD in #47172).
            if getattr(speculative_config, "use_heterogeneous_vocab", False):
                unsupported.append("heterogeneous-vocabulary draft models")

        if self.parallel_config.use_ubatching:
            unsupported.extend(self._get_dbo_unsupported_features())

        return unsupported

    def _get_v1_model_runner_unsupported_features(self) -> list[str]:
        unsupported: list[str] = []

        # PCP runtime support is implemented only by the V2 model runner.
        if self.parallel_config.prefill_context_parallel_size > 1:
            unsupported.append("prefill context parallel")

        # Note(arpera):
        # MRV1 + PP>1 + async sched + structured output
        # does not work in vLLM. For more info see:
        # https://github.com/vllm-project/vllm/issues/45014
        # Since recently MRV1 has been deprecated then
        # there was decided not to fix the issue but instead
        # to disallow such configuration.
        # At the same time MRV2 works fine in this case.
        if (
            self.parallel_config.pipeline_parallel_size > 1
            and self.scheduler_config.async_scheduling
        ):
            unsupported.append("pipeline parallelism with async scheduling")

        # DSpark is implemented only by the V2 GPU model runner.
        if self.speculative_config:
            if self.speculative_config.method == "dspark":
                unsupported.append("dspark speculative decoding")
            if self.speculative_config.enable_adaptive_verification:
                unsupported.append("adaptive draft verification")

        # Mixed sliding/full DFlash drafts need multiple KV groups (V2 only).
        if self._dflash_needs_multi_kv_group():
            unsupported.append("mixed sliding/full dflash drafts")

        # DFlash candidate heads exist only in the V2 speculator. On
        # V1 the same checkpoint drafts through DFlashProposer, which never
        # calls it, so the draft would degrade to DFlash1 silently.
        if self._is_dflash_candidate_draft():
            unsupported.append("DFlash candidate-head drafts")

        if self.model_config is not None and self.model_config.is_diffusion:
            unsupported.append("diffusion models")

        if self.parallel_config.enable_batch_sharded_sampling:
            unsupported.append("batch-sharded sampling")

        return unsupported

    def _validate_adaptive_verification(self) -> None:
        spec_config = self.speculative_config
        if not spec_config or not spec_config.enable_adaptive_verification:
            return

        if self.lora_config is not None:
            # The per-token LoRA mapping is built from CPU placeholder boundaries,
            # while the trimmed batch's true boundaries are decided on the GPU.
            raise ValueError(
                "Adaptive verification is not currently compatible with LoRA"
            )

        if not self.compilation_config.cudagraph_mode.has_full_cudagraphs():
            raise ValueError(
                "Adaptive verification requires full CUDA graphs. Use cudagraph "
                "mode FULL, FULL_DECODE_ONLY, or FULL_AND_PIECEWISE, or disable "
                "adaptive verification."
            )

        if self.parallel_config.pipeline_parallel_size > 1:
            # Cost curves and confidences currently only exist on the last PP rank;
            # earlier ranks would diverge on the trimmed batch shape.
            # TODO: we should be able to support adaptive verification with PP by
            # broadcasting the cost curves and confidences to all ranks.
            raise ValueError(
                "Adaptive verification is not currently compatible "
                "with pipeline parallelism"
            )

    def _validate_batch_sharded_sampling(self) -> None:
        """Validate `enable_batch_sharded_sampling` against the rest of the config."""
        if not self.parallel_config.enable_batch_sharded_sampling:
            # Default to False if not set.
            self.parallel_config.enable_batch_sharded_sampling = False
            return

        blockers: list[str] = []
        tp_size = self.parallel_config.tensor_parallel_size

        if tp_size <= 1:
            blockers.append("tensor_parallel_size is 1, so there is nothing to shard")
        elif self.scheduler_config.max_num_seqs < tp_size:
            # Requests are assigned to ranks whole, so fewer slots than ranks
            # leaves some ranks without work in every step.
            blockers.append(
                f"max_num_seqs ({self.scheduler_config.max_num_seqs}) is below "
                f"tensor_parallel_size ({tp_size})"
            )

        if self.model_config is not None and self.model_config.max_logprobs < 0:
            # max_logprobs == -1 allows vocab-size logprob requests, which the
            # fixed-width logprobs gather cannot reasonably size for.
            blockers.append("max_logprobs is -1, allowing vocab-size logprob requests")

        if self.model_config is not None and self.model_config.return_sampling_mask:
            # gather_sampler_output() drops SamplingMaskTensors: masks come back None.
            blockers.append(
                "return_sampling_mask is set and the batch-sharded gather does "
                "not forward sampling masks"
            )

        if (
            self.speculative_config is not None
            and self.speculative_config.enable_adaptive_verification
        ):
            # Adaptive verification picks the per-request draft split on the GPU,
            # so cu_num_logits_np is only an upper bound, while the shard plan is
            # built from that CPU array. The two disagree once the budget binds.
            # TODO(TheEpicDolphin): Support adaptive verification with batch-sharded
            # sampling.
            blockers.append(
                "it does not yet work with adaptive verification, which decides "
                "the per-request logits counts on the GPU, where the CPU-side "
                "shard plan cannot see them"
            )

        if blockers:
            raise ValueError(
                "Batch-sharded sampling was explicitly enabled via "
                "the --enable-batch-sharded-sampling flag, but is not supported "
                "in this configuration for the following reason(s): "
                f"{'; '.join(blockers)}."
            )

    def _get_dbo_unsupported_features(self) -> list[str]:
        """Collect what the V2 model runner cannot combine with DBO.

        The V2 runner microbatches a plain decoder forward pass. Anything that
        slices or replays the batch differently (drafting, adapters, pipeline
        stages, context parallelism, encoders) is not handled yet.
        """
        # TODO: DBO with model runner V2 is under development.
        # It should be enabled with explicit VLLM_USE_V2_MODEL_RUNNER environ.
        # Remove it when stable.
        if envs.VLLM_USE_V2_MODEL_RUNNER is None:
            return ["dual batch overlap"]

        unsupported: list[str] = []
        model_config = self.model_config
        parallel_config = self.parallel_config

        if self.lora_config is not None:
            unsupported.append("dual batch overlap with LoRA")
        if self.speculative_config is not None:
            unsupported.append("dual batch overlap with speculative decoding")
        if parallel_config.pipeline_parallel_size > 1:
            unsupported.append("dual batch overlap with pipeline parallelism")
        if (
            parallel_config.decode_context_parallel_size > 1
            or parallel_config.prefill_context_parallel_size > 1
        ):
            unsupported.append("dual batch overlap with context parallelism")
        if model_config is not None and (
            model_config.is_multimodal_model or model_config.is_encoder_decoder
        ):
            unsupported.append("dual batch overlap with multimodal models")
        if model_config is not None and model_config.is_hybrid:
            unsupported.append("dual batch overlap with hybrid models")
        if self.compilation_config.cudagraph_mode == CUDAGraphMode.PIECEWISE:
            # DBO captures FULL graphs only.
            unsupported.append("dual batch overlap with PIECEWISE CUDA graphs")
        if self.is_mm_encoder_only:
            unsupported.append("dual batch overlap with encoder only models")

        return unsupported

    def _validate_profiler_config(self) -> None:
        if self.profiler_config.profiler != "proton":
            return

        from vllm.platforms import current_platform

        if not current_platform.is_cuda():
            raise ValueError("The Proton profiler currently supports NVIDIA CUDA only")
        if (
            self.profiler_config.proton_graph_attribution
            and not self.use_v2_model_runner
        ):
            raise ValueError(
                "Proton CUDA graph attribution requires the V2 model runner."
            )

        has_cuda_graphs = (
            self.compilation_config.cudagraph_mode != CUDAGraphMode.NONE
            or (
                self.compilation_config.cudagraph_mm_encoder
                and not (self.model_config and self.model_config.enforce_eager)
            )
        )
        if not has_cuda_graphs:
            return
        mode = self.profiler_config.proton_mode
        if mode and mode.split(":", 1)[0].lower() == "pcsampling":
            raise ValueError(
                "Proton PC sampling requires CUDA graphs to be disabled. "
                "Use --enforce-eager."
            )
        if not self.profiler_config.proton_graph_attribution:
            raise ValueError(
                "Proton profiling with CUDA graphs requires "
                "proton_graph_attribution=True to capture replayed kernels. "
                "Enable attribution or use --enforce-eager."
            )

    def _disable_cudagraphs_for_v2_stock_torch_compile(self) -> None:
        """Run stock torch.compile without CUDA graphs in Model Runner V2.

        V1 never wraps a stock-compiled model in CUDAGraphWrapper, so it runs
        without CUDA graphs. V2's CUDA graph manager would otherwise capture
        FULL graphs around the stock-compiled model, so disable them to match.
        """
        compilation_config = self.compilation_config
        if (
            compilation_config.mode != CompilationMode.STOCK_TORCH_COMPILE
            or compilation_config.cudagraph_mode == CUDAGraphMode.NONE
        ):
            return
        logger.info_once(
            "CUDA graphs are not supported with stock torch.compile in Model "
            "Runner V2. Overriding cudagraph_mode %s to NONE.",
            compilation_config.cudagraph_mode.name,
        )
        compilation_config.cudagraph_mode = CUDAGraphMode.NONE
        compilation_config.max_cudagraph_capture_size = 0
        compilation_config.cudagraph_capture_sizes = []

    def _validate_v2_model_runner(self) -> None:
        """Check for features not yet supported by the V2 model runner."""
        if not HAS_TRITON:
            raise ValueError("Model Runner V2 requires Triton.")

        unsupported = self._get_v2_model_runner_unsupported_features()
        if unsupported:
            raise ValueError(
                f"Model Runner V2 does not yet support: {', '.join(unsupported)}"
            )

    def _validate_v1_model_runner(self) -> None:
        unsupported = self._get_v1_model_runner_unsupported_features()
        if unsupported:
            raise ValueError(
                f"Model Runner V1 does not support: {', '.join(unsupported)}"
            )

    def adjust_dcp_kv_cache_interleave_size(
        self, kv_cache_config: "KVCacheConfig"
    ) -> None:
        """Normalize DCP interleave size against block_size for NIXL P/D.

        Called by each worker (via ensure_kv_transfer_initialized), once it knows its
        own final block_size via kv_cache_config.
        """
        dcp_size = self.parallel_config.decode_context_parallel_size
        if dcp_size <= 1:
            return
        if self.parallel_config.dcp_kv_cache_interleave_size > 1 and (
            self.parallel_config.cp_kv_cache_interleave_size
            != self.parallel_config.dcp_kv_cache_interleave_size
        ):
            self.parallel_config.cp_kv_cache_interleave_size = (
                self.parallel_config.dcp_kv_cache_interleave_size
            )
            logger.warning_once(
                "cp_kv_cache_interleave_size is overridden by dcp_kv_cache"
                "_interleave_size. And dcp-kv-cache-interleave-size will be "
                "deprecated when PCP is fully supported."
            )

        if self.kv_transfer_config is None or not self.kv_transfer_config.has_connector(
            "NixlConnector"
        ):
            return
        if not self.parallel_config._allow_auto_resolve_cp_interleave_size:
            return

        # Get the kernel block_size, but don't use resolve_kv_cache_block_size to avoid
        # scaling by dcp_size (we need the local block_size here).
        local_block_size = min(
            g.kv_cache_spec.block_size for g in kv_cache_config.kv_cache_groups
        )
        if self.parallel_config.cp_kv_cache_interleave_size != local_block_size:
            interleave = self.parallel_config.cp_kv_cache_interleave_size
            self.parallel_config.cp_kv_cache_interleave_size = local_block_size
            logger.info_once(
                "When using PD disaggregation with DCP "
                "(decode_context_parallel_size=%d), "
                "cp_kv_cache_interleave_size is automatically adjusted "
                "from %d to block_size %d for block-level alignment.",
                dcp_size,
                interleave,
                local_block_size,
            )

    def validate_block_size(self) -> None:
        """Validate block_size against DCP and mamba constraints.

        Called after Platform.update_block_size_for_backend() has
        finalised block_size.
        """
        block_size = self.cache_config.block_size

        # Skip DCP interleave-size compatibility for NIXL P/D: the interleave
        # size is pinned to block_size by each worker.
        nixl_pd_active = (
            self.kv_transfer_config is not None
            and self.kv_transfer_config.has_connector("NixlConnector")
        )
        if self.parallel_config.decode_context_parallel_size > 1 and not nixl_pd_active:
            assert (
                self.parallel_config.cp_kv_cache_interleave_size <= block_size
                and block_size % self.parallel_config.cp_kv_cache_interleave_size == 0
            ), (
                f"Block_size({block_size}) should be greater "
                "than or equal to and divisible by cp_kv_cache_interleave_size "
                f"({self.parallel_config.cp_kv_cache_interleave_size})."
            )
        # Mamba cache align-mode constraints
        if self.cache_config.mamba_cache_mode == "align":
            assert not self.scheduler_config.disable_chunked_mm_input, (
                "Chunked MM input is required because we need the flexibility "
                "to schedule a multiple of block_size tokens even if they are "
                "in the middle of a mm input"
            )

    @model_validator(mode="after")
    def validate_nvfp4_kv_cache_with_mla(self) -> "VllmConfig":
        if self.model_config is None:
            return self
        # The ds_mla layouts are MLA-only by construction; the plain nvfp4
        # layout (head_size//2 + head_size//16) does not apply to MLA.
        if (
            self.cache_config.cache_dtype.startswith("nvfp4")
            and not self.cache_config.cache_dtype.endswith("_ds_mla")
            and self.model_config.use_mla
        ):
            raise ValueError(
                "nvfp4 KV cache is not supported with MLA (Multi-head Latent "
                "Attention) backends. Please use a different --kv-cache-dtype "
                "(e.g., 'fp8', 'auto', or 'nvfp4_ds_mla' with a sparse MLA "
                "backend) for MLA models such as DeepSeek."
            )
        return self

    @model_validator(mode="after")
    def validate_mamba_block_size(self) -> "VllmConfig":
        if self.model_config is None:
            return self
        mamba_block_size_is_set = (
            self.cache_config.mamba_block_size is not None
            and self.cache_config.mamba_block_size != self.model_config.max_model_len
        )
        if mamba_block_size_is_set and not self.cache_config.enable_prefix_caching:
            raise ValueError(
                "--mamba-block-size can only be set with --enable-prefix-caching"
            )
        return self

    @model_validator(mode="after")
    def validate_mamba_cached_kernel(self) -> "VllmConfig":
        if not self.cache_config.use_replayssm:
            self.cache_config.use_kda_recoverssm = False
            return self

        kda_architectures = (
            "KimiLinearForCausalLM",
            "KimiK3ForConditionalGeneration",
        )
        is_kda_model = (
            self.model_config is not None
            and self.model_config.architecture in kda_architectures
        )
        self.cache_config.use_kda_recoverssm = (
            self.num_speculative_tokens > 0 and is_kda_model
        )
        use_mamba_replayssm_spec = (
            self.num_speculative_tokens > 0 and not self.cache_config.use_kda_recoverssm
        )

        if self.model_config is not None and not self.model_config.supports_replayssm:
            raise ValueError(
                "--use-replayssm is not supported for architecture "
                f"{self.model_config.architecture!r}"
            )
        if (
            self.mamba_config.backend == MambaBackendEnum.FLASHINFER
            and self.cache_config.replayssm_buffer_len > 16
        ):
            raise ValueError(
                "FlashInfer ReplaySSM requires --replayssm-buffer-len <= 16"
            )
        if self.cache_config.use_kda_recoverssm:
            if self.mamba_config.enable_stochastic_rounding:
                raise ValueError(
                    "RecoverSSM supports bfloat16/float32 "
                    "SSM state caches, not --enable-mamba-cache-stochastic-"
                    "rounding, which requires an explicit float16 cache"
                )
            if (
                self.cache_config.mamba_cache_mode == "align"
                and not self.use_v2_model_runner
            ):
                raise ValueError(
                    "RecoverSSM with align mode requires VLLM_USE_V2_MODEL_RUNNER=1"
                )
            if self.parallel_config.pipeline_parallel_size > 1:
                raise ValueError(
                    "RecoverSSM currently requires pipeline_parallel_size=1"
                )
            if self.mamba_config.backend != MambaBackendEnum.TRITON:
                raise ValueError("RecoverSSM requires --mamba-backend triton")
        elif use_mamba_replayssm_spec:
            if self.cache_config.mamba_cache_mode != "none":
                raise ValueError(
                    "FlashInfer ReplaySSM speculative decoding requires "
                    "--mamba-cache-mode none"
                )
            query_len = 1 + self.num_speculative_tokens
            if self.cache_config.replayssm_buffer_len < query_len:
                raise ValueError(
                    "FlashInfer ReplaySSM speculative decoding requires "
                    "--replayssm-buffer-len >= 1 + num_speculative_tokens "
                    f"({query_len}); got "
                    f"{self.cache_config.replayssm_buffer_len}"
                )
            if self.mamba_config.backend != MambaBackendEnum.FLASHINFER:
                raise ValueError(
                    "Mamba2 ReplaySSM speculative decoding requires "
                    "--mamba-backend flashinfer"
                )
        elif self.mamba_config.backend == MambaBackendEnum.FLASHINFER:
            if self.cache_config.mamba_cache_mode == "align":
                raise ValueError(
                    "FlashInfer ReplaySSM does not support "
                    "--mamba-cache-mode align yet; use none"
                )
        elif self.mamba_config.backend != MambaBackendEnum.TRITON:
            raise ValueError(
                "--use-replayssm requires --mamba-backend triton or flashinfer"
            )
        elif self.use_v2_model_runner:
            raise ValueError(
                "Triton ReplaySSM requires Model Runner V1; use "
                "--mamba-backend flashinfer or Model Runner V1"
            )
        if (
            self.kv_transfer_config is not None
            and self.kv_transfer_config.is_kv_transfer_instance
        ):
            raise ValueError(
                "--use-replayssm is incompatible with KV connectors "
                "(P/D disaggregation, KV cache offload)"
            )
        return self

additional_config = Field(default_factory=dict) class-attribute instance-attribute

Additional config for specified platform. Different platforms may support different configs. Make sure the configs are valid for the platform you are using. Contents must be hashable.

attention_config = Field(default_factory=AttentionConfig) class-attribute instance-attribute

Attention configuration.

aux_output_config = Field(default_factory=AuxOutputConfig) class-attribute instance-attribute

Execution auxiliary output configuration.

cache_config = Field(default_factory=CacheConfig) class-attribute instance-attribute

Cache configuration.

compilation_config = Field(default_factory=CompilationConfig) class-attribute instance-attribute

torch.compile and cudagraph capture configuration for the model.

As a shorthand, one can append compilation arguments via -cc.parameter=argument such as -cc.mode=3 (same as -cc='{"mode":3}').

You can specify the full compilation config like so: {"mode": 3, "cudagraph_capture_sizes": [1, 2, 4, 8]}

device_config = Field(default_factory=DeviceConfig) class-attribute instance-attribute

Device configuration.

diffusion_config = None class-attribute instance-attribute

Diffusion LLM (dLLM) configuration.

ec_manager_config = Field(default_factory=EncoderCacheManagerConfig) class-attribute instance-attribute

The configurations for custom encoder cache manager.

ec_transfer_config = None class-attribute instance-attribute

The configurations for distributed EC cache transfer.

engram_config = None class-attribute instance-attribute

N-gram embedding storage and sharding settings.

instance_id = '' class-attribute instance-attribute

The ID of the vLLM instance.

kernel_config = Field(default_factory=KernelConfig) class-attribute instance-attribute

Kernel configuration.

kv_events_config = None class-attribute instance-attribute

The configurations for event publishing.

kv_transfer_config = None class-attribute instance-attribute

The configurations for distributed KV cache transfer.

load_config = Field(default_factory=LoadConfig) class-attribute instance-attribute

Load configuration.

logging_config = Field(default_factory=LoggingConfig) class-attribute instance-attribute

Logging configuration.

lora_config = None class-attribute instance-attribute

LoRA configuration.

mamba_config = Field(default_factory=MambaConfig) class-attribute instance-attribute

Mamba configuration.

model_config = None class-attribute instance-attribute

Model configuration.

needs_dp_coordinator property

Determine if the DPCoordinator process is needed.

The DPCoordinator is needed in two cases: 1. For MoE models with DP > 1: to handle wave coordination (even in external LB mode, since wave coordination runs in the coordinator) 2. For non-MoE models in internal/hybrid LB mode: to collect and publish queue stats for load balancing across DP ranks

Returns:

  • bool –

    True if DPCoordinator process is needed, False otherwise.

num_lookahead_tokens property

KV slots to reserve past the tokens the target model is scheduled for.

The drafter writes KV for positions beyond the target model's query range, so every component that reserves blocks must add this margin: the scheduler through allocate_slots, and the worker warmup, which builds its own SchedulerOutputs. Consumers must read this property rather than re-deriving their own per-method lookahead, so the scheduler and warmup cannot drift apart.

num_prefill_lookahead_tokens property

Prefill tokens past the computed range that the drafter reads.

Mid-prefill the drafter consumes tokens the target model has not been scheduled for yet, so every component that has to keep them available must apply this margin: the scheduler, which never ends a chunk within it and shifts encoder scheduling by it, and the KV cache manager, which treats the trailing this - 1 tokens as re-prefillable rather than finalized. Consumers must read this property rather than re-deriving their own per-method lookahead, so those components cannot drift apart.

observability_config = Field(default_factory=ObservabilityConfig) class-attribute instance-attribute

Observability configuration.

offload_config = Field(default_factory=OffloadConfig) class-attribute instance-attribute

Model weight offloading configuration.

optimization_level = OptimizationLevel.O2 class-attribute instance-attribute

The optimization level. These levels trade startup time cost for performance, with -O0 having the best startup time and -O3 having the best performance. -O2 is used by default. See OptimizationLevel for full description.

parallel_config = Field(default_factory=ParallelConfig) class-attribute instance-attribute

Parallel configuration.

performance_mode = 'balanced' class-attribute instance-attribute

Performance mode for runtime behavior, 'balanced' is the default. 'interactivity' favors low end-to-end per-request latency at small batch sizes (fine-grained CUDA graphs, latency-oriented kernels). 'throughput' favors aggregate tokens/sec at high concurrency (larger CUDA graphs, more aggressive batching, throughput-oriented kernels).

profiler_config = Field(default_factory=ProfilerConfig) class-attribute instance-attribute

Profiling configuration.

quant_config = None class-attribute instance-attribute

Quantization configuration.

reasoning_config = None class-attribute instance-attribute

The configurations for reasoning model.

scheduler_config = Field(default_factory=SchedulerConfig.default_factory) class-attribute instance-attribute

Scheduler configuration.

shutdown_timeout = Field(default=0, ge=0) class-attribute instance-attribute

Shutdown grace period for in-flight requests. Shutdown will be delayed for up to this amount of time to allow already-running requests to complete. Any remaining requests are aborted once the timeout is reached.

speculative_config = None class-attribute instance-attribute

Speculative decoding configuration.

structured_outputs_config = Field(default_factory=StructuredOutputsConfig) class-attribute instance-attribute

Structured outputs configuration.

uniform_decode_query_len property

Query length of every request in a uniform decode batch.

A decode step submits one query for the newly sampled token plus one for each draft token, so the widest uniform decode batch the scheduler can build is max_num_seqs * uniform_decode_query_len tokens. Anything that has to cover a decode batch reads this, so the sizing rule cannot drift between the places that apply it.

This deliberately does not derive from the KV slots a drafter reserves past the target's query range, which is a reservation contract rather than a query-length one. The two do not differ by a constant: DFlash reserves num_speculative_tokens + 1 slots yet still verifies 1 + num_speculative_tokens queries, while EAGLE reserves num_speculative_tokens and verifies the same 1 + n. Deriving one from the other would under-size EAGLE by a full request width, which is the failure this property exists to prevent.

use_cumem_cudagraph_pool property

Whether CUDA graphs go to the cuMem pool that sleep offloads.

watermark_config = None class-attribute instance-attribute

Text watermarking configuration.

weight_transfer_config = None class-attribute instance-attribute

The configurations for weight transfer during RL training.

__post_init__()

Verify configs are valid & consistent with each other.

Source code in vllm/config/vllm.py
1444
1445
1446
1447
1448
1449
1450
1451
1452
1453
1454
1455
1456
1457
1458
1459
1460
1461
1462
1463
1464
1465
1466
1467
1468
1469
1470
1471
1472
1473
1474
1475
1476
1477
1478
1479
1480
1481
1482
1483
1484
1485
1486
1487
1488
1489
1490
1491
1492
1493
1494
1495
1496
1497
1498
1499
1500
1501
1502
1503
1504
1505
1506
1507
1508
1509
1510
1511
1512
1513
1514
1515
1516
1517
1518
1519
1520
1521
1522
1523
1524
1525
1526
1527
1528
1529
1530
1531
1532
1533
1534
1535
1536
1537
1538
1539
1540
1541
1542
1543
1544
1545
1546
1547
1548
1549
1550
1551
1552
1553
1554
1555
1556
1557
1558
1559
1560
1561
1562
1563
1564
1565
1566
1567
1568
1569
1570
1571
1572
1573
1574
1575
1576
1577
1578
1579
1580
1581
1582
1583
1584
1585
1586
1587
1588
1589
1590
1591
1592
1593
1594
1595
1596
1597
1598
1599
1600
1601
1602
1603
1604
1605
1606
1607
1608
1609
1610
1611
1612
1613
1614
1615
1616
1617
1618
1619
1620
1621
1622
1623
1624
1625
1626
1627
1628
1629
1630
1631
1632
1633
1634
1635
1636
1637
1638
1639
1640
1641
1642
1643
1644
1645
1646
1647
1648
1649
1650
1651
1652
1653
1654
1655
1656
1657
1658
1659
1660
1661
1662
1663
1664
1665
1666
1667
1668
1669
1670
1671
1672
1673
1674
1675
1676
1677
1678
1679
1680
1681
1682
1683
1684
1685
1686
1687
1688
1689
1690
1691
1692
1693
1694
1695
1696
1697
1698
1699
1700
1701
1702
1703
1704
1705
1706
1707
1708
1709
1710
1711
1712
1713
1714
1715
1716
1717
1718
1719
1720
1721
1722
1723
1724
1725
1726
1727
1728
1729
1730
1731
1732
1733
1734
1735
1736
1737
1738
1739
1740
1741
1742
1743
1744
1745
1746
1747
1748
1749
1750
1751
1752
1753
1754
1755
1756
1757
1758
1759
1760
1761
1762
1763
1764
1765
1766
1767
1768
1769
1770
1771
1772
1773
1774
1775
1776
1777
1778
1779
1780
1781
1782
1783
1784
1785
1786
1787
1788
1789
1790
1791
1792
1793
1794
1795
1796
1797
1798
1799
1800
1801
1802
1803
1804
1805
1806
1807
1808
1809
1810
1811
1812
1813
1814
1815
1816
1817
1818
1819
1820
1821
1822
1823
1824
1825
1826
1827
1828
1829
1830
1831
1832
1833
1834
1835
1836
1837
1838
1839
1840
1841
1842
1843
1844
1845
1846
1847
1848
1849
1850
1851
1852
1853
1854
1855
1856
1857
1858
1859
1860
1861
1862
1863
1864
1865
1866
1867
1868
1869
1870
1871
1872
1873
1874
1875
1876
1877
1878
1879
1880
1881
1882
1883
1884
1885
1886
1887
1888
1889
1890
1891
1892
1893
1894
1895
1896
1897
1898
1899
1900
1901
1902
1903
1904
1905
1906
1907
1908
1909
1910
1911
1912
1913
1914
1915
1916
1917
1918
1919
1920
1921
1922
1923
1924
1925
1926
1927
1928
1929
1930
1931
1932
1933
1934
1935
1936
1937
1938
1939
1940
1941
1942
1943
1944
1945
1946
1947
1948
1949
1950
1951
1952
1953
1954
1955
1956
1957
1958
1959
1960
1961
1962
1963
1964
1965
1966
1967
1968
1969
1970
1971
1972
1973
1974
1975
1976
1977
1978
1979
1980
1981
1982
1983
1984
1985
1986
1987
1988
1989
1990
1991
1992
1993
1994
1995
1996
1997
1998
1999
2000
2001
2002
2003
2004
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
2027
2028
2029
2030
2031
2032
2033
2034
2035
2036
2037
2038
2039
2040
2041
2042
2043
2044
2045
2046
2047
2048
2049
2050
2051
2052
2053
2054
2055
2056
2057
2058
2059
2060
2061
2062
2063
2064
2065
2066
2067
2068
2069
2070
2071
2072
2073
2074
2075
2076
2077
2078
2079
2080
2081
2082
2083
2084
2085
2086
2087
2088
2089
2090
2091
2092
2093
2094
2095
2096
2097
2098
2099
2100
2101
2102
2103
2104
2105
2106
2107
2108
2109
2110
2111
2112
2113
2114
2115
2116
2117
2118
2119
2120
2121
2122
2123
2124
2125
2126
2127
2128
2129
2130
2131
2132
2133
2134
2135
2136
2137
2138
2139
2140
2141
2142
2143
2144
2145
2146
2147
2148
2149
2150
2151
2152
2153
2154
2155
2156
2157
2158
2159
2160
2161
2162
2163
2164
2165
2166
2167
2168
2169
2170
2171
2172
2173
2174
2175
2176
2177
2178
2179
2180
2181
2182
2183
2184
2185
2186
2187
2188
2189
2190
2191
2192
2193
2194
2195
2196
2197
2198
2199
2200
2201
2202
2203
2204
2205
2206
2207
2208
2209
2210
2211
2212
2213
2214
2215
2216
2217
2218
2219
2220
2221
2222
2223
2224
2225
2226
2227
2228
2229
2230
2231
2232
2233
2234
2235
2236
2237
2238
2239
2240
2241
2242
2243
2244
2245
2246
2247
2248
2249
2250
2251
2252
2253
2254
2255
2256
2257
2258
2259
2260
2261
2262
2263
2264
2265
2266
2267
2268
2269
2270
2271
2272
2273
2274
2275
2276
2277
2278
2279
2280
2281
2282
2283
2284
2285
2286
2287
2288
2289
2290
2291
2292
2293
2294
2295
2296
2297
2298
2299
2300
2301
2302
2303
2304
2305
2306
2307
def __post_init__(self):
    """Verify configs are valid & consistent with each other."""
    # To give each torch profile run a unique instance name.
    self.instance_id = f"{time.time_ns()}"

    if self.model_config is not None and self.model_config.is_submodel_config:
        # with_hf_config() view: the parent config was already validated,
        # and this view's empty architecture list makes the model-dependent checks
        # below unsafe (e.g. use_mla resolves the architecture registry).
        return

    self._resolve_mm_encoder_only()

    if self.is_mm_encoder_only and self.cache_config.enable_prefix_caching:
        # Such an instance publishes encoder embeddings and runs no language
        # model, so it holds no KV cache for prefix caching to reuse and its
        # coordinator would have no group to manage. Disable before
        # `try_verify_and_update_config` so model config hooks (e.g. the
        # hybrid mamba hook setting `mamba_block_size`) already see prefix
        # caching as disabled.
        logger.info(
            "Disabling prefix caching: this instance runs the "
            "multi-modal encoder only."
        )
        self.cache_config.enable_prefix_caching = False

    if self.performance_mode != "balanced":
        logger.info_once("Performance mode set to '%s'.", self.performance_mode)

    self.try_verify_and_update_config()
    self._resolve_and_verify_engram_config()

    self._check_supports_watermarking()
    # Models may have supplied their own DCP defaults above; anything still
    # unset falls back to the stock ones.
    self.parallel_config.set_dcp_defaults()

    if self.model_config is not None:
        self.model_config.verify_with_parallel_config(self.parallel_config)
        self.model_config.verify_dual_chunk_attention_config(self.load_config)

        self.parallel_config.is_moe_model = self.model_config.is_moe

    if (
        self.model_config is not None
        and self.model_config.multimodal_config is not None
        and self.model_config.multimodal_config.language_model_only
        and self.compilation_config.cudagraph_mm_encoder
    ):
        raise ValueError(
            "--language-model-only is incompatible with "
            "cudagraph_mm_encoder=True, since it disables all multimodal "
            "inputs and the multimodal encoder is never run. Please "
            "disable one of them."
        )

    self._verify_sampling_replay_config()
    self._verify_trace_replay_config()

    # A NIXL side is either fully replicated or fully DCP-sharded; MLA only.
    if (
        self.kv_transfer_config is not None
        and self.kv_transfer_config.has_connector("NixlConnector")
    ):
        dcp_size = self.parallel_config.decode_context_parallel_size
        transfer_tp_size = max(
            self.parallel_config.tensor_parallel_size,
            self.parallel_config.prefill_context_parallel_size,
        )
        assert dcp_size in (1, transfer_tp_size), (
            f"decode_context_parallel_size={dcp_size} must be 1 or equal "
            f"to the NIXL transfer parallel size={transfer_tp_size}."
        )
        if self.model_config is not None:
            assert self.model_config.use_mla or dcp_size == 1, (
                "PD with decode_context_parallel_size > 1 is only "
                "supported for MLA models."
            )
    if self.lora_config is not None:
        self.lora_config.verify_with_model_config(self.model_config)

    if (
        self.mamba_config.enable_stochastic_rounding
        and self.cache_config.mamba_ssm_cache_dtype != "float16"
    ):
        raise ValueError(
            "Stochastic rounding for Mamba cache requires "
            "the SSM cache to be float16. Please set it explicitly, "
            "by specifying `--mamba-ssm-cache-dtype float16`, or disable "
            "stochastic rounding by not specifying "
            "`--enable-mamba-cache-stochastic-rounding`."
        )

    if self.quant_config is None and self.model_config is not None:
        self.quant_config = VllmConfig._get_quantization_config(
            self.model_config, self.load_config
        )

    # "dummy" reads no weights at all, and the sharded formats read a vLLM
    # state dict, which stores tied word embeddings under the lm_head only.
    # Neither can tell us what the original checkpoint contained.
    if self.model_config is not None and self.load_config.load_format not in (
        "dummy",
        "sharded_state",
        "runai_streamer_sharded",
    ):
        self.model_config.maybe_untie_word_embeddings()

    if (
        self.quant_config is not None
        and self.model_config is not None
        and hasattr(self.quant_config, "use_deep_gemm")
        and self.quant_config.use_deep_gemm is None
    ):
        from vllm.utils.deep_gemm import should_auto_disable_deep_gemm

        model_type = getattr(self.model_config.hf_text_config, "model_type", None)
        if should_auto_disable_deep_gemm(model_type):
            self.quant_config.use_deep_gemm = False
            logger.warning_once(
                "Auto-disabled DeepGemm for model_type=%s on Blackwell. "
                "DeepGemm E8M0 scale format causes accuracy degradation "
                "for this architecture. Falling back to CUTLASS. "
                "To disable DeepGemm globally, set VLLM_USE_DEEP_GEMM=0.",
                model_type,
            )

    from vllm.platforms import current_platform
    from vllm.v1.executor.abstract import Executor

    executor_backend = self.parallel_config.distributed_executor_backend
    executor_class = Executor.get_class(self)
    executor_supports_async_sched = executor_class.supports_async_scheduling()
    uses_rocm_deepep_ht_dbo = (
        current_platform.is_rocm()
        and self.parallel_config.enable_dbo
        and self.parallel_config.all2all_backend == "deepep_high_throughput"
    )

    if self.scheduler_config.async_scheduling:
        # Async scheduling explicitly enabled, hard fail any incompatibilities.
        # Currently, async scheduling only support eagle speculative
        # decoding.
        if uses_rocm_deepep_ht_dbo:
            raise ValueError(
                "Async scheduling is not compatible with ROCm DeepEP "
                "high-throughput DBO. Please use --no-async-scheduling or "
                "select a different all2all backend."
            )
        if self.speculative_config is not None:
            if (
                self.speculative_config.method not in get_args(EagleModelTypes)
                and self.speculative_config.method not in get_args(NgramGPUTypes)
                and self.speculative_config.method != "draft_model"
                and self.speculative_config.method != "dspark"
                and self.speculative_config.method != "dflash"
            ):
                raise ValueError(
                    "Currently, async scheduling is only supported "
                    "with EAGLE/MTP/Draft Model/NGram GPU/DSpark/DFlash "
                    "kind of speculative decoding"
                )
            if self.speculative_config.disable_padded_drafter_batch:
                raise ValueError(
                    "Async scheduling is not compatible with "
                    "disable_padded_drafter_batch=True."
                )
        if not executor_supports_async_sched:
            raise ValueError(
                f"`{executor_backend}` does not support async scheduling yet."
            )
    elif self.scheduler_config.async_scheduling is None:
        # Enable async scheduling unless there is an incompatible option.
        if (
            self.model_config is not None
            and self.model_config.runner_type == "pooling"
        ):
            # The current implementation of asynchronous scheduling negatively
            # impacts performance of pooling models, so we disable by default.
            logger.debug(
                "Disabling asynchronous scheduling by default for pooling model."
            )
            self.scheduler_config.async_scheduling = False
        elif (
            self.speculative_config is not None
            and self.speculative_config.method not in get_args(EagleModelTypes)
            and self.speculative_config.method not in get_args(NgramGPUTypes)
            and self.speculative_config.method != "draft_model"
            and self.speculative_config.method != "dspark"
            and self.speculative_config.method != "dflash"
        ):
            logger.warning_once(
                "Async scheduling not supported with %s-based "
                "speculative decoding and will be disabled.",
                self.speculative_config.method,
            )
            self.scheduler_config.async_scheduling = False
        elif (
            self.speculative_config is not None
            and self.speculative_config.disable_padded_drafter_batch
        ):
            logger.warning_once(
                "Async scheduling is not compatible with "
                "disable_padded_drafter_batch=True and will be disabled.",
            )
            self.scheduler_config.async_scheduling = False
        elif not executor_supports_async_sched:
            logger.warning_once(
                "Async scheduling will be disabled because it is not supported "
                "with the `%s` distributed executor backend. ",
                executor_backend,
            )
            self.scheduler_config.async_scheduling = False
        elif uses_rocm_deepep_ht_dbo:
            logger.warning_once(
                "Async scheduling is disabled for ROCm DeepEP "
                "high-throughput DBO because that combination can corrupt "
                "DP+EP generation accuracy."
            )
            self.scheduler_config.async_scheduling = False
        elif (
            self.parallel_config.pipeline_parallel_size > 1
            and not self.use_v2_model_runner
        ):
            logger.warning_once(
                "Async scheduling is disabled because the V1 model runner "
                "does not support it with pipeline parallelism."
            )
            self.scheduler_config.async_scheduling = False
        else:
            self.scheduler_config.async_scheduling = True

    if self.parallel_config.disable_nccl_for_dp_synchronization is None:
        if self.scheduler_config.async_scheduling:
            if self.parallel_config.data_parallel_size > 1 and (
                self.model_config is None or self.model_config.is_moe
            ):
                logger.info_once(
                    "Disabling NCCL for DP synchronization "
                    "when using async scheduling.",
                )
            self.parallel_config.disable_nccl_for_dp_synchronization = True
        else:
            self.parallel_config.disable_nccl_for_dp_synchronization = False

    if (
        self.speculative_config is not None
        and self.scheduler_config.async_scheduling
        and self.model_config is not None
        and not self.model_config.disable_cascade_attn
    ):
        logger.warning_once(
            "Disabling cascade attention (not yet compatible with "
            "async speculative decoding).",
        )
        self.model_config.disable_cascade_attn = True

    if (
        self.observability_config.per_request_spec_decode_metrics != "none"
        and self.speculative_config is None
    ):
        raise ValueError(
            "--per-request-spec-decode-metrics requires speculative decoding "
            "to be enabled (via --speculative-config)."
        )

    if (
        self.model_config is not None
        and self.model_config.multimodal_config is not None
        and self.model_config.multimodal_config.mm_tensor_ipc == "torch_shm"
        and os.environ.get("VLLM_WORKER_MULTIPROC_METHOD") != "spawn"
    ):
        raise ValueError(
            "torch_shm is known to fail without "
            "VLLM_WORKER_MULTIPROC_METHOD set to spawn"
        )

    if (
        self.model_config is not None
        and self.scheduler_config.enable_chunked_prefill
        and self.model_config.dtype == torch.float32
        and current_platform.get_device_capability() == (7, 5)
    ):
        logger.warning_once(
            "Turing devices tensor cores do not support float32 matmul. "
            "To workaround this limitation, vLLM will set 'ieee' input "
            "precision for chunked prefill triton kernels."
        )

    if self.model_config is not None and self.model_config.enforce_eager:
        self.compilation_config.mode = CompilationMode.NONE
        self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE
        if self.parallel_config.enable_fault_tolerance:
            # Keep JIT warmup: in-inference Triton compilation latency
            # spikes can delay peer-fault detection past its deadline.
            logger.warning_once(
                "Enforce eager set, disabling torch.compile and CUDAGraphs. "
                "This is equivalent to setting -cc.mode=none "
                "-cc.cudagraph_mode=none"
            )
        else:
            logger.warning_once(
                "Enforce eager set, disabling torch.compile, CUDAGraphs, and "
                "JIT kernel warmup. This is equivalent to setting "
                "-cc.mode=none -cc.cudagraph_mode=none and "
                "--kernel_config.enable_jit_warmup=False"
            )
            self.kernel_config.enable_jit_warmup = False

    if os.environ.get("TORCH_COMPILE_DISABLE") == "1":
        logger.warning_once(
            "TORCH_COMPILE_DISABLE is set, disabling torch.compile. "
            "This is equivalent to setting -cc.mode=none"
        )
        self.compilation_config.mode = CompilationMode.NONE

    breakable_cudagraph_enabled = self._maybe_enable_breakable_cudagraph()

    if not breakable_cudagraph_enabled and (
        self.compilation_config.backend == "eager"
        or (
            self.compilation_config.mode is not None
            and self.compilation_config.mode != CompilationMode.VLLM_COMPILE
        )
    ):
        logger.warning_once(
            "Inductor compilation was disabled by user settings, "
            "optimizations settings that are only active during "
            "inductor compilation will be ignored."
        )

    def has_blocked_weights():
        if self.quant_config is not None:
            if hasattr(self.quant_config, "weight_block_size"):
                return self.quant_config.weight_block_size is not None
            elif hasattr(self.quant_config, "has_blocked_weights"):
                return self.quant_config.has_blocked_weights()
        return False

    # Enable quant_fp8 CUDA ops (TODO disable in follow up)
    # On H100 the CUDA kernel is faster than
    # native implementation
    # https://github.com/vllm-project/vllm/issues/25094
    if has_blocked_weights():
        custom_ops = self.compilation_config.custom_ops
        if "-quant_fp8" not in custom_ops:
            custom_ops.append("+quant_fp8")

    current_platform.apply_config_platform_defaults(self)

    if self.compilation_config.mode is None:
        if self.optimization_level > OptimizationLevel.O0:
            self.compilation_config.mode = CompilationMode.VLLM_COMPILE
        else:
            self.compilation_config.mode = CompilationMode.NONE

    # By default, enable torch wrapping only when using custom Inductor lowering
    if self.compilation_config.ir_enable_torch_wrap is None:
        self.compilation_config.ir_enable_torch_wrap = (
            self.compilation_config.mode == CompilationMode.VLLM_COMPILE
            and self.compilation_config.backend == "inductor"
        )

    if all(s not in self.compilation_config.custom_ops for s in ("all", "none")):
        if (
            self.compilation_config.backend == "inductor"
            and self.compilation_config.mode != CompilationMode.NONE
        ):
            self.compilation_config.custom_ops.append("none")
        else:
            self.compilation_config.custom_ops.append("all")

    # This populates IR op priorities,
    # must happen after compilation mode and backend are decided,
    # but before fusion defaults are applied as those may depend on op priority.
    self.kernel_config.set_platform_defaults(self)

    default_config = OPTIMIZATION_LEVEL_TO_CONFIG[self.optimization_level]
    self._apply_optimization_level_defaults(default_config)
    if self.kernel_config.enable_flashinfer_autotune is None:
        raise ValueError(
            "KernelConfig.enable_flashinfer_autotune must be set after applying "
            "optimization level defaults."
        )

    self._maybe_disable_dynamic_sd_for_data_parallel()
    self._maybe_override_dynamic_sd_cudagraph_mode()

    if (
        self.attention_config.hisparse_config is None
        and self.kv_transfer_config is not None
        and self.kv_transfer_config.has_connector("HiSparseConnector")
    ):
        self.attention_config.hisparse_config = HiSparseConfig()

    if self.attention_config.hisparse_config is not None:
        if not current_platform.is_cuda():
            raise ValueError("HiSparse currently requires NVIDIA CUDA.")
        if self.parallel_config.pipeline_parallel_size > 1:
            raise ValueError("HiSparse does not support pipeline parallelism.")
        if self.parallel_config.decode_context_parallel_size > 1:
            raise ValueError(
                "HiSparse does not support decode context parallelism."
            )
        if self.compilation_config.cudagraph_mode == CUDAGraphMode.FULL:
            raise ValueError(
                "HiSparse does not support cudagraph_mode=FULL; use "
                "FULL_AND_PIECEWISE (the default), which captures FULL graphs "
                "for decode batches."
            )
        if not self.scheduler_config.scheduler_reserve_full_isl:
            # Without it, async loads admitted against free host blocks can
            # each wait on host pages the others hold, and waiting requests
            # are never preempted to free them.
            raise ValueError(
                "HiSparse requires --scheduler-reserve-full-isl; remove "
                "--no-scheduler-reserve-full-isl."
            )
        if self.model_config is not None and not hasattr(
            self.model_config.hf_config, "index_topk"
        ):
            raise ValueError(
                "HiSparse is only supported for DSA models with index_topk."
            )
        if self.kv_transfer_config is not None and (
            self.kv_transfer_config.kv_connector
            not in (
                None,
                "NixlConnector",
                "MooncakeStoreConnector",
                "MultiConnector",
            )
        ):
            logger.warning(
                "HiSparse host-resident KV is configured with connector "
                "%s. NixlConnector (GPU-staged host imports) and "
                "MooncakeStoreConnector (shared-store offload) are the "
                "validated paths; other connectors are treated as "
                "debug/fallback paths.",
                self.kv_transfer_config.kv_connector,
            )

    self._normalize_piecewise_cudagraph_mode(
        breakable_cudagraph_enabled=breakable_cudagraph_enabled
    )

    from vllm.utils.torch_utils import HAS_OPAQUE_TYPE

    if HAS_OPAQUE_TYPE:
        # On torch >= 2.11 the hoisted OpaqueObject approach supersedes
        # fast_moe_cold_start, so force it off.
        self.compilation_config.fast_moe_cold_start = False
    elif self.compilation_config.fast_moe_cold_start is None:
        # resolve default behavior: try to be as safe as possible
        # this config is unsafe if any spec decoding draft model has a MOE.
        # We'll conservatively turn it off if we see spec decoding.
        self.compilation_config.fast_moe_cold_start = (
            self.speculative_config is None
        )

    self._set_max_num_scheduled_tokens()

    if current_platform.support_static_graph_mode():
        # if cudagraph_mode has full cudagraphs, we need to check support
        if model_config := self.model_config:
            if (
                self.compilation_config.cudagraph_mode.has_full_cudagraphs()
                and model_config.pooler_config is not None
            ):
                logger.warning_once(
                    "Pooling models do not support full cudagraphs. "
                    "Overriding cudagraph_mode to PIECEWISE."
                )
                self.compilation_config.cudagraph_mode = CUDAGraphMode.PIECEWISE
            elif (
                model_config.is_encoder_decoder
                and self.compilation_config.cudagraph_mode
                not in (CUDAGraphMode.NONE, CUDAGraphMode.FULL_DECODE_ONLY)
            ):
                logger.info_once(
                    "Encoder-decoder models do not support %s. "
                    "Overriding cudagraph_mode to FULL_DECODE_ONLY.",
                    self.compilation_config.cudagraph_mode.name,
                )
                self.compilation_config.cudagraph_mode = (
                    CUDAGraphMode.FULL_DECODE_ONLY
                )

        # Check if KV connector requires PIECEWISE mode for CUDA graphs
        if (
            self.kv_transfer_config is not None
            and self.kv_transfer_config.is_kv_transfer_instance
            and self.compilation_config.cudagraph_mode.has_full_cudagraphs()
        ):
            # Lazy import to avoid circular dependencies
            from vllm.distributed.kv_transfer.kv_connector.factory import (
                KVConnectorFactory,
            )

            connector_cls = KVConnectorFactory.get_connector_class(
                self.kv_transfer_config
            )
            if connector_cls.requires_piecewise_for_cudagraph(
                self.kv_transfer_config.kv_connector_extra_config
            ):
                logger.warning_once(
                    "KV connector %s requires PIECEWISE CUDA graph mode "
                    "due to layerwise async operations that cannot be "
                    "captured in CUDA graphs. "
                    "Overriding cudagraph_mode from %s to PIECEWISE.",
                    connector_cls.__name__,
                    self.compilation_config.cudagraph_mode.name,
                )
                self.compilation_config.cudagraph_mode = CUDAGraphMode.PIECEWISE

        # disable cudagraph when enforce eager execution
        if self.model_config is not None and self.model_config.enforce_eager:
            logger.info_once("Cudagraph is disabled under eager mode")
            self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE
            # override related settings when enforce eager
            self.compilation_config.max_cudagraph_capture_size = 0
            self.compilation_config.cudagraph_capture_sizes = []
        else:
            self.compilation_config.cudagraph_num_of_warmups = 1

        self._set_cudagraph_sizes()

    else:
        self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE

    if self.cache_config.kv_sharing_fast_prefill:
        if (
            self.speculative_config is not None
            and self.speculative_config.use_eagle()
        ):
            raise ValueError(
                "Fast prefill optimization for KV sharing is not "
                "compatible with EAGLE as EAGLE requires correct logits "
                "for all tokens while fast prefill gives incorrect logits "
                "for prompt tokens."
            )

        logger.warning_once(
            "--kv-sharing-fast-prefill requires changes on model side for "
            "correctness and to realize prefill savings."
        )

    if (
        self.model_config
        and self.model_config.architecture == "WhisperForConditionalGeneration"
        and os.environ.get("VLLM_WORKER_MULTIPROC_METHOD") != "spawn"
    ):
        logger.warning_once(
            "Whisper is known to have issues with "
            "forked workers. If startup is hanging, "
            "try setting 'VLLM_WORKER_MULTIPROC_METHOD' "
            "to 'spawn'."
        )

    if (
        self.kv_events_config is not None
        and self.kv_events_config.enable_kv_cache_events
        and not self.cache_config.enable_prefix_caching
    ):
        logger.warning_once(
            "KV cache events are on, but prefix caching is not enabled. "
            "Use --enable-prefix-caching to enable."
        )
    if (
        self.kv_events_config is not None
        and self.kv_events_config.publisher != "null"
        and not self.kv_events_config.enable_kv_cache_events
    ):
        logger.warning_once(
            "KV cache events are disabled, "
            "but the scheduler is configured to publish them. "
            "Modify KVEventsConfig.enable_kv_cache_events "
            "to True to enable."
        )
    current_platform.check_and_update_config(self)

    # After the platform hook, which has the last word on async scheduling.
    if (
        self.diffusion_config is not None
        and self.scheduler_config.scheduler_cls is None
    ):
        scheduler_name = (
            "DiffusionAsyncScheduler"
            if self.scheduler_config.async_scheduling
            else "DiffusionScheduler"
        )
        self.scheduler_config.scheduler_cls = (
            f"vllm.v1.core.sched.diffusion_scheduler.{scheduler_name}"
        )

    self._normalize_piecewise_cudagraph_mode(
        breakable_cudagraph_enabled=breakable_cudagraph_enabled
    )

    self._resolve_mm_embedding_inputs()
    self._resolve_mm_processor_device()
    self._resolve_mm_video_decode_device()
    self._validate_mm_processor_device()

    if self.use_v2_model_runner:
        self._disable_cudagraphs_for_v2_stock_torch_compile()
        self._validate_v2_model_runner()
    else:
        self._validate_v1_model_runner()

    self._validate_profiler_config()
    self._validate_batch_sharded_sampling()
    self._validate_adaptive_verification()

    # Re-compute compile ranges after platform-specific config updates
    # (e.g., XPU may lower max_num_batched_tokens when MLA is enabled)
    self._set_compile_ranges()

    if self.parallel_config.all2all_backend == "moonep":
        if (
            self.model_config is not None
            and self.model_config.quantization is not None
        ):
            raise ValueError(
                "The moonep all2all backend currently supports unquantized "
                "BF16 models only; got "
                f"quantization={self.model_config.quantization!r}. Use a "
                "different --all2all-backend for quantized models."
            )
        if (
            self.model_config is not None
            and self.model_config.dtype != torch.bfloat16
        ):
            raise ValueError(
                "The moonep all2all backend currently supports BF16 models "
                f"only; got dtype={self.model_config.dtype}. Use a "
                "different --all2all-backend or --dtype bfloat16."
            )
        if self.parallel_config.enable_eplb:
            raise ValueError(
                "The moonep all2all backend does not support EPLB yet: "
                "EPLB rearranges expert parameters in a layout MoonEP's "
                "replicated [E+B] weights do not follow. Disable "
                "--enable-eplb or use a different --all2all-backend."
            )
        if self.parallel_config.expert_placement_strategy != "linear":
            raise ValueError(
                "The moonep all2all backend requires linear expert "
                "placement: its load-time all-gather assumes each rank "
                "holds a contiguous chunk of the global expert range. Got "
                "--expert-placement-strategy "
                f"{self.parallel_config.expert_placement_strategy!r}."
            )
        # Enforced here rather than in set_splitting_ops_for_v1 so it
        # holds for every compilation mode, and keyed on use_all2all so
        # PCP/SP-only topologies are covered too: MoonEP dispatch/combine
        # are eager-only and must not be captured.
        if (
            self.parallel_config.use_all2all
            and self.compilation_config.cudagraph_mode != CUDAGraphMode.NONE
        ):
            logger.info(
                "MoonEP: Disabling CUDA Graphs since the MoonEP "
                "integration is currently eager-only."
            )
            self.compilation_config.cudagraph_mode = CUDAGraphMode.NONE

    # Do this after all the updates to compilation_config.mode
    effective_dp_size = (
        self.parallel_config.data_parallel_size
        if self.model_config is None or self.model_config.is_moe
        else 1
    )
    self.compilation_config.set_splitting_ops_for_v1(
        all2all_backend=self.parallel_config.all2all_backend,
        data_parallel_size=effective_dp_size,
    )

    # final check of cudagraph mode after all possible updates
    if current_platform.is_cuda_alike():
        if (
            self.compilation_config.cudagraph_mode.has_full_cudagraphs()
            and self.model_config is not None
            and not self.model_config.disable_cascade_attn
            and not self.compilation_config.cudagraph_mode.has_piecewise_cudagraphs()  # noqa: E501
        ):
            logger.warning_once(
                "No piecewise cudagraph for executing cascade attention. "
                "Will fall back to eager execution if a batch runs into "
                "cascade attentions."
            )

        if self.compilation_config.cudagraph_mode.requires_piecewise_compilation():
            assert (
                self.compilation_config.mode == CompilationMode.VLLM_COMPILE
                or envs.VLLM_USE_BREAKABLE_CUDAGRAPH
            ), (
                "Compilation mode should be CompilationMode.VLLM_COMPILE "
                "when cudagraph_mode piecewise cudagraphs is used, "
                f"cudagraph_mode={self.compilation_config.cudagraph_mode}"
            )
    if (
        self.model_config
        and envs.VLLM_BATCH_INVARIANT
        and not self.model_config.disable_cascade_attn
    ):
        self.model_config.disable_cascade_attn = True
        logger.warning_once(
            "Disabling cascade attention when VLLM_BATCH_INVARIANT is enabled.",
        )

    if self.parallel_config.use_ubatching:
        a2a_backend = self.parallel_config.all2all_backend
        assert a2a_backend in [
            "deepep_low_latency",
            "deepep_high_throughput",
            "nixl_ep",
        ], (
            "Microbatching currently only supports the deepep_low_latency, "
            "deepep_high_throughput, and nixl_ep all2all backends. "
            f"{a2a_backend} is not supported. To fix use "
            "--all2all-backend=deepep_low_latency, "
            "--all2all-backend=deepep_high_throughput, or "
            "--all2all-backend=nixl_ep and install the matching kernels."
        )

        if not self.model_config.disable_cascade_attn:
            self.model_config.disable_cascade_attn = True
            logger.warning_once("Disabling cascade attention when DBO is enabled.")

    if not self.instance_id:
        self.instance_id = random_uuid()[:5]

    if self.reasoning_config is not None and self.model_config is not None:
        self.reasoning_config.initialize_token_ids(self.model_config)
        if not self.reasoning_config.enabled:
            logger.warning_once(
                "Auto-initialization of reasoning token IDs failed. "
                "Please check whether your reasoning parser has implemented "
                "the `reasoning_start_str` and `reasoning_end_str`."
            )

    # Resolve kv_offloading-derived connector name into kv_transfer_config
    # before the HMA check below, which inspects the connector class.
    self._post_init_kv_transfer_config()
    self._verify_aux_output_compatibility()

    # Hybrid KV cache manager (HMA) runtime rules:
    # - Explicit enable (--no-disable-kv-cache-manager): error if runtime
    #   disables it
    # - No preference: auto-disable for unsupported features or connector configs
    # - Explicit disable (--disable-kv-cache-manager): always respect it
    need_disable_hybrid_kv_cache_manager = False
    # logger should only print warning message for hybrid models. As we
    # can't know whether the model is hybrid or not now, so we don't log
    # warning message here and will log it later.
    if not current_platform.support_hybrid_kv_cache():
        # Hybrid KV cache manager is not supported on non-GPU platforms.
        need_disable_hybrid_kv_cache_manager = True
    if (
        self.model_config is not None
        and self.model_config.attention_chunk_size is not None
    ):
        if (
            self.speculative_config is not None
            and self.speculative_config.use_eagle()
        ):
            # Hybrid KV cache manager is not yet supported with chunked
            # local attention + eagle.
            need_disable_hybrid_kv_cache_manager = True
        elif not envs.VLLM_ALLOW_CHUNKED_LOCAL_ATTN_WITH_HYBRID_KV_CACHE:
            logger.warning(
                "There is a latency regression when using chunked local"
                " attention with the hybrid KV cache manager. Disabling"
                " it, by default. To enable it, set the environment "
                "VLLM_ALLOW_CHUNKED_LOCAL_ATTN_WITH_HYBRID_KV_CACHE=1."
            )
            # Hybrid KV cache manager is not yet supported with chunked
            # local attention.
            need_disable_hybrid_kv_cache_manager = True

    if self.scheduler_config.disable_hybrid_kv_cache_manager is None:
        # Auto-disable HMA only when the connector config does not support it.
        if self.kv_transfer_config is not None:
            from vllm.distributed.kv_transfer.kv_connector.factory import (
                KVConnectorFactory,
            )

            if not KVConnectorFactory.supports_hma_config(self.kv_transfer_config):
                need_disable_hybrid_kv_cache_manager = True
                logger.warning(
                    "Turning off hybrid kv cache manager because "
                    "`--kv-transfer-config` selects a KV connector that "
                    "does not support it. Impact: hybrid SSM models "
                    "(e.g. Jamba, Bamba) require HMA and will fail at "
                    "startup without it; models with sliding window "
                    "attention will run with reduced performance. "
                    "To add HMA support to a KV connector, subclass "
                    "`SupportsHMA` defined in kv_connector/v1/base.py "
                    "(for MultiConnector, all child connectors must "
                    "support HMA)."
                )
        self.scheduler_config.disable_hybrid_kv_cache_manager = (
            need_disable_hybrid_kv_cache_manager
        )
    elif (
        self.scheduler_config.disable_hybrid_kv_cache_manager is False
        and need_disable_hybrid_kv_cache_manager
    ):
        raise ValueError(
            "Hybrid KV cache manager was explicitly enabled but is not "
            "supported in this configuration. Consider omitting the "
            "--no-disable-hybrid-kv-cache-manager flag to let vLLM decide"
            " automatically."
        )

    if self.scheduler_config.disable_hybrid_kv_cache_manager is None:
        # Default to enable HMA if not explicitly disabled by user or logic above.
        self.scheduler_config.disable_hybrid_kv_cache_manager = False

    if (
        self.attention_config.hisparse_config is not None
        and self.scheduler_config.disable_hybrid_kv_cache_manager
    ):
        raise ValueError(
            "HiSparse requires the hybrid KV cache manager; remove "
            "--disable-hybrid-kv-cache-manager or use connectors that "
            "support HMA."
        )

    if self.compilation_config.debug_dump_path:
        self.compilation_config.debug_dump_path = (
            self.compilation_config.debug_dump_path.absolute().expanduser()
        )
    if envs.VLLM_DEBUG_DUMP_PATH is not None:
        env_path = Path(envs.VLLM_DEBUG_DUMP_PATH).absolute().expanduser()
        if self.compilation_config.debug_dump_path:
            logger.warning(
                "Config-specified debug dump path is overridden"
                " by VLLM_DEBUG_DUMP_PATH to %s",
                env_path,
            )
        self.compilation_config.debug_dump_path = env_path

    # Enable quant_fp8 CUDA ops (TODO disable in follow up)
    # On H100 the CUDA kernel is faster than
    # native implementation
    # https://github.com/vllm-project/vllm/issues/25094
    if has_blocked_weights():
        custom_ops = self.compilation_config.custom_ops
        if "-quant_fp8" not in custom_ops:
            custom_ops.append("+quant_fp8")

    self._verify_kv_transfer_compat()
    if self.use_cumem_cudagraph_pool:
        # NCCL graph registration pins the offloaded pool; workers inherit this.
        value = os.environ.setdefault("NCCL_GRAPH_REGISTER", "0")
        if value != "0":
            logger.warning(
                "NCCL_GRAPH_REGISTER=%s pins the CUDA graph pool during sleep.",
                value,
            )
    # Log the custom passes that are enabled
    self.compilation_config.pass_config.log_enabled_passes()

_apply_optimization_level_defaults(defaults)

Apply optimization level defaults using self as root.

Recursively applies values from defaults into nested config objects. Only fields present in defaults are overwritten.

If the user configuration does not specify a value for a default field and if the default field is still None after all user selections are applied, then default values will be applied to the field. User specified fields will not be overridden by the default.

Parameters:

  • defaults

    (dict[str, Any]) –

    Dictionary of default values to apply.

Source code in vllm/config/vllm.py
def _apply_optimization_level_defaults(self, defaults: dict[str, Any]) -> None:
    """Apply optimization level defaults using self as root.

    Recursively applies values from defaults into nested config objects.
    Only fields present in defaults are overwritten.

    If the user configuration does not specify a value for a default field
    and if the default field is still None after all user selections are
    applied, then default values will be applied to the field. User specified
    fields will not be overridden by the default.

    Args:
        defaults: Dictionary of default values to apply.

    """

    def apply_recursive(config_obj: Any, config_defaults: dict[str, Any]) -> None:
        """Recursively apply defaults to config_obj, using self as root."""
        for key, value in config_defaults.items():
            if not hasattr(config_obj, key):
                continue

            current = getattr(config_obj, key)
            if isinstance(value, dict) and is_dataclass(current):
                apply_recursive(current, value)
            else:
                self._set_config_default(config_obj, key, value)

    apply_recursive(self, defaults)

_dflash_needs_multi_kv_group()

Whether a DFlash draft mixes sliding-window and full attention.

Source code in vllm/config/vllm.py
def _dflash_needs_multi_kv_group(self) -> bool:
    """Whether a DFlash draft mixes sliding-window and full attention."""
    spec = self.speculative_config
    if spec is None or spec.method != "dflash":
        return False
    draft_config = getattr(spec, "draft_model_config", None)
    if draft_config is None:
        return False
    layer_types = getattr(draft_config.hf_config, "layer_types", None) or []
    num_sliding = sum(lt == "sliding_attention" for lt in layer_types)
    return 0 < num_sliding < len(layer_types)

_disable_cudagraphs_for_v2_stock_torch_compile()

Run stock torch.compile without CUDA graphs in Model Runner V2.

V1 never wraps a stock-compiled model in CUDAGraphWrapper, so it runs without CUDA graphs. V2's CUDA graph manager would otherwise capture FULL graphs around the stock-compiled model, so disable them to match.

Source code in vllm/config/vllm.py
def _disable_cudagraphs_for_v2_stock_torch_compile(self) -> None:
    """Run stock torch.compile without CUDA graphs in Model Runner V2.

    V1 never wraps a stock-compiled model in CUDAGraphWrapper, so it runs
    without CUDA graphs. V2's CUDA graph manager would otherwise capture
    FULL graphs around the stock-compiled model, so disable them to match.
    """
    compilation_config = self.compilation_config
    if (
        compilation_config.mode != CompilationMode.STOCK_TORCH_COMPILE
        or compilation_config.cudagraph_mode == CUDAGraphMode.NONE
    ):
        return
    logger.info_once(
        "CUDA graphs are not supported with stock torch.compile in Model "
        "Runner V2. Overriding cudagraph_mode %s to NONE.",
        compilation_config.cudagraph_mode.name,
    )
    compilation_config.cudagraph_mode = CUDAGraphMode.NONE
    compilation_config.max_cudagraph_capture_size = 0
    compilation_config.cudagraph_capture_sizes = []

_get_dbo_unsupported_features()

Collect what the V2 model runner cannot combine with DBO.

The V2 runner microbatches a plain decoder forward pass. Anything that slices or replays the batch differently (drafting, adapters, pipeline stages, context parallelism, encoders) is not handled yet.

Source code in vllm/config/vllm.py
def _get_dbo_unsupported_features(self) -> list[str]:
    """Collect what the V2 model runner cannot combine with DBO.

    The V2 runner microbatches a plain decoder forward pass. Anything that
    slices or replays the batch differently (drafting, adapters, pipeline
    stages, context parallelism, encoders) is not handled yet.
    """
    # TODO: DBO with model runner V2 is under development.
    # It should be enabled with explicit VLLM_USE_V2_MODEL_RUNNER environ.
    # Remove it when stable.
    if envs.VLLM_USE_V2_MODEL_RUNNER is None:
        return ["dual batch overlap"]

    unsupported: list[str] = []
    model_config = self.model_config
    parallel_config = self.parallel_config

    if self.lora_config is not None:
        unsupported.append("dual batch overlap with LoRA")
    if self.speculative_config is not None:
        unsupported.append("dual batch overlap with speculative decoding")
    if parallel_config.pipeline_parallel_size > 1:
        unsupported.append("dual batch overlap with pipeline parallelism")
    if (
        parallel_config.decode_context_parallel_size > 1
        or parallel_config.prefill_context_parallel_size > 1
    ):
        unsupported.append("dual batch overlap with context parallelism")
    if model_config is not None and (
        model_config.is_multimodal_model or model_config.is_encoder_decoder
    ):
        unsupported.append("dual batch overlap with multimodal models")
    if model_config is not None and model_config.is_hybrid:
        unsupported.append("dual batch overlap with hybrid models")
    if self.compilation_config.cudagraph_mode == CUDAGraphMode.PIECEWISE:
        # DBO captures FULL graphs only.
        unsupported.append("dual batch overlap with PIECEWISE CUDA graphs")
    if self.is_mm_encoder_only:
        unsupported.append("dual batch overlap with encoder only models")

    return unsupported

_get_quantization_config(model_config, load_config) staticmethod

Get the quantization config.

Source code in vllm/config/vllm.py
@staticmethod
def _get_quantization_config(
    model_config: ModelConfig, load_config: LoadConfig
) -> QuantizationConfig | None:
    """Get the quantization config."""
    from vllm.platforms import current_platform

    if model_config.quantization is not None:
        from vllm.model_executor.model_loader.weight_utils import get_quant_config

        quant_config = get_quant_config(model_config, load_config)
        capability_tuple = current_platform.get_device_capability()

        if capability_tuple is not None:
            capability = capability_tuple.to_int()
            if capability < quant_config.get_min_capability():
                raise ValueError(
                    f"The quantization method {model_config.quantization} "
                    "is not supported for the current GPU. Minimum "
                    f"capability: {quant_config.get_min_capability()}. "
                    f"Current capability: {capability}."
                )
        supported_dtypes = quant_config.get_supported_act_dtypes()
        if model_config.dtype not in supported_dtypes:
            raise ValueError(
                f"{model_config.dtype} is not supported for quantization "
                f"method {model_config.quantization}. Supported dtypes: "
                f"{supported_dtypes}"
            )
        quant_config.maybe_update_config(
            model_config.model,
            hf_config=model_config.hf_config,
            revision=model_config.revision,
        )
        return quant_config
    return None

_get_v2_model_runner_unsupported_features()

Collect features not yet supported by the V2 model runner.

Source code in vllm/config/vllm.py
def _get_v2_model_runner_unsupported_features(self) -> list[str]:
    """Collect features not yet supported by the V2 model runner."""
    unsupported: list[str] = []
    speculative_config = self.speculative_config

    # V2 does not implement the external_launcher (torchrun) PP-output
    # broadcast that V1 uses to keep all ranks in sync (broadcast_pp_output).
    if (
        self.parallel_config.distributed_executor_backend == "external_launcher"
        and self.parallel_config.pipeline_parallel_size > 1
    ):
        unsupported.append("pipeline parallelism with external_launcher")

    if speculative_config is not None:
        if speculative_config.method in (
            "suffix",
            "medusa",
            "mlp_speculator",
            "custom_class",
        ):
            unsupported.append(f"speculative method '{speculative_config.method}'")

        # V2 EagleSpeculator does not support parallel_drafting (for P-Eagle).
        # DFlash and DSpark use parallel drafting natively in V2 via their
        # own speculators.
        if (
            speculative_config.parallel_drafting
            and speculative_config.method not in ("dflash", "dspark")
        ):
            unsupported.append("parallel drafting for EAGLE speculative decoding")

        # The V2 draft-model speculator has no token mapping between the
        # draft and target vocabularies (TLI is TBD in #47172).
        if getattr(speculative_config, "use_heterogeneous_vocab", False):
            unsupported.append("heterogeneous-vocabulary draft models")

    if self.parallel_config.use_ubatching:
        unsupported.extend(self._get_dbo_unsupported_features())

    return unsupported

_is_dflash_candidate_draft()

Whether the DFlash draft has a candidate head, by the architecture the speculator selects on (v1/worker/gpu/spec_decode/init.py).

Source code in vllm/config/vllm.py
def _is_dflash_candidate_draft(self) -> bool:
    """Whether the DFlash draft has a candidate head, by the architecture the
    speculator selects on (v1/worker/gpu/spec_decode/__init__.py)."""
    spec = self.speculative_config
    if spec is None or spec.method != "dflash":
        return False
    draft_config = getattr(spec, "draft_model_config", None)
    if draft_config is None:
        return False
    return bool(
        {"DFlash2DraftModel", "LiLiCorrDraftModel"}.intersection(
            draft_config.architectures or []
        )
    )

_post_init_kv_transfer_config()

Update KVTransferConfig based on top-level configs in VllmConfig.

Right now, this function reads the offloading settings from CacheConfig and configures the KVTransferConfig accordingly.

Source code in vllm/config/vllm.py
def _post_init_kv_transfer_config(self) -> None:
    """Update KVTransferConfig based on top-level configs in VllmConfig.

    Right now, this function reads the offloading settings from
    CacheConfig and configures the KVTransferConfig accordingly.
    """
    # KV offloading is only activated when kv_offloading_size is set.
    if (kv_offloading_size := self.cache_config.kv_offloading_size) is None:
        return

    kv_offloading_backend = self.cache_config.kv_offloading_backend

    # If no KVTransferConfig is provided, create a default one.
    if self.kv_transfer_config is None:
        self.kv_transfer_config = KVTransferConfig()

    if kv_offloading_backend == "native":
        if envs.VLLM_USE_SIMPLE_KV_OFFLOAD:
            config_connector = "SimpleCPUOffloadConnector"
        else:
            config_connector = "OffloadingConnector"
        self.kv_transfer_config.kv_connector = config_connector
        self.kv_transfer_config.kv_connector_extra_config.update(
            {"cpu_bytes_to_use": kv_offloading_size * (1 << 30)}
        )
    elif kv_offloading_backend == "lmcache":
        # Default to LMCache multi-process (MP) mode. The actual KV
        # storage capacity is managed by the standalone LMCache server
        # process, so ``kv_offloading_size`` is not propagated here.
        # ``LMCacheMPConnector`` falls back to ``tcp://localhost:5555``
        # when host/port are not provided via extra_config.
        self.kv_transfer_config.kv_connector = "LMCacheMPConnector"

    # This is the same for all backends
    self.kv_transfer_config.kv_role = "kv_both"

_resolve_and_verify_engram_config()

Resolve defaults and validate n-gram embedding settings.

Source code in vllm/config/vllm.py
def _resolve_and_verify_engram_config(self) -> None:
    """Resolve defaults and validate n-gram embedding settings."""
    model_config = self.model_config
    speculative_config = self.speculative_config
    # Draft configs inherit the target's communication groups and settings.
    # Validate the target because the draft may disable n-gram embeddings.
    if (
        speculative_config is not None
        and model_config is speculative_config.draft_model_config
    ):
        model_config = speculative_config.target_model_config
    if (
        model_config is not None
        and model_config.architecture == "DeepseekV41ForCausalLM"
        and getattr(model_config.hf_text_config, "engram_layer_ids", None)
        and self.parallel_config.use_ubatching
    ):
        raise ValueError(
            "DeepSeek V4.1 Engram does not support DBO or microbatching. "
            "Disable --enable-dbo and set --ubatch-size to 0."
        )
    if self.engram_config is None:
        if not model_has_engram_layers(model_config):
            return
        self.engram_config = EngramConfig()
    self.engram_config.verify_model_config(model_config)
    self.engram_config.resolve_dp_shared_memory(self.parallel_config)
    self.engram_config.verify_parallel_config(self.parallel_config)
    logger.info_once("Resolved Engram configuration: %s", str(self.engram_config))

_resolve_mm_embedding_inputs()

Accept embedding inputs, tensor optional, on disaggregated consumers.

An EC consumer loads embeddings from its connector. A KV consumer receives the prompt KV produced from those embeddings, so it does not need the tensors either. On every other deployment a missing tensor is a client error and must keep failing fast in the frontend.

Source code in vllm/config/vllm.py
def _resolve_mm_embedding_inputs(self) -> None:
    """Accept embedding inputs, tensor optional, on disaggregated consumers.

    An EC consumer loads embeddings from its connector. A KV consumer
    receives the prompt KV produced from those embeddings, so it does not
    need the tensors either. On every other deployment a missing tensor is
    a client error and must keep failing fast in the frontend.
    """
    model_config = self.model_config
    if model_config is None:
        return
    mm_config = model_config.multimodal_config
    if mm_config is None:
        return

    ec_config = self.ec_transfer_config
    kv_config = self.kv_transfer_config
    # Derived, so overwrite unconditionally rather than honouring a value
    # that was set by hand.
    mm_config.allow_missing_mm_embeddings = (
        ec_config is not None and ec_config.is_ec_consumer
    ) or (kv_config is not None and kv_config.is_kv_consumer)
    if not mm_config.allow_missing_mm_embeddings:
        return

    if not mm_config.enable_mm_embeds:
        # Allowing missing tensors still requires enabling embedding inputs
        # for the frontend to accept metadata-only requests.
        mm_config.enable_mm_embeds = True
        logger.info_once(
            "EC/KV consumer: accepting pre-computed-embedding inputs, "
            "which this role is sent by definition."
        )
    logger.info_once(
        "EC/KV consumer: pre-computed-embedding inputs may "
        "omit the embedding tensor."
    )

_resolve_mm_encoder_only()

Enable encoder-only mode for a dedicated EC producer.

Source code in vllm/config/vllm.py
def _resolve_mm_encoder_only(self) -> None:
    """Enable encoder-only mode for a dedicated EC producer."""
    ec_config = self.ec_transfer_config
    if ec_config is None or not ec_config.is_encode_only:
        return

    model_config = self.model_config
    mm_config = model_config.multimodal_config if model_config is not None else None
    if mm_config is None:
        raise ValueError(
            "An EC producer-only instance requires a multimodal model."
        )
    mm_config.mm_encoder_only = True

_resolve_mm_processor_device()

Settle --mm-processor-device=auto now that the EC role is known.

"auto" means "the accelerator, but only where the processor has it to itself and its output can be handed over without a copy back to host": an encode-only instance whose tensor transport carries device tensors. Every other deployment keeps the processor on CPU.

An explicit device -- from --mm-processor-device or straight from mm_processor_kwargs -- is already folded in by MultiModalConfig, so it is left alone here and validated by _validate_mm_processor_device.

Source code in vllm/config/vllm.py
def _resolve_mm_processor_device(self) -> None:
    """Settle `--mm-processor-device=auto` now that the EC role is known.

    "auto" means "the accelerator, but only where the processor has it to
    itself and its output can be handed over without a copy back to host":
    an encode-only instance whose tensor transport carries device tensors.
    Every other deployment keeps the processor on CPU.

    An explicit device -- from `--mm-processor-device` or straight from
    `mm_processor_kwargs` -- is already folded in by `MultiModalConfig`, so
    it is left alone here and validated by `_validate_mm_processor_device`.
    """
    model_config = self.model_config
    if model_config is None:
        return
    mm_config = model_config.multimodal_config
    if mm_config is None:
        return
    if mm_config.get_mm_processor_device_type() is not None:
        return

    from vllm.platforms import current_platform

    device_type = current_platform.device_type
    if device_type in ("", "cpu"):
        return

    ec_config = self.ec_transfer_config
    # An EC producer that is not also a consumer runs no forward pass and
    # allocates no KV cache, so frontend accelerator work has the device to
    # itself.
    if ec_config is None or not ec_config.is_encode_only:
        return

    if mm_config.mm_tensor_ipc != "torch_shm":
        # Any other transport serializes host bytes, so the output would be
        # copied back, and that copy costs more than running the transform
        # on device saves.
        logger.info_once(
            "EPD encoder instance: keeping the multi-modal processor on CPU "
            "because mm_tensor_ipc=%s cannot carry device tensors. Add "
            "--mm-tensor-ipc=torch_shm to run it on the accelerator.",
            mm_config.mm_tensor_ipc,
        )
        return

    mm_config.mm_processor_kwargs = {
        **(mm_config.mm_processor_kwargs or {}),
        "device": device_type,
    }
    logger.info_once(
        "EPD encoder instance: running the multi-modal processor on %s. "
        "Override with --mm-processor-device=cpu.",
        device_type,
    )

_resolve_mm_video_decode_device()

Default video decoding to torchcodec GPU backend for EPD encoder-only instance if the mm processor runs on CUDA.

The processor consumes the decoded frames on-device in that case, so keeping the frames on the GPU skips the host round-trip through the CPU media path. An explicit codec/backend choice in --media-io-kwargs is left alone, and the default is skipped where torchcodec (or its FFmpeg runtime) is unavailable.

Source code in vllm/config/vllm.py
def _resolve_mm_video_decode_device(self) -> None:
    """Default video decoding to torchcodec GPU backend for EPD encoder-only
    instance if the mm processor runs on CUDA.

    The processor consumes the decoded frames on-device in that case, so
    keeping the frames on the GPU skips the host round-trip through the
    CPU media path. An explicit codec/backend choice in
    `--media-io-kwargs` is left alone, and the default is skipped where
    torchcodec (or its FFmpeg runtime) is unavailable.
    """
    if self.model_config is None or self.model_config.multimodal_config is None:
        return
    mm_config = self.model_config.multimodal_config

    ec_config = self.ec_transfer_config
    # An EC producer that is not also a consumer runs no forward pass and
    # allocates no KV cache, so frontend accelerator work has the device to
    # itself.
    if ec_config is None or not ec_config.is_encode_only:
        return

    from vllm.platforms import current_platform

    device_type = current_platform.device_type
    if (
        device_type != "cuda"
        or mm_config.get_mm_processor_device_type() != device_type
    ):
        return

    # User set video backend or device explicitly
    video_kwargs = mm_config.media_io_kwargs.setdefault("video", {})
    if "backend" in video_kwargs or "device" in video_kwargs:
        return

    from vllm.utils.import_utils import check_torchcodec_available

    try:
        check_torchcodec_available()
    except (ImportError, RuntimeError):
        # torchcodec is not installed, or is installed without a usable
        # FFmpeg runtime (it raises rather than returning False).
        logger.info_once(
            "EPD encoder instance: keeping CPU video decoding because "
            "torchcodec is not available (needs a CUDA build with FFmpeg)."
        )
        return

    video_kwargs["backend"] = "torchcodec"
    video_kwargs["device"] = device_type
    logger.info_once(
        "EPD encoder instance: decoding video with NVDEC (torchcodec device=%s).",
        device_type,
    )

_set_compile_ranges()

Set the compile ranges for the compilation config.

Source code in vllm/config/vllm.py
def _set_compile_ranges(self):
    """Set the compile ranges for the compilation config."""
    compilation_config = self.compilation_config
    computed_compile_ranges_endpoints = []

    # The upper bound of the compile ranges is the max_num_batched_tokens.
    compile_range_end = self.scheduler_config.max_num_batched_tokens
    if compile_range_end is not None:
        computed_compile_ranges_endpoints.append(compile_range_end)

    # Add the compile ranges for flashinfer/aiter.
    if compilation_config.pass_config.fuse_allreduce_rms:
        tp_size = self.parallel_config.tensor_parallel_size
        from vllm._aiter_ops import rocm_aiter_ops

        max_size: int | None = None
        if rocm_aiter_ops.is_custom_all_reduce_enabled():
            from vllm.distributed.device_communicators.aiter_custom_all_reduce import (  # noqa: E501
                AiterCustomAllreduce,
            )

            max_size = AiterCustomAllreduce.effective_max_size()
        else:
            max_size = compilation_config.pass_config.flashinfer_max_size(tp_size)
        if max_size is not None and self.model_config is not None:
            assert isinstance(self.model_config.dtype, torch.dtype)
            max_token_num = max_size // (
                self.model_config.get_hidden_size()
                * self.model_config.dtype.itemsize
            )
            if compile_range_end is not None and max_token_num < compile_range_end:
                computed_compile_ranges_endpoints.append(max_token_num)
            else:
                logger.debug(
                    "Max num batched tokens below allreduce-rms fusion threshold, "
                    "allreduce-rms fusion will be enabled for all num_tokens."
                )

    if compilation_config.pass_config.fuse_rope_kvcache:
        max_token_num = (
            compilation_config.pass_config.rope_kvcache_fusion_max_token_num
        )
        if max_token_num is not None:
            if compile_range_end is not None and max_token_num < compile_range_end:
                computed_compile_ranges_endpoints.append(max_token_num)
            else:
                logger.debug(
                    "Max num batched tokens below rope+kvcache fusion threshold, "
                    "rope+kvcache fusion enabled for num_tokens <= %d.",
                    compile_range_end,
                )

    if compilation_config.pass_config.fuse_qk_norm_rope_kvcache:
        max_token_num = (
            compilation_config.pass_config.rope_kvcache_fusion_max_token_num
        )
        if max_token_num is not None:
            if compile_range_end is not None and max_token_num < compile_range_end:
                computed_compile_ranges_endpoints.append(max_token_num)
            else:
                logger.debug(
                    "Max num batched tokens below qk_norm+rope+kvcache "
                    "fusion threshold, fusion enabled for "
                    "num_tokens <= %d.",
                    compile_range_end,
                )

    if compilation_config.compile_ranges_endpoints is not None:
        for x in compilation_config.compile_ranges_endpoints:
            assert isinstance(x, int)
            assert x > 0, f"Invalid compile range endpoint: {x}"
            if compile_range_end is not None and x < compile_range_end and x > 1:
                computed_compile_ranges_endpoints.append(x)
    compilation_config.compile_ranges_endpoints = sorted(
        computed_compile_ranges_endpoints
    )

_set_config_default(config_obj, key, value)

Set config attribute to default if not already set by user.

Parameters:

  • config_obj

    (Any) –

    Configuration object to update.

  • key

    (str) –

    Attribute name.

  • value

    (Any) –

    Default value (static or callable).

Source code in vllm/config/vllm.py
def _set_config_default(self, config_obj: Any, key: str, value: Any) -> None:
    """Set config attribute to default if not already set by user.

    Args:
        config_obj: Configuration object to update.
        key: Attribute name.
        value: Default value (static or callable).

    """
    if getattr(config_obj, key) is None:
        # Some config values are known before initialization and are
        # hard coded.
        # Other values depend on the user given configuration, so they are
        # implemented with lambda functions and decided at run time.
        setattr(config_obj, key, value(self) if callable(value) else value)

_set_cudagraph_sizes()

VLLM defines the default candidate list of batch sizes for CUDA graph capture as:

```python default_max_graph_size = 1024 if is_data_center_blackwell else 512 decode_query_len = self.uniform_decode_query_len max_graph_size = min( max_num_seqs * decode_query_len * 2, default_max_graph_size )

1, 2, 4, then multiples of 8 up to 256 and then multiples of 16

up to max_graph_size

cudagraph_capture_sizes = [1, 2, 4] + list(range(8, 256, 8)) + list( range(256, max_graph_size + 1, 16))

max_num_batched_tokens is also appended to the list if it fits within max_cudagraph_capture_size, so the max batch size is captured even when off-stride. Uniform decode sizes are appended when they fit within the platform's default capture ceiling, since they need not land on an 8- or 16-token stride.

In the end, vllm_config.compilation_config.cudagraph_capture_sizes will be the final sizes to capture cudagraph (in ascending order).

These sizes are used to capture and reuse CUDA graphs for performance-critical paths (e.g., decoding). Capturing enables significantly faster kernel dispatch by avoiding Python overhead. The list is then filtered based on max_num_batched_tokens (e.g., 8192 on most GPUs), which controls the total allowed number of tokens in a batch. Since each sequence may have a variable number of tokens, the maximum usable batch size will depend on actual sequence lengths.

Example: With max_num_batched_tokens = 8192, and typical sequences averaging ~32 tokens, most practical batch sizes fall below 256. However, the system will still allow capture sizes up to the platform default if shape and memory permit.

Note: If users explicitly specify cudagraph capture sizes in the compilation config, those will override this default logic. At runtime:

- If batch size <= one of the `cudagraph_capture_sizes`, the closest
padded CUDA graph will be used.
- If batch size > largest `cudagraph_capture_sizes`, cudagraph will
not be used.
Source code in vllm/config/vllm.py
def _set_cudagraph_sizes(self):
    """VLLM defines the default candidate list of batch sizes for CUDA graph
    capture as:

    ```python
    default_max_graph_size = 1024 if is_data_center_blackwell else 512
    decode_query_len = self.uniform_decode_query_len
    max_graph_size = min(
        max_num_seqs * decode_query_len * 2, default_max_graph_size
    )
    # 1, 2, 4, then multiples of 8 up to 256 and then multiples of 16
    # up to max_graph_size
    cudagraph_capture_sizes = [1, 2, 4] + list(range(8, 256, 8)) + list(
        range(256, max_graph_size + 1, 16))

    `max_num_batched_tokens` is also appended to the list if it fits
    within `max_cudagraph_capture_size`, so the max batch size is captured
    even when off-stride. Uniform decode sizes are appended when they fit
    within the platform's default capture ceiling, since they need not land
    on an 8- or 16-token stride.

    In the end, `vllm_config.compilation_config.cudagraph_capture_sizes`
    will be the final sizes to capture cudagraph (in ascending order).

    These sizes are used to capture and reuse CUDA graphs for
    performance-critical paths (e.g., decoding). Capturing enables
    significantly faster kernel dispatch by avoiding Python overhead. The
    list is then filtered based on `max_num_batched_tokens` (e.g., 8192 on
    most GPUs), which controls the total allowed number of tokens in a
    batch. Since each sequence may have a variable number of tokens, the
    maximum usable batch size will depend on actual sequence lengths.

    Example:
        With `max_num_batched_tokens = 8192`, and typical sequences
        averaging ~32 tokens, most practical batch sizes fall below 256.
        However, the system will still allow capture sizes up to the
        platform default if shape and memory permit.

    Note:
        If users explicitly specify cudagraph capture sizes in the
        compilation config, those will override this default logic.
        At runtime:

        - If batch size <= one of the `cudagraph_capture_sizes`, the closest
        padded CUDA graph will be used.
        - If batch size > largest `cudagraph_capture_sizes`, cudagraph will
        not be used.

    """
    if (
        self.model_config is not None
        and not self.model_config.enforce_eager
        and self.compilation_config.cudagraph_mode != CUDAGraphMode.NONE
    ):
        # determine the initial max_cudagraph_capture_size
        max_cudagraph_capture_size = (
            self.compilation_config.max_cudagraph_capture_size
        )
        # Decode sizes to cover, in tokens. Populated only when the default
        # is computed here, so an explicit capture range is left exactly as
        # configured.
        uniform_decode_sizes: list[int] = []
        if max_cudagraph_capture_size is None:
            from vllm.platforms import current_platform

            default_max_graph_size = (
                1024 if current_platform.is_device_capability_family(100) else 512
            )
            decode_query_len = self.uniform_decode_query_len
            max_num_seqs = self.scheduler_config.max_num_seqs
            max_cudagraph_capture_size = min(
                max_num_seqs * decode_query_len * 2, default_max_graph_size
            )
            if decode_query_len > 1:
                # A uniform decode batch is decode_query_len tokens per
                # request, so the widest one is far outside this ceiling.
                # Coverage comes from appending the decode sizes rather than
                # extending the token-strided grid. Extending that grid to
                # the widest decode size produces 581 sizes at
                # max_num_seqs=512 and 16 draft tokens, versus 100 with the
                # request-count grid.
                #
                # The grid would not buy decode coverage anyway. Dispatch
                # requires an exact multiple of decode_query_len, so a
                # token-strided entry is only usable when it happens to be
                # one; at query length 17 a captured 560 rounds to 561 and
                # is rejected. Scaling a request-count grid keeps every
                # entry usable and the count comparable to the non-
                # speculative case.
                def request_counts(max_reqs: int) -> list[int]:
                    # At most the platform default number of requests,
                    # mirroring the one-token-per-request decode ceiling.
                    max_reqs = min(max_reqs, default_max_graph_size)
                    counts = [n for n in (1, 2, 4) if n <= max_reqs]
                    counts += list(range(8, min(max_reqs + 1, 256), 8))
                    counts += list(range(256, max_reqs + 1, 16))
                    return sorted(set(counts + [max_reqs]))

                # Dynamic speculative decoding picks the draft width from
                # the batch size, so a decode step is only uniform within a
                # tier and each tier needs its own sizes. Scaling by the
                # widest one alone leaves the narrower tiers short: the
                # manager rounds a capture size up to a multiple of the
                # tier's query length and drops it once the implied request
                # count exceeds max_num_seqs, so at query length 3 sizes
                # built from 17 stop covering at 227 of 256 requests.
                decode_tiers = [(decode_query_len, max_num_seqs)]
                speculative_config = self.speculative_config
                if (
                    speculative_config is not None
                    and speculative_config.uses_dynamic_speculative_decoding()
                ):
                    from vllm.v1.spec_decode.dynamic.utils import (
                        build_dynamic_sd_schedule_lookup,
                    )

                    schedule = (
                        speculative_config.num_speculative_tokens_per_batch_size
                    )
                    assert schedule is not None
                    # Read the tiers off the dense lookup the scheduler
                    # runs on, so the clamp against num_speculative_tokens
                    # and the carry-forward through gaps and the tail
                    # cannot drift from it. Validation lives elsewhere; an
                    # invalid schedule keeps the single-tier default.
                    try:
                        dense_schedule = build_dynamic_sd_schedule_lookup(
                            schedule,
                            vllm_max_batch_size=max_num_seqs,
                            vllm_num_speculative_tokens=self.num_speculative_tokens,
                        )
                    except ValueError:
                        pass
                    else:
                        # Ascending batch size, so the last write per
                        # query length is the widest batch running at it.
                        widest_batch: dict[int, int] = {}
                        for batch_size, num_spec in enumerate(
                            dense_schedule[1:], start=1
                        ):
                            widest_batch[num_spec + 1] = batch_size
                        decode_tiers = list(widest_batch.items())

                uniform_decode_sizes = sorted(
                    {
                        n * query_len
                        for query_len, tier_max_reqs in decode_tiers
                        for n in request_counts(tier_max_reqs)
                        if n * query_len <= max_cudagraph_capture_size
                    }
                )
            elif max_num_seqs <= max_cudagraph_capture_size:
                uniform_decode_sizes = [max_num_seqs]
        max_num_tokens = self.scheduler_config.max_num_batched_tokens
        max_cudagraph_capture_size = min(max_num_tokens, max_cudagraph_capture_size)

        assert max_cudagraph_capture_size >= 1, (
            "Maximum cudagraph size should be greater than or equal to 1 "
            "when using cuda graph."
        )

        # determine the cudagraph_capture_sizes
        if self.compilation_config.cudagraph_capture_sizes is not None:
            assert len(self.compilation_config.cudagraph_capture_sizes) > 0, (
                "cudagraph_capture_sizes should contain at least one element "
                "when using cuda graph."
            )
            # de-duplicate the sizes provided by the config
            dedup_sizes = list(set(self.compilation_config.cudagraph_capture_sizes))
            cudagraph_capture_sizes = [
                i for i in dedup_sizes if i <= max_num_tokens
            ]
            # sort to make sure the sizes are in ascending order
            cudagraph_capture_sizes.sort()
        else:
            if self.performance_mode == "interactivity":
                # Fine-grained CUDA graphs at small batch sizes
                # for minimal padding overhead
                interactivity_max = min(max_cudagraph_capture_size, 32)
                cudagraph_capture_sizes = list(range(1, interactivity_max + 1))
            else:
                cudagraph_capture_sizes = [
                    i for i in [1, 2, 4] if i <= max_cudagraph_capture_size
                ]
            if max_cudagraph_capture_size >= 8:
                # Step size 8 for small batch sizes, up to 256(not included)
                cudagraph_capture_sizes += list(
                    range(8, min(max_cudagraph_capture_size + 1, 256), 8)
                )
            if max_cudagraph_capture_size >= 256:
                # Step size 16 for larger batch sizes
                cudagraph_capture_sizes += list(
                    range(256, max_cudagraph_capture_size + 1, 16)
                )
            # ensure max_num_tokens is captured if within max capture size
            if (
                max_num_tokens <= max_cudagraph_capture_size
                and max_num_tokens not in cudagraph_capture_sizes
            ):
                cudagraph_capture_sizes.append(max_num_tokens)
            # Preserve the platform's default capture ceiling. Larger
            # uniform decode batches fall back to eager execution unless
            # users explicitly configure wider capture sizes.
            cudagraph_capture_sizes += [
                size for size in uniform_decode_sizes if size <= max_num_tokens
            ]
            # de-duplicate and sort the sizes
            cudagraph_capture_sizes = sorted(set(cudagraph_capture_sizes))

        # user-specific compilation_config.max_cudagraph_capture_size get
        # truncated to valid_max_size when they are inconsistent.
        valid_max_size = (
            cudagraph_capture_sizes[-1] if cudagraph_capture_sizes else 0
        )
        if (
            self.compilation_config.max_cudagraph_capture_size is not None
            and self.compilation_config.max_cudagraph_capture_size != valid_max_size
        ):
            # raise error only when both two flags are user-specified
            # and they are inconsistent with each other
            if self.compilation_config.cudagraph_capture_sizes is not None:
                raise ValueError(
                    "customized max_cudagraph_capture_size"
                    f"(={self.compilation_config.max_cudagraph_capture_size}) "
                    "should be consistent with the max value of "
                    f"cudagraph_capture_sizes(={valid_max_size})"
                )

            logger.warning(
                "Truncating max_cudagraph_capture_size to %d",
                valid_max_size,
            )
        # always set the final max_cudagraph_capture_size
        self.compilation_config.max_cudagraph_capture_size = valid_max_size

        if self.compilation_config.cudagraph_capture_sizes is not None and len(
            cudagraph_capture_sizes
        ) < len(self.compilation_config.cudagraph_capture_sizes):
            # If users have specified capture sizes, we only need to
            # compare the lens before and after modification since the modified
            # list is only the subset of the original list.
            logger.warning(
                (
                    "cudagraph_capture_sizes specified in compilation_config"
                    " %s is overridden by config %s"
                ),
                self.compilation_config.cudagraph_capture_sizes,
                cudagraph_capture_sizes,
            )
        # always write back the final sizes
        self.compilation_config.cudagraph_capture_sizes = cudagraph_capture_sizes

    else:
        # no cudagraph in use
        self.compilation_config.max_cudagraph_capture_size = 0
        self.compilation_config.cudagraph_capture_sizes = []

    # complete the remaining process.
    self.compilation_config.post_init_cudagraph_sizes()

_set_max_num_scheduled_tokens()

In most cases, the scheduler may schedule a batch with as many tokens as the worker is configured to handle.

Source code in vllm/config/vllm.py
def _set_max_num_scheduled_tokens(self):
    """In most cases, the scheduler may schedule a batch with as many tokens as the
    worker is configured to handle.
    """
    if self.speculative_config is not None:
        scheduled_token_delta = (
            self.speculative_config.max_num_new_slots_for_drafting
        )
        max_num_batched_tokens = self.scheduler_config.max_num_batched_tokens
        if self.scheduler_config.max_num_scheduled_tokens is None:
            self.scheduler_config.max_num_scheduled_tokens = max_num_batched_tokens

        if self.scheduler_config.max_num_scheduled_tokens <= 0:
            raise ValueError(
                "max_num_scheduled_tokens is set to"
                f" {self.scheduler_config.max_num_scheduled_tokens} based on"
                " the speculative decoding settings, which does not allow"
                " any tokens to be scheduled. Increase max_num_batched_tokens"
                " to accommodate the additional draft token slots, or decrease"
                " num_speculative_tokens."
            )
        if self.scheduler_config.max_num_scheduled_tokens < 8192:
            logger.warning_once(
                "max_num_scheduled_tokens is set to"
                f" {self.scheduler_config.max_num_scheduled_tokens} based on"
                " the speculative decoding settings. This may lead to suboptimal"
                " performance. Consider increasing max_num_batched_tokens to"
                " accommodate the additional draft token slots, or decrease"
                " num_speculative_tokens.",
            )

        if max_num_batched_tokens <= scheduled_token_delta:
            raise ValueError(
                "VllmConfig does not have enough slots to schedule a token and"
                " support the speculative decoding settings."
                f" Got {max_num_batched_tokens=} and {scheduled_token_delta=}."
            )

_uses_breakable_cudagraph_for_batch_invariance()

Avoid freezing runtime-M tile lookup in compiled forward (#54243). Breakable graphs look up tuned bf16, unquantized qkv/o/gate_up/down tiles at capture; lm_head runs outside compiled forward and does not benefit.

Source code in vllm/config/vllm.py
def _uses_breakable_cudagraph_for_batch_invariance(self) -> bool:
    """Avoid freezing runtime-M tile lookup in compiled forward (#54243).
    Breakable graphs look up tuned bf16, unquantized qkv/o/gate_up/down tiles
    at capture; lm_head runs outside compiled forward and does not benefit."""
    from vllm.model_executor.determinism import batch_invariant_configs as bi
    from vllm.platforms import current_platform

    model = self.model_config
    if (
        not envs.VLLM_BATCH_INVARIANT
        or model is None
        or model.enforce_eager
        or model.dtype != torch.bfloat16
        or model.quantization is not None
        or not current_platform.is_cuda()
    ):
        return False
    family = bi._get_tuned_matmul_arch_family(
        current_platform.get_device_capability()
    )
    if family is None or family not in bi._BATCH_INVARIANT_MATMUL_TUNED_CONFIGS:
        return False
    table = bi._BATCH_INVARIANT_MATMUL_TUNED_CONFIGS[family]
    parallel = self.parallel_config
    tp = parallel.tensor_parallel_size
    hidden = model.get_hidden_size()
    head = model.get_head_size()
    heads = model.get_num_attention_heads(parallel)
    kv_heads = model.get_num_kv_heads(parallel)
    shapes = [((heads + 2 * kv_heads) * head, hidden), (hidden, heads * head)]
    intermediate = getattr(model.hf_text_config, "intermediate_size", None)
    # Per-layer sizes (e.g. Gemma3n) are not modeled.
    if isinstance(intermediate, int):
        shapes += [(2 * intermediate // tp, hidden), (hidden, intermediate // tp)]
    return any(shape in table for shape in shapes)

_validate_batch_sharded_sampling()

Validate enable_batch_sharded_sampling against the rest of the config.

Source code in vllm/config/vllm.py
def _validate_batch_sharded_sampling(self) -> None:
    """Validate `enable_batch_sharded_sampling` against the rest of the config."""
    if not self.parallel_config.enable_batch_sharded_sampling:
        # Default to False if not set.
        self.parallel_config.enable_batch_sharded_sampling = False
        return

    blockers: list[str] = []
    tp_size = self.parallel_config.tensor_parallel_size

    if tp_size <= 1:
        blockers.append("tensor_parallel_size is 1, so there is nothing to shard")
    elif self.scheduler_config.max_num_seqs < tp_size:
        # Requests are assigned to ranks whole, so fewer slots than ranks
        # leaves some ranks without work in every step.
        blockers.append(
            f"max_num_seqs ({self.scheduler_config.max_num_seqs}) is below "
            f"tensor_parallel_size ({tp_size})"
        )

    if self.model_config is not None and self.model_config.max_logprobs < 0:
        # max_logprobs == -1 allows vocab-size logprob requests, which the
        # fixed-width logprobs gather cannot reasonably size for.
        blockers.append("max_logprobs is -1, allowing vocab-size logprob requests")

    if self.model_config is not None and self.model_config.return_sampling_mask:
        # gather_sampler_output() drops SamplingMaskTensors: masks come back None.
        blockers.append(
            "return_sampling_mask is set and the batch-sharded gather does "
            "not forward sampling masks"
        )

    if (
        self.speculative_config is not None
        and self.speculative_config.enable_adaptive_verification
    ):
        # Adaptive verification picks the per-request draft split on the GPU,
        # so cu_num_logits_np is only an upper bound, while the shard plan is
        # built from that CPU array. The two disagree once the budget binds.
        # TODO(TheEpicDolphin): Support adaptive verification with batch-sharded
        # sampling.
        blockers.append(
            "it does not yet work with adaptive verification, which decides "
            "the per-request logits counts on the GPU, where the CPU-side "
            "shard plan cannot see them"
        )

    if blockers:
        raise ValueError(
            "Batch-sharded sampling was explicitly enabled via "
            "the --enable-batch-sharded-sampling flag, but is not supported "
            "in this configuration for the following reason(s): "
            f"{'; '.join(blockers)}."
        )

_validate_mm_processor_device()

Hand the EC config to MultiModalConfig, which owns the rule.

Source code in vllm/config/vllm.py
def _validate_mm_processor_device(self) -> None:
    """Hand the EC config to `MultiModalConfig`, which owns the rule."""
    model_config = self.model_config
    if model_config is None:
        return
    mm_config = model_config.multimodal_config
    if mm_config is None:
        return

    mm_config.validate_mm_processor_device(self.ec_transfer_config)

_validate_v2_model_runner()

Check for features not yet supported by the V2 model runner.

Source code in vllm/config/vllm.py
def _validate_v2_model_runner(self) -> None:
    """Check for features not yet supported by the V2 model runner."""
    if not HAS_TRITON:
        raise ValueError("Model Runner V2 requires Triton.")

    unsupported = self._get_v2_model_runner_unsupported_features()
    if unsupported:
        raise ValueError(
            f"Model Runner V2 does not yet support: {', '.join(unsupported)}"
        )

_verify_aux_output_compatibility()

Reject configurations unsupported by enabled auxiliary outputs.

Source code in vllm/config/vllm.py
def _verify_aux_output_compatibility(self) -> None:
    """Reject configurations unsupported by enabled auxiliary outputs."""
    if not self.aux_output_config.enabled:
        return
    from vllm.platforms import current_platform

    # In-tree platforms only wire AuxOutput to MRV2. TPU and out-of-tree
    # platforms bring their own model runners and validate AuxOutput
    # support themselves.
    if not self.use_v2_model_runner and not (
        current_platform.is_tpu() or current_platform.is_out_of_tree()
    ):
        raise ValueError(
            "AuxOutput Connector requires Model Runner V2; set "
            "VLLM_USE_V2_MODEL_RUNNER=1."
        )
    if self.model_config.runner_type != "generate":
        raise ValueError("AuxOutput Connector only supports generate runners.")
    if not self.model_config.is_moe:
        raise ValueError("AuxOutput Connector only supports MoE models.")
    if not self.cache_config.enable_prefix_caching:
        raise ValueError("AuxOutput Connector requires prefix caching.")
    if (
        self.speculative_config is not None
        and self.speculative_config.enable_adaptive_verification
    ):
        raise ValueError(
            "--enable-return-routed-experts is incompatible with "
            "adaptive speculative verification."
        )
    if self.parallel_config.pipeline_parallel_size > 1:
        raise ValueError(
            "--enable-return-routed-experts is incompatible with "
            "pipeline parallelism (PP > 1)."
        )
    if (
        self.parallel_config.decode_context_parallel_size > 1
        or self.parallel_config.prefill_context_parallel_size > 1
    ):
        raise ValueError(
            "--enable-return-routed-experts is incompatible with "
            "context parallelism (DCP/PCP > 1)."
        )

    kv_transfer_config = self.kv_transfer_config
    if kv_transfer_config is not None:
        for connector_name in (
            "NixlConnector",
            "NixlPullConnector",
            "NixlPushConnector",
            "MoRIIOConnector",
            "MooncakeConnector",
        ):
            if kv_transfer_config.has_connector(connector_name):
                raise ValueError(
                    "--enable-return-routed-experts is incompatible with "
                    f"{connector_name}; PD auxiliary output is not supported."
                )

_verify_kv_transfer_compat()

Reject configurations that silently corrupt KV transfers.

Source code in vllm/config/vllm.py
def _verify_kv_transfer_compat(self) -> None:
    """Reject configurations that silently corrupt KV transfers."""
    if (
        self.kv_transfer_config is None
        or self.kv_transfer_config.kv_connector is None
    ):
        return

    # PyTorch's expandable_segments allocator uses CUDA VMM, which can
    # remap a virtual address range to different physical pages over the
    # engine's lifetime. KV connectors that pin KV cache memory (e.g.
    # NixlConnector via ibv_reg_mr, MooncakeConnector) end up with their
    # registrations pointing at stale physical pages after any remap,
    # producing RDMA failures like IBV_WC_REM_ACCESS_ERR /
    # NIXL_ERR_REMOTE_DISCONNECT at the first inter-node KV transfer.
    # We can't enumerate every in-tree and out-of-tree connector that
    # pins memory, so we conservatively reject the combination whenever
    # any KV connector is configured.
    #
    # CuMem allocator is exempt: CuMemAllocator.use_memory_pool toggles
    # expandable_segments off around its pool (see #40812), so the KV
    # cache allocated within that context lands on stable physical pages
    # even when the env var is set.
    if "expandable_segments:True" not in os.environ.get(
        "PYTORCH_CUDA_ALLOC_CONF", ""
    ):
        return
    if self.model_config is not None and (self.model_config.enable_cumem_allocator):
        return

    raise ValueError(
        f"KV connector {self.kv_transfer_config.kv_connector} is "
        "incompatible with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True "
        "unless enable_cumem_allocator is also enabled. PyTorch's CUDA VMM "
        "allocator can remap KV cache virtual addresses to different "
        "physical pages, invalidating any pinned/registered KV memory "
        "(e.g. IB memory regions registered by NIXL or Mooncake). Either "
        "unset expandable_segments:True or enable the cumem allocator "
        "(sleep mode does this automatically and also "
        "routes KV allocations through CuMemAllocator's pool, where "
        "expandable_segments is automatically disabled)."
    )

adjust_dcp_kv_cache_interleave_size(kv_cache_config)

Normalize DCP interleave size against block_size for NIXL P/D.

Called by each worker (via ensure_kv_transfer_initialized), once it knows its own final block_size via kv_cache_config.

Source code in vllm/config/vllm.py
def adjust_dcp_kv_cache_interleave_size(
    self, kv_cache_config: "KVCacheConfig"
) -> None:
    """Normalize DCP interleave size against block_size for NIXL P/D.

    Called by each worker (via ensure_kv_transfer_initialized), once it knows its
    own final block_size via kv_cache_config.
    """
    dcp_size = self.parallel_config.decode_context_parallel_size
    if dcp_size <= 1:
        return
    if self.parallel_config.dcp_kv_cache_interleave_size > 1 and (
        self.parallel_config.cp_kv_cache_interleave_size
        != self.parallel_config.dcp_kv_cache_interleave_size
    ):
        self.parallel_config.cp_kv_cache_interleave_size = (
            self.parallel_config.dcp_kv_cache_interleave_size
        )
        logger.warning_once(
            "cp_kv_cache_interleave_size is overridden by dcp_kv_cache"
            "_interleave_size. And dcp-kv-cache-interleave-size will be "
            "deprecated when PCP is fully supported."
        )

    if self.kv_transfer_config is None or not self.kv_transfer_config.has_connector(
        "NixlConnector"
    ):
        return
    if not self.parallel_config._allow_auto_resolve_cp_interleave_size:
        return

    # Get the kernel block_size, but don't use resolve_kv_cache_block_size to avoid
    # scaling by dcp_size (we need the local block_size here).
    local_block_size = min(
        g.kv_cache_spec.block_size for g in kv_cache_config.kv_cache_groups
    )
    if self.parallel_config.cp_kv_cache_interleave_size != local_block_size:
        interleave = self.parallel_config.cp_kv_cache_interleave_size
        self.parallel_config.cp_kv_cache_interleave_size = local_block_size
        logger.info_once(
            "When using PD disaggregation with DCP "
            "(decode_context_parallel_size=%d), "
            "cp_kv_cache_interleave_size is automatically adjusted "
            "from %d to block_size %d for block-level alignment.",
            dcp_size,
            interleave,
            local_block_size,
        )

compile_debug_dump_path()

Returns a rank-aware path for dumping torch.compile debug information.

Source code in vllm/config/vllm.py
def compile_debug_dump_path(self) -> Path | None:
    """Returns a rank-aware path for dumping
    torch.compile debug information.
    """
    if self.compilation_config.debug_dump_path is None:
        return None
    tp_rank = self.parallel_config.rank
    dp_rank = self.parallel_config.data_parallel_index
    append_path = f"rank_{tp_rank}_dp_{dp_rank}"
    path = self.compilation_config.debug_dump_path / append_path
    return path

compute_hash(include_version=True)

WARNING: Whenever a new field is added to this config, ensure that it is included in the factors list if it affects the computation graph.

Provide a hash that uniquely identifies all the configs that affect the structure of the computation graph from input ids/embeddings to the final hidden states, excluding anything before input ids/embeddings and after the final hidden states.

Parameters:

  • include_version

    (bool, default: True ) –

    Include the vLLM version in the hash.

Source code in vllm/config/vllm.py
def compute_hash(self, include_version: bool = True) -> str:
    """WARNING: Whenever a new field is added to this config,
    ensure that it is included in the factors list if
    it affects the computation graph.

    Provide a hash that uniquely identifies all the configs
    that affect the structure of the computation
    graph from input ids/embeddings to the final hidden states,
    excluding anything before input ids/embeddings and after
    the final hidden states.

    Args:
        include_version: Include the vLLM version in the hash.

    """
    factors: list[Any] = []

    # summarize vllm config
    vllm_factors: list[Any] = []
    if include_version:
        from vllm import __version__

        vllm_factors.append(__version__)
    if self.model_config:
        vllm_factors.append(self.model_config.compute_hash())
        if (
            self.compilation_config
            and getattr(self.compilation_config, "compile_mm_encoder", False)
            and self.model_config.multimodal_config
        ):
            vllm_factors.append(self.model_config.multimodal_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.cache_config:
        vllm_factors.append(self.cache_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.parallel_config:
        vllm_factors.append(self.parallel_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.scheduler_config:
        vllm_factors.append(self.scheduler_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.device_config:
        vllm_factors.append(self.device_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.load_config:
        vllm_factors.append(self.load_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.offload_config:
        vllm_factors.append(self.offload_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.attention_config:
        vllm_factors.append(self.attention_config.compute_hash())
    else:
        vllm_factors.append("None")
    vllm_factors.append(
        self.engram_config.compute_hash()
        if self.engram_config is not None
        else "None"
    )
    if self.lora_config:
        vllm_factors.append(self.lora_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.speculative_config:
        vllm_factors.append(self.speculative_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.structured_outputs_config:
        vllm_factors.append(self.structured_outputs_config.compute_hash())
    if self.profiler_config:
        vllm_factors.append(self.profiler_config.compute_hash())
    else:
        vllm_factors.append("None")
    vllm_factors.append(self.observability_config.compute_hash())
    if self.quant_config:
        pass  # should be captured by model_config.quantization
    if self.compilation_config:
        vllm_factors.append(self.compilation_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.kernel_config:
        vllm_factors.append(self.kernel_config.compute_hash())
    else:
        vllm_factors.append(None)
    if self.kv_transfer_config:
        vllm_factors.append(self.kv_transfer_config.compute_hash())
    else:
        vllm_factors.append("None")
    if self.ec_transfer_config:
        vllm_factors.append(self.ec_transfer_config.compute_hash())
    else:
        vllm_factors.append("None")
    vllm_factors.append(self.aux_output_config.compute_hash())
    if self.additional_config:
        if isinstance(additional_config := self.additional_config, dict):
            additional_config_hash = safe_hash(
                json.dumps(additional_config, sort_keys=True).encode(),
                usedforsecurity=False,
            ).hexdigest()
        else:
            additional_config_hash = additional_config.compute_hash()
        vllm_factors.append(additional_config_hash)
    else:
        vllm_factors.append("None")
    factors.append(vllm_factors)

    hash_str = safe_hash(str(factors).encode(), usedforsecurity=False).hexdigest()[
        :10
    ]
    return hash_str

enable_trace_function_call_for_thread()

Set up function tracing for the current thread, if enabled via the VLLM_TRACE_FUNCTION environment variable.

Source code in vllm/config/vllm.py
def enable_trace_function_call_for_thread(self) -> None:
    """Set up function tracing for the current thread,
    if enabled via the `VLLM_TRACE_FUNCTION` environment variable.
    """
    if envs.VLLM_TRACE_FUNCTION:
        tmp_dir = tempfile.gettempdir()
        # add username to tmp_dir to avoid permission issues
        tmp_dir = os.path.join(tmp_dir, getpass.getuser())
        filename = (
            f"VLLM_TRACE_FUNCTION_for_process_{os.getpid()}"
            f"_thread_{threading.get_ident()}_at_{datetime.now()}.log"
        ).replace(" ", "_")
        log_path = os.path.join(
            tmp_dir,
            "vllm",
            f"vllm-instance-{self.instance_id}",
            filename,
        )
        os.makedirs(os.path.dirname(log_path), exist_ok=True)
        enable_trace_function_call(log_path)

validate_block_size()

Validate block_size against DCP and mamba constraints.

Called after Platform.update_block_size_for_backend() has finalised block_size.

Source code in vllm/config/vllm.py
def validate_block_size(self) -> None:
    """Validate block_size against DCP and mamba constraints.

    Called after Platform.update_block_size_for_backend() has
    finalised block_size.
    """
    block_size = self.cache_config.block_size

    # Skip DCP interleave-size compatibility for NIXL P/D: the interleave
    # size is pinned to block_size by each worker.
    nixl_pd_active = (
        self.kv_transfer_config is not None
        and self.kv_transfer_config.has_connector("NixlConnector")
    )
    if self.parallel_config.decode_context_parallel_size > 1 and not nixl_pd_active:
        assert (
            self.parallel_config.cp_kv_cache_interleave_size <= block_size
            and block_size % self.parallel_config.cp_kv_cache_interleave_size == 0
        ), (
            f"Block_size({block_size}) should be greater "
            "than or equal to and divisible by cp_kv_cache_interleave_size "
            f"({self.parallel_config.cp_kv_cache_interleave_size})."
        )
    # Mamba cache align-mode constraints
    if self.cache_config.mamba_cache_mode == "align":
        assert not self.scheduler_config.disable_chunked_mm_input, (
            "Chunked MM input is required because we need the flexibility "
            "to schedule a multiple of block_size tokens even if they are "
            "in the middle of a mm input"
        )

default_breakable_cudagraph_architectures() cached

Architectures defaulting to breakable CUDA graphs on this platform.

Source code in vllm/config/vllm.py
@lru_cache
def default_breakable_cudagraph_architectures() -> frozenset[str]:
    """Architectures defaulting to breakable CUDA graphs on this platform."""
    from vllm.platforms import current_platform

    if current_platform.is_rocm():
        # Breakable CUDA graphs currently regress performance on ROCm for
        # models that can use torch.compile piecewise graphs instead. Do not
        # opt those in by default. Users can still force them with
        # VLLM_USE_BREAKABLE_CUDAGRAPH=1.
        #
        # DeepseekV41ForCausalLM cannot torch.compile, and the ROCm sparse
        # SWA backend only reports AttentionCGSupport.UNIFORM_BATCH. Default
        # FULL_AND_PIECEWISE then dies at capture unless breakable CUDA
        # graphs are on. Enable this architecture so the published AMD
        # recipe can start.
        return frozenset({"DeepseekV41ForCausalLM"})
    return DEFAULT_BREAKABLE_CUDAGRAPH_ARCHITECTURES

enable_act_fusion(cfg)

Enable if either SiLU+Mul or quant FP8 custom op is active; otherwise Inductor handles fusion. Also enable for FP4 models as FP4 quant is always custom so Inductor cannot fuse it.

Source code in vllm/config/vllm.py
def enable_act_fusion(cfg: "VllmConfig") -> bool:
    """Enable if either SiLU+Mul or quant FP8 custom op is active;
    otherwise Inductor handles fusion.
    Also enable for FP4 models as FP4 quant is always custom so Inductor cannot fuse it.
    """
    return (
        cfg.compilation_config.is_custom_op_enabled("silu_and_mul")
        or cfg.compilation_config.is_custom_op_enabled("quant_fp8")
        or (cfg.model_config is not None and cfg.model_config.is_nvfp4_quantized())
    )

enable_allreduce_rms_fusion(cfg)

Enable if TP > 1 and Hopper/Blackwell and flashinfer installed.

Source code in vllm/config/vllm.py
def enable_allreduce_rms_fusion(cfg: "VllmConfig") -> bool:
    """Enable if TP > 1 and Hopper/Blackwell and flashinfer installed."""
    from vllm.platforms import current_platform
    from vllm.utils.flashinfer import has_flashinfer

    # The fused all-reduce + RMSNorm path is not batch-invariant
    if envs.VLLM_BATCH_INVARIANT:
        return False

    if current_platform.is_rocm():
        from vllm._aiter_ops import rocm_aiter_ops

        return (
            rocm_aiter_ops.is_enabled() and cfg.parallel_config.tensor_parallel_size > 1
        )

    return (
        cfg.parallel_config.tensor_parallel_size > 1
        and current_platform.is_cuda()
        and has_flashinfer()
        and (
            current_platform.is_device_capability_family(100)
            or current_platform.is_device_capability(90)
        )
    )

enable_mla_dual_rms_norm_fusion(cfg)

Enable MLA dual RMS norm fusion on ROCm with AITER.

Source code in vllm/config/vllm.py
def enable_mla_dual_rms_norm_fusion(cfg: "VllmConfig") -> bool:
    """Enable MLA dual RMS norm fusion on ROCm with AITER."""
    from vllm._aiter_ops import rocm_aiter_ops

    return rocm_aiter_ops.is_enabled()

enable_norm_fusion(cfg)

Enable if either RMS norm or quant FP8 custom op is active; otherwise Inductor handles fusion.

Source code in vllm/config/vllm.py
def enable_norm_fusion(cfg: "VllmConfig") -> bool:
    """Enable if either RMS norm or quant FP8 custom op is active;
    otherwise Inductor handles fusion."""
    return (
        cfg.compilation_config.is_custom_op_enabled("rms_norm")
        or cfg.compilation_config.is_custom_op_enabled("quant_fp8")
        or cfg.kernel_config.ir_op_priority.rms_norm[0] != "native"
    )

enable_norm_pad_fusion(cfg)

Enable if using AITER RMSNorm and hidden size is 2880 i.e. gpt-oss.

Source code in vllm/config/vllm.py
def enable_norm_pad_fusion(cfg: "VllmConfig") -> bool:
    """Enable if using AITER RMSNorm and hidden size is 2880 i.e. gpt-oss."""
    return (
        cfg.kernel_config.ir_op_priority.fused_add_rms_norm[0] == "aiter"
        and cfg.model_config is not None
        and cfg.model_config.get_hidden_size() == 2880
    )

enable_qk_norm_rope_kvcache(cfg)

Enable fused QK-norm + RoPE/MRoPE + KV cache update with AITER.

Source code in vllm/config/vllm.py
def enable_qk_norm_rope_kvcache(cfg: "VllmConfig") -> bool:
    """Enable fused QK-norm + RoPE/MRoPE + KV cache update with AITER."""
    from vllm._aiter_ops import rocm_aiter_ops

    if not rocm_aiter_ops.is_enabled():
        return False
    return cfg.compilation_config.is_custom_op_enabled("rotary_embedding")

enable_rope_kvcache_fusion(cfg)

Enable if rotary embedding custom op is active and use_inductor_graph_partition is enabled.

Source code in vllm/config/vllm.py
def enable_rope_kvcache_fusion(cfg: "VllmConfig") -> bool:
    """Enable if rotary embedding custom op is active and
    use_inductor_graph_partition is enabled.
    """
    from vllm._aiter_ops import rocm_aiter_ops

    return (
        rocm_aiter_ops.is_enabled()
        and cfg.compilation_config.is_custom_op_enabled("rotary_embedding")
        and (
            cfg.compilation_config.use_inductor_graph_partition
            or not cfg.compilation_config.splitting_ops_contain_kv_cache_update()
        )
    )

enable_rope_kvcache_mla_fusion(cfg)

Enable if use_inductor_graph_partition is enabled.

Source code in vllm/config/vllm.py
def enable_rope_kvcache_mla_fusion(cfg: "VllmConfig") -> bool:
    """Enable if use_inductor_graph_partition is enabled."""
    return (
        cfg.compilation_config.use_inductor_graph_partition
        or not cfg.compilation_config.splitting_ops_contain_kv_cache_update()
    )

get_cached_compilation_config() cached

Cache config to avoid repeated calls to get_current_vllm_config()

Source code in vllm/config/vllm.py
@lru_cache(maxsize=1)
def get_cached_compilation_config():
    """Cache config to avoid repeated calls to get_current_vllm_config()"""
    return get_current_vllm_config().compilation_config

get_layers_from_vllm_config(vllm_config, layer_type, layer_names=None)

Get layers from the vLLM config.

Parameters:

  • vllm_config

    (VllmConfig) –

    The vLLM config.

  • layer_type

    (type[T]) –

    The type of the layer to get.

  • layer_names

    (Iterable[str] | None, default: None ) –

    The names of the layers to get. If None, return all layers.

Source code in vllm/config/vllm.py
def get_layers_from_vllm_config(
    vllm_config: VllmConfig,
    layer_type: type[T],
    layer_names: Iterable[str] | None = None,
) -> dict[str, T]:
    """Get layers from the vLLM config.

    Args:
        vllm_config: The vLLM config.
        layer_type: The type of the layer to get.
        layer_names: The names of the layers to get. If None, return all layers.

    """
    forward_context = vllm_config.compilation_config.static_forward_context
    if layer_names is None:
        layer_names = forward_context.keys()

    return {
        layer_name: layer
        for layer_name in layer_names
        if isinstance(layer := forward_context.get(layer_name), layer_type)
    }

set_current_vllm_config(vllm_config, check_compile=False, prefix=None)

Temporarily set the current vLLM config. Used during model initialization. We save the current vLLM config in a global variable, so that all modules can access it, e.g. custom ops can access the vLLM config to determine how to dispatch.

Source code in vllm/config/vllm.py
@contextmanager
def set_current_vllm_config(
    vllm_config: VllmConfig, check_compile=False, prefix: str | None = None
):
    """Temporarily set the current vLLM config.
    Used during model initialization.
    We save the current vLLM config in a global variable,
    so that all modules can access it, e.g. custom ops
    can access the vLLM config to determine how to dispatch.
    """
    global _current_vllm_config, _current_prefix
    old_vllm_config = _current_vllm_config
    old_prefix = _current_prefix
    from vllm.compilation.counter import compilation_counter

    num_models_seen = compilation_counter.num_models_seen
    try:
        # Clear the compilation config cache when context changes.
        # This is needed since the old config may have been accessed
        # and cached before the new config is set.
        get_cached_compilation_config.cache_clear()

        _current_vllm_config = vllm_config
        _current_prefix = prefix
        yield
    except Exception:
        raise
    else:
        if check_compile:
            vllm_config.compilation_config.custom_op_log_check()

        if (
            check_compile
            and vllm_config.compilation_config.mode == CompilationMode.VLLM_COMPILE
            and compilation_counter.num_models_seen == num_models_seen
        ):
            # If the model supports compilation,
            # compilation_counter.num_models_seen should be increased
            # by at least 1.
            # If it is not increased, it means the model does not support
            # compilation (does not have @support_torch_compile decorator).
            logger.warning(
                "`torch.compile` is turned on, but the model %s"
                " does not support it. Please open an issue on GitHub"
                " if you want it to be supported.",
                vllm_config.model_config.model,
            )
    finally:
        _current_vllm_config = old_vllm_config
        _current_prefix = old_prefix
        # Clear the compilation config cache when context changes
        get_cached_compilation_config.cache_clear()