Hardware-aware model optimisation
Research questionHow should model and runtime settings adapt to a specific GPU topology?
Target outcomeA reproducible configuration profile for each hardware class.
Measurable metricsThroughput, latency, utilisation, cost per token