Beyond a Performance Scalar: Evaluation Science for Evolutionary Computation

Evolutionary and nature-inspired algorithms are routinely compared through a single performance scalar: the best objective value reached at the end of a fixed computational budget, a mean or median aggregated over independent runs, a scalarized multi-objective score, or a position in a summary ranking table. Although convenient, collapsing an entire search process into one number conceals how performance was achieved, whether an apparent advantage is statistically and practically meaningful, which mechanisms caused it, and whether the resulting solutions are useful in practice.

This special session invites original research on evaluation methodologies that move beyond scalar summaries of algorithm performance. It connects three levels of evidence: theoretical and mechanism-level understanding, controlled benchmark evidence, and application-level utility. Submissions should introduce, validate, or critically compare analytical methods, metrics, statistical procedures, experimental designs, or evaluation frameworks. Papers whose principal contribution is a new optimizer evaluated only through conventional final-fitness tables fall outside the session’s central scope.

At the theoretical and mechanism level, the session welcomes analyses of convergence, runtime and sample complexity, stability, population dynamics, exploration and exploitation, diversity, search trajectories, adaptation, stagnation, and operator or component contributions.

At the benchmark level, it welcomes property- and distribution-aware evaluation across landscapes, dimensions, instances, budgets, targets, constraints, noise, dynamic changes, and fidelity levels.

At the application level, it welcomes evaluation that integrates solution quality with feasibility, reliability, uncertainty, interpretability, wall-clock cost, memory or energy consumption, scalability, and decision relevance.

Topics of Interest

Topics include, but are not limited to:

  • Anytime, time-to-target, attainment, and convergence-profile analysis
  • Behavioural, trajectory, population-motion, and decision-space-diversity measures
  • Mechanism attribution, ablation, sensitivity, and component-interaction analysis
  • Uncertainty quantification, effect sizes, statistical comparison, practical significance, and ranking stability
  • Distributional and multi-criteria alternatives to scalar aggregation, including performance profiles, run-time distributions, and Pareto-based reporting of cost and quality
  • Explainable performance analysis and data-driven characterization of algorithm behaviour
  • Robustness and generalization across problem properties, instances, budgets, fidelities, and distribution shifts
  • Evaluation of multimodal, constrained, dynamic, noisy, and multi-fidelity optimization
  • Cost-, resource-, reliability-, and decision-aware evaluation in real-world applications
  • Evaluation of evolutionary machine learning, AutoML, automated algorithm design, and LLM-assisted evolutionary computation
  • Frameworks linking theoretical predictions, benchmark observations, and practical outcomes

The session addresses three fundamental questions: what aspects of evolutionary search
should be evaluated, how should they be measured, and what evidence is required to support
conclusions in a particular context? It seeks methodological and empirically or theoretically
validated contributions rather than general opinion papers.

Organisers

  • Rohit Salgotra
    Faculty of Physics and Applied Computer Science / Centre of Excellence in Artificial Intelligence AGH University of Krakow, Poland
    rohits(at)agh.edu.pl
  • Tome Eftimov
    Computer Systems Department JoΕΎef Stefan Institute, Ljubljana, Slovenia
    tome.eftimov(at)ijs.si
  • Niki van Stein
    Leiden Institute of Advanced Computer Science, Leiden University, Leiden, The Netherlands
    n.van.stein(at)liacs.leidenuniv.nl
  • Anja Jankovic
    Chair of AI Methodology, RWTH Aachen University, Aachen, Germany
    jankovic(at)aim.rwth-aachen.de
  • Eva Tuba
    D. R. Semmes School of Science, Trinity University, San Antonio, TX, USA;
    Computer Systems Department, JoΕΎef Stefan Institute, Ljubljana, Slovenia
    eva.tuba(at)ijs.si