Monday 0.01 London FUZZ-IEEE Paper FUZZ 1 : Fuzzy data analysis, clustering and classifiers, pattern recognition, bio-informatics Session Chair: Thomas Runkler (Siemens AG, Technische Universtat Munchen) Possibilistic One Mean Clustering Using Particle Swarm Optimization Thomas Runkler (Siemens AG) Abstract Abstract Two popular methods for soft clustering are fuzzy clustering and possibilistic clustering. Possibilistic c-means (PCM) clustering is less sensitive to noise and outliers in data than fuzzy c-means clustering (FCM), but PCM may find non-distinct (duplicate) clusters. Both PCM and FCM can also be used to find single clusters, but FCM yields trivial clusters (the complete set with all memberships equal to one), whereas PCM can find meaningful single clusters, which is called possibilistic one-mean (P1M). P1M is also used for multiple clusters by sequentially finding one cluster after the other, the so-called sequential possibilistic one-means (SP1M), which can avoid non-distinct clusters, as an alternative to conventional PCM. P1M clustering can be done by alternating optimization (AO) of the P1M objective function, but P1M AO may converge to local minima. To the best of our knowledge this is the first paper that examines local minima of the P1M objective function. Moreover, we propose a particle swarm optimization (PSO) algorithm to avoid such local minima. Experiments with synthetic and real-world data indicate that PSO finds the global minimum of the P1M objective function significantly more often than AO, and also on average leads to lower values of the P1M objective function, which yields a higher probability of obtaining better clustering results, at the cost of longer run time, but still feasible even for larger data sets. Feature-weighted FKNN regression using fuzzy mutual information and Łukasiewicz similarity Mahinda Mailagaha Kumbure and Luukka Pasi (LUT University) Abstract Abstract This paper proposes a feature-weighted Minkowski distance-based fuzzy k-nearest neighbor regression method, called FWMD-FKNNreg. The key novelty of FWMD-FKNNreg is its integration of feature weighting directly into the Minkowski distance computation in FKNN regression, enabling adaptive control of each feature’s contribution. To achieve this, we introduce four new feature weighting schemes based on relevance, redundancy, and dependency, that are developed using fuzzy mutual information and Łukasiewicz similarity measure. An empirical evaluation identifies the most effective weighting strategy for FKNN regression, with relevance-based weighting outperforming others. Based on this finding, relevance-weighted FWMD-FKNNreg is evaluated on eight real-world datasets and compared with six competitive regression methods, including KNNreg, FKNNreg, Md-FKNNreg, support vector regression (SVR), LASSO, and multiple linear regression (MLR). Experimental results show that FWMD-FKNNreg achieves higher R2 and lower RMSE values in most cases, indicating that the proposed method offers a robust, simple, and effective alternative for regression problems. Hierarchical Clustering Based on Sequential Preference Stability Teresa González-Arteaga (University of Valladolid) and Rocío de Andrés Calle and José Manuel Cascón (Universty of Salamanca) Abstract Abstract This paper presents a novel hierarchical clustering methodology grounded in a refined formulation of a sequential preference stability measure. This measure captures the chance that two agents preserve identical opinions on an alternative across consecutive time points, while remaining well-defined in the presence of incomplete preference information. We formalize the measure, and derive a decomposition that separates within group and cross group stability contributions for arbitrary agent partitions. Leveraging these theoretical results, we introduce the Stability Hierarchical Clustering Algorithm. An efficient incremental update strategy substantially reduces computational overhead. The methodology is validated by means of two case studies: a synthetic temporal preference profile designed to illustrate its discriminative capability, and an empirical dataset concerning life-support decisions of cancer patients, which demonstrates its practical applicability. Hard Possibilistic C-Means: An Empirical Study of Noise, Overlap, and Rejection Paritosh Tiwari (Indian Institute of Science), James C. Bezdek (Emeritus Professor), Thomas Runkler (Siemens AG), and Punit Rathore (Indian Institute of Science) Abstract Abstract Hard Possibilistic C-Means (HPCM) is a prototype-based clustering method that replaces forced assignments with binary typicalities, allowing a point to be assigned to multiple clusters (overlap) or to none (rejection). This makes HPCM attractive in settings with ambiguity and background noise, but it also complicates evaluation and raises practical questions about when HPCM behaves differently from standard Hard C-Means (HCM). In this paper we present an empirical study of HPCM through controlled synthetic experiments and small real-world benchmarks. We visualize how HPCM differs from HCM in overlap and noise regimes, and we analyze the sensitivity of HPCM to the scale parameter that defines its acceptance thresholds, showing how uniform background noise can inflate it and change the effective rejection/overlap behaviour. Since HPCM does not output a classical partition, we evaluate it using set-valued/reject-style metrics that capture the accuracy--coverage and accuracy--set-size tradeoffs. Overall, the results highlight both the intended advantages of HPCM (explicit ambiguity and noise filtering) and the conditions under which its behaviour can degenerate due to scale estimation. Benchmarking Stress Detection: Fuzzy Logic vs. Deep Learning Amal Talbi (National School of Engineering of Sousse, University of Fribourg); Mehrez Abdellaoui (National School of Engineering of Sousse); and Edy Portmann (University of Fribourg) Abstract Abstract Detection systems for stress monitoring in airport settings must be precise, comprehensible, and deployable under practical limitations. A comparative benchmarking analysis of wearable physiological sensing-based stress detection techniques is presented in this paper, which contrasts cutting-edge deep learning and computer vision techniques with an interpretable fuzzy inference framework based on Design Science Research. The analysis takes into account typical wearable signals, such as electrocardiography (ECG), electrodermal activity (EDA), and heart rate variability (HRV), and assesses the consequences for real-time monitoring of both software and hardware deployment. Several benchmark datasets are used to test methods based on deployment-oriented characteristics, including robustness, interpretability, data requirements, computing cost, privacy, and sustainability. The results demonstrate that the fuzzy inference method offers transparent, stable, and privacy-preserving stress estimate with low computing overhead, while deep learning and vision-based models reach high accuracy in controlled circumstances. For ethical and sustainable airport applications, these results encourage the use of interpretable wearable-based stress monitoring. Monday 0.02 Berlin FUZZ-IEEE Paper FUZZ 2: FUZZ-IEEE SS06 Fairness and Trustworthiness in Intelligent Decision Support Systems & Main: Fuzzy web engineering, information retrieval, text mining and social network analysis Session Chair: Luis Martinez (University of Jaén) A Type-2 Fuzzy-Based Feature-Driven Rule Reduction Approach for Demand Forecasting within an Energy Smart Grid System Mahmoud Alfayan and Hani Hagras (university of essex) Abstract Abstract Abstract— Accurate short-term energy demand forecasting is essential for the stability and efficiency of modern energy smart grids, especially with the growing integration of renewable energy sources, which introduce significant uncertainty. Traditional machine learning models, such as neural networks and deep learning, are "black boxes" that deliver accurate predictions but lack transparency, limiting operator trust and actionable insights. To help users better understand the prediction model and increase their confidence in its outputs, it is necessary to employ eXplainable Artificial Intelligence (XAI), whose models and predicted results can be easily understood, analyzed, and improved by business users. In this paper, we present an Interval Type-2 Fuzzy Logic System (IT2 FLS)-based XAI system, designed to manage noise, non-stationarity, and uncertainty in hourly demand data. We utilize SHapley Additive Explanations (SHAP) and a Random Forest (RF) to enhance model interpretability and prediction reliability by reducing the number of inputs and rules in the IT2 FLS. In our approach, SHAP is used to rank features and combine their importance with RF-based significance for optimal feature selection and rule pruning in the IT2 FLS. This method reduces the root-mean-squared error (RMSE) on testing data by 55%, from 5.94 kilowatt-hours (with a full rule base of 2,781 rules) to 2.74 kilowatt-hours, using only three inputs and 16 rules—a 250-fold reduction in the number of rules. In comparison, the Type-1 FLS showed lower performance with an RMSE of 5.376, while the neural network achieved a lower RMSE of 1.410 but lacked interpretability. The proposed IT2 FLS-based XAI system thus offers a practical balance between accuracy and explainability, enabling real-time smart grid applications. This approach will be further applied within a smart grid to facilitate adaptive decision-making and optimization. A Fuzzy-Based Approach for Interpretable Spike Detection in Living Neural Biocomputers Adham Aboulkheir, Hani Hagras, and Michael Barros (University of Essex) Abstract Abstract The convergence of neuroscience and artificial intelligence has led to the development of "living biocomputers," where biological neural networks are cultured on Micro-Electrode Array (MEA) chips. This creates hybrid biocomputing platforms that blur the boundaries between biology and technology. A critical challenge in this domain is the accurate and interpretable detection of neural spikes from MEA recordings. While black-box models like deep neural networks have achieved good accuracy, their lack of transparency hinders scientific inquiry. Revealing spatial-temporal patterns across electrode arrays can inform electrode placement strategies, stimulation protocols, and our understanding of how biological neural networks process information. This approach could bridge the gap between understanding the organization of neurons, their activity and the computing task in hand and ultimately allow living neural biocomputers to have a solution to the current “black-box” biocomputing approaches. We adopt a fuzzy logic framework because it can handle the biological uncertainty and signal variability inherent in MEA recordings, enabling more faithful spike discrimination than rigid threshold-based or crisp machine learning approaches. This paper presents a fuzzy rule-based classifier that combines high performance with human-readable interpretability for spike detection in neural computing. Our approach integrates ANNIGMA (Artificial Neural Network for Input Gain and Measurement Analysis)-based feature selection to identify the most salient spike characteristics and a genetic algorithm to optimize the fuzzy rule base. We validated our model using multi-chip pooled validation protocol, achieving a peak F1 score on testing data of 97.74% on a two-chip experiment and 96.12% on a six-chip experiment. The resulting fuzzy rules are linguistically interpretable, allowing researchers to understand the underlying logic of spike detection. This work can be a step forward towards building trustworthy neural biocomputers with direct applications to drug discovery, disease modeling, and understanding the fundamentals of neural computation. Trustworthy and Explainable Neuro-Fuzzy Ensembles for Clinical Decision Support batyrkhan Sharipbay and Adnan Yazici (Nazarbayev University) Abstract Abstract We introduce an interpretable clinical classification framework that couples Adaptive Neuro-Fuzzy Inference Systems (ANFIS) with Fuzzy Random Forests (FRF) to achieve high performance while maintaining transparent decision logic. In the proposed design, ANFIS learns data-driven Gaussian membership functions and optimizes their mean and sigma parameters, which are then used to induce a robust FRF ensemble for decision-making. The end-to-end pipeline follows a leakage-aware methodology: (i) stratified train–test splitting with skewness-aware imputation computed strictly from the training set; (ii) Z-score normalization (zero mean, unit variance) fitted on training data only; (iii) mutual information–based feature selection via SelectKBest to obtain a compact, discriminative subset; (iv) SMOTE class balancing applied exclusively to the training set; (v) ANFIS-based optimization of Gaussian membership parameters; and (vi) FRF construction using ANFIS-optimized fuzzy partitions to produce interpretable predictions. Experiments on two UCI benchmarks (Cleveland Heart Disease and Wisconsin Diagnostic Breast Cancer) show that the ANFIS–FRF framework achieves AUC values of approximately 0.95 and 0.99, respectively, matching or exceeding strong baselines such as Random Forest and XGBoost. Importantly, the model yields a human-readable rule set, whose interpretability is further examined through an LLM-assisted evaluation. By decoupling fuzzy feature learning (ANFIS) from decision logic induction (FRF) and embedding mutual information-driven feature selection within rigorous preprocessing, the proposed approach provides an accurate, robust, and transparently explainable solution for clinical decision support. A Fairness-Aware Minimum Cost Consensus Model Bapi Dutta, Diego García-Zamora, Rosa M. Rodríguez, and Luis Martínez (Universidad de Jaén) Abstract Abstract Minimum cost consensus (MCC) models are widely used in group decision-making to achieve agreement with minimal opinion adjustment costs. Although recent studies have incorporated fairness into MCC frameworks, most existing approaches rely on aggregate social welfare or inequality-based axioms and remain insensitive to subgroup structures induced by protected attributes. As a result, consensus outcomes may still systematically favor certain subgroups despite being cost-optimal or globally fair. To address this limitation, this paper proposes a fairness-aware MCC framework that explicitly accounts for subgroup-level fairness. Fairness is defined through deviation-based measures that evaluate the parity between the final consensus outcome and the opinions of subgroups formed by protected attributes. Fairness is incorporated into MCC models via constraint-based formulations and a learning-inspired objective that interpolates between utilitarian and Rawlsian principles. The proposed models preserve convexity and computational tractability, while enabling a flexible balance between efficiency, consensus, and subgroup fairness in heterogeneous group decision-making environments. Improving Individual Fairness in Fuzzy Classifiers via Evolutionary Multi-Objective Optimization Takeru Konishi, Naoki Masuyama, and Yusuke Nojima (Osaka Metropolitan University) Abstract Abstract To address the ethical and societal risks of artificial intelligence, responsible artificial intelligence that ensures not only high accuracy but also transparency and fairness is increasingly demanded. Inherently interpretable models are valuable in scenarios where both transparency and fairness are important, as the model’s internal mechanisms can be inspected to support bias diagnosis and mitigation. Fuzzy systems are representative inherently interpretable models that enable flexible decisions considering real-world uncertainties. Multi-objective fuzzy genetics-based machine learning generates a diverse set of fuzzy classifiers that consider trade-offs among multiple objectives using an evolutionary multi-objective optimization algorithm. However, previous studies on fairness in fuzzy classifiers have focused exclusively on group fairness and have not considered individual fairness. In this study, we propose a multi-objective fuzzy genetics-based machine learning that simultaneously optimizes accuracy and individual fairness, and investigate how incorporating individual fairness as an objective affects the resulting set of fuzzy classifiers. Experimental results show the effectiveness of this approach and provide empirical insights into the relationships among accuracy, group fairness, and individual fairness, contributing to fair AI design. Source code is available at https://github.com/TakeruKonishi/MoFGBML Fairness. DPC-HOPE: A Fuzzy Density-Peak Approach to Overlapping Community Detection Swetha Balasubramanian and Pranab K. Muhuri (South Asian University) Abstract Abstract Overlapping community detection in complex networks remains a fundamental unresolved problem due to the fuzzy nature of node memberships and the computational bottleneck imposed by the O(n^2) complexity of conventional density-peak algorithms. Real-world networks are characterized by continuous membership distributions in which nodes belong to multiple functional groups. However, existing methods often compromise this structural realism by imposing discrete assignments or are computationally infeasible for networks at the million-node scale. In this paper, we propose DPC-HOPE, a framework that uses Density Peak Clustering (DPC) using a Structural Fuzzy Membership Propagation lens. We replace Euclidean distance with a High-Order Proximity Embedding (HOPE) and then introduce a Fuzzy Directed Acyclic Graph (F-DAG) propagation mechanism to model the gradient of membership flow from community peaks to the boundaries. This enables an explicit, non-heuristic determination of overlapping memberships with a computational complexity of O(m + n log n). Experimental results on large-scale LFR benchmarks and real-world networks demonstrate that DPC-HOPE consistently outperforms existing community detection algorithms in both accuracy and community quality. These results suggest that explicitly capturing topological fuzziness is essential to achieving both scalability and accuracy in large-scale network analysis. Monday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 1: Algorithms I Session Chair: Pauline Catriona Haddow (NTNU) QiSA: A Parameter-free Quantum-inspired Search Algorithm -- A Preliminary Study on Combinatorial Optimization Kuo-Chun Tseng and Wei-Chieh Lai (National Ilan University), Chen-Hsin Lu (National Tsing Hua University), Wei-Chun Huang (National Chung-Shan Institute of Science & Technology), Jen-Shin Hong (National Chi Nan University), and Hsin-Hung Cho (National Ilan University) Abstract Abstract This study is inspired by quantum algorithms and proposes the Quantum-inspired Search Algorithm (QiSA), an entropy-driven quantum-inspired metaheuristic for parameter-free combinatorial optimization. Guided by Occam's razor, QiSA simplifies the search process and reduces parameter dependence as much as possible in this preliminary study. It uses entropy as the core signal to unify parameter control, overcoming the traditional dilemma of balancing exploration and exploitation, and maps it to a confidence-like level to define an interpretable closed-loop termination criterion. QiSA also exhibits several observable quantum-like behaviors during the search, reflecting characteristic phenomena associated with quantum algorithms. Experiments on classic combinatorial optimization benchmarks show that QiSA delivers robust, consistent, and competitive performance, with the potential to stand alongside Differential Evolution in combinatorial optimization. Comparison of MCDE and MC-SHADE for Massively Parallel Optimization Rainer Storn (Rohde & Schwarz) Abstract Abstract Recently the multi-child differential evolution (MCDE) algorithm, in which each parent competes against more than one child, has been introduced. It has been demonstrated that MCDE achieves a wall-clock time reduction of up to two orders of magnitude for massively parallelized optimization. Results which were based on the classical DE algorithm DE/rand/1/bin showed that the speedup achieved by MCDE scales well with the number K of processors, even when K is orders of magnitude higher than the population size N. In this work the multi-child version of SHADE is compared with MCDE yielding promising results. Multi-Objective Multi-Agent Path Finding Based on Two-Stage Evolutionary Algorithms Haruto Ando, Yoshiki Nogata, and Tomohiro Harada (Saitama University) and Fumito Uwano (Okayama University) Abstract Abstract Multi-agent path finding (MAPF) is a critical problem in applications such as warehouse automation and autonomous driving. In real-world settings, multiple conflicting objectives, such as minimizing path length and ensuring safety, often must be optimized simultaneously, motivating the study of multi-objective MAPF (MOMAPF). However, existing MOMAPF approaches suffer from severe scalability limitations as the number of agents and the environment size increase. To address this challenge, we propose Two-Stage Evolutionary MOMAPF (TSE-MOMAPF), which optimizes a single agent's path via NSGA-II and refines the path set using an SMS-EMOA-based framework. Experiments on various grid maps show that the proposed method achieves superior scalability and provides broader Pareto front coverage in medium- and large-scale scenarios compared to an existing method. These results indicate that the proposed evolutionary framework effectively balances solution diversity and computational efficiency under strict time constraints. Self-Supervised State Representation for Reinforcement Learning-Assisted Evolutionary Algorithms Xiaotong Liu, Ye Tian, Hao Jiang, Langchun Si, Zimo Sheng, and Xingyi Zhang (Anhui University) Abstract Abstract Reinforcement learning-assisted evolutionary algorithms are attracting increasing attention, but existing methods rely on handcrafted state representations that are often problem-dependent, limited in expressiveness, and lack robustness to variations in population composition and distribution. To address these limitations, a self-supervised state representation learning framework is proposed for reinforcement learning-assisted evolutionary algorithms. This framework follows a two-stage learning paradigm consisting of representation learning and policy learning, where the state encoder is trained in the representation learning stage and subsequently used to generate state inputs for reinforcement learning based policy inference. Specifically, a set-aware attention-based encoder is designed and trained via self-supervised contrastive learning using population objective values, while an auxiliary guided learning strategy with dynamic weighting is employed to stabilize the representation learning process. To evaluate the effectiveness of the proposed framework, it is integrated into representative reinforcement learning-assisted evolutionary algorithms for single- and multi-objective optimization problems to assess its generalization capability, with additional ablation studies analyzing key components. The results demonstrate the effectiveness of the proposed framework in addressing the limitations of handcrafted designs. Exploring Island Genetic Algorithms and Unbounded Recognition Regions for Artificial Immune Systems Storm Jan Anton Visser and Pauline Catriona Haddow (NTNU) Abstract Abstract Artificial Immune Systems (AIS) have been extensively applied to the single-class anomaly detection problem, in optimization and adaptive control system applications. However, recent works are are finding ways to adapt AISs to address the needs of multi-class problems, at times achieving competitive advantages over traditional machine learning. Further, for higher dimensional problems such as Fake News classification, further extensions are needed. One such extension is Unbounded Recognition regions -- the area of the search space in which an antibody may detect an antigen (data). Exciting hybrid advances are also present in the field, combining AIS with, for example, Island Genetic Algorithms or Quantum computing. Basement Relief Inversion: A Comparative Study of the Dual Simplex and a Memetic Algorithm Arthur Anthony da Cunha Romão e Silva, Bruno Motta de Carvalho, Francisco Márcio Barboza, and Matheus da Silva Menezes (Federal University of Rio Grande do Norte) Abstract Abstract This paper addresses the inverse problem of reconstructing basement relief from gravity data by comparing two contrasting optimization strategies: a linearized Dual-Simplex (DS) method and a Memetic Algorithm (MA) combining global evolutionary operators with Gauss–Newton local refinement. The DS method solves a Linear Programming (LP) formulation based on $L_1$-norm data misfit and smoothness regularization, whereas the MA couples global exploration (BLX-$\alpha$ crossover, hybrid mutation, and tournament selection) with a weighted local update. Both methods are applied to synthetic and real 2D datasets—the Isolated Graben model and the Poem Bridge profile. In the synthetic case, characterized by smooth geometry and weak nonlinearity, DS achieves higher accuracy, including a 1\% relative error in 98\% of the observations. In contrast, in the more heterogeneous and nonlinear Poem Bridge dataset, the MA yields better cumulative relative data misfit error reduction, achieving a 5\% error in roughly 72\% of the data, compared to 58\% for DS. Overall, deterministic linearized methods prove efficient for well-behaved models, while evolutionary strategies offer greater robustness under complex real-world conditions. The results emphasize the complementary nature of both approaches and indicate that the optimal inversion strategy should be guided by the geological complexity and nonlinearity of the problem. Monday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 2 - Evolutionary Machine Learning I Session Chair: Lukas Sekanina (Brno University of Technology) Meta-Learning-Based Algorithm Selection for Multi-Objective Wind Farm Layout Optimization Gustavo Jorge Novaes Silva, Joao Gabriel Lofiego Sampaio Gomes Silva, and Islame Felipe da Costa Fernandes (Federal University of Bahia) Abstract Abstract Wind farm layout optimization involves multiple conflicting objectives and highly irregular search landscapes, making the selection of suitable multi-objective metaheuristics a challenging task. Although meta-learning has been successfully employed to select metaheuristics for single-objective optimization, its application to real-world multi-objective engineering problems remains unexplored. This paper proposes a meta-learning approach for automatically selecting metaheuristics for the Multi-Objective Wind Farm Layout Optimization Problem (MoWFLOP). The proposed approach uses fitness landscape meta-features to predict the performance of multi-objective state-of-the-art metaheuristics under different quality indicators. We evaluate 90 datasets with different sampling configurations from 516 wind farms. The analysis comprised the results of the metaheuristics, the impact of sampling configurations, and the meta-learning performance for metaheuristic selection. Results show that the meta-learning approach surpassed the baseline of selecting the same algorithm for all instances. Typed Linear Genetic Programming for Abstraction and Reasoning Corpus Zhixing Huang, Bing Xue, and Mengjie Zhang (Victoria University of Wellington) Abstract Abstract Abstraction and reasoning are essential abilities in human intelligence, which help humans understand the world and extend our knowledge to unseen cases. However, abstraction and reasoning are still big challenges for existing artificial intelligence models due to limited training instances and domain-specific concepts. Existing studies usually model the abstraction and reasoning tasks as program synthesis tasks. Genetic programming features in program synthesis. However, genetic programming is hardly applied to abstraction and reasoning tasks. To fill this gap and explore genetic programming's capability for abstraction and reasoning, this paper proposed a typed linear genetic programming method and evaluated its performance on the Abstraction and Reasoning Corpus. The results show that the typed linear genetic programming has a superior performance to many existing methods, implying a good potential to fulfill abstraction and reasoning in artificial intelligence. Recurrent Connections in Gene Expression Programming Neural Networks Yichen Yang and Jonathan Mwaura (Northeastern University) Abstract Abstract Gene Expression Programming Neural Networks (GEPNN) evolve neural network topologies through a fixed-length linear chromosome that decodes into expression trees. This encoding guarantees syntactic validity under genetic operations, but inherently produces acyclic structures that prevent GEPNN from solving temporal problems that require memory. This paper introduces the index terminal, a novel terminal type that references other nodes within the expression tree by their position on the chromosome. When a node references an earlier position, a recurrent connection is formed; when it references a later position, a skip connection is created. This mechanism enables arbitrary cycles in the computational graph while preserving fixed-length encoding and deterministic decoding properties found in Gene Expression Programming (GEP). The approach is validated on temporal XOR benchmarks that require one- and two-step memory. The results obtained confirm that evolution discovers recurrent connections through index terminals that maintain state across timesteps, enabling GEPNN to solve tasks previously inaccessible to the framework. Delayed regression experiments confirm that the mechanism generalizes to continuous-output tasks. This shows that index terminals support temporal sequence processing in both the classification and regression domains. These results establish that the proposed index terminals successfully extend GEPNN to temporal problems while maintaining the algorithmic simplicity that distinguishes GEP from variable-length neuroevolutionary encodings. Towards Improving Mutations in Neuroevolution Stav Bar-Sheshet and Amiram Moshaiov (Tel-Aviv University) and Adham Salih (Braude College of Engineering) Abstract Abstract The additive Gaussian noise mutation approach is commonly used in state-of-the-art neuro-evolution algorithms to optimize/train the weights of networks with fixed topology. In this paper, a novel modification to this mutation approach is proposed and studied. The proposed modification follows a theoretical analysis of mutations by the additive Gaussian noise. The well-known neuro-evolutionary algorithm of the Uber AI labs, which is based on additive Gaussian noise, is applied here as a reference algorithm. First, we modified the reference algorithm by replacing its mutation mechanism with the proposed mutation technique. Next, a comparison study is reported between the modified algorithm and the reference one. The comparison study is conducted on several well-known continuous control benchmark problems. The theoretical and computational studies indicate that the proposed modifications are not subjected to the constraints of the additive Gaussian noise approach, which appears to provide significantly better results with increasing problem dimensions. Evolutionary Multi-Objective Fusion of Deepfake Speech Detectors Vojtěch Staněk, Martin Perešíni, Lukáš Sekanina, Anton Firc, and Kamil Malinka (Faculty of Information Technology, Brno University of Technology) Abstract Abstract While deepfake speech detectors built on large self-supervised learning (SSL) models achieve high accuracy, employing standard ensemble fusion to further enhance robustness often results in oversized systems with diminishing returns. To address this, we propose an evolutionary multi-objective score fusion framework that jointly minimizes detection error and system complexity. We explore two encodings optimized by NSGA-II: binary-coded detector selection for score averaging and a real-valued scheme that optimizes detector weights for a weighted sum. Experiments on the ASVspoof 5 dataset with 36 SSL-based detectors show that the obtained Pareto fronts outperform simple averaging and logistic regression baselines. The real-valued variant achieves 2.37% EER (0.0684 minDCF) and identifies configurations that match state-of-the-art performance while significantly reducing system complexity, requiring only half the parameters. Our method also provides a diverse set of trade-off solutions, enabling deployment choices that balance accuracy and computational cost. Breaking Free from Hand-Crafted Rewards: A Genetic Programming Framework for End-Goal-Driven Reinforcement Learning Kamalesh Kumar (University of Massachusetts Amherst) and Jean-Alexis Delamer and James Hughes (St. Francis Xavier University) Abstract Abstract Reinforcement Learning (RL) enables the training of autonomous agents to accomplish various tasks through environmental feedback. This feedback, formalized as rewards, guides the agent toward optimal strategies for achieving desired objectives. However, a fundamental challenge exists: agents optimize for reward maximization rather than directly solving the intended tasks. The reward functions, typically designed manually, do not necessarily lead to optimal solutions and may fail to capture the true complexity of the problem domain. Even reward functions that appear intuitive to human designers often produce suboptimal or unexpected agent behaviors. Monday 0.10 Sydney IEEE CEC (Evolutionary Computation) CEC 3 - Optimization I Session Chair: John Sheppard (Montana State University) Incorporating Novelty Search Into Multi-Objective Cooperative Coevolution Shahnaj Mou, Asibul Islam, and John Sheppard (Montana State University) Abstract Abstract Multi-objective evolutionary algorithms (MOEAs) are widely used for solving complex optimization problems involving conflicting objectives. While NSGA-II (Non-dominated Sorting Genetic Algorithm II) remains a strong baseline, recent approaches such as Multi-Objective Factored Evolutionary Algorithms (MOFEA) and novelty-driven search aim to improve scalability and exploration in high-dimensional decision spaces. Based on a hypothesis that introducing novelty search into MOEAs should improve quality of returned Pareto sets, we present a systematic empirical comparison of four algorithms: NSGA-II, MOFEA, Novelty-MOEA (Multi-Objective Evolutionary Algorithm), and Novelty-MOFEA (Multi-Objective Factored Evolutionary Algorithms) on the ZDT test suite (excluding ZDT5) and all DTLZ problems. We evaluate their performance using standard quality indicators including Hypervolume (HV), $k$-nearest neighbor distance, Inverted Spacing (SP), and the UD unary quality indicator. Our results reveal that no single algorithm dominates across all problem types. Instead, performance depends strongly on problem characteristics such as modality, separability, and the presence of deceptive local optima. We analyze these trends and provide insights into when factored optimization and novelty mechanisms are most beneficial. SEHH: A Search Economics-Based Hyper-Heuristic for Single-Objective Real-Parameter Optimization Shao-Jhang Wu, Cheng-Hsun Chang, Yu-Xi Liu, and Chun-Wei Tsai (National Sun Yat-sen University) Abstract Abstract The basic idea of the hyper-heuristic (HH) algorithm is to integrate multiple search algorithms to leverage their strengths for optimization problems. One of the main challenges in HH design is selecting or generating low-level heuristics (LLHs) effectively within a limited number of evaluations. To address this challenge, this paper proposes an effective HH method based on the machine learning process for solving single-objective real-parameter optimization problems (SOPs). During the offline stage, a set of benchmark functions and an effective metaheuristic algorithm (i.e., search economics) will be used to automatically generate a “good” sequence of LLHs and a set of suitable hyperparameters for them. This good sequence can be regarded as a combination of LLH sets, which will construct a “new” metaheuristic algorithm for solving SOPs during the online stage. The experiments are conducted using the CEC 2024 and CEC 2025 benchmarks to evaluate the performance of the proposed method. The results conclusively demonstrate that this approach significantly outperforms the state-of-the-art algorithms. Solving Few-Shot Multiobjective Multitask Optimization via Iterative Sequential Transfer Tingyang Wei (Nanyang Technological University; Centre for Frontier AI Research (CFAR), A*STAR); Haofeng Wu (Nanyang Technological University); Ananda Phan Iman (Gwangju Institute of Science and Technology); Zhao Wei (Centre for Frontier AI Research (CFAR), A*STAR); Jiao Liu (Nanyang Technological University); and Yew-Soon Ong (Nanyang Technological University; Centre for Frontier AI Research (CFAR), A*STAR) Abstract Abstract Applying knowledge transfer across multiple optimization tasks, multitask optimization (MTO) emerges as a promising approach to solving synergistic optimization tasks simultaneously. However, the development of effective knowledge transfer mechanisms in MTO fundamentally relies on aligning elite solution distributions across tasks. This dependency creates a critical bottleneck in few-shot optimization regimes, as restricted evaluation budgets impede the identification of elite solution distributions required for beneficial transfer. This challenge is exacerbated in multiobjective multitask problems, where each optimizer must approximate a continuous Pareto manifold rather than a single optimal point. This paper introduces Iterative Sequential Transfer (IST) to circumvent this bottleneck. We model MTO as a sequence of sequential transfer optimization problems, concentrating evaluations on a single target per iteration. We propose a likelihood-informed task prioritization mechanism to maximize transfer utility by identifying the task most likely ready for knowledge integration. Empirical results on benchmark and real-world problems verify the effectiveness of the proposed method under tight budgets. Perceptor: Cross-Layer Adaptive Tuning Engine for LLM Inference Lei Li, Yubo Li, Zhipeng Shen, Anning Cai, Shuaiyi Zhang, Tao Yang, Chaoqun He, Yang Chu, and Bing Liu (Lenovo) Abstract Abstract On-site delivery and tuning of Large Language Model (LLM) inference face several challenges, like diverse environments, complex tuning parameters, and unpredictable workloads. We propose Perceptor, an end-to-end adaptive tuning engine to overcome these issues. Based on a perceive-adaption loop, Perceptor achieves lifetime-span adaptive performance tuning across various environments. It introduces a multi-stage optimization pipeline decoupling cross-layer key tuning parameters, and pioneers a collaborative strategy that combines pre-deployment (static) and post-deployment (online) optimizations to always keep optimal inference performance. Specifically in optimization pipeline, we have proposed novel adaptive optimization algorithms for parallel strategy, RoCE networks, and Prefill-Decode (PD) workload scheduling. Experiments demonstrate significant performance improvements achieved by Perceptor across various scenarios. Monday 0.11 Cape Town IEEE CEC (Evolutionary Computation) CEC 4 - Related Topics I Session Chair: Sanaz Mostaghim (Otto von Guericke University Magdeburg, Fraunhofer IVI Dresden) Diagnosing Optimizer-Landscape Interaction via Representation in Subset Selection Multitasking: Evidence from Piano Fingering Ananda Phan Iman (Gwangju Institute of Science and Technology); Tingyang Wei (Nanyang Technological University; Centre for Frontier AI Research (CFAR), A*STAR); and Chang Wook Ahn (Gwangju Institute of Science and Technology) Abstract Abstract Continuous relaxations are commonly used as an encoding scheme for evolutionary multitasking optimization (EMTO), allowing heterogeneous tasks to be optimized within a unified search space. However, these continuous representations fundamentally reshape the induced fitness landscape, and their impact on optimization behavior and knowledge transfer remains underexplored. In this work, we bridge the gap by analyzing the interaction between the representation-induced fitness landscape and optimizer behavior through fitness landscape analysis. Using piano fingering estimation as a representative case of the subset selection problem, we introduce several encoding variants and investigate how their induced landscapes influence search dynamics and solution quality under EMTO. Our analysis indicates that transfer efficacy is strongly associated with changes in representation-induced landscape characteristics, rather than algorithmic factors alone, highlighting the representation design as the key factor of effective knowledge transfer in multitasking optimization. Extending the Scalable Multi-Agent Pathfinding Problem to Modifiable Environments Carlo Nübel and Malte Speidel (Otto von Guericke University Magdeburg) and Sanaz Mostaghim (Otto von Guericke University Magdeburg, Fraunhofer IVI Dresden) Abstract Abstract In this paper, we extend the scalable multi-agent pathfinding problem to modifiable environments. Instead of only finding collision-free paths, agents collaboratively alter the environment to transform it to a predefined target state. Applications of this problem include smart warehousing, terrain leveling, autonomous mining, and automated construction. We introduce a scalable simulation environment for such tasks and propose an evolutionary optimization approach. Two experiments are conducted: (1) analyzing the effect of the number of agents on task performance and (2) comparing the performance of different crossover operators. We propose a new two-level crossover operator acting on different layers of the gene encoding and compare it with two one-level operators. Results indicate that more agents reduce the task completion time, but fewer agents yield more consistent solutions with better objective values due to the collision constraints. Moreover, the introduced two-level crossover operator achieves significantly better objective values than both one-level operators in most problem instances. Solving Electric Vehicle Routing by Iterative Instance Refinement Framework Tzu-Hao Lin and Ying-ping Chen (National Yang Ming Chiao Tung University) Abstract Abstract The Electric Vehicle Routing Problem (EVRP) requires the joint optimization of customer routing and energy feasibility under battery capacity constraints. Although many existing EVRP solution strategies implicitly follow a bilevel structure, routing and feasibility handling are commonly implemented within monolithic solvers, making algorithmic components tightly coupled and difficult to substitute or reuse. Adaptive Modular Differential Evolution Framework for Robotic Trajectory Optimization Jan Fiala, Martin Juříček, and Jakub Kůdela (Brno University of Technology) Abstract Abstract Optimization of robotic arm trajectories represents a complex computational challenge critical for efficient and safe automation. While traditional metaheuristic approaches have achieved notable success, they often require extensive manual parameter tuning and lack robustness when problem characteristics change. This work introduces a novel approach to robotic trajectory optimization based on the Modular Differential Evolution (MODDE) framework. Instead of designing a single, problem-specific algorithm, we employ a hyperheuristic framework that autonomously manages and configures the components of MODDE to identify optimal trajectories regarding execution time and motion smoothness. We demonstrate that employing this hyperheuristic control results in more robust and adaptive solutions compared to manually tuned metaheuristics, representing a significant advancement in the practical application of evolutionary computation. A Task-Similarity-Based Multitask Genetic Programming Approach to Symbolic Regression Ying Bi (Zhengzhou University, State Key Laboratory of Intelligent Agricultural Power Equipment); Yaxin Chang and Wenjing Li (Zhengzhou University); caitong yue (Zhengzhou University, State Key Laboratory of Intelligent Agricultural Power Equipment); and Jing Liang (Henan Institute of Technology; Zhengzhou University, State Key Laboratory of Intelligent Agricultural Power Equipment) Abstract Abstract Symbolic regression (SR) automatically derives interpretable mathematical expressions from data and has been widely applied in various fields. Many real-world SR tasks are intrinsically similar, such as fault prediction models for different engine types sharing common structures. However, most existing SR methods focus on independent single-task modeling and fail to exploit inter-task relationships to improve performance. To address this, a task-similarity-based multitask genetic programming (TSMTGP) approach is proposed for SR. TSMTGP adopts a novel two-tree individual representation, in which each solution consists of a shared tree across tasks and a task-specific tree tailored to each task, enabling effective knowledge transfer. Correspondingly, two fitness evaluation methods are designed for the shared tree and the two-tree representation. A task similarity measure strategy is developed by using Spearman’s rank correlation coefficient based on the performance of the shared tree on different tasks. Based on this measure, a similarity-aware evolutionary mechanism is introduced, in which the evolutionary process adaptively switched between high-similarity and low-similarity modes to improve knowledge transfer. TSMTGP is consistently evaluated under varying task similarities on both benchmark and real-world problems. TSMTGP consistently outperforms comparative methods, particularly on real-world datasets, demonstrating its effectiveness in dynamic knowledge transfer across different tasks. Forging Paths for DoS/DDoS Resilient Networks Kevin Olenic and Sheridan Houghten (Brock University) Abstract Abstract In today's technology focused environment, networks receive a constant multitude of requests from various user devices, leading to cases of high latency in the network if data transmissions are mismanaged. If a denial of service (DoS) or distributed denial of service (DDoS) attack is launched on a network with mismanaged data transmissions, the damage would be catastrophic. Thus, appropriate management of transmissions is paramount. Methods for managing communications exist, but they require costly computations and account for select variables such as energy consumption or the amount of data transmitted. Our approach utilizes genetic algorithms to generate network configurations with structures that reduce latency based on different factors: data transmitted over a connection, and energy consumption. Once designed, their resilience to DoS/DDoS attacks is evaluated. The results show the methodology designs networks resistant to attack and that the surviving networks can also be reconfigured to re-establish the fitness to a pre-attacked state. Monday 0.14 Singapore IEEE CEC (Evolutionary Computation) CEC 5 - SS01: Integrating Machine Learning Methods into Evolutionary Optimization Session Chair: Amir Gandomi (University of Technology Sydney) Adaptive Switching between Search and Model Refinement via Explainable Machine Learning in Surrogate-assisted Evolutionary Algorithms Kei Nishihara (Yokohama National University), Takahiro Sato (Muroran Institute of Technology), Kazuhiro Izui (Kyoto University), and Shinya Watanabe (Muroran Institute of Technology) Abstract Abstract Surrogate-assisted evolutionary algorithms (SAEAs) alternate expensive evaluations using surrogate models of objective functions, and thus have become a key approach for solving expensive optimization problems. Several efforts have been made to enhance the prediction accuracy of surrogate models. However, this does not necessarily lead to performance improvement of SAEAs, as solutions contributing to the update of the best objective value and model refinement are technically different from each other in nature. Thus, balancing the two aspects is crucial for accelerating the performance of SAEAs. Nevertheless, the resource allocation method to sample solutions for optimization or for model refinement remains unclear. Accordingly, this work provides a novel structure-aware adaptive switching methodology between evolutionary search and model refinement modes, grounded in explainable machine learning. Specifically, SHapley Additive exPlanations (SHAP) are employed to quantify the structural instability of the surrogate model. When the surrogate model is structurally unstable or inaccurate, inspired by active learning, the proposed SAEA framework generates solutions by anisotropically perturbing decision variables according to dimension-wise SHAP instability to refine the model. Otherwise, it switches to the evolutionary search mode to discover better solutions. The effectiveness of the proposed framework over state-of-the-art SAEAs is demonstrated through comparison on benchmark and engineering design problems. Online electric vehicle charging scheduling with commitment using surrogate-assisted optimization Abdennour Azerine (Université de Haute-Alsace); Mahmoud Golabi (Université de Haute-Alsace, IRIMAS UR 7499, 68100 Mulhouse, France); and Lhassane Idoumghar (IRIMAS, Université de Haute-Alsace, France; Université de Haute-Alsace, IRIMAS UR 7499, 68100 Mulhouse, France) Abstract Abstract The Electric Vehicle Charging Scheduling Problem (EVCSP) concerns the allocation of limited charging resources to heterogeneous vehicle requests under temporal and capacity constraints. Although advanced optimization models have been proposed for this problem, most assume an offline setting with complete knowledge of future arrivals, limiting their applicability in real charging stations. This paper introduces an online, preemptive EVCSP that preserves a high-fidelity problem formulation, including charger availability and grid capacity constraints, while enforcing a commitment policy under which admission decisions are irrevocable. We propose a hybrid solution approach that combines a genetic algorithm for online admission control and charger assignment with a mathematical optimization model for energy allocation. To meet real-time decision-making requirements, we embed a Gaussian-process surrogate within the evolutionary search to reduce the cost of repeatedly evaluating candidate schedules. Computational experiments across multiple instance sizes and booking modes show that surrogate assistance preserves objective values relative to exact evaluation while yielding statistically significant reductions in computational time, supporting practical online deployment of evolutionary optimization for real-time EV charging station operations. Dynamic Multi-Modal Particle Swarm Optimization for Training Neural Network Ensembles Under Concept Drift Chris Langeveldt and Andries Engelbrecht (Stellenbosch University) Abstract Abstract Artificial neural networks (NNs) often struggle to maintain performance when concept drift, a phenomenon where the underlying data distribution changes over time, occurs. While traditional training approaches based on backpropagation struggle to maintain performance under drift conditions, ensemble learning combined with dynamic multi-modal (DMM) optimization offers a promising alternative. This paper investigates the applicability of multi-quantum swarm optimization (MQSO), a DMM particle swarm optimization (PSO) variant, for training NN ensembles in dynamic environments. MQSO is evaluated against traditional gradient-based methods and standard quantum swarm optimization. MQSO achieves superior performance on lower-dimensional problems with simpler decision boundaries, effectively tracking multiple optima. However, gradient-based methods with retraining outperform MQSO on higher-dimensional, non-linear problems due to computational constraints when using MQSO. The findings further show that retraining mechanisms are essential for maintaining performance under concept drift. It is also clear that DMM PSO is particularly effective for tracking rapidly changing decision boundaries. A particle swarm optimisation maximum-margin classifier Sizalobuhle Ncube (Stellenbosch University), Jean-Pierre Van Zyl (Stellenbosch UniversityStellenbosch University), and Andries Engelbrecht (Stellenbosch University) Abstract Abstract This paper proposes a particle swarm optimisation maximum-margin classifier (PSO-MMC) that borrows core principles from support vector machines (SVMs). The task of identifying an optimal linear separating hyperplane is formulated as a constrained optimisation problem defined exclusively over training instances that form Tomek links. For non-linearly separable data, kernel mappings are incorporated in a manner analogous to the SVM kernel formulation. Prior to Tomek link identification, noise is removed from the training data using established filtering techniques. Experimental evaluation is conducted against classical SVMs, classification trees and 𝑘-nearest neighbour (KNN) across multiple benchmark data sets. The results show that the proposed classifier achieves improved classification accuracy and reduced optimisation complexity, particularly under conditions of class overlap and feature redundancy. The findings indicate that restricting optimisation to boundary-critical instances provides an efficient and scalable alternative model for maximum-margin classification. A Modular Multi-Tool Evolutionary Algorithm Framework for RNA Inverse Folding Kaiyu Nie and Nawwaf Kharma (Concordia University) Abstract Abstract Most RNA inverse folding methods rely on a single structure predictor, often resulting in sequences that overfit specific algorithmic biases. To address this, we present a predictor-agnostic Evolutionary Algorithm framework that decouples optimization logic from folding engines. On the Eterna100 benchmark, our framework achieves a 68% solve rate using single-tool modes, matching the performance of state-of-the-art reinforcement learning methods. Crucially, to mitigate the risk of model-specific errors, we introduce a consensus optimization strategy that seeks sequences validated by multiple, distinct folding models (e.g., thermodynamic and deep learning). By filtering out predictor-specific artifacts, this approach acts as a structural "safety net," effectively reducing the occurrence of extreme prediction outliers and yielding more robust candidates for downstream experimental testing. Finally, the framework features a modular "plug-and-play" architecture that supports granular, nucleotide-level constraint control, enabling precise customization for diverse design tasks. Monday 0.15 Washington IEEE CEC (Evolutionary Computation) CEC 6-SS10:Evolutionary Computation in Dynamic and Uncertain Environments Session Chair: Michalis Mavrovouniotis (Cyprus University of Technology) Dynamic Electric Vehicle Routing Problem using Population-Based Ant Colony Optimization Michalis Mavrovouniotis (ERATOSTHENES Centre of Excellence, Cyprus University of Technology); Changhe Li (Anhui University of Science and Technology); Charalambos Chrysostomou and Maria Anastasiadou (ERATOSTHENES Centre of Excellence, Cyprus University of Technology); and Shengxiang Yang (De Montfort University) Abstract Abstract The population-based ant colony optimization (P-ACO) algorithm has been proven effective in addressing the dynamic electric vehicle routing problem (DEVRP). In the DEVRP the travel time is subject to dynamic changes representing real-world traffic conditions. P-ACO maintains an archive of solutions that are used to update the pheromone trails. By default, the archive is updated in first in, first out fashion. In this study, we introduce a new update strategy based on the similarities of the solutions stored in the archive to maintain the diversity when updating the pheromone trails. The experimental results on different DEVRP test cases demonstrate the effectiveness of the proposed strategy in comparison with other existing ones. Dynamic Multi-Objective Optimization of Integrated Energy Systems in Steel Enterprises via Conditional Variational Autoencoder Min Chen (Northeastern University, China); Shengxiang Yang (De Montfort University); and Yanyan Zhang and Shengnan Zhao (Northeastern University, China) Abstract Abstract To overcome the limitations of conventional evolutionary algorithms in dynamic integrated energy systems for steel enterprises, this paper proposes a Conditional Variational Autoencoder (CVAE) guided multi-objective optimization framework. The method minimizes operational costs and external energy reliance by using online-trained CVAE to map dynamic environmental states (e.g., energy loads, price signals) to optimal solution distributions. Upon detecting changes, the model rapidly generates a state-conditioned initial population, providing a warm start for the evolutionary search. Experiments verify that this framework significantly accelerates convergence and improves solution stability and economic performance in dynamic environments. Multi-Point Search Towards Dynamic Multimodal Optimization with Frequent Solution Changes on their Locations, Moving speeds, and Numbers Shoei Fujita and Hiroyuki Sato (The University of Electro-Communications) and Keiki Takadama (The University of Tokyo) Abstract Abstract In the dynamic multimodal optimization problems where multiple optimal solutions change their locations and their number as time passes, enough number of individuals are indispensable to track the dynamically changing solutions. For this issue, many conventional methods take an approach of increasing the number of individuals becasue an optimal number of individuals are not known, but such an approach makes it hard to track the solutions when their speed become fast. To tackel this problem, this paper proposes (i) NCCO (sup+ul) which incorporates the suppression mechanism and an upper-limit mechanism into NMMSO- and CMA-ES-based Continuous Optimizer (NCCO) and (ii) NCCO (dec+ul) which replaces the suppression mechanism with a decrement mechanism that reduces the number of individuals as much as possible. The experiments on dynamic benchmark functions in the two and five dimensions based on the Moving Peaks Benchmark (MPB) demonstrate that NCCO (sup+ul) and NCCO (dec+ul) successfully track multiple optima in dynamic multimodal functions with lower errors than Multi-swarm Quantum Particle Swarm Optimization (mQSO) as the conventional method. Specifically, (1) NCCO (dec+ul), followed by NCCO (sup+ul) and NCCO, decreases Fitness Error and Distance Error, and (2) NCCO (dec+ul) also maintains the smallest number of individuals during the search. From Environmental Parameters to Pareto Optimal Prediction: LLM-Based Large-Scale Dynamic Multi-Objective Optimization Chenyang Li, Gary Yen, Lei Liu, and Zhenan He (Sichuan University) Abstract Abstract Existing prediction-based dynamic multi-objective evolution-ary algorithms (DMOEAs) face significant challenges in tracking dynamic Pareto optimal set (POS) for large-scale dynamic multi-objective optimization problems, including insufficient high-quality training data and the lack of prior knowledge about optimal solutions in the new environment. To address these issues, we propose a large language model (LLM)-based DMOEA. The proposed method jointly models environmental time-varying parameters and dynamic POS within a unified semantic space via LoRA-based fine-tuning. The fine-tuned LLM leverages parameter variation sequences from historical to new environments, together with historical POS sequences, as contextual information to accurately pre-dict the change directions of POS in a new environment. In doing so, this facilitates the evolution progress in finding the POS in the new environment. Additionally, we design an LLM reasoning-guided evolutionary search that narrows the search region from the large-scale search space to a promising local region, generating an initial population with both convergence and diversity for the new environment. Comprehensive exper-imental results on widely used dynamic benchmark problems demonstrate that LLM-DMOEA significantly outperforms state-of-the-art algorithms, particularly in large-scale scenar-ios. An Interval-Constrained MultiObjective Evolutionary Algorithm Integrating Interval Constrained Clustering and Transfer Learning Feimeng Wang, Dunwei Gong, Jing Sun, and Chunliang Zhao (Qingdao University of Science and Technology) Abstract Abstract Most practical engineering applications are often formulated as interval constrained multiobjective optimization problems (ICMOPs). In such problems, at least one objective or constraint contains interval parameters. With increasing constraints, the feasible region becomes narrow and fragmented. This greatly raises the difficulty of finding feasible solutions. Therefore, this paper proposes an interval-constrained multiobjective evolutionary algorithm integrating interval constraint clustering and transfer learning (ICCTL-ICMOEA). First, an interval constraint overlap degree is introduced to quantify constraint similarity, the Louvain algorithm is employed to cluster constraints into subclasses. Each constraint class is then combined with interval objectives to construct multiple subproblems. Next, within a multi-population evolutionary framework, a population state-aware search strategy is designed to independently evolve the main population and subpopulations. Following that, an interval constraint relaxation-guided knowledge transfer strategy is developed to migrate selected individuals from subpopulations to the main population, accelerating feasible region exploration. Finally, the proposed algorithm was tested on constructed benchmark ICMOPs and compared with a state-of-the-art interval-constrained multiobjective evolutionary algorithm and a constrained multiobjective evolutionary algorithm. Experimental results demonstrate the strong competitiveness of the proposed algorithm. Dual-Drive-ERL: A Dual-Population Interactive-Driven Evolutionary Reinforcement Learning Algorithm for Dynamic Traffic Assignment Jing-Yuan Chen, Feng-Feng Wei, Tai-You Chen, and Wei-Neng Chen (South China University of Technology) Abstract Abstract With the acceleration of urbanization, the demand for travel in urban areas and the number of private vehicles have risen sharply. However, the structure of the urban transportation network remains largely unchanged, and the road capacity is limited, unable to meet the increasing transportation demands, resulting in traffic congestion and further causing a series of negative impacts, such as time waste, aggravated environmental pollution, and frequent traffic accidents, ultimately leading to a low social benefit. Traditional travel planning usually only provides a few shortest path navigation suggestions, and the feedback of road conditions is often delayed, which easily leads travelers to make wrong traffic decisions, causing the traffic congestion to worsen. Reasonable traffic flow distribution can balance the traffic flow in the network and effectively alleviate traffic congestion. Therefore, the problem of traffic flow distribution has received widespread attention and research. This paper proposes a dual-population information interaction-driven evolutionary reinforcement learning algorithm, Dual-Drive-ERL, for solving the dynamic traffic flow distribution problem. This is the first time that the evolutionary reinforcement learning algorithm has been applied to the dynamic traffic flow distribution problem, and an innovative mechanism of dual-population information interaction-driven is proposed to enhance the exploration ability. The algorithm has shown superior performance to other comparison algorithms in experiments conducted in 45 different scenarios. Monday Brightlands Foyer IJCNN J2C Presentation, IJCNN Paper, IJCNN Position Paper, IJCNN Late Breaking Paper Poster Presentations IJCNN (1) Session Chair: Thorben Markmann (Bielefeld University), Aleksei Liuliakov (Bielefeld University) Optimistic Feasible Search for Closed-Loop Fair Threshold Decision-Making Wenzhang Du (Mahanakorn University of Technology International College (MUTIC)) Abstract Abstract Closed-loop decision-making systems (e.g., lending, screening, or recidivism risk assessment) often operate under fairness and service constraints while simultaneously inducing feedback effects: decisions change who appears in the future, yielding non-stationary data and potentially amplifying disparities. We study online learning of a one-dimensional threshold policy from bandit feedback under demographic parity (DP) and (optionally) service-rate constraints. The learner observes only a scalar score each round and selects a threshold; reward and constraint signals are revealed for the chosen threshold only. We propose Optimistic Feasible Search (OFS), a simple grid-based method that maintains confidence bounds for reward and constraint residuals for each candidate threshold. At each round, OFS selects the threshold that appears feasible under confidence bounds and, among those, maximizes optimistic reward; if no threshold appears feasible, OFS selects the threshold minimizing optimistic violation. This design directly targets feasible high-utility thresholds and is particularly effective for low-dimensional, interpretable policy classes where discretization is natural. STT-BP: Training Kernel-Learnable LIF for Efficient Neuromorphic Vision and Audio Recognition via Spatio-Temporal Tilted Backpropagation Haoran Gao, Xiang Fu, Yihang Chen, Xinyu Li, and Cong Shi (Chongqing University) Abstract Abstract Spiking Neural Networks (SNNs) offer a promising avenue for energy-efficient edge intelligence, yet their performance is often bottlenecked by the rigid, fixed dynamics of traditional neuron models like LIF and CuBa-LIF. These models typically rely on pre-defined exponential decay kernels, which fail to capture the diverse multi-scale temporal dependencies present in varying neuromorphic modalities. To overcome this limitation, we propose the Kernel-Learnable LIF (KL-LIF) neuron, a generalized model capable of end-to-end learning arbitrary temporal response kernels (Integrated and Reset kernels). To train deep networks composed of KL-LIF neurons efficiently, we derive the Spatio-Temporal Tilted Backpropagation (STT-BP) algorithm. By exploiting the duality between forward spike propagation and backward error flow, STT-BP establishes a tilted gradient path via kernelized error gradients, enabling unified optimization of synaptic weights and temporal kernels without complex state unrolling. Extensive experiments on neuromorphic vision (DVS-Gesture, DVS-CIFAR10) and spiking audio (SHD, SSC) datasets demonstrate that STT-BP significantly outperforms state-of-the-art methods. Notably, our approach reveals distinct optimal temporal scales for different modalities—shorter kernels for vision and longer kernels for audio—achieving top-tier accuracy. Probing the Functional Role of Muscle Synergies in Reinforcement Learning–Based Torso Balance via Neural Reconstruction and Synergy Ablation Siyuan Liu and Jiahao Chen (Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences) Abstract Abstract Muscle synergies are widely hypothesized as an organizational principle for controlling highly redundant musculoskeletal systems. However, most existing synergy analyses rely on post-hoc statistical decompositions of recorded muscle activity, making it difficult to directly relate synergy function to closed-loop control. In this work, we propose a reinforcement learning-based neuromuscular control framework that embeds muscle synergies as explicit, policy-structured components within the control architecture. A synergy decoder maps low-dimensional modulation signals to high-dimensional muscle excitations, enabling synergy structure to emerge as an intrinsic part of the learned controller. Based on this, we introduce a decoder-based synergy analysis method using neural reconstruction and synergy ablation. We evaluate the proposed approach on a high-dimensional musculoskeletal torso balance task. Synergy function is probed through both open-loop replay ablation and closed-loop online ablation. The results show that reconstruction-based importance and functional importance are related but not equivalent. It suggests a division of labor among synergy pathways, with different synergy channels contributing to feasibility, terminal stabilization, and precision shaping. Overall, this work provides a control-relevant framework for studying the functional roles of muscle synergies and bridges biomechanical motor control with reinforcement learning. The experiment video can be found in: https://sites.google.com/view/syn-decoder. Role Discovery with Dynamic Role Cardinality for Zero-Shot Scalable Cooperation Wanjun Jing and Zhen Liu (The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences, China); Zhiming Zhou (The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences); Keyan Guo and Jiebo Chen (Beijing Institute of Electrical Engineering, Beijing 100854, China); and Zixin Liu (The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences, China) Abstract Abstract Inspired by the natural emergence of roles in biological systems for efficient coordination, role-based multi-agent reinforcement learning (MARL) methods have been proposed to enhance policy heterogeneity and cooperative capability in complex tasks. However, existing approaches typically rely on a fixed number of roles, whereas in practical multi-agent systems, the required role cardinality often varies across task stages and agent numbers, causing learned cooperative strategies to transfer poorly under varying agent numbers. In this paper, we propose SRD-DRC, a scalable role discovery framework with dynamic role cardinality. The core idea is to dynamically aggregate a global context from the local interaction histories of all agents via an attention mechanism, and to adaptively estimate the number of roles required at each cooperation stage, thereby eliminating reliance on a predefined role cardinality. Based on this estimation, SRD-DRC employs contrastive learning to acquire discriminative role representations that capture behavior patterns evolving across cooperation stages. Moreover, role representations are integrated at both the policy and value levels by parameterizing role-conditioned heterogeneous layers and guiding centralized value decomposition to enable structured credit assignment. Experimental results on the StarCraft Multi-Agent Challenge (SMAC) benchmark demonstrate that SRD-DRC achieves superior performance across diverse cooperative scenarios and exhibits effective zero-shot scalability. DOFR: Diffusion-Model One-Iter Forgetting-Based Replay for Compute-Budgeted Continual Learning Taro Murayama (DENSO CORPORATION) and Ichiro Takeuchi (Nagoya University, RIKEN) Abstract Abstract Diffusion models achieve outstanding image-generation performance, yet practical deployment requires continual knowledge expansion under distribution shift. Most continual learning studies focus on memory constraints (limited access to old data). In practice, however, the primary bottleneck is often compute: GPU time and the number of allowable training iterations (optimizer steps) are limited. To improve the final generation quality with a limited number of iterations, we propose Diffusion-model One-iter Forgetting-based Replay (DOFR). Our idea comes from empirical observation that the loss increase after a single iteration on a new task (one-iteration forgetting) serves as a low-cost yet informative proxy for long-term forgetting. Based on this insight, DOFR estimates the forgetting risk at each diffusion timestep in a simple closed form, enabling replay budgets to be allocated to classes with higher forgetting risk. Experiments with Stable Diffusion v1.5 show that, under the same compute budget, DOFR improves overall generation quality over simple replay baselines, yielding a better compute-performance trade-off. The Geometry of Fragility: Benchmarking Robustness of Time Series Explanations via Curvature Maximization Yueshan Chen and Sihai Zhang (University of Science and Technology of China) Abstract Abstract Explainable Artificial Intelligence is critical for trustworthy time series forecasting, yet recent studies reveal its vulnerability to adversarial manipulations. Existing attacks, primarily designed for images, fail to preserve temporal semantics and often require white-box access, rendering them ineffective against non-differentiable methods like LIME and SHAP. To bridge this gap, we propose the Iso-surface Curvature Maximization Attack, a geometric attack method grounded in the insight that explanation instability is intrinsically governed by the curvature of the prediction manifold. Instead of targeting specific explainer outputs, ICMA directly maximizes the local curvature of the decision boundary under a spectro-temporal constraint. This constraint strictly confines perturbations to the high-frequency residual subspace, preserving the underlying trend and seasonality. Extensive experiments on real-world datasets across diverse architectures confirm that maximizing this geometric objective significantly destabilizes explanations, empirically verifying that high curvature is indeed the primary driver of explanation fragility. HSRNet: A Novel Hyperspectral Retentive Network for Hyperspectral Image Classification Junbo Zhou, Guoqiang Gao, Chengyu Zhang, Yujun Guo, Xueqin Zhang, Song Xiao, and Guangning Wu (Southwest Jiaotong University) Abstract Abstract Hyperspectral image (HSI) classification faces a persistent challenge: modeling long-range semantic context while preserving fine-grained local structures. Existing paradigms often struggle to balance these needs; Convolutional Neural Networks (CNNs) are limited by local receptive fields, whereas Transformer-style attention can mix irrelevant regions and overlook the sequential nature of spectral signatures. Moreover, treating HSI as a unified volume often blurs the distinct roles of the spectral and spatial axes. To address these issues, this paper proposes the Hyperspectral Retention Network (HSRNet), which factorizes feature learning into two stages: a spectral retention encoder that captures continuous band correlations in a pixel-wise spectral token sequence, followed by a spatial retention encoder that aggregates geometric context over the flattened spatial tokens. A distance-decayed retention operator with head-wise learnable decay factors controls the extent of contextual mixing, allowing different heads to emphasize short-range interactions or preserve per-token details, while a lightweight local context enhancement module injects explicit 2D neighborhood cues. Experiments on Pavia University and Houston 2013 datasets demonstrate the effectiveness of HSRNet, achieving Overall Accuracies of 98.9% and 98.3%, respectively, and producing more coherent classification maps than recent state-of-the-art baselines. TinySegPrune: Multi-Objective Pruning for Semantic Segmentation Networks on TinyML Hardware Zhuoran Xiong, Warren Gross, and Brett Meyer (McGill University) Abstract Abstract Network pruning is critical for supporting the efficient deployment of convolutional neural networks (CNNs) on resource-limited TinyML hardware. While pruning has proven effective in many cases, semantic segmentation models remain challenging to prune for deployment to TinyML devices due to their large output feature maps and consequent memory accesses, necessitating more fine-grained compression strategies. We propose TinySegPrune: a multi-objective pruning framework that parameterizes the pruning rate of each block in the network. We employ a latency predictor that estimates latency from block-wise pruning rates for fast hardware performance evaluation. We validate the effectiveness of our framework by pruning two different segmentation networks for a TinyML processor. Our framework finds a pruned model that reduces latency by 30.3% with only 1% accuracy reduction, outperforming benchmark methods that achieve only 14.7% latency reduction---a 2.1X larger improvement. Leveraging latency feedback, our framework effectively identifies critical blocks, reducing their latency by 30% to 50%, and significantly improving overall efficiency. AGAPE: Training-Free Audio-Guided Adaptive Prompt Ensembling for Few-Shot Audio Classification Kai Guo (Beijing Institute of Technology); Xiang Xie (Beijing Institute of Technology; Beijing Institute of Technology, Zhuhai); and Shangkai Zhao and Shu Wu (Beijing Institute of Technology) Abstract Abstract Although large-scale pre-trained models have revolutionized audio classification, adapting them to downstream few-shot tasks remains challenging. Existing methods face a dilemma: conventional few-shot adaptation requires resource-intensive gradient updates, while zero-shot inference often suffers from semantic misalignment in text prompts. To address these issues in deployment-constrained scenarios, we propose Audio-Guided Adaptive Prompt Ensembling (AGAPE), a training-free framework for few-shot audio classification. Specifically, AGAPE first utilizes limited support shots to construct optimal class-specific text prototypes, thereby enhancing text-prompt-based classification. Subsequently, the final classification is achieved by fusing the similarity scores of the test sample against both these optimized text prototypes and the audio prototypes. Furthermore, we introduce the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) for Test-Time Adaptation (TTA), enabling efficient model adaptation without gradient updates. Experiments on ESC-50 and FSD50K demonstrate that AGAPE significantly outperforms the baseline. Ablation studies further validate the effectiveness of our methods. RadProPoser: Probabilistic Radar Tensor Human Pose Estimation That Knows Its Limits Jonas Leo Mueller (Machine Learning and Data Analytics Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg; Munich Center for Machine Learning); Lukas Engel (Institute of Microwaves and Photonics, Friedrich-Alexander-Universität Erlangen-Nürnberg); Eva Dorschky and Daniel Krauss (Machine Learning and Data Analytics Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg); Ingrid Ullmann and Martin Vossiek (Institute of Microwaves and Photonics, Friedrich-Alexander-Universität Erlangen-Nürnberg); and Bjoern M. Eskofier (Machine Learning and Data Analytics Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg; Munich Center for Machine Learning) Abstract Abstract Radar-based human pose estimation enables privacy-preserving motion tracking for ambient intelligence, yet the noisy nature of radar sensing makes uncertainty quantification essential. We present RadProPoser, an end-to-end probabilistic framework that predicts three-dimensional body joints with per-joint uncertainties from raw radar tensor data. Using a variational encoder-decoder with spectral attention that fuses real and imaginary radar components across temporal frames, we model aleatoric uncertainty through learnable Gaussian and Laplace distributions. Trained on a new benchmark dataset with optical motion-capture ground truth, our method achieves 6.425 cm mean per-joint position error. The model outputs per-joint aleatoric uncertainties, and isotonic recalibration yields calibrated total uncertainty with expected calibration error of 0.027. Since spectral attention operates on individual radar tensor components, extending to multi-radar configurations requires only concatenating additional input streams. On the HuPR benchmark with dual orthogonal radars, this achieves 5.042 cm MPJPE. The framework runs at approximately 89 frames per second (FPS) on an NVIDIA RTX 3090, exceeding the 15 Hz radar frame rate. Biologically Informed Graph Attention Network for Interpretable Brain Connectivity Analysis Roger Zhu (Millburn High School) and Rishi Upadhyay (University of California, Los Angeles) Abstract Abstract Brain connectome analysis has long been a focus in neuroscience and neuroimaging research. However, many recent approaches focus heavily on fMRI and ROI-based modeling approaches, often overlooking key structural features and connectivity patterns of white matter tracts. In this work, we present a novel heterogeneous graph framework that incorporates structural features of white matter tracts, representing both white matter clusters and anatomical regions as nodes with edges capturing their spatial intersection. This approach allows us to model both structural integrity and connectivity, along with regional context, for more biologically meaningful patterns. Using GraphSAGE for feature aggregation and multi-head Graph Attention Networks (GAT) for interpretability, our proposed architecture achieves an AUC of 0.891 for predicting gender on the Human Connectome Project 1200 dataset, matching the performance of other existing models and validating our graph construction while using significantly fewer training data. Furthermore, we extract attention scores, identifying key white matter tracts and anatomical regions involved in sex differentiation, grounding model predictions in neuroanatomical relevance. Our results highlight the importance of structural information and pave the way for more biologically informed models. Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning Daisuke Niizumi (Tokyo Metropolitan University, NTT Inc.); Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen, and Noboru Harada (NTT Inc.); and Nobutaka Ono (Tokyo Metropolitan University) Abstract Abstract Since the introduction of Masked Autoencoders, various improvements to masking strategies have been explored. In this paper, we rethink masking strategies for audio representation learning using masked prediction-based self-supervised learning (SSL) on general audio spectrograms. While recent informed masking techniques have attracted attention, we observe that they incur substantial computational overhead. Motivated by this observation, we propose dispersion-weighted masking (DWM), a lightweight masking strategy that leverages the spectral sparsity inherent in the frequency structure of audio content. Our experiments show that inverse block masking, commonly used in recent SSL frameworks, improves audio event understanding performance while introducing a trade-off in generalization. The proposed DWM alleviates these limitations and computational complexity, leading to consistent performance improvements. This work provides practical guidance on masking strategy design for masked prediction-based audio representation learning. Same Accuracy, Different Geometry: Solver-Dependent Representation Learning in Third-Order Neural ODEs Gavin lee Goodship, Luis Miralles, and Stephen O'Sullivan (TU Dublin) Abstract Abstract Models with similar predictive accuracy are often assumed to learn comparable internal representations. We challenge this assumption in continuous-depth networks by showing that the numerical solver used to discretize Neural Ordinary Differential Equations (Neural ODEs) systematically alters the geometry of learned representations, even under matched predictive performance. Using a third-order Neural ODE framework, we study how Runge-Kutta (RK4) and extended-stability solvers (ESRK-15 and ESRK-21) shape latent trajectories in models trained within typical run-to-run accuracy variability ($\approx$2\%) on CIFAR-10 and CIFAR-100. Despite identical integration horizons and comparable validation accuracy, the resulting representations exhibit markedly different geometric structure, as quantified by curvature and jerk statistics derived directly from the latent dynamics. Our results show that solver stability properties and internal amplification act as implicit inductive biases that regulate representational smoothness and higher-order dynamics. These findings demonstrate that numerical integration is not merely an implementation detail, but a structural component of the model that influences how continuous-depth networks encode and regularize information. PP-Mixer: A Multi-Scale Pooled Patch MLP Framework for Long-Term Time Series Forecasting Shuo Zhang, Feihu Yan, Wenlong Xu, Yanliang Tan, Junlin An, and Xuelin Cheng (Zhejiang university) Abstract Abstract Long-term time series forecasting(LSTF) is crucial for applications such as energy management and traffic forecasting, yet remains challenging due to complex multi-scale temporal dependencies, including long-term trends and mixed periodic patterns. Recently patch-based methods improve local temporal modeling by partitioning sequences into subseries, but fixed-length patching struggles to adapt to diverse and varying periodic structures in real-world data. To address this issue, we propose PP-Mixer, a fully MLP-based framework for LTSF. PP-Mixer employs a Fourier-guided multi-scale pooled patch embedding to adaptively construct multi-scale representations aligned with dominant temporal periods. A multi-scale decomposition mixer module then models and integrates seasonal and trend components across scales, while a dual-dimension mixing block enables efficient temporal modeling and feature interaction within each scale. Experiments on eight real-world benchmarks show that PP-Mixer consistently outperforms Transformer-based and MLP-based baselines, with ablation and visualization analyses confirming the effectiveness of its period-aware and multi-scale design. EasyFSS: Dual-Polarity Matching and Multi-Prompt Guidance for Few-Shot Segmentation Yang Shi, Linsen Yu, and Guanglu Sun (Harbin University Of Science And Technology) Abstract Abstract Few-Shot Segmentation (FSS) aims to recognize and segment novel classes in images using only a limited number of annotated examples. Recently, increasing efforts have explored incorporating foundation models (e.g., DINOv2 and SAM) into FSS. However, existing pixel-wise matching approaches based on DINOv2 features often rely solely on foreground-constrained modeling, resulting in less discriminative similarity responses and suboptimal dense prompts. Moreover, relying solely on dense prompts remains inadequate for effectively guiding the Segment Anything Model (SAM) family in constraining object boundaries and low-confidence regions. To address these challenges, we propose EasyFSS, a SAM2-based few-shot segmentation framework that explicitly enhances dense prompts quality while introducing complementary sparse prompts. Specifically, based on multi-level DINOv2 features, we construct a Dual-Polarity Hierarchical Block (DPHB), extending foreground-constrained modeling to foreground–background polarity matching across semantic hierarchies. This design significantly enhances the discriminability of pixel-wise similarity responses and alleviates background-induced ambiguity. Furthermore, we design a Sparse Prompt Generation (SPG) module that leverages global prototypes and multi-scale spatial context to generate high-confidence sparse prompts, effectively constraining boundaries and low-confidence regions. Finally, dense and sparse prompts collaboratively guide the SAM2 decoding process, enabling accurate and reliable segmentation predictions. Extensive qualitative and quantitative experiments demonstrate that EasyFSS achieves state-of-the-art performance, with 1-shot mIoU improvements of 3.2% and 3.9% on PASCAL-5^i and COCO-20^i, respectively. Our code is publicly available at https://github.com/stoneoceam/EasyFSS. Energy-Efficient ECG Classification via Event-Based Sampling and In-Memory Computing Varun Raghuraman, Ryan Brinson, Rhimesh Lwagun, and Xiaoyue Hu (Kennesaw State University); Weiling Li (Dongguan University of Technology); and Yan Fang (Kennesaw State University, Georgia Institute of Technology) Abstract Abstract This paper presents an energy-efficient event-driven framework for cardiac arrhythmia classification that combines Level-Crossing Analog-to-Spike Converter (LC-ASC) sensing with a lightweight convolutional neural network (CNN) mapped onto a resistive compute-in-memory (CIM) system-on-chip. The proposed approach transforms the sparse, asynchronous output of the LC-ASC into compact spatio-temporal spike-frame representations, enabling effective feature extraction using standard CNN architectures while significantly reducing data volume. Evaluated on the MIT-BIH Arrhythmia Database, the proposed event-driven CNN achieves high classification accuracy and strong robustness to class imbalance, outperforming Nyquist-sampled CNN baselines in Macro-F1 score while requiring substantially fewer computations. To further improve efficiency, the network is deployed on an RRAM-based CIM processor comprising 288 parallel crossbar arrays, enabling highly parallel vector–matrix multiplication with minimal data movement. Hardware evaluation demonstrates sub-millisecond end-to-end latency per heartbeat and microjoule-level energy consumption, achieving orders-of-magnitude improvements compared to both desktop and edge GPU platforms. These results validate the proposed framework as a practical, real-time, and energy-efficient solution for near-sensor intelligence in wearable ECG monitoring systems. Reservoir Computing Combined with Optimizable Quantile-Based Subsampling for Small-Sample Time-Series Classification Ziqiang Li (Nagoya Institute of Technology); Yun Liu (Kyoto Institute of Technology, Institute of Science Tokyo); and Gouhei Tanaka (Nagoya Institute of Technology, International Research Center for Neurointelligence) Abstract Abstract Rapidly training reliable machine learning models from limited data remains highly desirable in real-world time-series classification (TSC) applications. Reservoir Computing (RC), originating from dynamical system modeling and characterized by fixed high-dimensional nonlinear projections with lightweight training, has emerged as a promising paradigm for small-sample TSC. In RC-based TSC models, a significant challenge lies in how the classifier can effectively utilize the rich temporal representations generated by the reservoir to produce accurate classification results. In this paper, we propose a novel subsampling strategy for reservoir states, termed Quantile-based Positive Proportion (QPP). Unlike existing subsampling strategies, QPP transforms continuous reservoir representations into binary representations using optimizable quantile thresholds and constructs features by computing the proportion of positive activations. Experiments are conducted on 13 small-sample TSC benchmark datasets from the UCR archive using two representative RC models. Comparative results against existing subsampling strategies show that the proposed QPP consistently achieves superior average rankings across datasets, which highlights the high effectiveness and generalization capability of QPP with respect to RC models in dealing with small-sample TSC tasks. Set Classification with Probabilistic Learning Vector Quantization on the Manifold Mohammad Mohammadi (Tilburg University), Sreejita Ghosh (University Medical Center Groningen), and Çiçek Güven (Tilburg University) Abstract Abstract Prototype-Based Networks (PBNs) provide intrinsic interpretability in image classification by learning characteristic traits of objects as evidence to support predictions. While most existing PBNs operate in Euclidean embedding spaces, recent work shows that representing images as linear subspaces provides robust and interpretable models. However, current subspace-based prototype methods are deterministic and do not explicitly account for (i) the manifold’s (in this case Grassmann) geometry during optimization; (ii) the uncertainty in the model's decisions. From Spatiotemporal Feature Alignment to Transferable Causal Discovery: Process Modeling for Interpretable Analysis of Hot-Rolled Surface Quality Mechanisms Yiquan An (University of Science and Technology Beijing, KU Leuven); Yihang Li (University of Science and Technology Beijing); Johannes De Smedt (KU Leuven); Xi Sun (Chinese Academy of Sciences); and Huaping Chen and Zhimin Lv (University of Science and Technology Beijing) Abstract Abstract To enable rigorous, interpretable analysis of hot-rolling surface-quality mechanisms, we propose an end-to-end framework for spatiotemporal registration and transferable causal discovery. It targets two central bottlenecks in industrial data-driven modeling: physically-inconsistent representations and regime-sensitive causal discovery under distribution shift. First, a physics-driven nonuniform spatiotemporal warping operator maps heterogeneous time series to a unified spatial domain, resolving spatiotemporal mismatches. Second, the Transferable Industrial Causal Engine (TiCE) formulates a progressive bilevel training paradigm that embeds mechanistic constraints and jointly activates an Independence-Constrained Domain-Invariant regularization with a Hierarchically Adaptive Fusion Weighting mechanism to resist structural drift under mixed regimes. Third, learned causal graphs guide Asymmetric SHAP attributions to preserve causal consistency in mechanistic explanations. Benchmarks and ablations on TE and UF-IMD demonstrate robustness to distributional shift and the necessity of modular synergy. In hot rolling, the framework reveals a binary antagonistic mechanism for surface defects: temperature-dominated oxidation kinetics versus multi-stage descaling suppression, thereby providing credible decision support for root-cause diagnosis in complex processes. Structure- and Style-Preserving Diffusion for High-Fidelity Underwater Crack Generation Haoyang Li, Yongsheng Ou, Kunmo Li, and Mingyang Liu (Dalian University of Technology) Abstract Abstract Bridge structure monitor relies on precise crack detection, yet current deep learning surveillance methods rely heavily on a large number of crack images as dataset, which are difficult to collect in complex and dangerous underwater environments. To alleviate this problem, we propose a training-free diffusion framework that enables efficient and high-quality underwater crack generation. Specifically, we first apply the denoising diffusion implicit models inversion to extract crack and underwater representations. We propose a structure preservation module that leverages global and local information from crack representations to enable the preservation of fine-grained crack details, mitigating structural loss. In parallel, we propose a style integration module to integrate underwater domain information and prevent content leakage issue. To better align the underwater domain style, we design a color adaptation module that dynamically performs adaptive instance normalization, which adjusts the color distribution of the translated result and further optimizes the visual realism of underwater crack images. Extensive experiments show that the proposed framework achieves a significant improvement on generating high fidelity underwater cracks. Code is available at https://github.com/LiHaoyang0517/High-Fidelity-Underwater-Crack-Generation The Effect of Pruning Order: A Case Study on CNN Filter Pruning Lorenzo Nikiforos, Luciano Prono, Fabio Pareschi, and Gianluca Setti (DET, Politecnico di Torino) Abstract Abstract Progressive structured pruning is a widely adopted strategy for compressing convolutional neural networks, where model capacity is gradually reduced through multiple prune–fine-tune stages. While most existing methods focus on defining effective importance criteria for convolutional filters, the order in which they are removed is typically fixed and the effect of pruning order is largely unexplored. In this work, we investigate different pruning order strategies for structured pruning of convolutional filters. Specifically, we introduce two alternative scheduling policies, tail-based and center-based, that modify the order of filter removal while preserving the same final pruning budget. Unlike classic approaches, which remove the least important filters first, the proposed scheduling strategies modify the pruning order of the components. In this work, we demonstrate the significant impact of pruning scheduling through extensive experiments with ResNet and ResNeXt architectures on CIFAR-100 and ImageNet-1K datasets. In particular, center-based scheduling consistently outperforms the standard progressive strategy, with larger gains at higher pruning rates and on more challenging datasets. Additional analyses show that these improvements are robust with a varying number of pruning stages and importance criteria. These results highlight how changing pruning scheduling is a simple yet effective method to improve structured pruning performance. Multi-Task Visual Perception Network with LLM Conditioning for Autonomous Navigation Praveen Kumar, K.R. Guruprasad, and Tushar Sandhan (IIT Kanpur) Abstract Abstract Long-term navigation for service robots faces critical challenges like the accumulation of odometry drift and sensor error, which progressively degrade 2D maps and renders traditional path planning algorithms (e.g., A*, RRT*, DiPPer, ViT-A*) ineffective over time. To address this, we propose a user-friendly, interactive framework that eliminates the reliance on globally consistent maps. Our approach integrates visual perception with Large Language Models (LLM) to interpret user commands via text or voice. Instead of relying on a drift-prone global map, the system generates a sequential action plan based on local visual cues and egocentric geometric instructions. These action plans are executed sequentially, allowing the robot to navigate known and unknown environments safely. By resetting localization relative to immediate targets, our framework effectively works with a minimum accumulation drift strategy, ensuring accurate, efficient, and collision-free navigation without the maintenance overhead of traditional mapping. Experiments on real-world and simulated data have shown significant improvements over other methods. Our code and dataset will be made public. LMABC: A Cost-Efficient Large Language Model Multi-Agent Framework for Automated Sequence Labeling in Building Codes Xingchen Hu (School of Computing and Information Technology, University of Wollongong); Zhenjun Ma (Sustainable Buildings Research Centre, University of Wollongong); Binbin Yong (School of Information Science and Engineering, Lanzhou University); and Jun Shen (School of Computing and Information Technology, University of Wollongong) Abstract Abstract Automated sequence labeling of regulatory documents is crucial for automated compliance checking (ACC) and domain knowledge graph (KG) creation, yet is hindered by a ``high-cost, low-resource'' dilemma. Public, high-quality labeled datasets for building codes are scarce and difficult to reuse across regions, making large-scale annotation expensive. Although large language models (LLMs) show strong few-shot capabilities, multi-round, clause-by-clause annotation with high-performance APIs remains prohibitively costly, whereas affordable models are less reliable. To resolve the cost-performance dilemma, we propose LMABC, a low-cost, low-resource LLM Multi-Agent framework for Automated Sequence Labeling in the Building Codes. Rather than relying on a single powerful model, LMABC coordinates multiple diverse, affordable LLM agents to achieve high-quality annotation through collective intelligence. The framework couples a task allocation mechanism with a two-phase labeling workflow consisting of parallel labeling and chain-based transfer labeling. Final results are obtained by a hybrid consensus ensemble that combines majority voting with dynamic weighting to support downstream KG creation. Experiments on building-code and general-domain benchmarks demonstrate that LMABC approaches high-performance models at a fraction of the cost, and achieves higher accuracy with fewer tokens than state-of-the-art (SOTA) methods on general tasks. QScheduler: Adaptive Gradient Sampling for Zeroth-Order On-Device Training on INT8 NPUs Victor Felipe Domingues do Amaral (STMicroelectronics; GeePs, University Paris-Saclay); Pierre Demaj (STMicroelectronics); Erwan Libessart (GeePs, University Paris-Saclay); Laurent Folliot (STMicroelectronics); and Anthony Kolar and Philippe Bénabès (GeePs, University Paris-Saclay) Abstract Abstract Zeroth-Order (ZO) optimization enables On-Device Learning (ODL) on NPU-equipped microcontrollers by estimating gradients through forward passes alone, bypassing the need for backpropagation primitives and reducing memory requirements. The number of gradient samples q critically affects training: insufficient samples produce noisy gradients that plateau early, while excessive samples consume more computational resources. However, finding an optimal q typically requires costly hyperparameter searches. This work introduces QScheduler, an adaptive algorithm that adjusts q based on training progress, and provides the first proof-of-concept of INT8 quantized on-device training on the STM32N6's Neural-ART NPU. Experiments on EuroSAT and STL-10 show that QScheduler matches well-tuned fixed-q configurations for both ResNet18 and MobileNetV2, without requiring prior q hyperparameter optimization. A PennyLane-Centric Dataset to Enhance LLM-based Quantum Code Generation using RAG Abdul Basit, Nouhaila Innan, Muhammad Haider Asif, Minghao Shao, Muhammad Kashif, Alberto Marchisio, and Muhammad Shafique (New York University (NYU) Abu Dhabi, UAE) Abstract Abstract Large Language Models (LLMs) offer powerful capabilities in code generation, natural language understanding, and domain-specific reasoning. Their application to quantum software development remains limited, in part because of the lack of high-quality datasets both for LLM training and as dependable knowledge sources. To bridge this gap, we introduce PennyLang, an off-the-shelf, high-quality dataset of 3,347 PennyLane-specific quantum code samples with contextual descriptions, curated from textbooks, official documentation, and open-source repositories. Our contributions are threefold: (1) the creation and open-source release of PennyLang, a purpose-built dataset for quantum programming with PennyLane; (2) a framework for automated quantum code dataset construction that systematizes curation, annotation, and formatting to maximize downstream LLM usability; and (3) a baseline evaluation of the dataset across multiple open-source and commercial models, including ablation studies, all conducted within a retrieval-augmented generation (RAG) pipeline. Using PennyLang with RAG substantially improves performance: for example, Qwen 7B’s success rate rises from 8.7% without retrieval to 41.7% with full-context augmentation, and LLaMa 4 improves from 78.8% to 84.8%, while also reducing hallucinations and enhancing quantum code correctness. Moving beyond Qiskit-focused studies, we bring LLM-based tools and reproducible methods to PennyLane for advancing AI-assisted quantum development. O2O-TP: An Offline-to-Online Transformer-Based Reinforcement Learning Approach for Portfolio Management Yuming Liao, Yuanxin Wei, Dan Huang, and Yutong Lu (School of Computer Science and Engineering, Sun Yat-Sen University) Abstract Abstract Portfolio management allocates wealth across assets to maximize returns. Reinforcement learning (RL) is promising, but it must balance risk-controlled decisions with rapid adaptation to changing markets. Offline RL trains reliably on fixed historical data using supervised-style objectives, yet its performance is limited by dataset quality and coverage. Online RL can adapt through exploration, but it requires costly real-time interaction with the environment. We introduce O2O-TP, an offline-to-online Transformer RL framework that pretrains on historical data and then refines online with low computational overhead. O2O-TP embeds a Dirichlet policy head to produce simplex-constrained portfolio weights, enabling deterministic or stochastic execution. Across four stock markets, O2O-TP achieves cumulative wealth up to 3.3 times higher than established finance and RL baselines, with low variance between random seeds. Online refinement keeps efficient, accounting for only 14% to 23% of total runtime. These results indicate that the proposed framework enables effective and computationally efficient adaptation in dynamic financial markets. RPMAT: Role-Guided Prioritized Multi-Agent Transformer for Dynamic Action Ordering Zixin Liu (The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences, China); Zhiming Zhou (The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences); Xiaohui Sun and Jiebo Chen (Beijing Institute of Electrical Engineering, Beijing 100854, China); and Wanjun Jing and Zhen Liu (The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences, China) Abstract Abstract Sequential decision making paradigms have recently emerged in Multi-Agent Reinforcement Learning (MARL) to address the coordination limitations of simultaneous execution. However, existing approaches typically employ fixed or observation-driven heuristic orderings, overlooking the intrinsic dependency of decision priority on agent roles. In this paper, we propose the Role-Guided Prioritized Multi-Agent Transformer (RPMAT), a novel framework that optimizes action generation order by leveraging explicit role semantics. Specifically, we design a Role-Guided Cross-Attention (RGCA) mechanism that queries global observation features using learned latent role representations to dynamically quantify the decision urgency of each agent. This mechanism allows high-contribution roles to preemptively broadcast their intentions, thereby guiding the team's sequential execution. Theoretically, we reinterpret this order optimization as anterior structural credit assignment, where agents are ranked based on their predicted potential contribution to the global value function. Extensive experiments on SMAC, Google Research Football, and Multi-Agent MuJoCo benchmarks demonstrate that RPMAT achieves consistent superiority over state-of-the-art baselines in both heterogeneous and homogeneous scenarios, validating the effectiveness of role-driven coordination. Distributed On-Orbit Privacy Preservation for Satellite Imagery via Lightweight Neural Inpainting Wenbo Wu, Cheng Tan, Zhishu Shen, and Di Huang (Wuhan University of Technology) and Qiushi Zheng and Jiong Jin (Swinburne University of Technology) Abstract Abstract Recent advances in Earth observation satellites have enabled high-resolution imagery for Internet of Remote Things (IoRT) applications, but such imagery often contains privacy-sensitive information that may be exposed during transmission or processing. Addressing privacy protection directly on orbit is challenging due to the strict computational, energy, and communication constraints of Low Earth Orbit (LEO) satellite systems and their limited ground-station contact windows. In this paper, we propose a Distributed On-orbit Privacy Protection framework DOPP that integrates lightweight neural detection and inpainting with inter-satellite collaborative computing. Each satellite employs a quantized YOLO-based detection model to identify privacy-sensitive regions and generate compact binary masks, which enable non-overlapping image slicing and parallel inpainting across neighboring satellites via inter-satellite links. A fine-tuned Mobile Inpainting Generative Adversarial Network (Mi-GAN) reconstructs masked regions with low computational overhead while preserving visual consistency. To ensure timely downlink transmission, a visibility-aware aggregation strategy dynamically selects the satellite with the most imminent ground-station access to assemble and transmit the processed imagery. Simulations on real-world remote sensing datasets demonstrate that DOPP achieves effective privacy preservation with high accuracy, while significantly reducing processing latency and energy consumption , resulting in a 15–20% lower Energy-Delay Product (EDP) compared with existing onboard cooperative baselines. These results highlight the feasibility of lightweight neural privacy protection for real-time, resource-constrained satellite imaging systems. SARE: Structure-Aware Reliability Evaluation and Processing for Event-Based Neural Systems Haiyu Li and Charith Abhayaratne (The University of Sheffield) Abstract Abstract Event-based neuromorphic sensors generate sparse and asynchronous streams whose reliability as neural inputs depends not only on global event statistics, but critically on the preservation of underlying spatio–temporal structure. In practice, noise corruption, localized signal loss, or aggressive preprocessing may disrupt structural continuity while leaving global statistics largely unchanged, yielding inputs that appear statistically plau- sible yet are structurally invalid. Such mismatches can lead to abrupt performance collapse or silent failure in neural systems. We propose a structure-preserving local–global graph processing framework for reliability diagnosis in event-based neural systems. Rather than serving as a denoising or enhancement module, the proposed processing stage is explicitly formulated as a di- agnostic intervention operator that applies controlled, structure- aware graph transformations to probe neural sensitivity under progressive structural degradation. The method jointly captures local neighborhood coherence and global geometric organization, enabling targeted structural interventions without directly opti- mizing task accuracy. To support systematic analysis, we introduce a structure- centric evaluation protocol based on geometry-driven objects, multi-scale motion settings, and standardized degradation mech- anisms, explicitly decoupling structural corruption from global event statistics. Experimental results demonstrate that structural integrity degrades significantly earlier than global statistical indi- cators and serves as a reliable predictor of abrupt failure in both spiking and non-spiking models. Compared with purely local or purely global strategies, the proposed local–global formulation exhibits superior robustness under severe structural degradation, revealing failure modes that remain invisible to conventional preprocessing and task-driven evaluation. Efficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems Sungmook Woo, Hyunku Kang, and Chanwoo Kim (Korea University) Abstract Abstract Punctuation restoration improves ASR (AutomaticSpeech Recognition) readability, however streaming ASR requires online decisions with limited future context. In streaming ASR, the system predicts punctuation incrementally, which makes generation-based approaches prone to latency and alignment failures under boundary-wise evaluation. This paper proposes a non-autoregressive scoring method (no free-form generation) that preserves the input transcript and makes a decision at each word boundary. Our method compares punctuation insertion hypotheses against a no-insertion baseline under a bounded K-subword-token lookahead, and calibrates decisions using a weight α and a validation-calibrated threshold τ (no parameter updates during inference). On IWSLT 2017, our scoring method achieves a 4-class macro F1 of 0.893 in the no-fine-tuning setting (validation-calibrated, K=2) and 0.937 after fine-tuning (K=2), outperforming the prompt-based baseline (0.566) and a fine-tuned ELECTRA baseline (0.913) under the same lookahead budget. We analyze the impact of the lookahead budget through ablation studies on K. A Neural Framework to Forecast Short-Term Waves for a Digital Twin of the Ocean Bernardo M. Chagas, Miguel S. E. Martins, João M. C. Sousa, and Susana Vieira (Instituto Superior Técnico) Abstract Abstract Accurate short-term wave forecasting is essential for ocean monitoring and for supporting stakeholder decision-making. This paper proposes a neural framework for one-hour-ahead prediction of significant wave height, mean wave period, and mean wave direction using the ERA5 dataset. The proposed approach uses a local 3 × 3 spatial grid to capture the essential neighbourhood dynamics of the wave and wind fields. The information from this grid, sampled over the two previous time steps, forms a 108-dimensional input vector that represents local spatio-temporal interactions. An autoencoder is then used to compress this representation into a 32-dimensional latent space, which is subsequently mapped to the target variables using a multilayer perceptron. The performance of the model is evaluated in a spatially disjoint test region. The results show high predictive accuracy, with coefficients of determination above 0.9959 for all target variables, and low absolute errors. The proposed framework provides a computationally efficient alternative for short-term wave forecasting suitable for integration into a digital twin of the ocean. Efficient Indoor Robotic Scene Analysis with Encoder-Centric Vision Transformers Söhnke Benedikt Fischedick, Benedict Stephan, Daniel Seichter, and Horst-Michael Gross (Ilmenau University of Technology) Abstract Abstract Indoor mobile robots require a comprehensive understanding of their surroundings to navigate and interact safely in cluttered environments. This typically involves multiple complementary perception tasks, such as scene classification, semantic segmentation, and panoptic segmentation, under real-time constraints. While recent mask-based segmentation methods already shift instance separation into the model, many existing approaches still remain computationally expensive and have mainly been studied in RGB-only settings. In this paper, we propose IRSAFormer, a Transformer-based approach for indoor scene analysis that follows the efficient Encoder-only Mask Transformer (EoMT) paradigm by injecting query tokens into a largely unmodified Vision Transformer (ViT) backbone and using lightweight task heads directly on top of it. To address the limitations of RGB-only perception, we analyze RGB-D incorporation strategies. Our results indicate that depth provides clear gains when pre-training is limited, whereas strong large-scale pre-training narrows the gap between RGB and RGB-D. Beyond closed-set prediction, we investigate distilling text-aligned embeddings as a more flexible output representation for language-driven downstream applications. We find that scene-level embeddings distill reliably, whereas query token-level embeddings remain more challenging. Overall, IRSAFormer achieves state-of-the-art performance on NYUv2, SUN RGB-D, and ScanNet V2 while maintaining high efficiency, reaching 26.8 FPS on an NVIDIA Jetson AGX Orin with a ViT-Base backbone. KD-MARL: Resource-Aware Knowledge Distillation in Multi-Agent Reinforcement Learning Monirul Islam Pavel and Muhammad Anwar Ma'sum (Adelaide University); Siyi Hu (Curtin University); Mahardhika Pratama and Zehong Jimmy Cao (Adelaide University); and Ryszard Kowalczyk (Adelaide University; Systems Research Institute, Polish Academy of Sciences) Abstract Abstract Real-world deployment of multi-agent reinforcement learning (MARL) systems is fundamentally constrained by limited compute, memory, and inference time. While expert policies achieve high performance, they rely on costly decision cycles and large-scale models that are impractical for edge devices or embedded platforms. Knowledge distillation (KD) offers a promising path toward resource-aware execution, but existing KD methods in MARL focus narrowly on action imitation, often neglecting coordination structure and assuming uniform agent capabilities. We propose resource-aware Knowledge Distillation for Multi-Agent Reinforcement Learning (KD-MARL), a two-stage framework that transfers coordinated behavior from a centralized expert to lightweight, decentralized student agents. The student policies are trained without critic, relying instead on distilled advantage signals and structured policy supervision to preserve coordination under heterogeneous and limited observations. Our approach transfers both action-level behavior and structural coordination patterns from expert policies with supporting heterogeneous student architectures, allowing each agent's model capacity to match its observation complexity, which is crucial for efficient execution under partial or limited observability along with limited onboard resources. Extensive experiments on SMAC and MPE benchmarks demonstrate that KD-MARL achieves high performance retention while substantially reducing computational cost, retaining over 90% of expert performance while reducing computational cost by up to 28.6× FLOPs. The proposed approach preserves expert-level coordination through structured distillation, enabling practical MARL deployment across resource-constrained onboard platforms. A hybrid framework for multimodal feature learning in chromatic aberration detection Jaroslaw Bernacki (Wroclaw University of Science and Technology) and Rafal Scherer (AGH University of Krakow) Abstract Abstract Chromatic aberration (CA) remains a common optical artifact in digital imaging, causing spatial misalignments between color channels and resulting in visible color fringing. In this paper, we propose TriadNet, a novel hybrid convolutional-transformer architecture that formulates CA detection as a multimodal feature learning problem for accurate analysis of real-world photographs. TriadNet integrates three complementary modality-specific pathways: an edge-centric pathway that captures structural, high-frequency information where CA manifests most prominently, a channel-aware attention mechanism that models chromatic inter-channel discrepancies, and an edge-modulated feature processing strategy that facilitates effective multimodal feature fusion. Experimental evaluation on a custom dataset of real-world images demonstrates that TriadNet outperforms several baselines, achieving $F_1$-scores exceeding 0.98 and robust generalization across varying aberration strengths. The proposed method offers a lightweight yet effective solution for automatic CA detection within modern image processing pipelines. Visuomotor Control for Single Task Robotic Manipulation via Text-Conditioned Flow Matching Somdeb Saha, Vighnesh Vatsal, and Kaushik Das (Tata Consultancy Services) Abstract Abstract This paper introduces Text-Conditioned Flow Matching Policy (TC-FMP), a continuous-time generative imitation learning framework that leverages natural language as a lightweight yet effective guidance signal for visuomotor control. TC-FMP learns an instruction-conditioned velocity field that guarantees the smooth evolution of trajectories from noise toward the demonstrated action distribution. To mitigate mode collapse and improve conditional fidelity, classifier-free guidance is incorporated by randomly dropping textual conditioning during training. At inference time, guided and unguided vector fields are combined to produce a scaled velocity field, which is integrated using a second-order Heun solver to generate action sequences. The resulting action trajectories are executed in a receding-horizon control scheme, enabling closed-loop execution over long horizons. Language is fused with visual and proprioceptive inputs via FiLM modulation, introducing negligible architectural overhead while allowing natural language prompts to steer task-relevant behaviors. TC-FMP is first evaluated on the planar PushT benchmark, and then to demonstrate its robustness and practical utility; on a long-horizon, contact-rich bimanual peg-in-hole insertion task under varying noise conditions. Experimental results show that increasingly specific language instructions consistently improve both average and best-case returns compared to diffusion-based and flow-matching baselines. These findings highlight that even single-task visuomotor control can substantially benefit from language-driven guidance, improving robustness, controllability, and performance. CMInet: Automated Assessment of Physiological Cerebellar Mobility in Upright and Supine MRI Jiamin Wang (Hong Kong University of Science and Technology); Shiying Ke, Fuhua Jia, and Xiaoying Yang (University of Nottingham Ningbo China); Kun Yan (Department of Radiology, Ningbo No. 2 Hospital); and Tianxiang Cui, Yulin Wang, and Chengbo Wang (University of Nottingham Ningbo China) Abstract Abstract Chiari type I malformation (CM-I) is characterized by the caudal herniation of cerebellar tonsils through the foramen magnum. Diagnosis traditionally relies on measuring tonsillar position relative to the McRae line, a manual process prone to inter-observer variability. While AI has advanced skeletal landmark prediction in CT, applying these methods to MRI remains challenging due to low osseous contrast, creating a dependency on radiation-based imaging. In this work, we present CMInet, the first fully MRI-based framework integrating cerebellar segmentation with direct osseous landmark regression. This approach eliminates the need for supplementary CT scans, streamlining clinical workflows. Leveraging a rare, hardware-limited paired dataset from specialized upright MRI, our model effectively captures subtle gravity-dependent structural changes. Our framework delivers an objective tool for anatomical assessment, reducing manual workload while establishing a critical normative baseline to distinguish physiological mobility from pathological herniation. Federated Learning for Binary Neural Networks with Secure Aggregation Yinuo Wang, Zheng Huang, Jie Guo, and Weidong Qiu (Shanghai Jiao Tong University) Abstract Abstract Federated learning (FL) faces key challenges in communication efficiency and privacy protection. Most existing studies primarily address one of these aspects, while comprehensive security mechanisms often conflict with efficiency requirements. Binary neural networks (BNNs) can substantially reduce computation and communication costs in FL, but aggressive binarization typically leads to severe performance degradation. In this work, we propose BiSec-FL, a secure and efficient FL framework that integrates binary aggregation with homomorphic encryption (HE). Specifically, we introduce Scale and Residual Compensated Binary Aggregation (SR-BinAgg) to mitigate quantization loss in binary aggregation, and adopt a selective secure aggregation scheme based on multi-key HE to protect full-precision parameters. Experiments show that BiSec-FL achieves stable convergence across multiple datasets and settings, reducing single-round communication costs by over 96% on average compared with FedAvg, while incurring significantly lower cryptographic overhead than fully encrypted FL systems. Improving the Robustness of Control of Chaotic Convective Flows with Domain-Informed Reinforcement Learning Michiel Straat and Thorben Markmann (Bielefeld University), Sebastian Peitz (TU Dortmund), and Barbara Hammer (Bielefeld University) Abstract Abstract Chaotic convective flows arise in many real-world systems, such as microfluidic devices and chemical reactors. Stabilizing these flows is highly desirable but remains challenging, particularly in chaotic regimes where conventional control methods often fail. In this work, we improve the practical feasibility of controlling chaotic flows by Reinforcement Learning (RL), focusing on Rayleigh-Bénard Convection (RBC), a canonical model for convective heat transport. In particular, our goal is to robustly reduce the convective heat transfer. Although RL has shown promise for control in laminar flow settings, its ability to generalize and remain robust under chaotic and turbulent dynamics is not well explored, despite being critical for real-world deployment. Hence, to enhance generalization and sample efficiency, we introduce domain-informed RL agents that are trained using Proximal Policy Optimization across diverse initial conditions and flow regimes. We incorporate domain knowledge in the reward function via a term that encourages Bénard cell merging, as an example of a desirable macroscopic property. In laminar flow regimes, the domain-informed RL agents reduce convective heat transport by up to 33%. In chaotic flow regimes, the agents still achieve a significant 10% reduction, which is markedly better than the conventional controllers used in practice. We compare the domain-informed to uninformed agents: Our results show that the domain-informed reward design results in steady flows, faster convergence during training, and generalization across flow regimes without retraining. Our work demonstrates that elegant domain-informed priors can greatly enhance the robustness of RL-based control of chaotic flows, bringing real-world deployment closer. CodeFORGE: Framework for Orchestrated Rare Programming Language Generation and Evaluation Zhuoyan Yu (NYU Shanghai); Minghao Shao (NYU Tandon School of Engineering, NYU Abu Dhabi); Yijun Shen and Yichen Zhao (NYU Shanghai); Boyuan Chen (NYU Tandon School of Engineering, NYU Abu Dhabi); Yik-Cheung Tam (NYU Shanghai); and Muhammad Shafique (NYU Abu Dhabi) Abstract Abstract Large language models (LLMs) have shown strong code generation ability, yet most benchmarks concentrate on high-resource languages such as Python, Java, and C++, leaving a gap in evaluating LLM performance on low-resource and legacy languages. Languages including COBOL, Fortran, Ada, and Prolog remain in financial infrastructure, aerospace systems, scientific computing, and symbolic reasoning, but are supported by far less training data and lack widely accepted benchmarks. To fill this gap, we present CodeFORGE, an automated benchmark synthesis framework tailored to data-scarce programming languages. Our method uses LLMs to generate programming tasks together with corresponding test suites through a designed pipeline. CodeFORGE adopts a model-cascading verification strategy: multiple LLMs attempt each generated task in sequence, leveraging iterative feedback from earlier failures to enable cross-validation and ensure benchmark reliability. Only tasks that satisfy automated checks inside isolated Docker environments are included in the final benchmark. With this framework, we build CodeFORGE-Bench, a comprehensive benchmark spanning 12 low-resource languages across diverse paradigms, including procedural, functional, and logic programming. Extensive experiments evaluating state-of-the-art LLMs reveal performance gaps relative to high-resource languages, offering insights into current limitations when LLMs face linguistically diverse and data-scarce code generation settings. Hybrid Deep Learning for Traceability and Classification of Industrial Slate Tiles Soren Antebi, Stefan Eickeler, and Sandra Halscheidt (Fraunhofer IAIS); Rene Schmitz and Michael Müllers (Rathscheck Schiefer und Dachsysteme); and Dirk Hecker and Rafet Sifa (Fraunhofer IAIS) Abstract Abstract Applying deep learning to instance-aware re- identification of slate tiles and extraction site classification can improve production efficiency and quality control in the slate tile industry. These tasks are particularly important for handling natural materials where visual variability can make manual inspection costly and error-prone. We present a lightweight, hybrid deep learning approach that combines image matching and classification within a single framework. The system integrates a feature-matching branch based on XFeat with a MobileNetV3- based classification branch. The XFeat branch, combined with a LightGlue matching head, improves instance matching performance by +15.4% AUC. For classification, features from both backbones are shared and fused, resulting in a +10.9% accuracy improvement over a standard MobileNetV3 model. Our approach is evaluated on a newly created industrial dataset consisting of 2,610 slate tile images from six extraction sites. The results demonstrate the effectiveness of the proposed approach for object re-identification and classification in an industrial setting. PATH-NM: Learning Orthogonal N:M Sparse Patterns via Differentiable Transformation for Efficient Heterogeneous MARL Zhenfeng Su, Zhen Liu, and Zhiming Zhou (CASIA) Abstract Abstract Semi-structured $N\!:\!M$ sparsity enables hardware acceleration on modern GPUs, but its use in heterogeneous multi-agent reinforcement learning (MARL) remains largely unexplored. Existing parameter-sharing methods either ignore hardware constraints or encourage policy homogenization, while strict $N\!:\!M$ sparsity introduces additional challenges due to non-stationary activations, discrete constraints, and agent heterogeneity. We propose PATH-NM (Pattern-Aware Transformation for Hardware-Efficient $N\!:\!M$ Sparse MARL), a sparsity-aware framework that learns agent-specific $N\!:\!M$ sparse subnetworks on top of a shared backbone. PATH-NM reformulates pruning as pattern distribution learning and combines dynamic activation-aware scoring, differentiable pattern transformation, and orthogonal diversity regularization. This design supports end-to-end optimization, adapts to non-stationary RL dynamics, and explicitly promotes structural diversity across agents. Experiments on MPE and SMACv2 show that PATH-NM consistently outperforms representative parameter-sharing baselines in both sample efficiency and final performance, while remaining fully compatible with hardware-accelerated $N\!:\!M$ sparse inference. Multi-stage Dynamic Selection for Cross-Project Defect Prediction Juscimara Gomes Avelino, Juscelino Sebastião Avelino Júnior, and George Darmiton da Cunha Cavalcanti (Universidade Federal de Pernambuco) and Rafael Menelau Oliveira Cruz (École de Technologie Supérieure) Abstract Abstract Cross-Project Defect Prediction (CPDP) involves building models using data from external projects, called training projects, to predict modules from the target project. However, traditional CPDP methods suffer from the distribution shift between training and target projects that affects the model's performance. This paper proposes a novel CPDP framework that addresses this issue by proposing a two-stage multiple classifier system (MCS) selection scheme: one working at the project level and another at the module level. In the first stage, the framework evaluates multiple possible MCS configurations to find one that covers and generalizes well across multiple training projects. Consequently, the proposal is likely to obtain a diverse set of classifiers, each specialized in tackling software modules with distinct characteristics. The second selection stage operates at test time, selecting the most competent classifiers to predict each new module in the target project. Unlike previous approaches that apply the same classifiers to the entire target project, the proposed framework performs module-level model selection. This way, the system is more robust to changes in distributions between training and target projects because the selected set of classifiers is module-dependent. Our experimental results using 82 projects from four different CPDP benchmark datasets demonstrate that the proposed approach outperforms the state-of-the-art CPDP methods in most scenarios. The code, dataset, and further details about the proposed method are publicly available at https://github.com/jsaj/Multi_DES. Edge-Cell Graph Encoding and Fixed-Point Arithmetic for FPGA Growing Neural Gas Teuku Zikri Fatahillah, Raditya Artha Rochmanto, Achmad Fahrul Aji, Anhar Risnumawan, and Naoyuki Kubota (Tokyo Metropolitan University) Abstract Abstract We present an FPGA implementation of Growing Neural Gas (GNG) on a Sipeed Tang Nano~9K (Gowin GW1NR-9C) as a bounded-memory hardware substrate for online self-organizing clustering. The design uses an FSM datapath with BRAM-based storage, integer fixed-point arithmetic (Q8/Q16), and a compact \emph{edge-cell graph encoding} that stores the upper-triangular adjacency as 8-bit age+1 values, avoiding dense adjacency matrices and dynamic memory allocation. The complete design is implemented without any soft-core processor or floating-point unit, and is evaluated on Two Moons and Concentric Circles benchmarks. The platform establishes a hardware foundation for deploying continual self-organizing clustering on resource-constrained edge devices. Dual Path Neural CDE for Probabilistic Stepwise Forecasting Shengnan Di, Zhidong Li, Bin Liang, and Fang Chen (University of Technology Sydney) Abstract Abstract Probabilistic forecasting for multivariate stepwise time series is difficult because observations are dominated by long stable regimes, while predictive changes occur sparsely and asynchronously across variables. We propose a dual-path neural controlled differential equation framework that disentangles persistent trend structure from event-driven updates in continuous time. A trend pathway summarizes regime-level dynamics via smoothing, whereas an event pathway sparsifies the primary channel by retaining endpoints and statistically significant change windows to form an irregular event sequence. The two representations are fused by a learned time-varying gate and mapped to predictive distribution parameters. We study Gaussian, Laplace, and Student-t likelihoods and evaluate forecast quality using the continuous ranked probability score together with calibration and sharpness metrics. Experiments show consistent improvements over classical and neural baselines, and indicate that heavy-tailed likelihoods improve distributional accuracy under regime shifts, while calibrated interval reliability remains subject to a sharpness trade-off. Trust Fusion with Choquet-Based Interaction Aggregation in Federated Learning Alicja Rachwał, Albert Rachwał, and Paweł Karczmarek (Lublin University of Technology) Abstract Abstract Federated learning (FL) enables collaborative model training without centralizing raw data, but its performance can degrade under non-identically distributed (non-IID) client data and unreliable client updates. This paper introduces Φ-FL, a server-side aggregation strategy that augments Federated Averaging with trust fusion based on multiple reliability signals computed on a small server proxy set, robust magnitude control via interquartile-range (IQR) norm clipping, and interaction-sensitive reweighting using a 2-additive Choquet-inspired mechanism to promote complementary trusted updates. We evaluate Φ-FL on a suite of benchmark tabular classification datasets under Dirichlet-partitioned non-IID settings, using repeated paired runs. Compared to FedAvg, Φ-FL improves the mean test accuracy from 69.56% to 74.74%, achieving higher mean accuracy on 9 out of 12 datasets. In relation to more advanced baseline model, FLTrust, the proposed method Φ-FL achieves average accuracy higher by 1.7 percentage point, while the median accuracy increases by 3.7 percentage point. The significance of the obtained results is confirmed by statistical tests. Experience Constrained Hierarchical Federated Reinforcement Learning for Large-scale UAV Teams in Hazardous Environments Qinwei Huang and Rui Zuo (Syracuse University), Simon Khan (Air Force Research Laboratory), and Qinru Qiu (Syracuse University) Abstract Abstract Conventional federated learning assumes that greater learner participation improves training performance, by leveraging abundant, independently generated local data. However, in federated reinforcement learning (FRL) for unmanned aerial vehicle (UAV) teams in hazardous environments where experience generation is severely constrained by safety considerations, energy limitations, and mission duration, this assumption may break. This work introduces Experience-Constrained Hierarchical Federated Reinforcement Learning (EC-HFRL), a framework in which clusters act as federated learning agents, while multiple intra-cluster learners represent parallel learning resources that reuse a shared experience pool. We show that increasing participation does not necessarily improve learning performance. Instead, learning performance is strongly associated with experience reuse strategy and the dominance of key analytically identified gradient transition experiences within a cluster. In particular, minibatch size primarily determines effective replay exposure, while higher intra-cluster participation increases reuse level. Empirical results demonstrate that the commonly observed performance regimes are strongly associated with the structure of the learning signal, rather than federated aggregation effects, clarifying the limited and secondary role of learner participation in experience-constrained FRL. Reformulating the Forward-forward Algorithm with Layer-wise Similarity-based Objectives James Gong, Raymond Luo, Emma Wang, and Leon Ge (University of Auckland); Bruce Li (Imperial College London); and Felix Marattukalam and Waleed Abdulla (University of Auckland) Abstract Abstract Backpropagation underpins the success of artificial neural networks, yet it has critical limitations such as backward locking and global error propagation. The Forward-Forward algorithm has been proposed as a more biologically plausible method that replaces the backward pass with an additional forward pass. Nevertheless, the Forward-Forward algorithm significantly trails backpropagation in accuracy, and its optimal form requires multiple forward passes at inference time, resulting in low inference efficiency. In this work, the Forward-Forward algorithm is reformulated within a similarity-learning paradigm. This proposed algorithm is named Forward-forward Algorithm Unified with Similarity-based Training (FAUST). FAUST incorporates similarity-based objectives and, by caching estimates of class centroids, achieves single-pass inference. On matched MLP and CNN architectures, FAUST yields improved accuracy over Forward-Forward baselines and narrows the gap to backpropagation across MNIST, Fashion-MNIST, and CIFAR-10. Applying FAUST to ResNet-18 yields competitive results on ImageNette. AwareNet: A Topography-Informed Spatio-Temporal Network for Motor Imagery EEG Classification Mingjie Gao (Yunnan University/School of Information Science) and Zifeng Yao (Sichuan Normal University/Institute of Brain and Psychological Sciences) Abstract Abstract Accurate decoding of motor imagery electroencephalography (MI-EEG) is fundamental to brain–computer interface applications, yet remains challenging due to low signal-to-noise ratio and amplitude fluctuations. Existing approaches often underutilize spatial relationships and struggle to model long-range temporal dependencies. We propose AwareNet, an end-to-end spatio-temporal EEG decoding framework that employ a Bidirectional Temporal Convolution layer to capture both past and future contextual information, enriching temporal feature representations. Next, we design a lightweight Coordinate-aware Spatial Gating module that leverages prior knowledge of actual electrode positions to model spatial relationships. Furthermore, we introduce the Channel-aware Transformer module to capture global dependencies. CAT is an eeg-friendly attention mechanism which replaces conventional dot-product attention with cosine similarity and incorporates per-head learnable temperature scaling, enabling more stable and adaptive temporal modeling under noisy, small-sample conditions. Extensive experiments demonstrate that AwareNet consistently outperforms state-of-the-art baselines on multiple datasets. Visualization analyses further confirm that AwareNet yields more discriminative and neurophysiologically consistent representations, with our CAT module surpassing conventional MSA in training stability, decoding performance, and interpretability. Overall, AwareNet provides accurate, stable, and interpretable MI-EEG decoding, demonstrating strong potential for practical BCI deployment. CGSSN: A Coupled Graph Spiking Skeleton Network for Human Action Recognition Using Neuromorphic Vision Sensors Lihui Wu and Charith Abhayaratne (The University of Sheffield) Abstract Abstract Abstract—Event cameras provide sparse and asynchronous event streams with high temporal resolution, enabling human action recognition under fast motion and challenging illumina- tion conditions. However, effective modeling of event streams for event-based human action recognition remains challenging due to polarity imbalance, timestamp jitter, and the inherent mismatch between structural reasoning and transient temporal dynamics. To address these challenges, we propose CGSSN, a coupled graph–spiking dual-path framework that jointly models structure-aware spatial relations and spike-driven temporal dy- namics. CGSSN introduces a unified polarity-aware Poisson tem- poral encoding (UPPTE) module to enhance spike stability and robustness against temporal jitter, and adopts parallel graph and spiking pathways to capture complementary structural and tran- sient cues. In addition, a lightweight cross-branch alignment and gated fusion strategy is incorporated to improve representation consistency across heterogeneous pathways. Experimental results on five public event-based human action recognition benchmarks demonstrate that CGSSN consistently outperforms existing CNN- based, Transformer-based, and spiking-based methods, achieving state-of-the-art performance on multiple datasets. Comprehensive Ablation studies further verify the effectiveness of each proposed components. Keep the Teacher Non-Fine-Tuned: Knowledge Distillation for Stable Molecular Dynamics with Neural Network Potentials Naoki Matsumura, Yuta Yoshimoto, Yuto Iwasaki, Meguru Yamazaki, and Yasufumi Sakai (Fujitsu) Abstract Abstract Neural network potentials (NNPs) accelerate molecular dynamics (MD), but stable NNP-MD requires training data with broad configurational coverage, including rare high-energy states. Conventional knowledge distillation pipelines fine-tune a foundation teacher on density functional theory (DFT) and then use the fine-tuned teacher to generate training data via NNP-MD simulations for training a lightweight student. We show this order can harm MD: fine-tuning steepens the teacher potential energy surface, traps trajectories in low-energy basins, and collapses high-energy coverage. We therefore reverse the workflow. We keep the teacher non-fine-tuned as a generalist model to sample diverse NNP-MD configurations labeled with teacher energies and forces, distill a lightweight student on these soft targets, and then fine-tune only the student on a small DFT-labeled subset. On polyethylene glycol (PEG) and Li10GeP2S12 (LGPS) materials using MatterSim as the generalist teacher, the student achieves near-experimental property accuracy while reducing expensive DFT labeling by 10× (PEG) and 5× (LGPS) relative to representative active learning baselines. It also yields large inference gains (e.g., 82× speedup for a 3,100-atom PEG system) and avoids teacher memory bottlenecks. In PEG, We further show that a fine-tuned-teacher pipeline can match static validation error yet fail in MD due to missing high-energy coverage, showing that configurational coverage is as important as static accuracy for NNP‑MD. 2DGS-Room: Seed-Guided 2D Gaussian Splatting with Geometric Constraints for High-Fidelity Indoor Scene Reconstruction Wanting Zhang, Haodong Xiang, Zhichao Liao, Xiansong Lai, Xinghui Li, and Long Zeng (Tsinghua University) Abstract Abstract The reconstruction of indoor scenes is a critical task for embodied robotics, where accurate perception of geometry is essential for navigation, scene understanding, and spatial reasoning. Although recent 3D Gaussian Splatting (3DGS) methods show promising results in novel view synthesis, their performance degrades in textureless regions and complex indoor layouts, resulting in incomplete and noisy reconstructions. In this paper, we introduce 2DGS-Room, a novel method leveraging 2D Gaussian Splatting for high-fidelity indoor scene reconstruction. Specifically, we introduce a seed-guided strategy to control the distribution of 2D Gaussians through dynamic growth and pruning, while depth and normal priors enhance geometry in both detailed and low-texture regions. Furthermore, multi-view consistency constraints improve structural robustness across views. Extensive experiments demonstrate that our method achieves state-of-the-art performance in indoor scene reconstruction. PRISM-Med: A Perception, Routing, Inspection, and Self-Correction Multi-agent Collaboration System for Medical VQA 杰 周 (East China Normal University) Abstract Abstract Medical Visual Question Answering (Med VQA) presents a unique challenge demanding rigorous precision and zero hallucinations. Unlike general VQA tasks, interpreting medical imagery requires rigorous multi-scale feature understanding and multi-step diagnostic reasoning. Existing approaches, predominantly relying on single-agent Multimodal Large Language Models, often suffer from “tunnel vision”, focusing on salient features while missing subtle yet critical anomalies and lack robust self-verification mechanisms. Furthermore, the static inference strategies employed by current models lack the flexibility to adapt to varying question complexities, leading to inefficient resource scheduling. Inspired by how a prism decomposes light to reveal spectral details, we propose PRISM-Med (Perception, Routing, Inspection, Self-correction Multi-agent Collaboration System for Medical VQA), a modular framework. Our system features three key innovations: (1) a Two-Stage Diagnostic Architecture mimicking clinical “Triage-to-Consultation” with Organizer-based dynamic routing; (2) a Clinically Grounded Multi-Agent Framework with four specialized experts to mitigate tunnel vision; (3) a Dynamic Inference Optimization Strategy based on Proximal Policy Optimization (PPO) for adaptive generation configuration. PRISM-Med achieves state-of-the-art performance among open-source models on PMC-VQA, MedXQA, MMMU and OmniMedVQA benchmarks, outperforming baselines by approximately 3.68% on average. The code is available at https://gitee.com/zhou-jie021125/med.git. Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation Ligong Lei, Wenwen Lu, Xudong Pang, Zaokere Kadeer, and Aishan Wumaier (Xinjiang University) Abstract Abstract Automatic speech recognition (ASR) systems often degrade on accented speech because acoustic-phonetic and prosodic shifts induce a mismatch to training data, making labeled accent adaptation costly. Text-centric pseudo-label filters (e.g., perplexity) can favor fluent hypotheses that remain acoustically misaligned; using them as supervision during fine-tuning amplifies errors. We present a reference-free selection pipeline for accent adaptation under a transductive, label-free protocol: unlabeled target speech drives pool reduction and percentile thresholds without oracle transcripts. We optionally apply target-aware FLMI preselection, then decode several hypotheses per utterance and score them with speech–text alignment in a frozen SONAR embedding space and predicted word error rate (WER), treating multimodal consistency as a proxy for pseudo-label quality under accent-induced mismatch. A simple percentile rule selects the training subset. In-domain, selecting ∼1.5k utterances from a 30k pool achieves 10.91% WER, close to 10.45% with 30k supervised labels. With a mismatched cross-domain candidate pool, consistency-filtered subsets avoid degradation from unfiltered pseudo-labels under strong accent shift; matched-hour experiments on a stronger ASR backbone further confirm gains over random sampling and recent selection baselines. CoSTr: Fully Sparse Transformer with Mutual Information for Pragmatic Collaborative Perception Chen Xia, Qianxin Qu, and Ziyi Song (Tsinghua University); Guipeng Zhang (Institute of Computing Technology); and Sheng Zhou and Zhisheng Niu (Tsinghua University) Abstract Abstract Collaborative perception faces critical pragmatic challenges under bandwidth constraints, particularly in guaranteeing perception performance while overcoming spatial-temporal imperfections such as sensor misalignment and asynchrony. Although recent fully sparse approaches have improved computational efficiency, they remain constrained by local receptive fields that limit long-range reasoning, lacking intelligent communication redundancy reduction, and still exhibit limited robustness to real-world noise. To address these challenges, we propose CoSTr, a Collaborative Sparse Transformer framework that operates natively on sparse feature representations through integrated sparse convolution and transformer. Our approach introduces a robust spatial-temporal attention mechanism that explicitly compensates for pose errors and communication delays during feature fusion. To minimize communication overhead, we propose a mutual information-based criterion that operates directly on sparse features to select and transmit only the most critical information. Extensive experiments on OPV2V and V2XSet demonstrate that CoSTr not only achieves state-of-the-art perception performance and communication efficiency, but also significantly improves the robustness against spatial-temporal disturbances, establishing a new benchmark for pragmatic collaborative perception. Divide and Conquer Strategy on Monocular Depth Estimation of Semiconductor SEM Images Hao Sun (School of Microelectronics, University of Science and Technology of China); Peijin Cai (Institute of Advanced Technology, University of Science and Technology of China); and Sihai Zhang (School of Microelectronics, University of Science and Technology of China) Abstract Abstract Single-image scanning electron microscope (SEM) depth estimation plays a critical role in semiconductor three-dimensional metrology. However, existing end-to-end approaches often struggle to maintain stable generalization performance in complex structural scenarios composed of multiple geometric primitives. To address this issue, we propose a Divide-and-Conquer style structured modeling approach for single-image SEM depth estimation. The proposed method explicitly decomposes complex scenes into multiple local structural regions and shares depth prediction parameters across regions, thereby guiding the model to learn reusable local structure--depth mapping relationships and reducing its reliance on global layout statistics. The proposed approach is systematically validated on a simulation-generated SEM dataset. The training data include both basic structural primitives and their compositions forming SRAM-like structures, while the testing phase further evaluates the model’s prediction stability under variations in structural compositions. Experimental results demonstrate that, compared with conventional end-to-end models, the proposed method achieves more stable and reliable depth estimation performance in complex compositional scenarios, validating the effectiveness of structured modeling strategies in improving the generalization capability of single-image SEM depth estimation. Retrieval-Augmented Recommendation Explanation Generation with Hierarchical Aggregation Bangcheng Sun, Yazhe Chen, Jilin Yang, Xiaodong Li, and Hui Li (Xiamen University) Abstract Abstract Explainable Recommender System (ExRec) provides transparency to the recommendation process, increasing users' trust and boosting the operation of online services. With the rise of large language models (LLMs), whose extensive world knowledge and nuanced language understanding enable the generation of human-like, contextually grounded explanations, LLM-powered ExRec has gained great momentum. However, existing LLM-based ExRec models suffer from profile deviation and high retrieval overhead, hindering their deployment. To address these issues, we propose Retrieval-Augmented Recommendation Explanation Generation with Hierarchical Aggregation (REXHA). Specifically, we design a hierarchical aggregation based profiling module that comprehensively considers user and item review information, hierarchically summarizing and constructing holistic profiles. Furthermore, we introduce an efficient retrieval module using two types of pseudo-document queries to retrieve relevant reviews to enhance the generation of recommendation explanations, effectively reducing retrieval latency and improving the recall of relevant reviews. Extensive experiments demonstrate that our method outperforms existing approaches by up to 12.6% w.r.t. the explanation quality while achieving high retrieval efficiency. Geometry-Aligned EEG Scalpograms: Graph Harmonics Enable Accurate, Calibrated Dementia Subtyping Siddhant Ujjain, Ahmad Siraj Hashmi, Pooja Singh, Ekta Srivastava, Sandeep Kumar, and Tapan Kumar Gandhi (IIT Delhi) Abstract Abstract Dementia subtyping from routine scalp EEG could enable low-cost decision support, but deep models often ignore 10–20 electrode geometry and can overfit small clinical cohorts. We introduce graph-harmonic scalpograms, a geometry-aligned representation that projects 19-channel EEG onto low-frequency scalp graph harmonics and converts harmonic activity into compact time-frequency images. Using Open-Neuro ds004504 (N = 88, eyes-closed; AD/FTD/CN), recordings are segmented into 8 s windows (75% overlap), projected onto the first K = 12 non-DC graph modes, transformed over 1–45 Hz using F = 60 log-spaced bins, and encoded as Band-RGB scalpograms (δ+θ, α, β+γ) under strict subject-wise 80/10/10 splits to prevent leakage. Models are trained with class-weighted cross-entropy, geometry-preserving spectro-temporal augmentations, and post-hoc temperature scaling to improve probability calibration; evaluation is performed at the subject level by averaging window posteriors with 95% BCa bootstrap confidence intervals. A lightweight CNN trained from scratch achieves the best subject-level performance (balanced accuracy 0.83 ± 0.07,macro-F1 0.847 ± 0.06, AUROC 0.95 ± 0.05), outperforming transfer baselines such as ResNet-34 (0.65 ± 0.05 /0.667 ± 0.04 / 0.90 ± 0.03). Overall, aligning the model inductive bias with 10–20 sensor geometry supports accurate, well-calibrated subject-level dementia subtyping using lightweight CNNs that are promising for small clinical EEG cohorts. Domain-Specialized Object Detection via Model-Level Mixtures of Experts Svetlana Pavlitska (FZI Research Center for Information Technology), Malte Stüven and Beyza Keskin (Karlsruhe Institute of Technology (KIT)), and J. Marius Zöllner (FZI Research Center for Information Technology) Abstract Abstract Mixture-of-Experts (MoE) models provide a structured approach to combining specialized neural networks and offer greater interpretability than conventional ensembles. While MoEs have been successfully applied to image classification and semantic segmentation, their use in object detection remains limited due to challenges in merging dense and structured predictions. In this work, we investigate model-level mixtures of object detectors and analyze their suitability for improving performance and interpretability in object detection. We propose an MoE architecture that combines YOLO-based detectors trained on semantically disjoint data subsets, with a learned gating network that dynamically weights expert contributions. We study different strategies for fusing detection outputs and for training the gating mechanism, including balancing losses to prevent expert collapse. Experiments on the BDD100K dataset demonstrate that the proposed MoE consistently outperforms standard ensemble approaches and provides insights into expert specialization across domains, highlighting model-level MoEs as a viable alternative to traditional ensembling for object detection. Our code is available at \url{https://github.com/KASTEL-MobilityLab/mixtures-of-experts/}. Research on Green Wave Guidance Strategy Under Cross-Network Collaboration Architecture Lingqiu Zeng, Yuqiao Liu, and Qingwen Han (Chongqing University); Jianmei Lei (State Key Laboratory of Intelligent Vehicle Safety Technology); and Yufei Yan, Lei Ye, and Ruilong Yang (Chongqing University) Abstract Abstract Green wave guidance is an important application in intelligent transportation. This paper proposes a green wave guidance strategy based on cross-network collaboration. First, a cross-network collaboration framework is introduced to optimize multi-intersection green wave guidance, incorporating an improved Ising model for road pass-through rate assessment. Then, an improved ant colony algorithm based on road pass-through rate distribution is developed to generate vehicle trajectory guidance and speed recommendations, with a dynamic trajectory update strategy to enhance information usability. Experimental results show that this method effectively evaluates traffic conditions and improves vehicle passage efficiency. AutoEval: A Practical Framework for Autonomous Evaluation of Mobile Agents Jiahui Sun and Zhichao Hua (Shanghai Jiao Tong University) Abstract Abstract Accurate and systematic evaluation of mobile agents can significantly advance their development and real-world applicability. However, existing benchmarks for mobile agents lack practicality and scalability due to the extensive manual effort required to define task reward signals and implement corresponding evaluation codes. To this end, we propose AutoEval, an autonomous agent evaluation framework that tests a mobile agent without any manual effort. First, we design a Structured Substate Representation to describe the UI state changes while agent execution, such that task reward signals can be automatically generated. Second, we utilize a Judge System that can autonomously evaluate agents’ performance given the automatically generated task reward signals. By providing only a task description, our framework evaluates agents with fine-grained performance feedback to that task without any extra manual effort. We implement a prototype of our framework and validate the automatically generated task reward signals, finding over 93% coverage to human-annotated reward signals. Moreover, to prove the effectiveness of our autonomous Judge System, we manually verify its judge results and demonstrate that it achieves 94% accuracy. Finally, we evaluate the state-of-the-art mobile agents using our framework, providing detailed insights into their performance characteristics and limitations. The Trichromatic Strong Lottery Ticket Hypothesis: A Unifying View of Supermask-Based Learning Ángel López García-Arias (NTT, Inc.); Yasuyuki Okoshi and Hikari Otsuka (Institute of Science Tokyo); Daiki Chijiwa, Yasuhiro Fujiwara, and Susumu Takeuchi (NTT, Inc.); and Masato Motomura (Institute of Science Tokyo) Abstract Abstract The Strong Lottery Ticket Hypothesis (SLTH) posits that high-performing subnetworks can be extracted from randomly initialized neural networks using binary pruning masks, or supermasks. This paper introduces the Trichromatic Strong Lottery Ticket Hypothesis (T-SLTH), a unified generalization of the SLTH that decomposes supermasks into three conceptually orthogonal components: connectivity, sign, and magnitude. This trichromatic decomposition unifies existing SLTH variants, enables new configurations, and clarifies the connection between supermask-based training and quantization-aware training. The proposed framework is evaluated on image classification benchmarks with ResNet architectures, achieving state-of-the-art accuracy-compression tradeoffs among supermask-based models. ResNet-50 is compressed 38x on CIFAR-100 and 25x on ImageNet without training any weights. Additional experiments across datasets, architectures, and modalities demonstrate that trichromatic supermasks act as robust structural priors. These results position partially random networks trained via supermasks as a compelling alternative to conventional weight optimization for efficient inference systems. DMR-Seg: Endowing Multimodal Large Language Models with Segmentation Capability via Discretized Mask Representations Jiru Deng, Wenhao Wu, and Zhiheng Li (Tsinghua University) Abstract Abstract Referring Image Segmentation (RIS) aims to segment objects based on natural language expressions. Existing methods leverage Multimodal Large Language Models (MLLMs) to interpret text descriptions and localize objects, while mask generation follows two paradigms. The first approach represents segmentation masks using polygon coordinates, which inherently suffers from precision errors. The alternative approach integrates MLLMs with Segmentation Foundation Models (SFMs). Combining two models with differing output mechanisms can lead to optimization inefficiencies and hinder the effective propagation of information. In this work, we propose DMR-Seg, an image segmentation method that integrates mask generation within MLLM. We use the modified VQVAE to discretize binary segmentation masks into sequences of tokens. By training the MLLM to predict these sequences autoregressively, the model is endowed with the capability to achieve image segmentation. To reduce decoding dependency caused by autoregressive generation, we further designed a training strategy that proactively injects token noise. Experimental results on multiple datasets demonstrate that our model outperforms existing segmentation methods. Furthermore, ablation studies confirm the effectiveness of the proposed training strategy. GP-Adapter: Gaussian Process CLIP-Adapter for Few-Shot Out-of-Distribution Detection Taisei Saito, Koretaka Ogata, and Takafumi Hiroi (Ricoh Company, Ltd.) Abstract Abstract We propose GP-Adapter, a training-free framework that augments CLIP (Contrastive Language-Image Pre-training) with Gaussian Process (GP) uncertainty modeling for few-shot classification and out-of-distribution (OOD) detection. While CLIP achieves strong zero-shot recognition, it yields deterministic similarity scores and offers limited uncertainty information, which is critical under distribution shift and data scarcity. GP-Adapter constructs modality-specific, class-wise one-class GPs on top of frozen CLIP embeddings using an RBF kernel for image features and a linear kernel for text prompts and fuses their predictive statistics to produce a variance-aware confidence score for OOD detection. The method requires no fine-tuning of the CLIP backbone and relies only on a small K-shot cache and lightweight hyperparameter selection, with memory cost scaling as O(CK^2) for C classes and K shots. Experiments on ImageNet and multiple OOD benchmarks show that GP-Adapter provides competitive few-shot performance and consistently improves OOD detection when combined with prompt-learning baselines, highlighting the complementarity between GP-based uncertainty modeling and prompt learning. Overall, our results suggest that integrating probabilistic inference with large pre-trained vision-language models can improve reliability in low-data and distribution-shifted settings. Code is available at https://github.com/tms-byte/GP-Adapter Q-SARIMA: A Hybrid Quantum–Classical Extension of SARIMA for Time Series Forecasting in Precision Agriculture Lucas Grogenski Meloca (State University of Maringá); Rodrigo Clemente Thom de Souza (Federal University of Paraná, State University of Maringá); Linnyer Ruiz Aylon (State University of Maringá); and Viviana Cocco Mariani (Federal University of Paraná) Abstract Abstract This paper introduces Q-SARIMA, a hybrid quantum-classical architecture for 10-day meteorological forecasting in precision agriculture. By integrating SARIMA structural robustness with Variational Quantum Circuits (VQC), the model maps qubit expectation values to statistical coefficients using the gradient-free COBYLA optimizer. Evaluation on real-world Brazilian datasets via walk-forward validation shows that while classical SARIMA provides superior stability, Q-SARIMA is competitive for high-volatility series like wind speed. These results highlight NISQ-era challenges, such as convergence instability and cost surface roughness. This study establishes a methodological baseline for hybrid quantum-classical integration, offering critical insights into the trade-offs between quantum expressivity and statistical reliability in agricultural monitoring systems. Importance-Aware Scheduling for High-Dimensional Hyperparameter Optimization Ruinan Wang, Ian Nabney, and Mohammad Golbabaee (University of Bristol) Abstract Abstract Hyperparameter Optimization (HPO) is essential for building high-performing ML/DL models, yet conventional optimizers often struggle in high-dimensional spaces where evaluations are costly and progress is diluted across many low-impact variables. We propose Greedy Importance First (GIF), an importance-aware scheduling strategy that uses a small-sample warm start to estimate hyperparameter importance, forms importance-based groups, allocates trials proportionally, and retains a full-space fallback. We evaluate GIF under fixed evaluation budgets on five anisotropic analytic functions ($d\!\in\!\{5,10,30,50\}$), Bayesmark, and NAS-Bench-301 (33D). On the higher-dimensional benchmarks, GIF reaches better incumbents with faster convergence than TPE, BOHB, Random Search, and Sequential Grouping. On Bayesmark, where the effective dimensionality is smaller, GIF remains competitive but the margins are smaller. Ablation studies show that importance estimation, proportional allocation, and the fallback step all contribute to the gains. We also verify that the HIA component recovers the intended anisotropy on the analytic benchmarks. These results suggest that GIF is a simple and plug-compatible way to improve sample efficiency in high-dimensional HPO. Hardware-aware Calibrated Clustered Attention for Efficient Visual Geometric Transformers Weitian Wang (Robert Bosch GmbH, Ruhr University Bochum); Shubham Rai and Cecilia De la Parra (Robert Bosch GmbH); and Akash Kumar (Ruhr University Bochum) Abstract Abstract The Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruction, as it is the first model that directly infers all key 3D attributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism requires global attention layers with extremely long sequences that causes a significant latency bottleneck. In this paper, we propose blockwise clustered attention (BC attention) to accelerate the global attention layers in VGGT. By limiting the clustering within HW-friendly neighborhood blocks, BC attention reduces the computation overhead of query clustering as well as the costly data movement between on- and off-chip memory. This enables BC attention to scale to long sequences and deliver practical latency improvements on GPUs. Moreover, we introduce a hashing hyperplane calibration method and a threshold-based error compensation method to reduce clustering errors efficiently, which is a bottleneck in the current clustered attention mechanism. Overall, our experiments on GPU demonstrate that calibrated BC attention accelerates the global attention layers by 2.10-2.63× and the whole backbone by 1.77-2.35× with negligible loss (1%) for large scenes. With a small performance loss (< 5%), calibrated BC attention further achieves a 2.26-2.87× latency improvement on the global attention layers and a 1.90-2.55× improvement on the backbone. Graph Neural Networks for Human Activity Recognition in Smart Environments David-Dumitru Ţigău and Ali Mahmoudi (Eindhoven University of Technology), Fetze Pijlman (Signify), and Tanir Ozcelebi (Eindhoven University of Technology) Abstract Abstract Human Activity Recognition (HAR) in smart environments enables assisted living, health monitoring, and home automation. Traditional HAR methods often rely on handcrafted features or shallow models that do not generalize across different sensor layouts. Sequence-based deep learning models, such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory Networks (LSTMs), improve temporal modeling but typically treat sensor streams as independent channels. As a result, these models do not explicitly encode spatial relationships between sensors or the structural constraints imposed by the environment, limiting their ability to model sensor interactions driven by the environment's physical structure. We study Graph Convolutional Networks (GNNs) for HAR using ambient sensor data, modeling sensors as nodes and their spatio-temporal interactions as edges. Our graph construction uses domain-informed room adjacency and temporal co-activation, with shared node and edge encoders. We evaluate a Graph Convolutional Network (GCN), a Graph Attention Network (GAT), and a hybrid GAT–Long Short-Term Memory model (GAT-LSTM) on two CASAS smart homes (HH111 and HH101). Sparse, domain-informed graphs consistently outperform fully connected graphs, and longer temporal windows improve robustness. GAT-LSTM attains the best classification performance, while GCN provides the most efficient baseline. Overall, combining spatial attention with temporal modeling yields a strong framework for HAR while highlighting remaining challenges for multi-resident activities and data-driven room adjacency. Advanced Recombinant Transformer: Instrument-Level Audio Compatibility Scoring with Self-Supervised Acoustic Embeddings Ahmad Hammoudeh, Muhammad Taimoor Haseeb, and Gus Xia (MBZUAI) Abstract Abstract Vertical loop compatibility—predicting the perceptual compatibility of two music loops layered together in a mix—has seen limited progress due to the absence of a large-scale, high-fidelity, labeled dataset. We introduce the Advanced Recombinant Transformer (ART). To train ART, we curate a large dataset of aligned positive pairs by rendering studio-quality stems from multi-instrument MIDI. We construct matched-condition negatives using score-based mining with calibrated thresholds based on a blinded musician audit. For representation, we replace log-mel features with frozen embeddings from the Music Understanding model with large-scale self-supervised training (MERT), which capture timbral nuances and provide cues for higher-level musical structure. An order-invariant head scores pairwise compatibility. Experimental results show a 10.1% improvement in compatibility classification over the previous state-of-the-art. Extensive subjective tests (N = 107) on production-grade, human-composed audio loops find seamlessness and creativity ratings comparable to human-curated pairings. Samples are available at: https://loopcompatibility.github.io/demo GenSAM: One-Shot Medical Image Segmentation via Diffusion-Based Test-Time Adaptation Zhihao Mao, Bangpu Chen, and Xuehai Chen (China University of Geosciences) Abstract Abstract Medical image segmentation is fundamental for clinical diagnosis, yet developing deep learning models remains challenging due to the scarcity of pixel-level annotations. Few-shot segmentation offers a promising solution by leveraging only a few labeled examples, but existing methods struggle with the support-query misalignment problem—a single support image often fails to represent the diverse appearances of target organs across patients. In this paper, we propose GenSAM, a novel training-free framework that addresses this challenge by leveraging diffusion models for test-time adaptation. Our key insight is that query-guided style injection can synthesize support views that are visually closer to the query domain while maintaining anatomical consistency, effectively bridging the support-query gap without additional annotations. GenSAM comprises three components: (1) Generative Support Expansion (GSE), which synthesizes diverse augmented views using query-guided diffusion-based style transfer; (2) Multi-center Prompting, which generates positive and negative point prompts from similarity maps to guide SAM2; and (3) Consistency-Aware Ensemble, which aggregates predictions weighted by confidence scores to filter unreliable results. Experiments on three medical imaging benchmarks demonstrate that GenSAM significantly outperforms existing training-free methods and achieves competitive performance with training-based approaches, highlighting the potential of combining generative and discriminative foundation models for medical image analysis. EventPoseFormer: Event-Based Real-time 3D Human Pose Estimation Melica Omer Ali (University of Genova, Istituto Italiano di Technologia); Francesca Odone (University of Genova); Arren Glover and Chiara Bartolozzi (Istituto Italiano di Technologia); and Gaurvi Goyal (Maastricht University) Abstract Abstract Human Pose Estimation (HPE) provides a structured representation of human motion and has become a key component in a wide range of applications, spanning from medical monitoring and rehabilitation to sports biomechanics and consumer-oriented applications such as dance and yoga instruction. While 2D pose estimation is sufficient for some of these use cases, others require 3D pose information to be effective. Despite significant recent progress in HPE, achieving real-time 3D estimation with low latency is still a challenging problem. This challenge may be addressed through the use of event cameras — novel, bio-inspired, visual sensors which react to changes in the intensity values in the scene, asynchronously, at pixel level. This sensing modality is extremely efficient, as it removes redundancy at the sensor level, and is well suited for the design of high-speed, robust artificial vision algorithms, including HPE. In this paper, we combine MoveEnet, a HPE system based on event cameras which performs real-time high-frequency 2D pose estimation, with PoseFormerV2, a transformer-based 2D-to-3D pose lifting model to develop an end-to-end system for monocular real-time 3D HPE. Our method not only performs more accurately and faster than existing HPE models for event cameras, but is also competitive with RGB models while being substantially faster. Our final system has an end-to-end latency of 30ms, a real-time frequency of 38Hz and a 3D MPJPE of 85.7mm on event-Human 3.6M dataset. Code for the end-to-end real-time pipeline can be found at https://github.com/event-driven-robotics/HPE-event-poseformer. Self-Supervised Transferability-Sensitive Multimodal Heterogeneous Graph Learning Khalil BACHIRI and Maria Malek (CY Cergy Paris Université) and Ali YAHYAOUY (Sidi Mohamed Ben Abdellah University) Abstract Abstract Multimodal heterogeneous graphs are increasingly used to model complex systems combining diverse entity types, relations, and data modalities. While multimodal graph learning has shown strong empirical performance, existing methods often rely on uniform cross-modal alignment or indiscriminate fusion, which can induce negative transfer when modalities are imbalanced, sparse, or semantically incompatible. In this paper, we propose a self-supervised, transferability-sensitive framework for multimodal heterogeneous graph neural networks (S$^2$T-MMHGNN). Instead of enforcing symmetric alignment, our approach leverages directional predictive self-supervision to explicitly estimate inter-modality transferability and uses this signal to regulate multimodal fusion. This design selectively amplifies complementary modalities while suppressing unreliable cross-modal interactions. We provide theoretical analysis establishing collapse avoidance and monotone fusion behavior under increasing modality incompatibility. Extensive experiments on five real-world benchmarks demonstrate consistent improvements over state-of-the-art heterogeneous, multimodal, and self-supervised baselines, particularly under noisy and sparse multimodal settings. CogKT: Cognitive Ability-based Knowledge Tracing towards Interpretable Knowledge Mastery Representation in Personalized Learning Peiran Zhang (Tsinghua University), Nan He (Beijing University Of Technology), Binbin Qi (South China Normal University), and Lifeng Sun (Tsinghua University) Abstract Abstract In Knowledge Tracing (KT) tasks, to address the black-box characteristic of knowledge mastery, recent work has proposed some interpretable methods to represent knowledge mastery. Educational theories suggest that learners apply knowledge concepts across different cognitive levels. Therefore, cognitive levels should be represented in the unit of knowledge mastery, which is overlooked by existing studies. We propose the concept of "cognitive ability" as the unit of knowledge mastery to represent cognitive levels, and propose a Cognitive Ability-based Knowledge Tracing (CogKT) method. In CogKT, the relationships among cognitive abilities are modeled using cognitive ability graphs and Graph Neural Network (GNN), while the temporal dynamics of cognitive abilities is captured using LSTM network enhanced by cognitive level embedding. Evaluated on a real-world dataset comprising 2393 students, CogKT outperforms baseline models on performance metrics AUC and ACC. In addition, cognitive abilities exhibit a significant influence on test outcomes, indicating that they contribute to enhancing the interpretability of knowledge mastery representation. Stability Regions Control Gradient Flow: Why Neural ODEs Need Extended-Stability Solvers Gavin lee Goodship, Luis Miralles, and Stephen O'Sullivan (TU Dublin) Abstract Abstract Neural ordinary differential equations (ODEs) provide a principled continuous-depth formulation, but their practical use is often limited by slow and numerically fragile training due to the weaknesses of standard ODE solvers. We introduce extended-stability Runge--Kutta methods, explicit fixed-step solvers that admit much larger step sizes without numerical blow-up in negative-real-axis dominated regimes than classical schemes such as RK4, or adaptive methods such as Dormand--Prince. RAP: Retrieve, Adapt, and Prompt-Fit for Training-Free Few-Shot Medical Image Segmentation Bangpu Chen and Zhihao Mao (China University of Geoscience) Abstract Abstract Few-shot medical image segmentation (FSMIS) has achieved notable progress, yet most existing methods mainly rely on semantic correspondences from scarce annotations while under-utilizing a key property of medical imagery: anatomical targets exhibit repeatable high-frequency morphology (e.g., boundary geometry and spatial layout) across patients and acquisitions. We propose RAP, a training-free framework that Retrieves, Adapts, and Prompt-Fits Segment Anything Model 2 (SAM2) for FSMIS. First, RAP retrieves morphologically compatible supports from an archive using DINOv3 features to reduce brittleness to single-support choice. Second, it adapts the retrieved support mask to the query by fitting boundary-aware structural cues, yielding an anatomy-consistent pre-mask under domain shifts. Third, RAP converts the pre-mask into prompts by sampling positive points via Voronoi partitioning and negative points via sector-based sampling, and feeds them into SAM2 for final refinement—without any fine-tuning. Extensive experiments on multiple medical segmentation benchmarks show that RAP consistently surpasses prior FSMIS baselines and achieves stateof-the-art performance. Overall, RAP demonstrates that explicit structural fitting combined with retrieval-augmented prompting offers a simple and effective route to robust training-free few-shot medical segmentation. Concurrent Clinical Predictions via Dual-Space Message Passing and Ontology Alignment Shervin Mehryar and Michel Dumontier (Maastricht University) Abstract Abstract The growing availability of electronic health records (EHR) presents increasing opportunities for data-driven applications in healthcare. These records, however, are represented using disparate coding systems and often suffer from fragmentation and structural inconsistencies. Such data fragmentation can lead to substantial performance gaps across different predictive tasks. In this work, a framework is proposed in order to mitigate these artefacts and enable multi-task predictions. To address performance gaps, we introduce a concurrent loss that guides the learning process over fragmented entities. To address structural misalignment, we propose a neural message-passing algorithm that enables the integration of ontological information. Experiments conducted on cardiovascular disease (CVD) patients demonstrate that the proposed system overcomes prediction gaps across multiple clinical tasks, including laboratory tests, diagnoses, medications, and mortality prediction, performed concurrently. The results highlight the potential of ontology-aligned representation learning for high-precision decision-making in healthcare systems. AdversaRiskQA: An Adversarial Factuality Benchmark for High-Risk Domains Adam Szelestey, Sofie van Engelen, Tianhao Huang, Justin Snelders, Qintao Zeng, and Songgaojun Deng (Eindhoven University of Technology) Abstract Abstract Hallucination in large language models (LLMs) contributes to the spread of misinformation and diminished public trust, particularly in high-risk domains. Among hallucination types, factuality is crucial, as it concerns a model's alignment with established world knowledge. Adversarial factuality, defined as the deliberate insertion of misinformation into prompts, tests a model's ability to detect and resist confidently framed falsehoods. However, existing work lacks high-quality, domain-specific resources for assessing model robustness under such adversarial conditions, and no prior research has examined the impact of injected misinformation on long-form text factuality. ReloDepth: Reprojection Loss Filtering for Self-Supervised Depth Estimation in Dynamic Environments Siyu Chen and Hong Liu (Peking University); Wenhao Li (Nanyang Technological University); and Ying Zhu, Jianbing Wu, and Guoquan Wang (Peking University) Abstract Abstract Depth estimation plays a pivotal role in robotics applications, where self-supervised approaches have emerged as promising solutions by leveraging abundant unlabelled real-world data. While existing methods achieve notable performance under static scene assumptions, their effectiveness significantly degrades in dynamic environments due to moving objects. To address this issue, we present ReloDepth, a novel method for self-supervised depth estimation in dynamic scenes. It tackles the challenge of dynamic objects from two key perspectives. First, within the self-supervised learning framework, we introduce a reprojection loss mask to explicitly localize regions likely contaminated by dynamic objects, thereby mitigating their corrupting influence on the reprojection-based supervisory signal at the loss level. Second, for multi-frame depth estimation, we propose a photometric gated cost volume strategy to mask redundant regions prior to cost volume, improving depth estimation robustness by steering the attention of the network toward photometrically informative regions. Furthermore, we propose a spectral entropy uncertainty module that incorporates spectral entropy to guide uncertainty estimation during depth fusion, effectively addressing issues arising from cost volume computation in dynamic environments. Extensive experiments on KITTI and Cityscapes datasets demonstrate that the proposed method consistently outperforms existing self-supervised depth estimation baselines. Geodesic-SVDD Latent Mixup for One-Shot Distributed Tabular Learning Andrea Moleri (Honda Research Institute Europe GmbH, University of Milan-Bicocca); Christian Internò (University of Bielefeld); Markus Olhofer (Honda Research Institute Europe GmbH); Barbara Hammer (University of Bielefeld); and Ali Raza (Honda Research Institute Europe GmbH) Abstract Abstract Distributed Learning on structured data is a key and widely adopted approach in privacy-sensitive sectors, ranging from financial fraud prevention to federated healthcare diagnostics. While Gradient Boosted Decision Trees are widely used in centralized settings due to their efficiency, they are incompatible with standard gradient-based federated learning protocols because of their non-differentiability. Conversely, standard deep classifiers that are amenable to federated optimization often suffer from catastrophic forgetting under label distribution skew. Together, these limitations constitute the Federated Tabular Gap. In this work, to address this gap we propose Geodesic-SVDD, a framework for Distributed Spherical Prototype Learning that aggregates probabilistic manifolds rather than model weights. Specifically, for each client, we employ a Residual Spherical Encoder and Geodesic Latent Mixup that enforces strict hyperspherical constraints (∥ϕ(x)∥ ≡ 1), preventing the manifold collapse observed in linear approximations. Furthermore, leveraging this insight, we propose a novel One-Shot “Product of Experts” protocol that enables privacy-preserving aggregation without iterative synchronization, effectively handling disjoint class distributions. Finally, we mechanistically analyze the trade-off between interpolation strategies, demonstrating that while linear mixup offers an enticing computational shortcut, it destabilizes optimization gradients by orders of magnitude (109×). More broadly, our work showcases how differentiable manifold constraints can be leveraged to outperform state-of-the-art Federated methods by ≈ 20% on non-IID tabular benchmarks. DiF3CON: Diffusion Forgetting via Continual Unlearning Cătălin Ciocîrlan, Vlad Vasilescu, and Ana Nicolae (National University of Science and Technology Politehnica Bucharest) Abstract Abstract Diffusion models have emerged as the dominant framework for image generation, offering high fidelity samples and flexible conditioning. However, they may pose risks with regard to generating copyright-sensitive or inappropriate content that may be exploited by malicious actors. Machine unlearning emerged as a counteraction to this phenomenon, seeking to efficiently remove the influence of specific training data without sacrificing the general performance. Prior work on diffusion unlearning mostly targets text-to-image or class-conditional models, while image-to-image settings such as inpainting remain largely unexplored. In this work, we introduce DiF3CON (DiFFusion Forgetting via CONtinual unlearning), a continual unlearning framework for diffusion-based inpainting models. Inspired by regularization and stability mechanisms from multi-task learning, DiF3CON simulates iterative, efficient class-wise data removal, without compromising the performance of the base diffusion model. Extensive experiments on the Places365 dataset show that DiF3CON achieves unlearning stability over significantly larger forgetting data volumes, advancing this relatively underexplored area of machine unlearning for diffusion-based image-to-image models. Our code is available at https://github.com/Cata400/dif3con. Monday 0.01 London FUZZ-IEEE Paper FUZZ 3: FUZZ-IEEE SS11 Software for Soft Computing Session Chair: Giovanni Acampora (University of Naples Federico II), James Keller (University of Missouri) JWA–PyTorch: an Open-Source PyTorch-Based Toolkit for the Joint Weighted Average Operator Youssef Abdulghani and Christian Wagner (University of Nottingham) and Stephen Broomell, Sean Conway, and Ethan Guthrie (Purdue University) Abstract Abstract Aggregation operators form the fundamental set of tools used in information aggregation and fusion. The Joint Weighted Average (JWA) was recently proposed as a new aggregation operator that blends two different weighted aggregation approaches; source-based and evidence-based. Specifically, the JWA enables the systematic, joint consideration of the well-known linear weighted average (LWA) and the ordered weighted average (OWA) operators--and their associated weighting strategies. Here, we present a PyTorch-based toolkit to facilitate adoption of JWA aggregation in practice as well as to learn JWA weights from data. The latter is critical, not only for machine learning applications, but in particular for leveraging the explanatory power of the underlying weights, for example, in social science research. We integrate the operator learning approach with the established supervised learning framework that PyTorch provides, offering access to powerful optimization. We showcase and demonstrate the functionality of the toolkit. The open-source PyTorch-based toolkit opens the door for further development and accessible experimentation with the JWA framework and offers the only freely available approach for supervised learning in this setting. The toolkit is lightweight, extensible, and easy to use, allowing researchers and practitioners from other disciplines to leverage and advance JWA, associated techniques and research. ISurvey: An R Toolkit for Interval-Valued Survey Data Analysis Yu Zhao, Christian Wagner, Brendan Ryan, and Direnc Pekaslan (University of Nottingham) Abstract Abstract Recent years have seen growing interest in interval-valued (IV) survey response formats as an alternative to conventional single-valued (SV) scales, as they allow respondents to express uncertainty and imprecision as ranges. While IV data analysis methods have advanced, software support for applied IV survey analysis remains limited. To address this gap, this paper presents ISurvey, a free and open-source R toolkit for IV survey data analysis, designed for researchers and practitioners with diverse backgrounds. ISurvey supports datasets containing IV, SV, and single-choice items. It provides descriptive statistics under the Fréchet framework, visual descriptives, and survey psychometric assessment methods such as IV correlation for convergent validity and IV Cronbach's α for reliability. The toolkit also implements IV extensions of diagnostic methods used for actionable prioritization and benchmarking in survey research, including gap analysis, importance-performance analysis, and the customer satisfaction index. For illustration, we present a complete data analysis workflow using a UK rail customer survey, in which IV and SV responses were collected via a between-subject design. This provides a step-by-step tutorial that guides users from raw exports (from DECSYS, an open-source IV survey platform) to analysis, ready for reporting and decision-making. f-EVOVAQ: A GPU-based Framework for Evolutionary Training of Variational Quantum Algorithms Angela Chiatto, Autilia Vitiello, and Giovanni Acampora (University of Naples Federico II) Abstract Abstract In the field of quantum computing, Variational Quantum Algorithms (VQAs) combine quantum circuits with classical optimization to perform several tasks including classification and regression. Although gradient-based methods are commonly used for their training, these methods suffer from barren plateaus, noise sensitivity, and high computational costs. In this scenario, Evolutionary Algorithms (EAs) provide a gradient-free alternative for training VQAs enabling the discovery of high-quality solutions while requiring limited hardware resources. Unfortunately, the number of software tools aimed to the application of EAs for training of VQAs is limited and, moreover, they rarely exploits GPU acceleration. In order to bridge this gap, this work presents f-EVOVAQ, a Python package that extends the EVOVAQ framework by implementing vectorized evolutionary operators on GPU using CuPy arrays. The package enables efficient population-level optimization while remaining user-friendly for non-experts in evolutionary computation. Experiments on quantum classification tasks demonstrate significant reductions in optimization time and improved scalability, with limitations arising from GPU–CPU data transfers. Linguistic Pattern–Enriched Transformers for Prompt Injection Detection in LLM Systems Carmen De Maio (University of Salerno), Fei Hao (School of Artificial Intelligence and Computer Science), and Sabrina Senatore (University of Salerno) Abstract Abstract Nowadays, Large Language Models (LLMs) are widely included in many daily applications and services, where they play a crucial role in supporting users for automated decision-making tasks. As their adoption grows, so does their vulnerability to prompt injection attacks, which use natural language instructions to bypass safety constraints and also manipulate model behavior. To address these issues, the paper proposes a hybrid framework for prompt injection detection that combines domain-adaptive pretraining (DAPT) of RoBERTa encoder with a set of rule-based text patterns to detect suspicious adversarial interaction modes within prompt inputs. The proposed approach is evaluated through an empirical analysis across multiple prompt injection datasets. Experimental results show that the hybrid model outperforms text-only baselines and maintains reasonable performance under distributional shift. An ablation study is conducted to evaluate the individual component contribution. FuzzyLinguistics: A Python Package for Generating, Simplifying, and Comparing Linguistic Summaries of Data Brendan Alvey, James Keller, Brendan Young, and Derek Anderson (University of Missouri) Abstract Abstract Large Language Models (LLMs) can generate polished text, but they are not guaranteed to remain faithful to the underlying numbers. Fuzzy Linguistic Summarization (FLS) converts data into graded protoform statements with explicit fuzzy semantics and quality measures. We present FuzzyLinguistics, an open-source Python package for generating, simplifying, and comparing fuzzy linguistic summaries of tabular data. It offers a configuration-driven, Pydantic-validated workflow, vectorized protoform enumeration and scoring, graph-based redundancy reduction with optional compression, and differential summaries for comparing two datasets or model outputs under a shared linguistic vocabulary. Experiments on three minimal examples demonstrate deterministic outputs, strong redundancy reduction, and dataset comparisons, positioning FuzzyLinguistics as a practical grounding layer for explainable analysis and LLM-based reporting. A Python Library for Contextual SVM Explanations Juan Pisco-Jordan, Luis Ramos-Pozo, and Ana Tapia-Rosero (ESPOL Polytechnic University) Abstract Abstract Support Vector Machines (SVMs) are powerful classification models, but their black box nature often limits their adoption in high-stakes domains where interpretability is crucial. The contextualized SVM framework provides a path towards explainability by training specialized models for specific data contexts, but a practical software implementation has been missing. This paper introduces XSVM@ctx-Lib as an open-source Python library that implements the contextualized SVM architecture. The library is built to be computationally efficient by parallelizing both the model training and evaluation processes. XSVM@ctx-Lib was validated against standard SVM implementations using multiple datasets, demonstrating significant performance advantages. The library achieved reductions in training and classification times. Furthermore, the contextual approach improved classification precision and accuracy by up to 9%. XSVM@ctx-Lib provides researchers and practitioners with a powerful tool that enhances SVMs by simultaneously improving transparency, computational speed, and predictive accuracy. Monday 0.04 Brussels IJCNN Paper Neural Learning and Optimization I Session Chair: Jose Principe (University of Florida), Manos Kirtas (Centre for Nanosciences and Nanotechnologies, Aristotle University Of Thessaloniki) Proxy-Driven Estimation of Distribution Algorithm for Efficient Neural Architecture Search André Silva, Mario Haddad-Neto, Rayol Mendonca-Neto, Vanessa Camara, Moyses Lima, Luis Uebel, and Luiz Cordovil-Jr (Sidia Science and Technology Institute) Abstract Abstract Automating the design of Deep Neural Networks for specific applications using Neural Architecture Search (NAS) has become increasingly attractive. However, traditional NAS methods are often hindered by the high cost of evaluating numerous candidate networks, rendering them impractical for large-scale searches. In this work, we assessed a proxy-driven Estimation of Distribution Algorithm (EDA) for optimization of topological and quantitative hyperparameters. By leveraging the ideas of designing network design spaces and iterated sampling algorithms with zero-cost proxies, our approach reduces the search cost significantly. Furthermore, the proxy driven EDA achieves near-optimal results in the NATS-Bench topology and size search spaces on CIFAR-10, CIFAR-100, and ImageNet16-120 datasets for image classification. The results show that a proxy-driven EDA can reliably explore large, mixed-type design spaces with negligible computational overhead, opening the door to rapid, resource-friendly architecture discovery. PAVE: Parallel Adaptive View Excitation for Scalable Spectral Learning on Tabular Data Matthew Fried and Mohammad Alshibli (SUNY Farmingdale) Abstract Abstract We introduce Parallel Adaptive View Excitation (PAVE), a framework for parallel selection and training across sample-view and feature-view graph Laplacians. Each view independently decomposes the data manifold, exposing complementary spectral structure for downstream learning. Unlike unified-graph pipelines that merge all relations into a single Laplacian with inherent serial bottlenecks, PAVE exposes coarse-grained task parallelism by evaluating sample-view and feature-view manifolds concurrently and selecting the view with higher validation accuracy, while retaining the alternate view for interpretability. Across 16 benchmark datasets, PAVE matches or exceeds strong gradient-boosted baselines on most tasks, with performance degradations typically small when present, while maintaining comparable or improved F1 and AUC. On modern multi-core systems, PAVE achieves 1.5--1.9$\times$ training speedup via view-level parallelism, with approximately 60\% strong-scaling efficiency up to 32 workers and near-ideal efficiency under weak scaling. These results show that dual-view spectral excitation provides a lightweight, interpretable, and scalable alternative to conventional feature selection and manifold learning. Maintaining Difficulty: A Margin Scheduler for Triplet Loss in Siamese Networks Training Roberto Sprengel Minozzo Tomchak (Universidade Federal do Paraná); Oge Marques (Florida Atlantic University); and Lucas Garcia Pedroso, Luiz Eduardo Oliveira, and Paulo Lisboa de Almeida (Universidade Federal do Paraná) Abstract Abstract The Triplet Margin Ranking Loss is one of the most widely used loss functions in Siamese Networks for solving Distance Metric Learning (DML) problems. This loss function depends on a margin parameter $\mu$, which defines the minimum distance that should separate positive and negative pairs during training. In this work, we show that, during training, the effective margin of many triplets often exceeds the predefined value of $\mu$, provided that a sufficient number of triplets violating this margin is observed. This behavior indicates that fixing the margin throughout training may limit the learning process. Based on this observation, we propose a margin scheduler that adjusts the value of $\mu$ according to the proportion of easy triplets observed at each epoch, with the goal of maintaining training difficulty over time. We show that the proposed strategy leads to improved performance when compared to both a constant margin and a monotonically increasing margin scheme. Experimental results on four different datasets show consistent gains in verification performance. Our trained models and source code are available at github.com/robertotomchak/maintaining-difficulty. Regime Change Hypothesis: Foundations for Decoupled Dynamics in Neural Network Training Cristian Pérez-Corral, Alberto Fernández-Hernández, and Jose I. Mestre (Universitat Politècnica de València); Manuel F. Dolz (Universitat Jaume I); Enrique S. Quintana-Ortí (Universitat Politècnica de València); and José Duato (Openchip & Software Technologies) Abstract Abstract Despite the empirical success of DNN, their internal training dynamics remain difficult to characterize. In ReLU-based models, the activation pattern induced by a given input determines the piecewise-linear region in which the network behaves affinely. Motivated by this geometry, we investigate whether training exhibits a two-timescale behavior: an early stage with substantial changes in activation patterns and a later stage where weight updates predominantly refine the model within largely stable activation regimes. We first prove a local stability property: outside measure-zero sets of parameters and inputs, sufficiently small parameter perturbations preserve the activation pattern of a fixed input, implying locally affine behavior within activation regions. We then empirically track per-iteration changes in weights and activation patterns across fully-connected and convolutional architectures, as well as Transformer-based models, where activation patterns are recorded in the ReLU feed-forward (MLP/FFN) submodules, using fixed validation subsets. Across the evaluated settings, activation-pattern changes decay 3 times earlier than weight-update magnitudes, showing that late-stage training often proceeds within relatively stable activation regimes. These findings provide a concrete, architecture-agnostic instrument for monitoring training dynamics and motivate further study of decoupled optimization strategies for piecewise-linear networks. For reproducibility, code and experiment configurations will be released upon acceptance. Distribution-Free Pretraining of Classification Losses via Evolutionary Dynamics Xiang Meng and Yan Pei (The University of Aizu, Japan) Abstract Abstract We propose Evolutionary Dynamic Loss (EDL), a framework that learns a transferable classification loss in the probability space using unlimited synthetic prediction-label pairs, without accessing real samples during the main loss pretraining stage. EDL parameterizes the loss as a lightweight network and is trained with a semantics-free ranking-consistency objective that assigns larger penalties for more erroneous predictions. To robustly explore the space of loss functions, we optimize EDL via an evolutionary strategy and introduce chaotic mutation to improve exploration under noisy fitness evaluations. Experiments on CIFAR-10 with ResNet backbones show that EDL can serve as a drop-in replacement for cross-entropy and achieves competitive or improved accuracy, while ablation studies confirm that chaotic mutation yields faster convergence and better synthetic pretraining metrics than standard Gaussian mutation. On the Generalization Gain of Transfer Learning Chenming Cao (The Hong Kong Polytechnic University), Xiaoming Xue (China University of Petroleum (East China)), Liang Feng (Chongqing University), and Yao Hu and Kay Chen Tan (The Hong Kong Polytechnic University) Abstract Abstract Transfer learning (TL) has proven highly effective at improving generalization, yet theoretical insights lag behind its rapid methodological advances. This study attempts to reduce this research gap by establishing a unified theoretical framework for analogy-based TL, grounded in a commonly used concept known as \emph{similarity}. We systematically map the three core subprocesses, i.e., retrieval, adaptation and evaluation, to the fundamental challenges in TL: what, how and when to transfer. Building upon this framework, we derive two theorems governing the gain in generalization performance of analogy-based TL: (1) the theorem of unconditionally non-negative gain, which ensures safety against negative transfer through accurate evaluation; and (2) the theorem of conditionally positive gain, contingent upon either large-scale source retrieval or effective adaptation. Furthermore, we discuss the compatibility of this framework with the No Free Lunch theorem, clarifying that the efficacy of TL arises from the alignment of analogy-based inductive biases with the non-uniform distribution of problems, rather than a violation of the theorem. These insights theoretically rationalize the paradigm shift from seeking universal transferability metrics to exploiting specialized inductive biases. Monday 0.05 Paris IJCNN Paper Brain-Inspired and Cognitive Neural Systems Session Chair: Thomas Trappenberg (Dalhousie University), Punit Rathore (Indian Institute of Science) Using Disinhibition versus Direct Control in a Spiking Neural Model of Dopamine-Driven Reinforcement Learning Roberto Sautto, Nicolas Cuperlier, Thanos Manos, and Marwen Belkaid (ETIS UMR 8051, CY Cergy Paris Université, ENSEA, CNRS, F95000 Cergy, France) Abstract Abstract Dopaminergic signalling is central to value learning and decision making. It has been observed that multiple pathways with different patterns of connectivity project to midbrain dopaminergic neurons, some involving direct excitatory projections while others involve disinhibition. However, the respective contributions of these patterns to dopamine control, and their computational and functional advantages remain unclear. In the current work we simulate and evaluate two fully spiking neural models of dopaminergic control, based either solely on disinhibition, or solely on direct inhibitory and excitatory projections. We compare these models in terms of their engineering properties, their resulting spiking profiles, and their ability to successfully acquire representations of expected value in a 3-armed bandit task. We find that both models are able to operate at an asynchronous-irregular firing regime, but that the firing profile of the direct integration model is less resilient to disruption and more sensitive to incoming signals. In addition, the disinhibition model performs better in the learning task. We conclude that while the direct model is more parsimonious, disinhibition-based control remains advantageous in the operational context. Our results have implications for the study of decision-making brain circuits as well as for the design of brain-inspired systems. Emergence of Spiral Waves in Locally Connected Neural Networks Jinhang Li (Zhejiang University); Guiyang Lv (Taizhou University); and Zongli Wang, Ping Zhu, and Guoguang He (Zhejiang University) Abstract Abstract Spiral waves are a representative class of spatiotemporal activity patterns widely observed in biological nervous systems and are closely related to neural information propagation. In this work, we investigate the emergence of spiral waves in a locally connected two-dimensional neural network composed of Aihara chaotic neurons. By systematically varying key parameters, including the local connection range, refractory-related parameters, feedback strength, network size, and boundary conditions, we examine dynamical behavior of the network under the different conditions. Numerical simulations show that spiral wave formation depends sensitively on the interplay between local connectivity and network parameters, and only emerges within specific parameter regions. Changes in network structures and parameters can significantly alter the spatiotemporal evolution of activity patterns. Our results clarify how refractoriness interact with local connectivity to regulate the emergence and stability of spiral wave patterns in neural networks. The Role of Grid Cells in Reducing Spatial Aliasing in Hippocampal Place Representations Alexander Johnson, Obadah Ghizawi, and Ali A. Minai (University of Cincinnati) Abstract Abstract Spatial aliasing occurs when two or more distinct locations produce highly similar place-cell representations, primarily due to environmental symmetry or repetitive structures. This issue is most pronounced when place representations are constructed solely from boundary vector cell (BVC) inputs, because symmetric or repetitive structures can yield indistinguishable sensory patterns across multiple locations in an environment. This work introduces grid cell signals to mitigate spatial aliasing in such settings. Because grid cells contribute periodic, internally generated spatial signals that vary independently of environmental geometry, they play a key role in disambiguating perceptually identical locations. We integrate multiple modules of analytically constructed grid cells with BVC-driven place cells and show that this leads to a 94--99\% reduction in spatial aliasing relative to a BVC-only baseline across three environments: an open environment without obstacles; an environment with a cross-shaped central obstacle creating high visual symmetry; and a maze environment. The greatest improvement occurs in the environment with the highest visual symmetry. These results indicate that grid cells provide information complementary to boundary-based inputs, yielding more reliable place representations in geometrically ambiguous environments. Dual Range Adaptation Combines Experiential Scales to Learn Subjective Values through Efficient Coding Marwen Belkaid, Uma Navare, and Akram El Wathek Meraghni (ETIS UMR 8051, CY Cergy Paris Université, ENSEA, CNRS, F95000 Cergy, France) Abstract Abstract Reward magnitudes can vary substantially from a context to another. Value coding thus requires adapting neural responses to input changes to make use of their full dynamic range. Two models have recently been proposed aiming to provide a theoretical account for context-sensitive choice valuation in the framework of reinforcement learning. One model implements range adaptation relative to the minimum and maximum achievable outcomes, building on the notions of efficient coding and signal normalization. Another model borrows concepts from intrinsic motivation and posits that valuation integrates extrinsic and intrinsic rewards. Both models are able to fit human data well, but the latter was found to perform better under specific experimental conditions. In this paper, we propose that dual range adaptation over locally and globally experienced reward ranges combines the best of the two worlds. Our model benefits from the theoretical foundations and computational advantages of range adaptation. It also displays choice behaviors qualitatively similar to the model integrating intrinsic rewards. Model fitting suggests that data previously favoring the second model are best explained by our model. Our dual range adaptation model thus offers a new mechanistic account as well as testable predictions related to the representation of value for decision-making. Monday 0.10 Sydney IJCNN Paper LLM Agents and Neural Reasoning Session Chair: Chris Yakopcic (University of Dayton, OH, USA), Snehasis Banerjee (TCS Research) Integrating LLM Agents with Heterogeneous Systems via the Web of Things Seyyedeh Ensiye Kiyamousavi (Eindhoven University of Technology, GATE-AI Institute); Boris Kraychev (GATE-AI Institute); Jan Bosch (Eindhoven University of Technology); and Helena Holmström Olsson (Malmö University) Abstract Abstract Current tool-calling protocols such as the Model Context Protocol (MCP) enable AI agents to invoke external services, but their rigid transport assumptions and lack of mandatory security can be limiting in heterogeneous, event-driven IoT environments underpinning edge intelligence. This work analyzes MCP’s limitations and introduces a Web-of-Things (WoT)–based integration layer that uses standardized semantic descriptions and multiple protocol bindings to support dynamic tool discovery and subscription-based events, with security requirements declared via Thing Descriptions. We design a reference architecture around WoT and validate it using a prototype system comprising a web-based front end (implemented with Cesium map) and heterogeneous back-end tools, services, and sensor streams exposed as WoT Things. In our testbed, the prototype sustains high-frequency event interactions, restricts access to protected endpoints in automated security scanning, and completes multi-step tasks with low additional integration overhead. CADRL: Aligning LLMs with Context-Aware Disentangled Reward Learning Bo Xu (university of toronto) Abstract Abstract Reinforcement Learning from Human Feedback (RLHF) has emerged as the dominant paradigm for aligning large language models with human preferences. However, existing approaches suffer from a fundamental limitation: collapsing multifaceted human preferences into a single scalar reward signal leads to suboptimal credit assignment and reward hacking behaviors. This work introduces Context-Aware Disentangled Reward Learning (CADRL), a novel framework that decomposes reward signals into four interpretable latent factors---accuracy, coherence, style, and safety---and dynamically weighs their importance based on task context. Through a variational autoencoder architecture constrained for interpretability, the proposed meta-reward model learns disentangled factor representations while a context encoder predicts factor-specific importance weights conditioned on prompt semantics. A hierarchical policy architecture operates at two levels: high-level factor selection guided by the meta-reward model, and low-level token generation with factor-decomposed value functions. Extensive experiments on Reddit TL;DR summarization and Anthropic HH-RLHF dialogue demonstrate that CADRL achieves 6.7\% and 6.3\% improvements over strong baselines while requiring 15.2\% less training time. The factor-level advantage estimation reduces gradient variance by 40\%, enabling more stable optimization and superior sample efficiency. Process Rewards for Outcome-Guided Reasoning Steps in LLM Reinforcement Learning Mohammad Rezaei Ravari and Jens Lehmann (Technische Universität Dresden) and Sahar Vahdati (Leibniz Universität Hannover) Abstract Abstract Mathematical reasoning in large language models has improved substantially with reinforcement learning using verifiable rewards, where final answers can be checked automatically and converted into reliable training signals. Most such pipelines optimize outcome correctness only, which yields sparse feedback for long, multi-step solutions and offers limited guidance on intermediate reasoning errors. Recent work therefore introduces process reward models (PRMs) to score intermediate steps and provide denser supervision. In practice, PRM scores are often imperfectly aligned with final correctness and can reward locally fluent reasoning that still ends in an incorrect answer. When optimized as absolute rewards, such signals can amplify fluent failure modes, induce reward hacking, and destabilize policy updates. Existing PRM-based methods improve PRM quality, filter trajectories, or modify reward aggregation, but they do not directly constrain how process rewards interact with outcome correctness during optimization. We propose PROGRS, a framework that leverages PRMs while keeping outcome correctness dominant. PROGRS treats process rewards as relative preferences within outcome groups rather than absolute targets. We introduce outcome-conditioned centering, which shifts PRM scores of incorrect trajectories to have zero mean within each prompt group. It is removing systematic bias while preserving informative rankings. To stabilize process signals, PROGRS combines a frozen quantile-regression PRM with a multi-scale coherence evaluator that penalizes short-window confidence volatility. We integrate the resulting centered process bonus into Group Relative Policy Optimization (GRPO) without auxiliary objectives or additional trainable components. Across MATH-500, AMC, AIME, MinervaMath, and OlympiadBench, PROGRS consistently improves Pass@1 over outcome-only baselines (e.g., 74.9% vs.69.7% on MATH-500; 59.0% vs.52.0% on AMC-2023) and achieves stronger performance with fewer rollouts. These results show that outcome-conditioned centering enables safe and effective use of process rewards for mathematical reasoning. A Task-Aware Rank-Based Trained LLM Surrogate for Cross-Task Neural Architecture Search Kai-Chun Wu, Hao-Qun Huang, and Hsin-Yu Wang (National Sun Yat-sen University) and Chun-Wei Tsai (National Sun Yat-sen University, Department of Computer Science and Engineering) Abstract Abstract Using large language models (LLMs) to improve generalization across unseen out-of-domain tasks is promising in neural architecture search (NAS), yet it still struggles to solve regression problems precisely when performance distributions shift across domains. To address the issue, the proposed framework strategically integrates several existing methods: (1) employing rank-based contrastive loss, which focuses on learning-to-rank by the candidates relative comparisons, (2) take advantage of hierarchical pairwise sampling to emphasize fine-grained distinctions among top-performing architectures but still maintain a global ordering and (3) utilizing decoder-based large language models to enable potential inference ability gains. To verify the performance of the framework, it is evaluated on the NAS-Bench-201 and TransNAS-Bench-101 benchmarks. Experimental results provide strong evidence that the proposed strategic integration enables LLMs to adapt to regression problems and demonstrates significantly more effective transferability, achieving superior zero-shot performance compared to conventional strong predictors. Active In-Context Learning for Tabular Foundation Models Wilailuck Treerath and Fabrizio Pittorino (Politecnico di Milano) Abstract Abstract Active learning (AL) reduces labeling cost by querying informative samples, but in tabular settings its cold-start gains are often limited because uncertainty estimates are unreliable when models are trained on very few labels. Tabular foundation models such as TabPFN provide calibrated probabilistic predictions via in-context learning (ICL), i.e., without task-specific weight updates, enabling an AL regime in which the labeled context - rather than parameters - is iteratively optimized. We formalize Tabular Active In-Context Learning (Tab-AICL) and instantiate it with four acquisition rules: uncertainty (TabPFN-Margin), diversity (TabPFN-Coreset), an uncertainty-diversity hybrid (TabPFN-Hybrid), and a scalable two-stage method (TabPFN-Proxy-Hybrid) that shortlists candidates using a lightweight linear proxy before TabPFN-based selection. Across 20 classification benchmarks, Tab-AICL improves cold-start sample efficiency over retrained gradient-boosting baselines (CatBoost-Margin and XGBoost-Margin), measured by normalized AULC up to 100 labeled samples. When to ASK: Uncertainty-Gated Language Assistance for Reinforcement Learning Juarez Monteiro (Kunumi Institute), Nathan Gavenski (King's College London), and Adriano Veloso and Gianlucca Zuin (Kunumi Institute) Abstract Abstract Reinforcement learning (RL) agents often struggle with out-of-distribution (OOD) scenarios, leading to high uncertainty and random behavior. While language models (LMs) contain valuable world knowledge, larger ones incur high computational costs, hindering real-time use, and exhibit limitations in autonomous planning. We introduce Adaptive Safety through Knowledge (ASK), which combines smaller LMs with trained RL policies to enhance OOD generalization without retraining. ASK employs Monte Carlo Dropout to assess uncertainty and queries the LM for action suggestions only when uncertainty exceeds a set threshold. This selective use preserves the efficiency of existing policies while leveraging the language model’s reasoning in uncertain situations. In experiments on the FrozenLake environment, ASK shows no improvement in-domain, but demonstrates robust navigation in transfer tasks, achieving a reward of 0.95. Our findings indicate that effective neuro-symbolic integration requires careful orchestration rather than simple combination, highlighting the need for sufficient model scale and effective hybridization mechanisms for successful OOD generalization. Monday 0.11 Cape Town IJCNN Paper Robust and Adversarial Machine Learning Session Chair: Rafael de Santiago (Universidade Federal de Santa Catarina, Departamento de Informática e Estatística), Pranjala Kolapwar (SGGS Institute of Engineering and Technology) Perturbing the Phase: Analyzing Adversarial Robustness of Complex-Valued Neural Networks Florian Eilers, Christof Duhme, and Xiaoyi Jiang (University of Münster) Abstract Abstract Complex-valued neural networks (CVNNs) are rising in popularity for all kinds of applications. To safely use CVNNs in practice, analyzing their robustness against outliers is crucial. One well known technique to understand the behavior of deep neural networks is to investigate their behavior under adversarial attacks, which can be seen as worst case minimal perturbations. We design Phase Attacks, a kind of attack specifically targeting the phase information of complex-valued inputs. Additionally, we derive complex-valued versions of commonly used adversarial attacks. We show that in some scenarios CVNNs are more robust than RVNNs and that both are very susceptible to phase changes with the Phase Attacks decreasing the model performance more, than equally strong regular attacks, which can attack both phase and magnitude. Investigating the Robustness of Subtask Distillation under Spurious Correlation Pattarawat Chormai (Technische Universität Berlin); Klaus-Robert Müller (Technische Universität Berlin, Berlin Institute for the Foundations of Learning and Data); and Grégoire Montavon (Charité Universitätsmedizin Berlin, Berlin Institute for the Foundations of Learning and Data) Abstract Abstract Subtask distillation is an emerging paradigm in which compact, specialized models are extracted from large, general-purpose 'foundation models' for deployment in environments with limited resources or in standalone computer systems. Although distillation uses a teacher model, it still relies on a dataset that is often limited in size and may lack representativeness or exhibit spurious correlations. In this paper, we evaluate established distillation methods, as well as the recent SubDistill method, when using data with spurious correlations for distillation. As the strength of the correlations increases, we observe a widening gap between advanced methods, such as SubDistill, which remain fairly robust, and some baseline methods, which degrade to near-random performance. Overall, our study underscores the challenges of knowledge distillation when applied to imperfect, real-world datasets, particularly those with spurious correlations. Robustness of Fuzzy ARTMAP to Adversarial Attacks and Progressive Adversarial Training for Streaming Learning Shane Cairns (Missouri University of Science and Technology); Leonardo Enzo Brito da Silva (Instituto Metrópole Digital, Universidade Federal do Rio Grande do Norte); and Sasha Petrenko, Donald Wunsch, and Jian Liu (Missouri University of Science and Technology) Abstract Abstract Incremental learners deployed on streaming data must remain robust to evolving adversarial perturbations, yet most adversarial-robustness studies assume offline multi-epoch training with repeated access to historical data. We investigate adversarial robustness in Fuzzy ARTMAP, a prototype-based Adaptive Resonance Theory (ART) model that supports single-pass learning without replay. To enable strong adaptive evaluation despite ARTMAP’s non-differentiable winner-take-all dynamics, we propose WB-Softmax, a differentiable softmax relaxation that aggregates category-level activations into class-level scores for gradient-based attacks. Across USPS, MNIST, and Fashion-MNIST, WB-Softmax Projected Gradient Descent (PGD) achieves 89–100% attack success on vanilla models, exceeding transfer and query-based baselines at matched budgets and satisfying anti-gradient-masking sanity checks. We then study adversarial training under true streaming constraints by comparing offline versus online adversarial example generation and standard versus selective updates. Offline adversarial training consistently collapses robustness despite substantial category growth, whereas online training is dataset dependent. We introduce progressive two-stage selective training that combines selective filtering with ϵ scheduling, achieving the best overall robustness (AURAC) on all three datasets (28.2%, 64.5%, and 41.3%). Finally, leveraging ART’s explicit category geometry, we monitor incremental cluster validity indices to diagnose separation collapse and design a separation-aware absorption rule that improves high-ϵ robustness on USPS (11.0±1.4% at ϵ=0.30). Generation and Mitigation of Precision-Based Evasion Attacks in NNs Verification Andrea Gimelli, Stefano Demarchi, Davide Anguita, Armando Tacchella, and Luca Oneto (University of Genoa) Abstract Abstract Adversarial examples exploit small perturbations in the input data to manipulate the decisions of neural network classifiers. To address this challenge, verification tools have been developed to certify a range of perturbations that cannot alter the prediction of a classifier. Nevertheless, classifiers can be implemented using different precisions and arithmetic due to computational or hardware constraints. Recent studies reveal that discrepancies between the implemented and the verified classifier can be exploited to invalidate existing guarantees, since verifiers predominantly rely on floating-point arithmetic. In this work, we show that state-of-the-art verifiers (i.e., the podium of the VNN-COMP) remain vulnerable to this class of attacks by introducing a new, computationally efficient, precision-based evasion attack. We then investigate how augmenting a verifier with interval-based arithmetic can effectively mitigate these attacks, thereby enabling precision-aware verification. As our empirical evaluation confirms, our approach ensures consistent guarantees and successful mitigation. Exploring Sparsity and Smoothness of Arbitrary Lp Norms in Adversarial Attacks Christof Duhme, Florian Eilers, and Xiaoyi Jiang (University of Münster) Abstract Abstract Adversarial attacks against deep neural networks are commonly constructed under Lp norm constraints, most often using p = 1, p = 2 or p = infinity, and potentially regularized for specific demands such as sparsity or smoothness. These choices are typically made without a systematic investigation of how the norm parameter p influences the structural and perceptual properties of adversarial perturbations. In this work, we study how the choice of p affects sparsity and smoothness of adversarial attacks generated under Lp norm constraints for values of p in [1, 2]. To enable a quantitative analysis, we adopt two established sparsity measures from the literature and introduce three smoothness measures. In particular, we propose a general framework for deriving smoothness measures based on smoothing operations and additionally introduce a smoothness measure based on first-order Taylor approximations. Using these measures, we conduct a comprehensive empirical evaluation across multiple real-world image datasets and a diverse set of model architectures, including both convolutional and transformer-based networks. We show that the choice of L1 or L2 is suboptimal in most cases and the optimal p value is dependent on the specific task. In our experiments, using Lp norms with p in [1.3, 1.5] yields the best trade-off between sparse and smooth attacks. These findings highlight the importance of principled norm selection when designing and evaluating adversarial attacks. TsallisPGD: Adaptive Gradient Weighting for Adversarial Attacks on Semantic Segmentation Alexander Matyasko, Xin Lou, Indriyati Atmosukarto, and Wei Zhang (Singapore Institute of Technology) Abstract Abstract Attacking semantic segmentation models is significantly harder than image classification models because an attacker must flip thousands of pixel predictions simultaneously. Standard pixel-wise cross-entropy (CE) is ill-suited to this setting: it tends to overemphasize already-misclassified pixels, which slows optimization and overstates model robustness. To address these issues, we introduce TsallisPGD, an adversarial attack built on the Tsallis cross-entropy, a generalization of CE parameterized by $q$, which adaptively reshapes the gradient landscape by controlling gradient concentration across pixels. By varying $q$, we steer the attack toward pixels at different confidence levels. We first show that no single fixed-$q$ is universally optimal, as its effectiveness depends on the dataset, model architecture, and perturbation budget. Motivated by this, we propose a dynamic $q$-schedule that sweeps $q$ during optimization. Extensive experiments on Cityscapes, Pascal VOC, and ADE20K show that TsallisPGD, using a single validation-selected schedule, achieves the best average attack rank across all evaluated settings and improves over CEPGD, SegPGD, CosPGD, JSPGD, and MaskedPGD in reducing accuracy and mIoU on both standard and robust models. Monday 0.14 Singapore IJCNN Paper Scientific ML and Bio-Physical Sensing Session Chair: Eyad Elyan (Robert Gordon University), Mathieu Vandwalle (X-FAB) Decoding Functional Multiplicity: Graph Learning Approaches to Multifunctional Proteins Mattia Cervellini ("Sapienza" University of Rome, LUISS University); Enrico De Santis and Antonello Rizzi ("Sapienza" University of Rome); and Alessio Martino (LUISS University) Abstract Abstract Proteins are central to virtually all biological processes, and their structural diversity underlies a wide range of functions. Among them, multifunctional proteins (i.e., those associated with multiple enzymatic activities) pose unique challenges for computational analysis. In this work, we apply a suite of machine learning and deep learning techniques to predict multifunctionality roles starting from protein 3D structures. Proteins are represented as residue contact networks, enabling the use of graph machine learning approaches. We explore feature-based methods grounded in simplicial complexes analysis alongside end-to-end Graph Neural Networks implementing recent message-passing schemes. Our models address the task of multi-label classification of the first level Enzyme Commission numbers of multifunctional proteins. Evaluation across a strictly multifunctional subset of the human proteome associated with repeated stratified validation demonstrates that structural graph representations effectively capture signals of multifunctionality, shedding light on how protein architecture encodes a diverse range of biochemical roles. Neutrino Oscillation Parameter Estimation Using Structured Hierarchical Transformers Giorgio Morales (University of Caen Normandy), Gregory Lehaut and Antonin Vacheret (LPC Caen), and Frederic Jurie and Jalal Fadili (University of Caen Normandy) Abstract Abstract Neutrino oscillations encode fundamental information about neutrino masses and mixing parameters, offering a unique window into physics beyond the Standard Model. Estimating these parameters from oscillation probability maps is, however, computationally challenging due to the maps’ high dimensionality and nonlinear dependence on the underlying physics. Traditional inference methods, such as likelihood-based or Monte Carlo sampling approaches, require extensive simulations to explore the parameter space, creating major bottlenecks for large-scale analyses. In this work, we introduce a data-driven framework that reformulates atmospheric neutrino oscillation parameter inference as a supervised regression task over structured oscillation maps. We propose a hierarchical transformer architecture that explicitly models the two-dimensional structure of these maps, capturing angular dependencies at fixed energies and global correlations across the energy spectrum. To improve physical consistency, the model is trained using a surrogate simulation constraint that enforces agreement between the predicted parameters and the reconstructed oscillation patterns. Furthermore, we introduce a neural network-based uncertainty quantification mechanism that produces distribution-free prediction intervals with formal coverage guarantees. Experiments on simulated oscillation maps under Earth-matter conditions demonstrate that the proposed method is comparable to a Markov Chain Monte Carlo baseline in estimation accuracy, with substantial improvements in computational cost (around 240$\times$ fewer FLOPs and 33$\times$ faster in average processing time). Moreover, the conformally calibrated prediction intervals remain narrow while achieving the target nominal coverage of 90\%, confirming both the reliability and efficiency of our approach. Score-Based Log-Density Estimation for Anomaly Detection in Astronomical Surveys Sebastián Guzmán-Olave and Pablo A. Estévez (University of Chile) Abstract Abstract Astronomical brokers must filter millions of alerts to prioritize follow-up in surveys such as the Zwicky Transient Facility (ZTF) and the Vera C. Rubin Observatory’s Legacy Survey of Space and Time (LSST). To address this challenge we present Score-based Log-Density Network (SLDNet), an unsupervised generative model that estimates the gradient of the data log-density in a multiscale manner. Trained with denoising score matching, SLDNet defines an anomaly score using a conservative potential as a proxy for the negative log-density, avoiding the calculation of the normalization constant. Hyperspectral Image Classification with Imbalanced Data Based on a Semi-Supervised Learning Algorithm Nian Zhang and Paul Cotae (University of the District of Columbia) Abstract Abstract Hyperspectral image classification is often challenged by severe class imbalance and limited labeled samples, which significantly degrade the performance of conventional supervised learning methods. To address these issues, this paper proposes a semi-supervised learning framework that iteratively incorporates informative unlabeled samples based on neighborhood consistency to improve minority-class representation. Unlike synthetic oversampling techniques such as SMOTE, the proposed method leverages the intrinsic data distribution without generating artificial samples. Audio-to-Image Bird Species Retrieval without Audio-Image Pairs via Text Distillation Ilyass Moummad (Inria); Marius Miron (Earth Species Project); Lukas Rauch (University of Kassel); David Robinson (Earth Species Project); Alexis Joly (Inria); and Olivier Pietquin, Emmanuel Chemla, and Matthieu Geist (Earth Species Project) Abstract Abstract Audio-to-image retrieval offers an interpretable alternative to audio-only classification for bioacoustic species recognition, but learning aligned audio–image representations is challenging due to the scarcity of paired audio–image data. We propose a simple and data-efficient approach that enables audio-to-image retrieval without any audio–image supervision. Our proposed method uses text as a semantic intermediary: we distill the text embedding space of a pretrained image–text model (BioCLIP-2), which encodes rich visual and taxonomic structure, into a pretrained audio–text model (BioLingual) by fine-tuning its audio encoder with a contrastive objective. This distillation transfers visually grounded semantics into the audio representation, inducing emergent alignment between audio and image embeddings without using images during training. We evaluate the resulting model on multiple bioacoustic benchmarks. The distilled audio encoder preserves audio discriminative power while substantially improving audio–text alignment on focal recordings and soundscape datasets. Most importantly, on the SSW60 benchmark, the proposed approach achieves strong audio-to-image retrieval performance exceeding baselines based on zero-shot model combinations or learned mappings between text embeddings, despite not training on paired audio–image data. These results demonstrate that indirect semantic transfer through text is sufficient to induce meaningful audio–image alignment, providing a practical solution for visually grounded species recognition in data-scarce bioacoustic settings. Physics-Based modeling for Synthetic Urban Flood Rendering from Monocular Video MD Fayaz Bin Hossen, Ahmed Temtam, Kwame AMPOFO, and M. Iftekharuddin Khan (Old Dominion University) Abstract Abstract Reliable reconstruction of urban flooding is important for urban planning, flood monitoring, and risk analysis; however, progress is hindered by the limited availability of datasets with usable scene geometry and flood-depth annotations. Existing approaches often rely on LiDAR scanning, 3D reconstruction, or purely generative models, which can be computationally expensive and may not enforce physically realistic fluid dynamics over time. Moreover, obtaining labeled ground truth data —particularly flood depth —is costly and difficult to scale. To address these limitations, we propose a monocularvideo- driven framework that combines open-vocabulary semantic segmentation, monocular depth estimation, and a lightweight physics-based simulation model using Shallow Water Equations (SWE). The simulation operates on scene-derived terrain and dynamic obstacles to produce temporally consistent flooding evolution, which is then used to synthesize flooding imagery and corresponding dense supervision signals. As a proof of concept, we also demonstrate a physics-informed neural surrogate model that regularizes predictions with SWE residuals. Overall, the proposed framework enables efficient generation of physically grounded synthetic data for training and benchmarking flood segmentation and depth estimation models, supporting data-driven urban resilience and disaster preparedness. Monday 0.15 Washington IJCNN Paper IJCNN SS17 Generative Foundation Models for Robotics: From Language and Vision to Embodied Action Session Chair: Erdi Sayar (Paderborn University), Alper Yegenoglu (Paderborn University), Van Huyen Dang (Paderborn University), Erdal Kayacan (Paderborn University, Germany) Beyond Geometry: Leveraging Pose and Visual Scene Understanding for Indoor Room Segmentation in Mobile Robots Guilherme Goes Zanetti and Elio David Triana Rodriguez (Universidade Federal do Espírito Santo); Thiago Paixão (Instituto Federal do Espírito Santo); and Filipe Mutz, Claudine Badue, Alberto F. De Souza, Anselmo Frizera, and Thiago Oliveira-Santos (Universidade Federal do Espírito Santo) Abstract Abstract Accurate residential indoor room segmentation is crucial for mobile robots, but traditional geometric methods often fail in cluttered environments, while vision-based approaches are still dependent on geometric analysis, and therefore limited to simple environments with clear geometric boundaries. We propose a framework that can handle complex room boundaries by using a robot's visual input to simultaneously segment and classify rooms. This is achieved by projecting image-based room predictions onto a 2D occupancy map using spatial aggregation techniques. By leveraging the vision cone of the robot and powerful image classification models such as CNNs, ViTs and VLMs, the map is segmented into rooms with a higher Intersection over Union (IoU) than traditional methods while handling environments with complex geometry. The results demonstrate that our method using VLMs and vision cone aggregation outperforms other approaches, achieving up to 76.22\% classification accuracy on ArchitectThor room segmentation without any task-specific training, offering superior accuracy and flexibility for robotic semantic mapping, while overcoming the geometry-dependency of traditional methods. We also provide the code and improved dataset. BTGenBot-2: Efficient Behavior Tree Generation with Small Language Models Riccardo Andrea Izzo, Gianluca Bardaro, and Matteo Matteucci (Politecnico di Milano) Abstract Abstract Recent advances in robot learning increasingly rely on LLM-based task planning, leveraging their ability to bridge natural language with executable actions. While prior works showcased great performances, the widespread adoption of these models in robotics has been challenging as 1) existing methods are often closed-source or computationally intensive, neglecting the actual deployment on real-world physical systems, and 2) there is no universally accepted, plug-and-play representation for robotic task generation. Addressing these challenges, we propose BTGenBot-2, a 1B-parameter open-source small language model that directly converts natural language task descriptions and a list of robot action primitives into executable behavior trees in XML. Unlike prior approaches, BTGenBot-2 enables zero-shot BT generation, error recovery at inference and runtime, while remaining lightweight enough for resource-constrained robots. We further introduce the first standardized benchmark for LLM-based BT generation, covering 52 navigation and manipulation tasks in NVIDIA Isaac Sim. Extensive evaluations demonstrate that BTGenBot-2 consistently outperforms GPT-5, Claude Opus 4.1, and larger open-source models across both functional and non-functional metrics, achieving average success rates of 90.38% in zero-shot and 98.07% in one-shot, while delivering up to 16× faster inference compared to the previous BTGenBot. SAFER: Incorporating Visual Context for Multitask Failure Detection in VLA Models Puren Tap, Rafi Putra Abdurrahman, and Sanem Sarıel (Istanbul Technical University) and Sinan Kalkan (Middle East Technical University) Abstract Abstract As Vision-Language-Action (VLA) models are increasingly used in complex, multi-task robotic setups, automatic detection of execution failures has become critical for safety and reliability. Recent studies have demonstrated that latent action embeddings of a VLA can be leveraged to predict task success. However, relying solely on action embeddings neglects the environmental context that often precipitates failure. In this paper, we address this gap by first demonstrating that the distributions of image embeddings correlate with failure scenarios. Motivated by this finding, we propose SAFER, a fusion approach that integrates both action and image embeddings into a dedicated failure prediction module. Through extensive experiments, we show that image embeddings provide complementary information to action embeddings, providing up to 8% point improvement in seen and 2% point improvement on unseen tasks with 5.7x times faster inference over the state-of-the-art for detecting failures. Overall, SAFER underscores the necessity of visual context in monitoring the internal reasoning of VLA models and provides a more robust framework for autonomous failure detection. Code and trained models are available at: https://github.com/purentap/safer Palette Inpainting Diffusion Curriculum Reinforcement Learning (PIDCRL) Faruk Oruç (Technical University of Munich), Erdi Sayar (Paderborn University), Giovanni Iacca (University of Trento), Alois Knoll (Technical University of Munich), and Erdal Kayacan (Paderborn University) Abstract Abstract Curriculum reinforcement learning (CRL) aims to speed up agent learning by organizing tasks in a progressively harder sequence. However, many existing CRL methods struggle to guide agents toward meaningful goals, especially when domain knowledge is limited or unavailable. To overcome this limitation, we introduce PIDCRL, a diffusion-based CRL framework that leverages a mask-conditional, image-to-image pretrained diffusion model to automatically generate curriculum goals from agent trajectory heatmaps. Given a heatmap trajectory image as input, the model produces candidate curriculum goals, which we then filter and select using several strategies, including averaging, Q- value scoring, and a trainable reward-based mechanism. These strategies identify curriculum goals that are both achievable and appropriately challenging, enabling effective curricula without expert-designed heuristics. We validate PIDCRL across three maze environments, where it matches or outperforms ten state-of-the-art CRL baselines. Using Large Language Models to Guide Purpose-Driven Grounded Learning in Autonomous Robots Sergio Martínez-Alonso, Emanuel Fallas-Hernández, Alejandro Romero, Jose Antonio Becerra, and Richard J. Duro (Universidade da Coruña) Abstract Abstract Large Language Models (LLMs) are increasingly used in autonomous robotics, often in end-to-end frameworks that map natural language instructions directly to actions. While effective in constrained settings, such approaches suffer from limited grounding, poor explainability, and weak knowledge reuse. We propose an alternative integration strategy in which LLMs act as a high-level cultural knowledge repository that informs, but does not control, a robotic cognitive architecture. All perception, learning, and action selection remain grounded in embodied interaction. We show that leveraging LLM-provided priors during exploration in previously unknown domains significantly accelerates the acquisition of grounded goals compared to intrinsically motivated exploration. Experiments in simulation and on a real robot demonstrate improved early learning efficiency, goal discovery, and interpretability. Our results suggest that embedding LLMs within hybrid cognitive architectures enables scalable, explainable, and open-ended robotic learning. Fuzzy Logic Theory-based Adaptive Reward Shaping for Robust Reinforcement Learning (FARS) Hürkan Sahin, Van Huyen Dang, Erdi Sayar, Alper Yegenoglu, and Erdal Kayacan (Paderborn University) Abstract Abstract Reinforcement learning (RL) often struggles in real-world tasks with high-dimensional state spaces and long horizons, where sparse or fixed rewards severely slow down exploration and cause agents to get trapped in local optima. This paper presents a fuzzy-logic-based reward shaping method that integrates human intuition into RL reward design. By encoding expert knowledge into adaptive, interpretable terms, fuzzy rules promote stable learning and reduce sensitivity to hyperparameters. The proposed method leverages these properties to adapt reward contributions based on the agent’s state, enabling smoother transitions between fast motion and precise control in challenging navigation tasks. The extensive simulation results on autonomous drone racing benchmarks show stable learning behavior and consistent task performance across scenarios of increasing difficulty. The proposed method achieves faster convergence and reduced performance variability across training seeds in more challenging environments, with success rates improving by up to approximately 5% compared to non-fuzzy reward formulations. Monday 2.2 Zambezi IJCNN Paper IJCNN SS09 Multimodal Deep Learning in Applications Session Chair: Stefania Tomasiello (University of Salerno), Dawid Połap (Silesian University of Technology, Poland) Aligning Hierarchical Vision Transformers as Multi-Scale Dense Feature Extractors Rebanta Dey, Tribikram Dhar, and Samrat Dutta (TCS Research) Abstract Abstract Dense prediction requires representations that are multi-scale, semantically aligned, and resolution stable, yet existing hierarchical vision transformers are task-specific, optimization-sensitive, and inefficient under scale variation. We propose MAHViT, a self-supervised hierarchical vision transformer designed as a multi-scale dense feature extractor. MAHViT introduces a dual-Layer Normalization data-flow that stabilizes deep hierarchical attention by jointly preserving gradient flow and feature diversity, enabling scalable self-supervised pre-training. To efficiently connect hierarchical encoders with dense decodes, we propose a Feature Diversity Sampling (FDS) neck that suppresses redundant tokens while retaining semantically distinct features. Pre-trained using a BYOL-based objective with resolution-adaptive alignment, MAHViT encoder learns scale-consistent representations. The encoder is evaluated on semantic segmentation task by pairing a Progressive Upsampling decoder with the FDS Neck. Through extensive evaluations, MAHViT achieves state-of-the-art or competitive multi-scale mIoU on COCO-2017, ADE20K, and Cityscapes across multiple model scales, while maintaining favorable FLOPs–accuracy trade-offs suitable for edge deployment. Data-lightweight particulate matter forecasting using temporal attention and Kolmogorov-Arnold Networks Jaroslaw Bernacki (Wroclaw University of Science and Technology) and Rafal Scherer (Czestochowa University of Technology, AGH University of Krakow) Abstract Abstract Air pollution, particularly particulate matter (PM) concentration, poses a serious threat to public health and the environment. Accurate short-term forecasting of PM levels based on historical concentration data is essential for timely decision-making, public warning systems, and effective air quality management. In this study, we propose an interpretable and data-lightweight approach for predicting PM2.5 and PM10 concentrations using a model that combines temporal attention with a Kolmogorov-Arnold Network. The proposed model is evaluated using real-world air quality data collected without using meteorological or other auxiliary data. Its performance is assessed using standard error metrics, including mean absolute error (MAE), mean squared error (MSE), and root mean squared error (RMSE). The results demonstrate that the proposed model consistently outperforms state-of-the-art deep learning and statistical methods. AutoML-Based Fake Review Detection: A Unified Textual and Multimodal Learning Framework Tayasan Milinda Hewadaunda Gedara (IMT School for Advanced Studies Lucca, University of Salerno) and Rocco Loffredo and Nadeeka Malkanthi Kiringoda Arachchige (University of Salerno) Abstract Abstract Fake reviews undermine the credibility of online platforms, yet most existing detection methods rely on manually designed pipelines with extensive feature engineering and heuristic model selection, limiting robustness and cross-domain generalization. To address these challenges, this paper proposes a unified AutoML-based framework for fake review detection that automates model selection, feature optimization, and hyperparameter tuning in an end-to-end manner, evaluated in both text-only and multimodal settings. The framework is evaluated across multiple AutoML systems and five benchmark datasets, including Amazon, FRD, Yelp-NYC, TripAdvisor, and a synthetic E-Commerce dataset, and integrates textual, numerical, and visual information during the multimodal stage. Experimental results show that multimodal AutoML consistently outperforms traditional machine learning models and remains competitive with state-of-the-art deep learning approaches, achieving up to 10% accuracy increment and 9.1% F1-score, demonstrating AutoML’s effectiveness as a scalable and deployment-ready solution. Parameter-Efficient Score Calibration for Text-Image Retrieval: Interpretable Thresholding with Vision-Language Models Taku Fujitomi, Naoya Sogi, Takashi Shibata, and Makoto Terao (NEC Corporation) Abstract Abstract Embedding-based retrieval with Vision-Language Models (VLMs) enables high-performance text-image retrieval by projecting queries and candidates into a shared vector space and ranking them via cosine similarity. However, setting relevance thresholds manually requires effort and lacks adaptability to differences across users, queries, and models. To overcome this, we introduce a fast, lightweight score calibration method that normalizes cosine similarity scores using a formula based on the mean of estimated non-relevant candidates. This approach lessens sensitivity to candidate set size and relevance ratio. By interpreting the calibrated scores as relevance probabilities, we can estimate precision and recall at any threshold, making it possible for users to intuitively specify preferred metrics and directly control results. Experiments on diverse datasets show that our calibration method improves text-image relevance classification at global thresholds and delivers more stable retrieval quality according to user-specified criteria, allowing intuitive and consistent control over the relation between threshold and retrieval performance. Monday 0.04 Brussels IEEE CEC (Evolutionary Computation), CEC Position Paper CEC 7 - Large Language Models and Neuromorphic EC Session Chair: Mengjie Zhang (Victoria University of Wellington, Centre for Data Science and Artificial Intelligence) Position Paper: Suggestions from Large Language Models about Experimental Settings for Benchmarking Evolutionary Multi-objective Optimization Algorithms Lie Meng Pang and Hisao Ishibuchi (Southern University of Science and Technology) Abstract Abstract Benchmarking is essential for the evaluation of evolutionary multi-objective optimization (EMO) algorithms, yet benchmarking results are sensitive to experimental settings such as test problems, compared algorithms, performance indicators, and their parameter specifications. Recently, large language models (LLMs) have been increasingly used in EMO research and academic activities (such as assisting with paper reviews). As a result, LLMs may influence how benchmarking studies are conducted, especially for newcomers to the EMO field. This position paper examines the reliability of benchmarking suggestions generated by two LLMs by comparing their responses with those reported in earlier studies at IEEE CEC 2025. Using the same prompts and experimental settings, we analyze similarities and differences in the suggested algorithms, test problems, and performance indicators for benchmarking EMO algorithms between the two LLMs and also between the two years. In addition, we report benchmarking suggestions from the LLMs for constrained multi-objective optimization. The observations indicate that, while many LLM-generated suggestions remain consistent with widely used benchmarking practices, variations and inaccuracies are also observed. These findings highlight the need for careful interpretation of LLM-generated suggestions and emphasize the importance of benchmarking guidelines when evaluating EMO algorithms. Towards Neuromorphic Evolutionary Computation: A Position Paper Jorge Mario Cruz-Duarte and El-Ghazali Talbi (Université de Lille) Abstract Abstract Neuromorphic Computing (NC) and Evolutionary Computation (EC) are mature, yet largely fragmented. This position paper argues that their convergence is timely and technically plausible. We define Neuromorphic Evolutionary Computation (NEC) as a paradigm that bridges NC and EC. We review the current landscape and its limitations, with an emphasis on the prevalence of hybrid execution and incomplete energy-latency reporting. We then present a probable Cortico-Basal-Thalamic Loop (CBTL)-inspired conceptual architecture for operator selection and outline what must be implemented on-chip to achieve an NEC approach, including operator embodiment, associative memory, and state estimation. We summarise open challenges spanning numerical precision, scalability, learning rules, benchmarking, theory, and tooling, including the need for convergence analysis framed in terms of the best-so-far process. We close with a phased roadmap and concrete reporting requirements towards a reproducible NEC ecosystem. Fast and Adaptive Multi-Objective Feature Selection for Classification Fei Ming (Westlake University), Bing Xue and Mengjie Zhang (Victoria University of Wellington), and Yaochu Jin (Westlake University) Abstract Abstract Identifying high-quality feature subsets for decision-makers by wrapper-based multi-objective feature selection (MOFS) has been attracting increasing attention. As a mainstream approach, evolutionary methods offer distinct advantages but also struggle with increasingly complex application scenarios and high-dimensional data, mainly in terms of time efficiency, exponentially expanding search space, and adaptive determination of a suitable classifier. To overcome these challenges, this paper proposes two simple yet effective methods: Fast Initialization (FI) and one-generation Adaptive K-Nearest Neighbor (AK). FI leverages mutual information and tournament selection to locate high-quality initial feature subsets computationally efficiently. AK verifies that, with theoretical proof, using a single generation can determine the most suitable KNN for different data to improve feature selection performance with very little time overhead and without any data analysis or assumption. Experiments on 20 real-world high-dimensional datasets demonstrate the superior performance of FI and AK to advanced initialization and KNN methods for MOFS. We also validated that the obtained feature subsets generalize well to an LLM for tabular data, enabling it to be seamlessly applied to high-dimensional data and achieve superior performance. Monday 0.10 Sydney IJCNN Position Paper Foundations of Neural and Scientific ML (Position Track) Session Chair: Donald Wunsch (Missouri University of Science and Technology), Alessio Martino (Dept. of AI, Data and Decision Sciences, LUISS University) Position Paper: Neurotransmitters as a Missing Dimension in Artificial Neural Networks Yupei Li (Imperial College London, Technical University of Munich); Manuel Milling (Technical University of Munich); Berrak Sisman (Johns Hopkins University); and Björn W. schuller (Technical University of Munich, Imperial College London) Abstract Abstract Artificial neural networks (ANNs), as core components of modern deep learning (DL) systems, lack the adaptive flexibility and long-term stability exhibited by biological systems. This limitation largely stems from the fact that conventional ANNs rely on uniform, local, and gradient-based parameter updates, while neglecting internal learning principles that are biological mechanisms such as neurotransmitters signalling or neuroplasticity. Consequently, many existing approaches focus on architectural expansion or mathematical fine-tuning techniques such as regularisation or parameter isolation. Inspired by the superior adaptability and plasticity of mammalian brains, we posit that neuromodulation with neurotransmitters constitutes a third axis of learning, complementary to neural activity and synaptic plasticity, and should be explicitly modelled in artificial neural networks. In this positional paper, we argue that incorporating neuromodulatory principles into ANN design represents a promising and underexplored research direction, and we advocate for greater attention to this perspective in the development of adaptive and continual learning systems. Position Paper: Deep Learning in Hilbert Lattices Vassilis Kaburlasos (International Hellenic University) Abstract Abstract This work proposes two necessary and sufficient conditions for neural computing namely, first, a Hilbert space H and, second, non-linear transformations in H. The conventional deep learning (neural computing) paradigm is extended hierarchically from the Euclidean (Hilbert) space R^N to the Hilbert space G_1^N, where G_1 is a quotient of equivalence classes of Generalized Intervals' Numbers (GINs). The convex cone F_1 of G_1 is considered, where an element of F_1, namely Intervals' Number (IN), is an (information) granule that may represent a probability/possibility distribution including real numbers. It turns out that F_1 is a sublattice of lattice G_1, where partial order represents semantics; moreover, axiomatic logic can be applied in F_1. Because all the required mathematical instruments are available, neural computing can be pursued in F_1^N including logic/reasoning toward eXplainable Artificial Intelligence (XAI). The far-reaching potential of proposed tools in practical applications is discussed. Position Paper: Unlocking the Potential of AI Researchers in Scientific Discovery—What Is Missing? Hengjie Yu, Shuya Liu, Haiyun Yang, and Yuping Yan (Westlake University); Maozhen Qu (National University of Singapore); and Yaochu Jin (Westlake University) Abstract Abstract The potential of AI researchers in scientific discovery remains largely untapped. Over the past decade, AI for Science (AI4Science) publications in 145 Nature Index journals have increased fifteen-fold, yet they still account for less than 3% of the total publications. Drawing upon the Diffusion of Innovation theory, we project AI4Science’s share of total publications to rise from 2.72% in 2024 to approximately 20% by 2050. Achieving this shift requires fully harnessing the potential of AI researchers, as nearly 95% of AI-driven research in these journals is led by experimental scientists. To facilitate this, we propose structured workflows and strategic interventions to position AI researchers at the forefront of scientific discovery. Specifically, we identify three critical pathways: equipping experimental scientists with accessible AI tools to amplify the impact of AI researchers, bridging cognitive and methodological gaps to enable more direct involvement in scientific discovery, and proactively fostering a thriving AI-driven scientific ecosystem. By addressing these challenges, we aim to empower AI researchers as key drivers of future scientific breakthroughs. Monday 0.14 Singapore IJCNN Paper Deep learning for computer vision Session Chair: Alper Yegenoglu (Paderborn University), Stefania Tomasiello (University of Salerno) EEG-MFTNet: An Enhanced EEGNet Architecture with Multi-Scale Temporal Convolutions and Transformer Fusion for Cross-Session Motor Imagery Decoding Panagiotis Andrikopoulos and Siamak Mehrkanoon (Utrecht University) Abstract Abstract Brain-computer interfaces (BCIs) enable direct communication between the brain and external devices, providing critical support for individuals with motor impairments. However, accurate motor imagery (MI) decoding from electroencephalography (EEG) remains challenging due to noise and cross-session variability. This study introduces EEG-MFTNet, a novel deep learning model based on the EEGNet architecture, enhanced with multi-scale temporal convolutions and a Transformer encoder stream. These components are designed to capture both short and long-range temporal dependencies in EEG signals. The model is evaluated on the SHU dataset using a subject-dependent cross-session setup, outperforming baseline models, including EEGNet and its recent derivatives. EEG-MFTNet achieves an average classification accuracy of 58.9% while maintaining low computational complexity and inference latency. The results highlight the model’s potential for real-time BCI applications and underscore the importance of architectural innovations in improving MI decoding. This work contributes to the development of more robust and adaptive BCI systems, with implications for assistive technologies and neurorehabilitation. SCoRe: Clean Image Generation from Diffusion Models Trained on Noisy Images Yuta Matsuzaki, Seiichi Uchida, and Shumpei Takezaki (Kyushu University) Abstract Abstract Diffusion models trained on noisy datasets often reproduce high-frequency training artifacts, significantly degrading generation quality. To address this, we propose SCoRe (Spectral Cutoff Regeneration), a training-free, generation-time spectral regeneration method for clean image generation from diffusion models trained on noisy images. Leveraging the spectral bias of diffusion models, which infer high-frequency details from low-frequency cues, SCoRe suppresses corrupted high-frequency components of a generated image via a frequency cutoff and regenerates them via SDEdit. Crucially, we derive a theoretical mapping between the cutoff frequency and the SDEdit initialization timestep based on Radially Averaged Power Spectral Density (RAPSD), which prevents excessive noise injection during regeneration. Experiments on synthetic (CIFAR-10) and real-world (SIDD) noisy datasets demonstrate that SCoRe substantially outperforms post-processing and noise-robust baselines, restoring samples closer to clean image distributions without any retraining or fine-tuning. Monday 0.15 Washington IJCNN Paper IJCNN SS06 Collaborative Learning of Trustworthy Computational Intelligence Systems (CLOTHES 2026) - Third Edition Session Chair: Pietro Ducange (University of Pisa, Italy) Anticipating Narrative Impact through Collaborative LLM-based Multi-Agent System Maria Di Gisi (IMT School for Advanced Studies Lucca); Giuseppe Fenza, Domenico Furno, and Mariacristina Gallo (University of Salerno); and Pio Pasquale Trotta (IMT School for Advanced Studies Lucca) Abstract Abstract The online information ecosystem is shaped by collaborative and adaptive processes in which narratives evolve through collective interpretation and reframing. Anticipating these dynamics is essential for developing trustworthy and human-centered AI systems that can provide proactive, interpretable support to human decision-makers before problematic narratives reach critical diffusion. However, most existing approaches rely on post-hoc analyses, offering limited utility for early intervention and pre-bunking strategies. This paper presents a collaborative multi-agent simulation framework based on Large Language Models (LLMs) to study the evolution of news narratives. Heterogeneous agents interact within a controlled network to iteratively generate and reframe news headlines, producing structured trajectories of narrative transformation. A notion of simulated Narrative Impact to quantify a news item’s capacity to organize, sustain, and attract collective narrative activity over time, is formulated. Building on this formulation, an early alerting task predicts the structural impact of a narrative from partial observations of the simulation, enabling the identification of high-impact narratives at early stages. Experimental results, obtained using an XGBoost-based predictive model, demonstrate that the Narrative Impact can be accurately identified (R2 = 0.71) at early stages, even in the absence of intentional misinformation. Multi-perspective Poisoning Detection using Graph Convolutional Network Mario Luca Bernardi (University of Sannio) and Marta Cimitile and Anna Vacca (UnitelmaSapienza) Abstract Abstract Federated Learning (FL) enables collaborative model training across distributed nodes while preserving data privacy. However, its decentralized architecture introduces significant security vulnerabilities, particularly susceptibility to poisoning attacks where malicious participants inject misleading information into the training process. In this paper, we propose a Multi-perspective Poisoning Detection approach that based on a Graph Convolutional Network architecture that integrates both temporal and spatial perspectives within a unified framework. This integration enables more robust and comprehensive drift analysis by applying simultaneous spatio-temporal evaluation to all detected drift instances. We evaluate the proposed approach on a new dataset generated starting from a public dataset. Experimental results demonstrate that the integrated hierarchical approach achieves superior detection accuracy with respect to the plain transformer-based autoencoder baseline and reduced false positive rates, enhancing the resilience of federated learning systems against sophisticated poisoning attacks while maintaining adaptability to legitimate concept drift. FedE2A: Large Language Model Agents for Energy-Aware Robust Federated Learning Muhammad Usman and Mario Luca Bernardi (University of Sannio), Marta Cimitile (Unitelm sapienza), and Samiksha Bisht (University of Sannio) Abstract Abstract Federated learning in resource-constrained environ-ments poses a complex multi-objective optimization problemcharacterized by extreme data heterogeneity, stochastic client availability, and fluctuating energy constraints. Traditional clientselection mechanisms, relying predominantly on static heuristicsor single-objective optimization, lack the flexibility to navigate thedynamic Pareto frontier between data utility, system robustness,and energy efficiency. We propose FedE2A, a framework that reframes client selection as a context-aware decision process governed by a Large Language Model (LLM). In this architecture, clients are abstracted into structured Quality Profile Vectors (QPV) that encapsulate strictly local states—including epistemic uncertainty via Monte Carlo Dropout, system health, and historical reliability. The server-side LLM acts as a reasoning engine, interpreting these high-dimensional profiles to dynamically arameterize the selection policy based on the current training phase and convergence trajectory. Unlike rigid rule based systems, this neuro-symbolic approach allows for semantic interpretation of system anomalies and adaptive balancing of competing objectives. Extensive experiments on CIFAR-10 under non-IID settings and label noise demonstrate that FedE2A yields significant improvements in convergence stability and energy efficiency compared to state-of-the-art baselines, establishing LLM-based governance as a principled methodology for robust orchestration in heterogeneous federated systems. Monday 2.18 Mekong IJCNN Paper SS30 Computational Intelligence and AI Applications for Sustainable Energy Management in Smart Grids and Energy Communities Session Chair: Enrico De Santis (University of Rome "La Sapienza") Black-Box Adversarial Attacks on Smart Grid Stability Score Prediction and a Stochastic Quantile Smoothing Defense Emad Efatinasab, Mirco Rampazzo, and Alessandro Brighente (University of Padua); Chuadhry Mujeeb Ahmed (Newcastle University); and Mauro Conti (University of Padua, Örebro University) Abstract Abstract Regression-based stability prediction is a core component of decentralized smart grid control, yet its security under adversarial data-injection attacks remains largely unexplored. Unlike prior work that focuses on binary stability classification, many operational systems rely on continuous stability scores that directly influence control decisions. In this paper, we study black-box adversarial attacks against regression-based stability prediction under realistic query and injection budget constraints. We adapt and evaluate multiple query-based derivative-free attacks designed to directly maximize regression output distortion rather than misclassification. We further introduce an uncertainty-aware sampling strategy that allows an attacker to concentrate a limited injection budget on the most vulnerable inputs, significantly amplifying attack impact. A Physics-Guided Framework for Solar-Aware Load Disaggregation in PV-Integrated Households Muhammad Affan Khan, Giulia Tanoni, Enrik Xhani, Stefano Squartini, and Emanuele Principi (Università Politecnica delle Marche) Abstract Abstract The ever increasing adoption of residential photovoltaic (PV) systems fundamentally alters true household consumption and therefore challenges the applicability of conventional Non-Intrusive Load Monitoring (NILM) techniques. In this paper, we propose a physics-guided neural network for the disaggregation of household appliances, as well as behind the meter solar generation using the net load signal and local irradiance measurements. Unlike conventional methods that regress solar power directly, our method introduces a latent Correction Factor (CF) that is estimated alongside appliance power profiles. This predicted CF scales the irradiance input to reconstruct solar generation, strictly enforcing physical constraints (e.g. zero output under no irradiance) directly into the learning process. This formulation enables the CF to implicitly capture PV system efficiency variations due to environmental and operational factors without explicit supervision. Experimental evaluation on real world datasets demonstrate that the proposed approach achieves improved disaggregation accuracy and enhanced consistency compared to unconstrained solar regression baselines in effectively decoupling household consumption from distributed generation in PV-integrated environments. Reconfiguration-Robust Spatiotemporal Learning for Dynamic EV Hosting Capacity Estimation in Distribution Feeders Md Touhidul Haque and Petr Musilek (University of Alberta) Abstract Abstract The rapid growth of electric vehicle charging loads is creating new operational stress in power distribution systems, where additional demand can trigger voltage violations and equipment overloads. To manage this risk, utilities use hosting capacity analysis to assess how much new charging load can be accommodated at each location without violating operating limits. Existing learning-based estimators can provide fast predictions, but many implicitly assume a fixed network configuration and therefore their accuracy degrades when switching operations alter feeder topology during maintenance, restoration, or load balancing. This paper introduces a topology-aware spatiotemporal learning framework that predicts time-varying hosting capacity at the bus level while remaining robust to feeder reconfiguration. The method represents the feeder as a graph with physically meaningful edge descriptors and performs joint spatial-temporal reasoning over a short history window using an edge-conditioned attention mechanism. Experiments on a real-world, smart-meter-informed distribution system under multiple reconfiguration scenarios show that the proposed approach maintains high accuracy on unseen topologies and consistently outperforms topology-agnostic baselines. These results indicate that explicitly modeling feeder topology and equipment limits in data-driven frameworks is essential for reliable, dynamic electric vehicle hosting capacity assessment in adaptive distribution networks. Monday 2.2 Zambezi IJCNN Paper IJCNN SS26 Brain Machine Intelligence: Models, Systems, and Translational Applications Session Chair: Marwen Belkaid (ETIS UMR 8051, CY Cergy Paris Université, ENSEA, CNRS, F95000 Cergy, France), PANTEA KEIKHOSROKIANI (University of Oulu) Cost-Efficient Multi-Scale Fovea for Semantic-Based Visual Search Attention João Luzio, Alexandre Bernardino, and Plinio Moreno (Institute for Systems and Robotics, Instituto Superior Técnico) Abstract Abstract Semantics are one of the primary sources of top-down preattentive information. Modern deep object detectors excel at extracting such valuable semantic cues from complex visual scenes. However, the size of the visual input to be processed by these detectors can become a bottleneck, particularly in terms of time costs, affecting an artificial attention system's biological plausibility and real-time deployability. Inspired by classical exponential density roll-off topologies, we apply a new artificial foveation module to our novel attention prediction pipeline: the Semantic-based Bayesian Attention (SemBA) framework. We aim at reducing detection-related computational costs without compromising visual task accuracy, thereby making SemBA more biologically plausible. The proposed multi-scale pyramidal field-of-view retains maximum acuity at an innermost level, around a focal point, while gradually increasing distortion for outer levels to mimic peripheral uncertainty via downsampling. In this work we evaluate the performance of our novel Multi-Scale Fovea, incorporated into SemBA, on target-present visual search. We also compare it against other artificial foveal systems, and conduct ablation studies with different deep object detection models to assess the impact of the new topology in terms of computational costs. We experimentally demonstrate that including the new Multi-Scale Fovea module effectively reduces inherent processing costs while improving SemBA's scanpath prediction accuracy. Remarkably, we show that SemBA closely approximates human consistency while retaining the actual human fovea's proportions. Enhancing Efficiency and Performance in Deepfake Audio Detection through Neuron-level Dropin & Neuroplasticity Mechanisms Yupei Li (Imperial College London, Technical University of Munich); Shuaijie Shao (University College London); Manuel Milling (Technical University of Munich); and Björn Schuller (Technical University of Munich, Imperial College London) Abstract Abstract Current audio deepfake detection has achieved remarkable performance using diverse deep learning architectures such as ResNet, and has seen further improvements with the introduction of large models (LMs) like Wav2Vec. The success of large language models (LLMs) further demonstrates the benefits of scaling model parameters, but also highlights one bottleneck where performance gains are constrained by parameter counts. Simply stacking additional layers, as done in current LLMs, is computationally expensive and requires full retraining. Furthermore, existing low-rank adaptation methods are primarily applied to attention-based architectures, which limits their scope. Inspired by the neuronal plasticity observed in mammalian brains, we propose novel algorithms, dropin and further plasticity, that dynamically adjust the number of neurons in certain layers to flexibly modulate model parameters. We evaluate these algorithms on multiple architectures, including ResNet, Gated Recurrent Neural Networks, and Wav2Vec. Experimental results using the widely recognised ASVSpoof2019 LA, PA, and FakeorReal dataset demonstrate consistent improvements in computational efficiency with the dropin approach and a maximum of around 39% and 66% relative reduction in Equal Error Rate with the dropin and plasticity approach among these dataset, respectively. The code and supplementary material are available at Anonymous link. Monday 2.1 Volga IJCNN Paper Explainable AI Session Chair: Naoyuki Kubota (Tokyo Metropolitan University), Imen Jdey (REGIM Lab, Sfax University), Simona Casini (University of Pisa) Hyperparameter Transfer Using Item Response Theory Takeaki Sakabe and Yuko Sakurai (Nagoya Institute of Technology); Hisashi Kashima (Kyoto University); and Satoshi Oyama (Nagoya City University, RIKEN Center for Advanced Intelligence Project) Abstract Abstract Hyperparameter optimization (HPO) is computationally expensive due to large search spaces and dataset-dependent optimal configurations. To reduce this cost, transfer-based HPO methods have been proposed. However, existing methods rely solely on observed performance metrics, failing to disentangle the latent capability of hyperparameter configurations from dataset characteristics, which can cause biased transfer. In this paper, we propose a hyperparameter transfer method using Item Response Theory (IRT), a testing theory that enables the independent estimation of problem characteristics and examinee capability. By associating hyperparameter configurations with examinees and individual data samples with test items, the proposed method estimates the latent capability of configurations on source datasets. This is then combined with the dataset characteristics estimated from a small number of evaluations on a target dataset, enabling efficient hyperparameter transfer. Experimental results on multiple datasets demonstrate that the proposed method can select effective hyperparameter configurations with fewer evaluation trials compared to an existing method. Diagnosing Neural Convergence with Topological Alignment Spectra Tiago Fernandes Tavares and Fabio José Ayres (Insper) and Paris Smaragdis (MIT) Abstract Abstract Representational similarity in neural networks is inherently scale-dependent, yet widely used metrics such as Centered Kernel Alignment (CKA) and Procrustes analysis provide only global scalar estimates. These scalars often fail to distinguish micro-scale geometric jitter (local noise) from macro-scale semantic reorganization, compressing multi-scale structural relationships into a single uninformative value. We introduce the Topological Alignment Spectrum (TAS), a multi-scale diagnostic tool that sweeps normalized mean Jaccard similarity over varying neighborhood sizes. By normalizing the metric over an analytically-derived expected range (from expected overlap under randomness to perfect alignment), TAS yields a dimension-invariant metric over a spectrum of scales, where one indicates perfect structural alignment, zero reflects chance-level agreement, and negative values signal active anti-alignment at specific scales. Experiments on synthetic point clouds demonstrate that TAS allows the recognition of distinct types of alignment perturbation: local jitter harms fine-grained neighborhoods but preserves cluster-level structure, while cluster-center shuffling preserves local similarity but disrupts global alignment -- phenomena that remain invisible or conflated under global, single-scalar metrics. Applying TAS to the MultiBERTs collection reveals that fine-tuning induces comprehensive topological reorganization across scales, challenging the view of task adaptation as merely conservative or localized. While models from different random seeds remain locally divergent, semantic clusters emerge as the dominant scale of alignment. TAS thus offers a granular, topology-aware alternative for diagnosing convergence and representational stability in deep networks. A Physics-Informed Cascade Surrogate for Likelihood-Aligned Rare-Event Monte Carlo of SPAD Avalanche Breakdown Mathieu Vandwalle (X-FAB) Abstract Abstract Monte Carlo carrier-transport simulations used to estimate the spatial avalanche breakdown probability in single-photon avalanche diodes (SPADs) are essential for accurate photon detection probability (PDP) modeling, yet remain computationally prohibitive due to stochastic transport and rare impact-ionization events. A physics-informed cascade neural surrogate is introduced that emulates the Monte Carlo avalanche simulation workflow by decomposing it into three physically ordered stages: spatial electric-field reconstruction, excess-bias inference, and breakdown-probability prediction. The final stage is trained directly on Monte Carlo injection counts using a binomial negative log-likelihood, yielding calibrated spatial probability maps rather than mean-based regression outputs. Across variations in process, geometry, temperature, and bias, converged Monte Carlo breakdown-probability maps are accurately reproduced while runtime is reduced by by up to six orders of magnitude. More broadly, the cascade and likelihood-aligned training provide a reusable template for accelerating rare-event Monte Carlo in semiconductor transport and reliability, including high-voltage semiconductor device avalanche breakdown and radiation-induced degradation. Monday Virtual Room 1 IJCNN Paper SS09 Multimodal Deep Learning in Applications I Session Chair: Fang Yang (Hebei university), Yinfeng Yu (Xinjiang University) Adaptive Global-Local Contrastive Learning and Information Enhancement for Multimodal Entity and Relation Extraction Tiantian Chen, Meiling Wang, and Fang Yang (Hebei University) Abstract Abstract Multimodal Named Entity Recognition (MNER) and Multimodal Relation Extraction (MRE) extract entities and relations from text by fusing complex visual data. Despite notable advances, existing methods suffer from two core challenges: The first issue is the key information loss, where noise suppression eliminates entity-related visual cues in MNER and relation-critical details in MRE, degrading performance; The second issue is insufficient original information, which triggers entity misjudgment in MNER and relation determination deviations in MRE. This paper proposes a multimodal joint task network (AGLCLIE) to optimize MNER and MRE simultaneously. Specifically, an adaptive global-local contrastive learning module is designed to address the issue of key information loss: it constructs a Consistency-Aware Factor to quantify the semantic consistency between the textual and visual modalities and achieves adaptive and dynamic adjustment of temperature parameters in the two dimensions of "word-image patch" (global) and "phrase-region" (local). In addition, an information enhancement module is designed to address the problem of insufficient original information: it guides the Large Language Model to generate accurate explanations of entities and nouns through prompts, so as to supplement semantics and eliminate ambiguities. Experimental results show that our model achieves F1-scores of 75.97%, 87.82% and 87.27% on the Twitter15, Twitter17 and MNRE datasets, exceeding the prior state-of-the-art methods by 0.65%, 0.95% and 4.14%, respectively. These results confirm that AGLCLIE outperforms current SOTA methods on both MNER and MRE tasks. Spatial-Aware Conditioned Fusion for Audio-Visual Navigation Shaohang Wu and Yinfeng Yu (Xinjiang University) Abstract Abstract Audio-visual navigation tasks require agents to locate and navigate toward continuously vocalizing targets using only visual observations and acoustic cues. However, existing methods mainly rely on simple feature concatenation or late fusion, and lack an explicit discrete representation of the target’s relative position, which limits learning efficiency and generalization. We propose Spatial-Aware Conditioned Fusion (SACF). SACF first discretizes the target’s relative direction and distance from audio-visual cues, predicts their distributions, and encodes them as a compact descriptor for policy conditioning and state modeling. Then, SACF uses audio embeddings and spatial descriptors to generate channel-wise scaling and bias to modulate visual features via conditional linear transformation, producing target-oriented fused representations. SACF improves navigation efficiency with lower computational overhead and generalizes well to unheard target sounds. Reliability-Aware Geometric Fusion for Robust Audio-Visual Navigation Teng Liu and Yinfeng Yu (Xinjiang University) Abstract Abstract Audio-Visual Navigation (AVN) requires an em bodied agent to move toward a sounding target using both visual observations and binaural audio signals. However, in complex indoor environments, binaural cues are not always reliable. This problem becomes more noticeable when the agent encounters sound categories that were not seen during train ing. In this work, we present RAVN (Reliability-Aware Audio Visual Navigation), a framework that adjusts cross-modal fusion according to the reliability of audio observations. Instead of assuming that audio information is always trustworthy, our method estimates its reliability and uses this estimate to guide how audio and visual features are combined. To achieve this, we design an Acoustic Geometry Reasoner (AGR) trained with geometric proxy supervision. AGR learns observation-dependent uncertainty through a heteroscedastic Gaussian NLL objective. The learned uncertainty serves as a practical reliability signal and does not require geometric labels at inference time. Leveraging this signal, our Reliability-Aware Geometric Modulation (RAGM) module applies a soft gate to the incoming visual features. This explicitly suppresses inter-modal conflicts at the fusion stage. To test the architecture, we benchmarked RAVN within the SoundSpaces framework across Replica and Matterport3D layouts. Our comprehensive evaluations show consistent im provements over previous baselines, both in heard and unheard scenarios. Audio Spatially-Guided Fusion for Audio-Visual Navigation Xinyu Zhou and Yinfeng Yu (Xinjiang University) Abstract Abstract Audio-visual Navigation refers to an agent utilizing visual and auditory information in complex 3D environments to accomplish target localization and path planning, thereby achieving autonomous navigation. The core challenge of this task lies in the following: how the agent can break free from the dependence on training data and achieve autonomous navigation with good generalization performance when facing changes in environments and sound sources. To address this challenge, we propose an Audio Spatially-Guided Fusion for Audio-Visual Navigation method. First, we design an audio spatial feature encoder, which adaptively extracts target-related spatial state information through an audio intensity attention mechanism; based on this, we introduce an Audio Spatial State Guided Fusion (ASGF) to achieve dynamic alignment and adaptive fusion of multimodal features, effectively alleviating noise interference caused by perceptual uncertainty. Experimental results on the Replica and Matterport3D datasets indicate that our method is particularly effective on unheard tasks, demonstrating improved generalization under unknown sound source distributions. Monday Virtual Room 2 IJCNN Paper SS09 Multimodal Deep Learning in Applications II Session Chair: Zhiyi Zhu (Communication University of China), yu song (East China Normal University) BCMIRec: Behavior Co-occurrence Enhanced Multi-Interest Dual-Graph Learning for Multimodal Recommendation yu song (Shanghai Key Laboratory of Trustworthy Computing, East China Normal University, Shanghai 200062, China) Abstract Abstract Multimodal recommendation leverages item contents such as images and texts to alleviate interaction sparsity, yet it still faces noisy implicit feedback, unreliable content similarity, and entangled user intents. To address these issues, we propose BCMIRec, a behavior co-occurrence enhanced multi-interest dual-graph framework for multimodal recommendation. BCMIRec first constructs an item co-occurrence graph from co-selected behaviors and expands each user’s neighborhood to provide more behavior-consistent and clustered evidence. It then performs dual-graph propagation on the user--item bipartite graph and the co-occurrence graph, and fuses the resulting representations with an adaptive gate to balance global collaboration and local consistency. On top of the enhanced representations, BCMIRec learns multiple user intent vectors through temperature-controlled soft routing over the expanded neighborhood, and produces ranking scores via a LogSumExp-based mixture prediction with a diversity regularization to reduce interest collapse. Experiments on three Amazon benchmarks demonstrate that BCMIRec consistently outperforms strong multimodal baselines across Recall and NDCG, with clear gains on long-tail items. FHFusion: Frequency Heterogeneity Network for Infrared and Visible Image Fusion Zhou Fu and Min Li (Xinjiang University, School of Computer Science and Technology); chen chen (Xinjiang University, College of Software); Xiaoyi Lv (Xinjiang University, School of Computer Science and Technology); and cheng chen (Xinjiang University, College of Software) Abstract Abstract Infrared and visible image fusion aims to integrate complementary information from multiple modalities. Recently, frequency-domain-based fusion methods have gained attention due to their advantages in edge texture recognition and global information integration. However, existing frequency-domain fusion methods fail to fully utilize the rich information embedded within frequency representations and the cross-modal complementarity among different spectral components. Therefore, we propose FHFusion. Specifically, our approach leverages the semantic and stylistic heterogeneity inherent in the frequency domain of infrared and visible images. Through spectral decoupling mechanisms, we achieve differentiated information fusion that effectively exploits cross-modal frequency domain complementary information. Further, we design an Amplitude Selection Module (ASM) and a Phase Discrimination Module (PDM) adaptively fusing energy distribution differences in the amplitude spectrum and structural complementary information in the phase spectrum, respectively. Experiments on M3FD, LLVIP and TNO dataset show FHFusion outperforms state-of-the-art methods in subjective and objective evaluations. FocusBEV: Jointly Focusing on 2D Semantic Guidance and Holistic LiDAR Representation Sixian Chan, Yaohui Li, Jian Tao, and Xinggang Fan (Zhejiang University of Technology) Abstract Abstract Accurate 3D object detection stands as a cornerstone of autonomous driving systems. Among various methodologies, the Bird's-Eye-View (BEV) paradigm, which fuses LiDAR point clouds and surround-view images, has emerged as the predominant approach. However, existing frameworks struggle to fully exploit the semantic potential of images and construct contextually integral LiDAR representations. To address these challenges, we propose FocusBEV, a "dual-focus" framework. First, we introduce the Semantic-Focus module, a 2D semantic guidance mechanism. Leveraging a lightweight 2D detector as a "semantic prior", this module proactively enhances foregrounds on feature maps prior to view transformation, thereby improving semantic fidelity. Second, we design the Long-range Interaction Aggregator (LIA) that uses shifted-window self-attention to reconstruct coherent global dependencies from sparse LiDAR features, overcoming the locality of CNNs. Finally, extensive experiments conducted on nuScenes demonstrate that FocusBEV achieves 72.3 mAP and 74.6 NDS, explicitly validating the effectiveness of our FocusBEV. MLVTG: Mamba-Based Feature Alignment and LLM-Driven Purification for multi-modal Video Temporal Grounding Zhiyi Zhu, Xiaoyu Wu, Zihao Liu, and Linlin Yang (Communication University of China) Abstract Abstract Video Temporal Grounding (VTG), which aims to localize video clips corresponding to natural language queries, is a fundamental yet challenging task in video understanding. Existing Transformer-based methods often suffer from redundant attention and suboptimal multi-modal alignment. To address these limitations, we propose MLVTG, a novel framework comprising two key designed modules: MambaAligner and LLMRefiner. MambaAligner uses the bidirectional scanning and gated filtering strategy to model temporal dependencies and extract robust video representations for multi-modal alignment. LLMRefiner leverages the specific frozen layer of a pre-trained Large Language Model (LLM) to implicitly transfer semantic priors, enhancing multi-modal alignment without fine-tuning. This dual alignment strategy, temporal modeling via structured state-space dynamics and semantic purification via textual priors, enables more precise localization. Extensive experiments on QVHighlights, Charades-STA, and TVSum demonstrate that MLVTG achieves highly competitive performance. Monday Virtual Room 3 IJCNN Paper SS09 Multimodal Deep Learning in Applications III Session Chair: Haowen Zhu (Southeast University), Alberto Lopez Casanova (Future Connections, R&D Dept.) MIRAGE: Modality-Aware 3D Vision-Language Model for Liver Lesion Diagnosis in MRI via Volumetric-Text Alignment Haowen Zhu, Jingyang Zhang, and Chunfeng Yang (Southeast University) Abstract Abstract The accurate diagnosis of liver lesions from 3D multimodal Magnetic Resonance Imaging (MRI) is a critical yet challenging task in oncological decision-making. Existing deep learning approaches often struggle due to the scarcity of annotated volumetric medical data and the inherent heterogeneity of MRI sequences, where distinct physical parameters (e.g., T2-weighted vs. DWI) create significant distributional shifts. Furthermore, current medical Vision-Language Models (VLMs) adapted for 3D tasks often rely on \textbf{static early fusion strategies}, treating diverse MRI modalities as homogeneous input channels. This indiscriminate integration leads to feature interference and fails to replicate the \textbf{dynamic, semantic-driven} nature of radiological interpretation. To address these limitations, we propose \textbf{MIRAGE} (\textbf{M}ultimodal MRI \textbf{R}epresentation \textbf{A}ggregation with \textbf{G}ated \textbf{E}xperts). We introduce a library of modality-aware expert encoders, adopting a \textbf{decoupled learning paradigm}: experts are first contrastively pretrained to align specific physical imaging features with LLM-synthesized radiology semantics, ensuring robust representation learning prior to diagnosis. To mimic the selective attention of radiologists, we design a \textbf{Modality Routing Gate (MRG)} that dynamically prioritizes the most discriminative expert path for each patient, reducing redundancy and gradient conflicts. Additionally, a Cross-modal Interactive Feature Extraction and Refinement (CIFER) module is employed to fuse these selective features with textual knowledge. Extensive evaluations on the large-scale LLD-MMRI dataset demonstrate that MIRAGE achieves state-of-the-art performance with 91.57\% accuracy, significantly outperforming existing medical VLMs while maintaining computational efficiency suitable for clinical deployment. Enabling Visual Reasoning in Text-Only LLMs via Symbolic Scene Descriptions Wenzhuo Lei, Mufan Cao, Chaorong Ye, Yi Xu, and Cheng Chen (Tongji University) Abstract Abstract While modern multi-modal models rely on dense visual-language alignment, the extent to which pure text-only Large Language Models (LLMs) can perform visual reasoning remains underexplored. We investigate this by presenting DeepYOLO, a neuro-symbolic framework that converts images into compact, grammar-constrained symbolic scene reports using an off-the-shelf object detector. This allows a text-only LLM (e.g., DeepSeek-V3) to answer visual questions without any multi-modal pre-training. Our experiments on a VQAv2 validation subset show that DeepYOLO achieves 42.3% accuracy, significantly outperforming an explicit Scene Graph baseline (38.2%) while requiring 2.5× fewer tokens. Crucially, replacing detector outputs with ground-truth objects (Oracle) yields 44.5%, demonstrating that the symbolic interface is sufficient for the task and that the primary bottleneck remains the object detector rather than the LLM’s reasoning capability. DeepYOLO offers a modular, auditable, and highly token-efficient alternative for visual reasoning, particularly suitable for privacy-sensitive or edge-compute scenarios. VAPrompt: Variational Prompting for Weakly Supervised Video Anomaly Detection Wenbo Pang (Jiangnan University); tao zhang (Jiangnan University, Central South University); and gaoe qin (Jiangsu Huaying Intelligent Technology Co) Abstract Abstract Weakly supervised video anomaly detection (WSVAD) aims to identify abnormal events in long, untrimmed videos using only video-level labels, which is both practical and challenging. Existing methods often suffer from semantic ambiguity, caused by the diverse and context-dependent nature of anomalies, and attention dispersion, where models fail to focus on subtle anomalous cues. To address these issues, we propose VAPrompt, a semantic distribution-aware framework that explicitly models anomaly uncertainty and enhances spatiotemporal localization. The core component, Variational Text Prompt Refinement (VTPR), represents textual prompts as stochastic latent distributions rather than fixed embeddings, enabling the model to capture diverse anomaly semantics and improve generalization. In addition, a Prompt-Infused Spatiotemporal Saliency Aggregation (PISSA) module integrates cross-modal attention to highlight anomaly-relevant regions while suppressing background noise. Furthermore, a Multi-Anchor Semantic Alignment Objective (MASAO) is introduced to enforce fine-grained semantic separation under weak supervision. Extensive experiments on the UCF-Crime and XD-Violence benchmarks demonstrate that VAPrompt consistently outperforms state-of-the-art methods. Specifically, our approach achieves 89.01% AUC and 72.32% AnoAUC on UCF-Crime, and 86.16% AP on XD-Violence, along with substantial improvements in temporal localization accuracy. These results indicate that modeling semantic uncertainty and integrating prompt-guided saliency are effective for robust weakly supervised video anomaly detection, making VAPrompt a practical and extensible solution for complex real-world surveillance scenarios. Deep Reinforcement Learning for Downlink Power Control in 5G/6G Networks Alberto Lopez Casanova (Future Connections), Rocío Pérez de Prado (University of Jaén), Miguel A. Regueira Caumel (Future Connections), Carlos Hidalgo-Luque (University of Jaén), and Chuan Heng Foh (University of Surrey) Abstract Abstract From the creation of early mobile network standards to current 5G and beyond, interference management has been extensively researched as it is crucial for improving network performance and service quality. Among the different techniques, downlink power control, which consists of dynamically adjusting the power with which base stations transmit to users, has demonstrated to be critical to mitigate inter-cell interference. In this context, through the years, traditional optimization methods have been applied, achieving good performance but facing challenges in current standards characterized by high complexity in terms of service requirements. In this paper, we propose a deep reinforcement learning based approach for downlink power control, motivated by the results achieved in recent years by data-driven approaches to radio resource management optimization. We train an agent based on the double dueling deep Q-network algorithm, relying on a centralized single-agent reinforcement learning architecture. The OMNeT++ discrete-event simulator with the Simu5G library is chosen to simulate a 5G Standalone scenario where the agent is trained with the goal of maximizing the network-wide signal-to-interference-plus-noise ratio. We tested the trained agent in this scenario against the baseline approach native to the simulator, based on fixed power transmission. Validation results show that the proposed approach improves the signal-to-interference-plus noise ratio while achieving a reduction in physical resource block usage compared with the baseline approach. While the work presented in this paper is based on scalar network key performance indicators, it establishes a baseline for future improvements to incorporate multimodal data, enabling the use of multimodal architectures to achieve a deeper understanding of the network state. Monday Virtual Room 4 IJCNN Paper SS09 Multimodal Deep Learning in Applications IV Session Chair: mingchao zhang (inner Mongolia University of Technology), Zhikui Chen (Dalian University of Technology) Dustformer: A PM10 Concentration Prediction Model Integrating Spatiotemporal Modeling and Mixture-of-Experts Mechanism Mingchao Zhang and Yongsheng Wang (Inner Mongolia University of Technology), Delong Zhang (Meteorological Data Center of Inner Mongolia Autonomous Region), and Guolin Zhang (Inner Mongolia University of Technology) Abstract Abstract Elevated PM10 concentrations, driven by both anthropogenic emissions and frequent sand and dust events, pose serious threats to public health and infrastructure; accurate PM10 forecasting is therefore essential for pollution control and resilient energy operations. This study proposes Dustformer, a spatiotemporal forecasting framework designed to fuse ground observations, meteorological data, satellite remote sensing data, and topography data. The proposed architecture integrates a Mixture-of-Experts (MoE) Transformer with dynamic routing to capture complex multimodal temporal patterns, coupled with an attention-weighted Dynamic Graph Neural Network (DGNN) to explicitly model evolving inter-station pollutant diffusion paths. Additionally, a robust sample screening mechanism is incorporated to mitigate the impact of data noise. Extensive experiments on real-world datasets demonstrate that Dustformer outperforms several representative baseline methods in terms of predictive accuracy and robustness. Furthermore, systematic ablation studies confirm the indispensability of each module, revealing that adaptive temporal expert routing and dynamic spatial dependency learning are critical for minimizing forecasting errors. D2TNet: A Decoupled Dual-Teacher Framework for Robust Infrared and Visible Image Fusion in Diverse Degraded Scenes Yujie Liu (Nanchang Hangkong University), Peng Liu (Beihang University), Congxuan Zhang (Nanchang Hangkong University), Zhen Chen (Beihang University), and Feng Lu (Nanchang Hangkong University) Abstract Abstract Real-world image degradations pose severe challenges for infrared-visible image fusion (IVIF), as most existing methods rely on the implicit assumption of pristine inputs. Conventional single-teacher frameworks struggle to balance image restoration and detail preservation, as denoising objectives inevitably suppress the high-frequency textures critical for fusion. Similarly, cascaded two-stage approaches suffer from error propagation and limited end-to-end efficacy. To resolve this fundamental conflict, we propose a novel Decoupled Dual-Teacher (D2T) framework that disentangles restoration and fusion into dedicated expert networks. A lightweight student network distills this knowledge via a Quadruple-Supervision Strategy, aligning its intermediate features with robust representations from both teachers to achieve simultaneous restoration and fusion. Notably, our model operates in a fully blind manner, handling unknown degradations without requiring manual priors while maintaining high inference efficiency. Extensive experiments on challenging degraded datasets show that D2T consistently outperforms state-of-the-art (SOTA) fusion methods, producing fused images with superior clarity and structural fidelity while effectively suppressing noise and artifacts from unknown degradations. CoherentDrive: Conflict-Aware World-State Grounding for Hallucination-Resistant Autonomous Driving Reasoning Shuo Liu, Lei Shi, Yufei Gao, and Yucheng Shi (Zhengzhou University) Abstract Abstract Large language models (LLMs) are increasingly used to support language-based interaction in autonomous driving, including scene question answering and risk explanation. Yet a critical reliability failure persists: when prompted with long, redundant, and mutually inconsistent perception logs from heterogeneous sensors, LLMs may produce fluent but unsupported statements, hallucinate non-existent traffic agents, and amplify false premises. We address this problem from an interface perspective. We present CoherentDrive, a prompt-ready grounding framework that compiles multi-source perception facts into a compact, conflict-resolved Coherent World State (CWS) with entity identifiers, provenance, and uncertainty flags. CoherentDrive further enforces a grounding-first answering protocol that restricts responses to CWS-supported entities and relations, enabling explicit refusal when evidence is missing. On nuScenes-QA under a zero-shot setting with multiple LLM backbones, CoherentDrive consistently improves answer correctness while substantially reducing hallucinated entity mentions, and in a single-pass CWS prompting setting it also shortens prompt length and reduces inference latency. Extensive experiments demonstrate that conflict-resolved world-state prompting provides an effective and reliable foundation for hallucination-resistant driving question answering. Self-supervised Graph Contrastive Learning for Incomplete Multi-view Clustering Anni Chen (University of Wollongong) and Shan Jin, Zhikui Chen, Yanfan Li, and Shuo Yu (Dalian University of Technology) Abstract Abstract Improving clustering performance on incomplete multi-view data is crucial for intelligent systems. However, most existing incomplete multi-view clustering methods assume consistent missing patterns across views, which rarely holds in real-world applications with heterogeneous and view-specific missingness. To address this issue, we propose SIGMA (Self-supervised Graph Contrastive Learning for Incomplete Multi-view Data Alignment), a method that enables effective multi-view representations under inconsistent missing patterns. Compared with existing studies, SIGMA extracts view-specific structures while independently inferring missing information for each view. Specifically, SIGMA initializes view-specific graphs and a global consistent graph using $k$-nearest neighbors, and dynamically propagates reliable structural information from the global graph to incomplete views. Furthermore, a weighted graph contrastive learning strategy is introduced to preserve both local discriminability and global structural consistency across recovered views. Extensive experiments on multiple benchmark datasets demonstrate that SIGMA consistently outperforms state-of-the-art incomplete multi-view clustering methods. Monday Virtual Room 5 IJCNN Paper SS09 Multimodal Deep Learning in Applications V Session Chair: Mingyong Li (Chongqing Normal University), Dongxun Jiang (Tongji University) Noise-Robust Cross-Modal retrieval via Noisy Label Re-Matching and Pseudo-Classification Yukai Wang and Mingyong Li (Chongqing Normal University) Abstract Abstract Cross-modal matching relies on large-scale data, which researchers often collect through crowd-sourcing or web crawling to reduce annotation costs. However, this process inevitably introduces mismatched pairs, leading to the noisy correspondence problem. Since noisy correspondences are difficult to identify and correct, noise-robust cross-modal learning becomes a challenging task. To achieve robust cross-modal retrieval, we propose a noisy label re-matching and pseudo-classification framework. The noisy label re-matching module identifies and corrects mismatched pairs by estimating correspondence confidence in the learned representation space and refining alignment with a pre-trained vision–language model. The pseudo-classification module treats titles as category labels and generates pseudo-labels from clean samples to promote confident and balanced predictions. In addition, a confidence-based weighting strategy is introduced to adaptively suppress harmful samples during training. Extensive experiments on Flickr30K, MS-COCO, and CC120K demonstrate that our method outperforms state-of-the-art approaches. O2SMatch: Cross-modality Knowledge Transferring Induced Optical and SAR Image Matching jiahang Song and ganggang Dong (Xidian University) Abstract Abstract Accurate matching between SAR and optical images remains challenging due to the significant discrepancy in imaging mechanisms and visual characteristics.To address this issue, this paper proposes a cross-modality knowledge transferring framework for SAR--optical image matching.First, an adversarially accelerated diffusion model is introduced to construct an intermediate domain between SAR and optical images, effectively bridging the modality gap.Based on the translated representations, a cascaded cross-sensor matching architecture is developed using a densely connected convolutional network equipped with hierarchical attention mechanisms, including cross-channel, cross-tensor, and cross-kernel attention.Extensive experiments on measured datasets demonstrate that the proposed method significantly outperforms competitive approaches. Vil-unet: beyond Transformers for 3d Medical Image Segmentation with Vision Xlstm Yong Wu (South China Normal University); Wenjun Huang (Sun Yat-Sen University); Ziyu Hu (Aalto University); Yiqi Ma (Shanghai Jiaotong University); and Pi Fang, Yuechen Yin, and Qinghua Zhong (South China Normal University) Abstract Abstract Efficient segmentation of medical images has evolved from relying solely on Convolutional Neural Networks (CNNs) to exploring hybrid models that integrate CNNs with Vision Transformers (ViTs). Despite the advantages of transformers in capturing global dependencies, their high computational and storage demands pose challenges. This research introduces the novel ViL-UNet, which combines CNNs with Vision Extended Long Short-Term Memory (Vision-xLSTM), aiming to balance performance and computational efficiency. The Vision-xLSTM blocks capture long-range spatial and global relationships, both at the network’s bottleneck and within the skip connections to refine multi-scale feature fusion. Our primary goal is to demonstrate that Vision-xLSTM provides a suitable backbone for medical image segmentation, offering superior performance with reduced computational costs. The ViL-UNet achieves state-of-the-art results on the publicly available Synapse and ACDC datasets. Index Terms— Medical Image Segmentation, xLSTM, 3D Imaging. Lung Cancer Detection Model Integrating Orthogonal Views and Blood Analysis Indicators for Patients with Pleural Effusion Dongxun Jiang and Dongdong Zhang (Tongji University) Abstract Abstract Pleural effusion, as a highly prevalent clinical symptom, has garnered wide-spread attention. However, most models are designed for lung cancer detection, but the detection of tumor lesions within pleural effusion still requires improvement. This paper proposes a lung cancer detection model integrating orthogonal views and blood analysis indicators for patients with pleural effusion(OVBI-LCDM), in which feature enhancement network based on view complementarity is designed to more accurately depict hidden lesions in orthographic views. Blood indicators analysis module is introduced to analyze correlation and importance of six blood indicators and determine their weights. Feature fusion module based on analysis results is proposed to dynamically re-evaluate the contributions of orthogonal view images and blood indicators, thereby achieving mutual enhancement of these features. The experimental results show that the model achieved good classification performance, with an accuracy of 82.76%, outperforming other models. Monday Virtual Room 6 IJCNN Paper SS09 Multimodal Deep Learning in Applications VI Session Chair: Jinhui Yu (Zhejiang Provincial Museum), yiqiang he (zhejiang university of technology) EGF-CHITR: An Entity-Guided Framework for Cultural Heritage Cross-Modal Image–Text Retrieval Gang Xiao and Yiqiang He (zhejiang university of technology); Fengyi Ye (Hangzhou Tianyuan Textile Co., Ltd); Jinhui Yu (Zhejiang Provincial Museum); and Jun Xu (zhejiang university of technology) Abstract Abstract Chinese cultural heritage artifacts are rich in multimodal signals, making image-text cross-modal retrieval crucial for digital preservation and reuse. Yet reliable image-text alignment remains challenging due to (i) background-heavy professional descriptions that introduce redundant semantics and (ii) CLS-level global matching that can obscure locally discriminative visual evidence. We propose EGF-CHITR, a lightweight framework that integrates entity-prior gating with bidirectional local alignment to enhance cross-modal discriminability. Specifically, entity-prior gating injects domain knowledge to reweight textual tokens, while bidirectional local alignment models hierarchical patch--token interactions to suppress redundant semantics and strengthen locally discriminative correspondence under global semantic consistency. Experiments on the ceramics dataset PorTi and the public cultural-heritage dataset CulTi demonstrate that EGF-CHITR consistently outperforms mainstream methods across multiple metrics in both retrieval directions. Noise- and Occlusion-Robust Mask-Guided Adaptive Fusion for Audio-Visual Person Verification Juntao Wang (Xinjiang University), Qiuming Zhao (Tsinghua University), Xiaolong Wu (Naval Aviation University), Mingxing Xu (Tsinghua University), Askar Hamdulla (Xinjiang University), and Thomas Fang Zheng (Tsinghua University) Abstract Abstract Audio-visual multimodal person verification has attracted increasing attention for improving robustness in challenging conditions by leveraging the complementarity between speech and face cues. Despite recent progress, existing feature-level fusion methods---e.g., soft-attention and gating---typically infer modality reliability implicitly from identity embeddings, without dedicated degradation-aware cues for acoustic noise or facial occlusion. Consequently, the learned modality weights can be unreliable when degradations are severe or asymmetric, yielding suboptimal fusion representations. In this work, we propose a mask-guided adaptive fusion framework (MGAF) that explicitly estimates modality reliability from degradation-aware masks and uses this signal to guide fusion. Specifically, MGAF predicts a time--frequency mask to characterize noise-corrupted speech and learns a feature-level visual mask on convolutional representations to enhance reliable responses; the masks are aggregated into explicit reliability cues, based on which MGAF performs mask-driven re-weighting of unimodal embeddings and constructs the fused representation via weighted concatenation. This design allows the model to emphasize the more reliable modality under strong corruption while retaining complementary information when both modalities are informative. We train MGAF on VoxCeleb2 and evaluate it on VoxCeleb1 under systematically simulated noise and occlusion settings, including asymmetric degradations. On the standard VoxCeleb1 test sets, MGAF achieves EERs of 0.207\%, 0.212\%, and 0.433\% on VoxCeleb1-O, VoxCeleb1-E, and VoxCeleb1-H, respectively, and achieves consistent gains over representative fusion baselines under degraded scenarios---especially under asymmetric degradations---without over-suppressing either modality when both are reliable. AdapDeFormer: Adaptive Deformable Attention Aggregation with Hierarchical Query Learning for Multi-Modal Object Re-identification Xingan Ma (University of Bonn), Yuhao Wang (Dalian University of Technology), and Jinhui Yi and Juergen Gall (University of Bonn) Abstract Abstract Multi-modal object Re-Identification (ReID) exploits complementary cues from heterogeneous sensors to recognize identities under challenging environmental conditions. However, existing methods suffer from three key limitations. First, modality reliability can vary substantially across instances and environments, rendering static fusion suboptimal. Second, sensor-induced spatial misalignment often corrupts cross-modal correspondences and degrades feature fusion. Third, naive feature aggregation lacks a progressive mechanism to integrate identity cues from fine-grained details to high-level semantics. To address these issues, we propose AdapDeFormer, a unified framework that combines Enhanced Deformable Attention (EDA) with a Hierarchical Adaptive Query Network (HAQN) for robust multi-modal representation learning. Specifically, EDA reweights modalities via Adaptive Modal Weighting (AMW) and corrects spatial discrepancies through Multi-Scale Offset Fusion (MSOF) to enable alignment-aware fusion. Building on the enhanced features, HAQN performs progressive identity feature extraction with hierarchical query learning, where cross-layer query propagation and modality-adaptive aggregation facilitate discriminative representation learning. Extensive experiments on three multi-modal object ReID benchmarks demonstrate the effectiveness and robustness of AdapDeFormer. Monday Virtual Room 7 IJCNN Paper SS09 Multimodal Deep Learning in Applications VII Session Chair: 树兰 张 (四川师范大学, College of Computer Science), 斌 王 (East China University of Science and Technology) ChartMVC: A General Multi-View Contrastive Learning Framework for Robust Chart Question Answering Bin Wang and JianHua Li (East China University of Science and Technology) Abstract Abstract Chart Question Answering (ChartQA) has recently gained significant attention in visual data understanding. Despite achieving promising results on standard benchmarks, existing vision-language models—ranging from lightweight architectures to recent Large Multimodal Models (LMMs)—often degrade significantly under diverse chart styles and visual disturbances, such as variations in color, font, and layout. To improve robustness, we propose ChartMVC(Chart Multi-View Consistency Contrastive learning), a general Multi-View Contrastive Learning framework designed to enhance visual stability in chart understanding. ChartMVC forces the model to learn semantically consistent representations across diverse visual styles through a queuebased contrastive mechanism. To validate the effectiveness of our framework, we instantiate it using the UniChart architecture. Extensive experiments on the ChartQA benchmark show that our approach significantly outperforms the original baseline in relaxed accuracy metrics. Furthermore, the model exhibits superior robustness against various chart perturbations, demonstrating the effectiveness of the ChartMVC framework for real-world chart understanding tasks. UniVerse: A Unified Voxel-Semantic Ensemble Framework for Multi-Modal 3D Object Detection Yichen Yao, Xiang Qiu, and Sixian Chan (Zhejiang University of Technology) Abstract Abstract As a core component of autonomous driving perception systems, 3D object detection faces a critical bottleneck: how to achieve effective joint representation of geometric features and semantic information across heterogeneous modalities (LiDAR point clouds and RGB images). Prevailing multi-sensor fusion methods are prone to cascading errors stemming from sequential optimization processes spanning geometric modeling, cross-modal alignment, and semantic integration. These errors ultimately lead to systematic failures in semantic understanding and undermine robustness. To address this limitation, we propose the UniVerse framework, which systematically cascades a progressive optimization paradigm through "feature disentanglement, alignment enhancement, and deep aggregation." Specifically, UniVerse introduces a novel triple synergy mechanism: the unified dual-stream perception encoder (GeSe-Encoder) ensures high-purity geometric and semantic features at the source; the cross-modal hierarchical feature fusion module (CHFM) achieves precise alignment through bidirectional enhancement; and the semantic-aware global enhancement module (SAGE) enables robust semantic understanding through multi-source interaction, ultimately supporting effective joint modeling of geometric structure and semantic context. Extensive experiments on the KITTI dataset validate the effectiveness of our method, achieving 65.63% 3D mAP and outperforming comparable baselines by 3.52%, confirming the value of the proposed cascade optimization strategy. A Social Media Emotion Detection Model Based on Implicit Emotion Multi-Hop Memory and Knowledge Embedding Liu Meiling, Li Ming, and Xi Xinlei (Northeast Forestry University) Abstract Abstract Social media platforms have become a primary source for mining public sentiment and analyzing user behavior. While explicit emotions are easily detectable, a significant portion of online expression is implicit—conveyed through subtle metaphors, irony, or context—posing a severe challenge for traditional sentiment analysis models. Existing approaches typically rely on surface-level semantic features, often failing to capture these underlying emotional cues due to a lack of background knowledge and reasoning capabilities. To address this, we propose a novel framework: a Social Media Emotion Detection Model based on Implicit Emotion Multi-Hop Memory and Knowledge Embedding. Unlike standard methods, our model incorporates a Knowledge-Embedded Attention Mechanism to retrieve external commonsense knowledge, bridging the gap between literal text and implicit intent. Furthermore, we design an Implicit Emotion Multi-Hop Memory Network to capture long-range contextual dependencies and learn the prior distribution of emotional states. Experiments on the fine-grained GoEmotions dataset demonstrate that our proposed method achieves state-of-the-art performance, significantly outperforming strong baselines in detecting implicit and complex emotion categories.Experiments on the fine-grained GoEmotions dataset demonstrate that our proposed method achieves state-of-the-art performance, yielding a Precision of 54.66%, Recall of 53.80%, and Macro-F1 score of 52.34%. These results significantly outperform strong baselines such as BERT (Macro-F1: 46%) and BiDLSTM (Macro-F1: 50.4%). This superior performance is attributed to our model's unique ability to bridge implicit semantic gaps via external knowledge embeddings and capture long-range emotional dependencies through the multi-hop memory network. EDARFusion: Entropy-Difference-Aware Dynamic Registration for Unaligned Multimodal Medical Image Fusion Shulan Zhang, Guihua Liao, and Renbin Fang (四川师范大学, College of Computer Science) Abstract Abstract Unaligned multimodal medical image fusion faces dual challenges of spatial misalignment and cross-modal information inconsistency. Existing collaborative registration–fusion frameworks often enforce global alignment, causing unreliable regions such as low-texture areas and noise to severely interfere with deformation estimation; they are also prone to inducing erroneous deformations and ghosting in regions with large modality discrepancies. To address these challenges, from the perspective of distinctly handling anatomical misalignment that should be aligned and modality-specific differences that should be preserved, we propose EDARFusion: Entropy-Difference-Aware Dynamic Registration for Unaligned Multimodal Medical Image Fusion. Specifically, the Feature-Reliability Guided Pre-Registration Adapter (FRG-PreReg) predicts a feature-level reliability prior map to suppress the interference of unreliable regions in deformation learning; the Entropy-Difference-Aware Flow Registration (EDA-FlowReg) generates a gating-weight map from local information entropy differences and imposes soft constraints on incremental deformation updates in multi-scale estimation, thereby enabling discrepancy-aware dynamic alignment and reducing ghosting and fusion artifacts caused by misregistration; in the fusion stage, an alignment-consistency constraint is introduced to dynamically integrate structural and modality-specific information on the basis of reliability-aware alignment. Extensive experiments and visualizations of the local information entropy-difference map, the gating-weight map, and the reliability map demonstrate that the proposed method achieves superior registration and fusion performance compared with multiple representative approaches, with interpretable decision-making. Monday Virtual Room 8 IJCNN Paper SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models I Session Chair: Ruwang Jiao (Soochow University), XINGBANG DU (Hokkaido University) Embodied Multimodal Neural Architecture Search System for Low-Resolution Image Classification Xinyi Yuan and Nan Li (School of Computer and Information Technology, Shanxi University); Ruwang Jiao (School of Future Science and Engineering, Soochow University); and Yayu Zhang and Zhifang Wei (School of Computer and Information Technology, Shanxi University) Abstract Abstract With the rapid development of embodied intelligence, robust image classification from low-resolution multimodal sensory inputs has become a critical objective for embodied agents. Due to hardware limitations and environmental interference, image inputs captured by embodied agents often suffer from severe resolution degradation, and manually designed architectures for low-resolution classification lack adaptability to embodied multimodal perception. To address these challenges, we propose an \emph{Embodied Multimodal Neural Architecture Search} (EM-NAS) system, which automates the design of network architectures for low-resolution image classification in embodied scenarios. Specifically, we design an embodied perception-oriented NAS search space that integrates super-resolution (SR) as a feature enhancement component and incorporates cross-modal fusion operators, enabling the automatic discovery of architectures tailored to embodied classification demands. We further propose a Context-Aware Multimodal Fusion (ECMCF) module, which is optimized via NAS to dynamically fuse visual and environmental modal features for robust feature representation. Additionally, an embodied joint loss function is introduced to guide NAS search toward architectures with both high classification accuracy and environmental robustness. Experiments on extended Embodied-CIFAR10 and Tiny-ImageNet datasets demonstrate that the EM-NAS-discovered architecture outperforms state-of-the-art baselines, achieving a 4.2\% improvement in classification accuracy (with auxiliary SR gains of 2.5 dB in PSNR and 0.027 in SSIM). This work introduces a NAS-driven paradigm for embodied low-resolution image classification, providing an automated architecture design solution for embodied perception systems. Conditional Neural Architecture Search Tailored for Reactor 3D Power Distribution Prediction Minxiao Zhong (Sichuan University, State Key Laboratory of Advanced Nuclear Energy Technology); Qing Li (State Key Laboratory of Advanced Nuclear Energy Technology); and Yanan Sun (Sichuan University) Abstract Abstract The accurate prediction of three-dimensional power distribution is critical for ensuring the safety of nuclear reactors. Traditional prediction methods struggle in computational efficiency and prediction accuracy. Recent deep learning approaches offer a promising alternative but frequently employ generic architectures due to a lack of domain-specific knowledge guiding the architecture design. They fail to capture the inherent anisotropy, geometric symmetry, and state-dependent nonlinearity of reactor behavior. To overcome each limitations, this paper proposes a conditional neural architecture search method. It comprises three essential components: (1) an anisotropic convolution block is designed to decouple radial and axial feature learning, thereby matching the anisotropy of physical principles; (2) a geometry-aware attention block is proposed to embed rotational symmetry priors and adaptively focus on critical regions; and (3) a lightweight conditioner is developed to dynamically activate the architecture for specific reactor states, effectively modeling state-dependent nonlinearities. Experiments show that our automatically searched architecture significantly outperforms manual baselines. For instance, it achieves a mean square error of 0.0109 with only 873K parameters, reducing error by 44.1% compared to a V-Net baseline. The results demonstrate that our method enables highly accurate, efficient, and adaptive prediction suitable for reactor safety evaluation. SmartDyGNN: Dynamic Sparse Training for GNNs via Structural and Gradient Fusion Boran Hu and Nan Li (Shanxi University), Ruwang Jiao (Soochow University), Yayu Zhang (Shanxi University), Cuie Yang (Northeastern University), and Zhifang Wei (Shanxi University) Abstract Abstract Standard graph neural networks (GNNs) typically perform message passing over all nodes during training, incurring substantial computational redundancy. To solve the problem, this paper proposes SmartDyGNN, a dynamic sparse training algorithm that adaptively activates critical nodes by jointly exploiting structural redundancy in graph-structured data, embedding dynamics, and task-gradient sensitivity. Embedding dynamics capture the convergence trends of node representations throughout training, while task-gradient sensitivity reflects how much each node’s representation influences the final task loss. Together, these signals guide the algorithm in dynamically identifying and retaining the most optimization-relevant nodes at each training iteration. On benchmark datasets Cora, Citeseer, and PubMed, the proposed method stably maintains the proportion of active nodes at approximately 73%,thereby reducing message-passing operations by 25%–30%, while achieving a relative improvement of 0.20%–0.37% in accuracy over the full-graph baseline. Moreover, SmartDyGNN generalizes well to other architectures, including GCN and GAT. G-ICSO-NAS: Shifting Gears between Gradient and Swarm for Robust Neural Architecture Search XINGBANG DU (Graduate School of Information Science and Technology, Hokkaido University); ENZHI ZHANG and RUI ZHONG (Information Initiative Center, Hokkaido University); YANG CAO (Graduate School of Information Science and Technology, Hokkaido University); and Masaharu Munetomo (Information Initiative Center, Hokkaido University) Abstract Abstract Neural Architecture Search (NAS) has become a pivotal technique in automated machine learning. Evolutionary Algorithm (EA)-based methods demonstrate superior search quality but suffer from prohibitive computational costs, while gradient-based approaches like DARTS offer high efficiency but are prone to premature convergence and performance collapse. To bridge this gap, we propose G-ICSO-NAS, a hybrid framework implementing a three-stage optimization strategy. The Warm-up Phase pre-trains supernet weights w via differentiable methods while architecture parameters α remain frozen. The Exploration Phase adopts a hybrid co-optimization mechanism: an Improved Competitive Swarm Optimizer (ICSO) with diversity-aware fitness navigates the architecture space to update α, while gradient descent concurrently updates w. The Stability Phase employs fine-grained gradient-based search with early stopping to converge to the optimal architecture. By synergizing ICSO's global navigation capability with differentiable methods' efficiency, G-ICSO-NAS achieves remarkable performance with minimal cost. In the context of the DARTS search space, an accuracy of 97.46% is achieved on CIFAR-10 with a computational budget of just 0.15 GPU-Days. The method also exhibits strong transfer potential, recording accuracies of 83.1% (CIFAR-100) and 75.02% (ImageNet). Furthermore, regarding the NAS-Bench-201 benchmark, G-ICSO-NAS is shown to deliver state-of-the-art results across all evaluated datasets. Monday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V1 Session Chair: Bin Cao (Hebei University of Technology) Entropy-Guided Maximin-Fitness MOPSO with Simulated Annealing for Multi-Objective Dual-Resource Flexible Job Shop Scheduling Jiajie Fan (South China Normal University), Kailai Zhuang (Tiangong University), and Rundong Gao (Shenyang University of Chemical Technology) Abstract Abstract In modern manufacturing systems, dual-resource scheduling has emerged as a critical challenge, where each operation must be simultaneously assigned to both an eligible machine and a qualified worker. This joint allocation requirement introduces tight resource coupling, which significantly expands the combinatorial search space while severely constraining the feasible solution region and resulting in complex multi-dimensional trade-offs between completion time and operational costs. This paper investigates a Dual-Resource Constrained Flexible Job Shop Scheduling Problem (DRC-FJSP) that minimizes the makespan and a comprehensive production cost integrating raw-material holding, work-in-process holding, direct processing costs, and earliness/tardiness penalties. To effectively approximate Pareto-optimal solutions, we propose EM-MFF-MOPSO, a discrete multi-objective framework. The algorithm features an entropy-guided mechanism for stage-based parameter adaptation, a Maximin Fitness Function (MFF) strategy for leader selection, and a Simulated Annealing (SA) local search to refine solutions. Experiments on reproducible benchmark instances of increasing scales demonstrate that EM-MFF-MOPSO consistently produces more competitive and better-distributed Pareto fronts than representative baselines. Scale-Adaptive Surrogate Model Management for Expensive Optimization Xing-Yu Wang and Jian-Yu Li (Nankai University), Jun Zhang (Hanyang University), and Zhi-Hui Zhan (Nankai University) Abstract Abstract Data-driven evolutionary algorithms (DDEAs) have been widely used for solving expensive optimization problems by replacing expensive fitness evaluations with surrogate models constructed from limited evaluated data. However, although various surrogate construction and selection strategies have been investigated, the evaluation scale of surrogate-assisted fitness estimation is usually predefined and remains unchanged throughout the evolutionary process. Such a static evaluation scale may be inefficient or even misleading, as the evaluation requirements vary across different evolutionary stages. To alleviate this issue, this paper proposes a scale-adaptive model management (SAMM)-based DDEA, named SAMM-DDEA, together with three contributions. First, a SAMM strategy is proposed to dynamically adjust the evaluation scale of surrogate-assisted fitness estimation according to evolutionary feedback. In addition, a Full-Coverage Sampling (FCS) strategy is proposed to enhance ensemble surrogate diversity under limited data conditions, so as to enhance algorithm performance. Third, the SAMM-DDEA is developed and validated on widely-used expensive optimization problems with 10 to 100 dimensions, demonstrating superior or highly competitive performance compared with several state-of-the-art DDEAs. A Surrogate-Assisted Evolutionary Algorithm with Geometry-Guided Initialization and Memory-Guided Variation for Expensive Multi-Objective Problems Zitong Su and Min-Rong Chen (SOUTH CHINA NORMAL UNIVERSITY) Abstract Abstract In recent years, numerous surrogate-assisted evolutionary algorithms (SAEAs) have been developed for expensive multi-objective optimization problems (EMOPs). However, under a stringent evaluation budget, many existing SAEAs are hindered by low-quality initialization, inefficient use of training data for surrogate modeling, and limited exploitation of historical high-quality solutions. This paper proposes an SAEA with geometry-guided initialization and memory-guided variation (GIMV-SAEA). Specifically, a geometry-guided initialization scheme combines boundary probing with a center anchor and early local elite sampling to generate high-quality initial solutions while maintaining uniform exploration. An evolutionary-stage-aware adaptive training set construction strategy is further designed to balance surrogate training cost and predictive accuracy. In addition, a memory-guided variation operator preserves historical elite solutions and applies stronger perturbations to poorly converged decision variables. Experimental results on the DTLZ, WFG, and MaF test suites, as well as offset DTLZ instances, demonstrate that GIMV-SAEA achieves competitive performance on expensive multi- and many-objective optimization. Multi-stage Large-scale Multiobjective Evolutionary Algorithm Integrating Multidirectional Sampling and Branch Search Bin Cao (State Key Laboratory of Intelligent Power Distribution Equipment and System,Hebei University of Technology) and Zhenyu Wang (School of Artificial Intelligence,Hebei University of Technology) Abstract Abstract Most existing algorithms for large-scale multiobjective optimization problems (LSMOPs) suffer from monotonous and fixed search strategies and evolutionary stages. Consequently, the solutions generated by these algorithms struggle to achieve rapid convergence toward the Pareto front. To address this issue, this paper proposes a Multi-stage Large-scale Multiobjective Evolutionary Algorithm integrating Multi-directional sampling and Branch Search (LMOEA-MBMS). Specifically, a breadth multi-directional sampling strategy is first introduced to explore five distinct directions correlated with the sample solutions. Subsequently, a depth branch search strategy is implemented to further explore the non-dominated solutions of these samples along the sampling directions, thereby enhancing the algorithm’s convergence. Finally, a multi-stage environmental selection framework is designed to achieve the specific goals of the algorithm across different evolutionary phases. The proposed algorithm was evaluated against six state-of-the-art algorithms on eighteen LSMOPs and two practical problems of Time varying Ratio Error Estimation (TREE). Experimental results demonstrate that the performance of the proposed algorithm outperforms other state-of-the-art methods, effectively improving both convergence and solution diversity, with significant advancements in both theory and application. Monday Virtual Room 1 IJCNN Paper SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models II Session Chair: Ruwang Jiao (Soochow University), Lianbo Ma (College of Software, Northeastern University, China; Foshan Graduate School of Innovation, Northeastern University) GFNAS: Towards Gradient-Friendly Neural Network Architecture Search From Sharpness-Aware Perspective Jianlun Ma and Xinxin Xu (Northeastern University) and Lianbo Ma (Northeastern University; Foshan Graduate School of Innovation, Northeastern University) Abstract Abstract Neural architecture search (NAS) aims to automatically discover high-performing network architectures for specific tasks. In recent years, NAS methods that leverage sharpness signals have shown promising potential for improving model generalization. However, NAS still falls short in robustness under out-of-distribution (OOD) shifts. Recent studies have incorporated Sharpness-Aware Minimization (SAM) into the search process, but under weight-sharing and short training-budget settings, the standard SAM perturbation direction may entangle deterministic gradient trends with mini-batch stochastic noise, thereby biasing the sharpness proxy. To address this issue, we propose GFNAS, a gradient-friendly sharpness-aware NAS framework that estimates the deterministic gradient trend via an exponential moving average (EMA) and suppresses it when constructing perturbations, yielding a noise-consistent direction. Without increasing the search cost or modifying the search space, this strategy provides a more robust robustness-oriented evaluation signal. Experiments on CIFAR-10/100 and their corrupted variants demonstrate that GFNAS improves both accuracy and corruption robustness (mCE), reduces search cost, and alleviates sensitivity to sharpness hyperparameters. ADG-Former: Adaptive Dynamic Graph and Gated Fusion Transformer for Traffic Flow Prediction Chunyang Shi, Dongmin Chao, and Lingling Zhang (School of Computer Science and Technology, Beihua Universit) and Jianhui Lv (The First Affiliated Hospital of Jinzhou Medical University, Multi-modal Data Fusion and Precision Medicine Laboratory) Abstract Abstract Traffic flow prediction has been recognized as a pivotal technology in the development of smart cities, facilitating data support for resource scheduling and traffic management. By leveraging an efficient spatiotemporal graph attention mechanism, outstanding performance in balancing prediction accuracy and computational efficiency has been demonstrated by the STGformer model. However, a static adjacency matrix fails to accommodate dynamic changes in road connectivity within actual traffic scenarios, such as road closures or peak-hour congestion, consequently limiting prediction performance under complex conditions. To address these limitations, an enhanced STGformer model, termed ADG-Former, is introduced based on adaptive optimization of dynamic graph structures. Initially, a dynamic adjacency matrix is constructed using spatiotemporal correlations of traffic flow and speed, wherein effective connectivity relationships between nodes are updated in real time through an adaptive threshold strategy. Subsequently, a gated fusion module is employed to dynamically balance the contribution weights of static road topology and dynamic traffic dependencies, enhancing adaptability to complex scenarios. This study presents a feasible optimization scheme for traffic flow prediction in dynamic scenarios and demonstrates engineering applicability. Auto-HFL: Automated Sub-Network Discovery for Heterogeneous Federated Learning Ying Qian (College of software, Northeastern University) and Lianbo Ma (College of software, Northeastern University; Foshan Graduate School of Innovation, Northeastern University) Abstract Abstract The deployment of Large Machine Learning Models (LMLMs) on edge devices is hindered by the conflict between the colossal size of modern architectures and the limited, heterogeneous resources of edge clients. Federated Learning (FL) enables privacy-preserving collaborative training, yet standard FL protocols typically assume model homogeneity, making them unsuitable for diverse edge environments. In this paper, we propose Auto-HFL, a framework for Automated Sub-Network Discovery. Unlike traditional Neural Architecture Search (NAS) which incurs high computational costs, Auto-HFL introduces a lightweight, resource-aware discovery mechanism. It allows clients to automatically extract and train optimal sub-networks from a global supernet based on their real-time hardware constraints (e.g., memory and bandwidth). We further design a specialized aggregation strategy to handle the structural heterogeneity of the discovered models. Extensive experiments on [Insert Dataset, e.g., CIFAR-10/100] demonstrate that Auto-HFL significantly reduces communication and computation overhead while maintaining accuracy comparable to full-model training. This work paves the way for automating model deployment in the era of large-scale edge intelligence. Adaptive Bi-Directional Asymmetric Flip and Multi-Criteria Knowledge Transfer Strategy in Multitask Framework for High-Dimensional Feature Selection yikang zhou and nan li (Shanxi University) and ruwang jiao (Soochow University) Abstract Abstract Evolutionary multi-task optimization (EMTO) has been widely applied to address the challenges of high-dimensional feature selection. However, existing EMTO algorithms still suffer from two key deficiencies: in the feature search process, the flipping strategy uses fixed probabilities and thresholds, failing to adapt to population dynamics during evolution; and in the knowledge transfer stage, the adaptive aggregation strategy randomly selects sources, ignoring particle value and task correlation, which may lead to negative transfer. To overcome these limitations, this paper proposes AAMCSO-AP, a dynamic adaptive aggregative multi-task competitive particle swarm optimization algorithm. It introduces two novel mechanisms: a population entropy and feature distribution-aware dynamic update strategy for winners, and a value assessment and relationship-aware knowledge transfer strategy for losers, enabling adaptive parameter adjustment and intelligent knowledge transfer. Experiments on seven real high-dimensional datasets (with up to 10,509 features) demonstrate that AAMCSO-AP achieves significantly superior classification accuracy and feature reduction compared to the original AAMCSO and other state-of-the-art algorithms, confirming the effectiveness of the proposed dynamic mechanisms. Monday Virtual Room 2 IJCNN Paper SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models III Session Chair: Meng Wang (Liaoning Technical University), Zhiguo Hu (Shanxi University) Learning Graph Structures via Sparse Training with Adaptive Edge Growth sisi chen and nan li (Shanxi University) and ruwang jiao (Soochow University) Abstract Abstract Graph Neural Networks (GNNs) have achieved remarkable success in processing graph-structured data, yet they continue to face two major limitations: high computational cost and the presence of structural noise in real-world graphs, such as missing edges or spurious connections. Existing sparse training approaches effectively reduce computational overhead by pruning model weights and node features. However, they operate strictly on the initial graph topology, and the resulting trained graph is merely a subgraph of the original one, leaving the structural optimization problem unaddressed. By contrast, Graph Structure Learning (GSL) methods are capable of dynamically optimizing graph topology during training. However, these approaches typically rely on global pairwise similarity computation, resulting in a quadratic time complexity of O(N^2), which severely limits their scalability to large-scale graphs. Although anchor-based variants of Iterative Deep Graph Learning (IDGL) [1] partially alleviate this issue, they depend on low-rank approximations that inevitably discard local structural information of the original graph and may further introduce additional optimization instability. Therefore, an efficient graph structure optimization method that preserves local structural details is highly desirable. To address these limitations, we propose a dual-sparsity framework that integrates the principles of IDGL with the Comprehensive Graph Gradual Pruning (CGP) [2] paradigm, enabling the joint optimization of model parameters and graph topology. Specifically, we compute cosine similarity between node pairs to measure semantic relevance, which guides both the addition of necessary edges and the removal of redundant ones. Meanwhile, to overcome the high computational cost inherent in standard IDGL, we employ a hybrid sampling strategy that combines 2-hop neighborhood search with random negative sampling, reducing the computational complexity from O(N^2) to O(E). This strategy also mitigates the risk of the optimization process falling into local minima. Experiments on the Cora, Citeseer, and Pubmed datasets demonstrate that the designed framework achieves good performance while maintaining dual sparsity in both the model and graph structure, achieving comparable or superior performance static sparse baselines and exhibiting particularly notable generalization improvements on the larger Pubmed dataset. Complementarity-Guided Neural Architecture Search for Encrypted Traffic Classification Zhiguo Hu and Yadong Lu (Shanxi University), Xi Ma (Unit 96901 PLA), Lifeng Guo and Guoqing Liu (Shanxi University), and Kaikai Yang (Shanxi University of Finance and Economics) Abstract Abstract Deep learning-based encrypted traffic classification often relies on fixed network architectures with simple fusion strategies for multi-modal traffic features. As a result, feature complementarity and redundancy are not explicitly considered, which can limit model efficiency and generalization. Neural Architecture Search (NAS) provides a promising approach to automatic architecture optimization, but its application to large heterogeneous search spaces remains computationally expensive and slow to converge. In this work, we propose a complementarity-guided monte carlo tree search with evolutionary fine-tuning framework (CG-MCTS-NAS). The framework defines a structured search space over five heterogeneous traffic features, including packet length, inter-arrival time, raw bytes, byte entropy, and flow-level statistics. A complementarity matrix derived from feature correlations is incorporated into the PUCT strategy of MCTS, biasing the search toward more complementary feature combinations and reducing redundant exploration. Based on this design, a two-stage optimization strategy is employed, combining global architecture search with evolutionary fine-tuning of hyperparameters. Experiments on multiple real-world datasets show that CG-MCTS-NAS consistently outperforms recent state-of-the-art methods in classification performance, while reducing the required MCTS search iterations by approximately 55.6\%. The complete code for the proposed method is available at https://github.com/yadonglu7-hue/CG-MCTS-NAS. Global-Local nnUNet with Anatomical-Metabolic Statistical Priors for Improved PET/CT based Whole Body Lymphoma Segmentation Meng Wang, Yarong Feng, Yongwei Tang, Man Li, and Yuxin Liang (Liaoning Technical University); Chao Lv (China Medical University); and Tianyu Shi (Shenyang Ligong University) Abstract Abstract Automatic segmentation methods for lymphoma based on whole-body positron emission tomography (PET) and computed tomography (CT) still face significant challenges due to metabolic overlap between physiologically highly metabolizing and excreting organs (sFEPUs) and lesions. To address this, this paper proposes a whole-body lymphoma segmentation method that integrates anatomical and metabolic statistical priors. First, the CT-based whole-body organ segmentation method TotalSegmentator is employed to automatically segment 107 tissue/organ types across 300 PET/CT datasets. The metabolic intensity distribution of each tissue/organ is statistically analyzed to construct a population-level human metabolic statistical model. Building upon this foundation, an organ-specific metabolic suppression strategy is introduced for physiologically hypermetabolic regions, thereby reducing the complexity of the whole-body lymphoma segmentation task. Furthermore, a two-stage global-to-local segmentation framework is established: nnUNet performs coarse segmentation across the entire body, followed by local PET/CT-based fine segmentation using nnUNet to refine segmentation based on the probability map generated in the first stage, which locates potential lesion regions. Experimental results on the AutoPET dataset demonstrate that this method effectively reduces false positives caused by physiologically hypermetabolic organs and outperforms state-of-the-art medical image segmentation methods in overall segmentation performance. MCA-UNet: A Multi-Center Adaptive U-Net for Improved Tumor Segmentation in Multi-Center PET/CT Images Meng Wang, Yongwei Tang, Yarong Feng, and Yuxin Liang (Liaoning Technical University); Tianyu Shi (Shenyang Ligong University); and Chao Lv (China Medical University) Abstract Abstract The performance of medical image segmentation models faces the core challenge of data heterogeneity in multi-center collaborative training. Existing methods typically process mixed multi-center data directly while ignoring center-specific characteristics, resulting in suboptimal performance. To address this issue, this study proposes a Multi-Center Adaptive U-Net (MCA-UNet) for PET/CT images. The proposed method introduces a hierarchical center adaptation mechanism: center-specific adapters are employed in shallow layers to handle low-level variations, while multi-branch cross-center blocks in deeper layers balance general features and center-specific representations. In addition, MCA-UNet incorporates a low-rank dual-path fusion strategy to further enhance overall performance. Extensive experiments were conducted on the publicly available HECKTOR 2021 head and neck tumor segmentation dataset. Compared with other methods, MCA-UNet achieves superior performance, with a Dice score of 81.50% and an HD95 of 5.48 mm, with limited parameter growth (less than 5%). These results demonstrate that the proposed method can significantly improve overall segmentation performance on multi-center PET/CT images. Monday Virtual Room 3 IJCNN Paper SS43 Computational Intelligence and Software Engineering I Session Chair: Priyaranjan Pattnayak (Oracle Cloud AI), Xiaobing Xiong (Key Laboratory of Cyberspace Security, Ministry of Education; Information Engineering University) CausalVul: Robust Vulnerability Detection via Dual-Granularity Semantic Fusion and Causal Structural Learning Guoyu Huo, Xiaobing Xiong, Qi Han, and Fei Kang (Key Laboratory of Cyberspace Security, Ministry of Education; Information Engineering University) Abstract Abstract Software vulnerability detection is a critical task in cybersecurity. Existing hybrid methods combining Graph Neural Networks (GNNs) and Pre-trained Code Models (PCMs) face two fundamental limitations: structural noise in Code Property Graphs (CPGs) introduces spurious correlations, and the granularity mismatch between global semantics and local features results in insufficient fusion. To address these issues, we propose CausalVul, a robust framework based on dual-granularity semantic fusion and causally-motivated structural pruning. First, we design a curriculum learning-driven dynamic edge pruning mechanism that gradually shifts focus from global exploration to vulnerability-relevant causal subgraphs. Second, we construct a dual-granularity encoding architecture to extract function-level sequence semantics and fine-grained graph structural features in parallel, facilitating deep cross-modal fusion via gating mechanisms. Finally, we propose a multi-view optimization objective combining node-level Multiple Instance Learning, class center regularization, and supervised contrastive learning to enhance discriminative power on imbalanced data. Experiments on three widely used benchmarks (Devign, ReVeal, and Big-Vul) demonstrate that CausalVul achieves F1-scores of 65.05%, 51.01%, and 54.55%, respectively. Notably, on the highly imbalanced Big-Vul dataset, our method achieves competitive performance, demonstrating the effectiveness of the proposed approach on challenging real-world scenarios. CL-ICE: A White-Box Framework for Cross-Language Code Translation Evaluation Qi Han, Guoyu Huo, Yan Guang, and Fei Kang (Key Laboratory of Cyberspace Security, Ministry of Education; Information Engineering University) Abstract Abstract Evaluating cross-language code translation remains a fundamental challenge due to the inherent diversity of functionally equivalent implementations. Traditional reference-based metrics fail to account for valid solutions that diverge in syntax, while execution-based evaluation is often difficult to scale because it depends on runnable environments, test cases, and external dependencies. We propose CL-ICE (Cross-Language Internal Concept Evaluation), a white-box framework that probes the internal representations of Code Large Language Models (Code LLMs) to assess translation quality without execution or reference comparison at inference time. Our approach leverages two key insights from recent interpretability research: (1) intermediate layers of Code LLMs form language-agnostic concept layers—where abstract, universal representations of algorithmic logic are encoded regardless of surface syntax, and (2) LLMs encode an internal “correctness” signal in their activation space that can be extracted through Linear Artificial Tomography (LAT). CL-ICE includes two complementary components: concept-layer profiling for semantic invariance and language-specific LAT for latent correctness probing. Experiments demonstrate that CL-ICE achieves correlation with functional correctness (r = 0.582) that approaches that of GPT-4-based oracle evaluation (r = 0.615), while maintaining deterministic operation and low computational cost. SAD-Gen: Structure-Aware Unit Test Generation via Dual-Encoder Fusion and Latent Space Disentanglement Jingqi Gao, Haoran Yan, Kun Yang, Nan Liang, Wentao Qiu, Pengfei Ding, Jingan Chen, and Qingqiang Wu (Xiamen University) Abstract Abstract Large Language Models (LLMs) have demonstrated impressive capabilities in automated unit test generation. How- ever, they predominantly rely on textual co-occurrence patterns, often neglecting the explicit control flow and dependency struc- tures inherent in source code. This limitation frequently results in test cases that suffer from logical hallucinations and insufficient boundary coverage. While Graph Neural Networks (GNNs) excel at modeling non-Euclidean code structures, effectively aligning structural features with the textual semantics of LLMs remains a significant challenge. To address these issues, we propose SAD- Gen, a Structure-Aware Disentangled Generation framework. First, we present a dual-modal encoding architecture that lever- ages a textual encoder to capture sequential semantics alongside Graph Attention Networks (GAT) for extracting structural fea- tures from Abstract Syntax Trees (ASTs). These representations are dynamically fused via an adaptive gated mechanism. Second, to maximize generation diversity, we propose a novel Latent Space Disentanglement Strategy that decouples the program rep- resentation into immutable Logic Latents and perturbable Data Latents. By injecting Gaussian noise into the data latent space during generation, our model actively explores diverse boundary conditions while strictly preserving the underlying program logic. Extensive experiments on the Methods2Test and MBPP datasets demonstrate that SAD-Gen consistently outperforms state-of-the- art baselines in terms of syntax pass rate, branch coverage, and mutation score, validating the effectiveness and robustness of our proposed approach. Automated Postmortem Action Generation (PAG) using LLM Priyaranjan Pattnayak and Sanchari Chowdhuri (Oracle Cloud AI) Abstract Abstract We present Postmortem Action Generation (PAG), a novel NLP framework for intelligently generating structured Corrective and Preventive Measures (CAPMs) from post-incident analyses in large-scale cloud systems. Unlike prior Root Cause Analysis (RCA) work that identifies why failures occur, PAG focuses on what to do next producing actionable, role-specific remediation steps. Our approach combines hybrid retrieval of similar incidents, persona-conditioned prompting, and LLM-based validation to ensure correctness, actionability, and role fit. Evaluated on over 8,000 real-world cloud incidents, PAG achieves strong performance with factual consistency (0.83), groundedness (85.3%), and expert-rated correctness (4.3/5), surpassing fine-tuned and rule-based baselines across automated and human evaluations. The architecture is lightweight, scalable, and production-ready frozen LLMs paired with retrieval and validation loops delivering reliable, large-scale post-incident decision support. Monday Virtual Room 4 IJCNN Paper SS43 Computational Intelligence and Software Engineering II Session Chair: Yunpeng Wang (Taiyuan University of Technology), Chengxiao Zhao (Qilu University of Technology (Shandong Academy of Sciences), Shandong Computer Science Center (National Supercomputer Center in Jinan)) A Vulnerability Detection Method with Semantic-Sensitive Contrastive Learning and Token-Level Graph Representation Dawei Zhao, Chengxiao Zhao, Xin Li, Lijuan Xu, and Fenghua Tong (Qilu University of Technology (Shandong Academy of Sciences), Shandong Computer Science Center (National Supercomputer Center in Jinan)) and Haipeng Peng (Information Security Center, State Key Laboratory of Networking and Switching Technology, and the National Engineering Laboratory for Disaster Backup and Recovery, Beijing University of Posts and Telecommunications) Abstract Abstract As information systems become increasingly complex, vulnerability detection has emerged as a critical means of identifying and mitigating security risks. Artificial intelligence-driven methods offer unique advantages in terms of automation and intelligence. However, the training process of current intelligent detection models tends to learn shallow patterns from the training data, overlooking the underlying semantic features of vulnerabilities. To address this challenge, we propose a novel vulnerability detection method named CoGRVD, which employs semantic-sensitive contrastive learning and token-level graph representation learning to capture critical semantic and structural features. Specifically, we design semantically equivalent yet vulnerability-aware code transformation methods that guide the model to focus on capturing meaningful vulnerability-related semantics. In addition, we transform the input into a token-level co-occurrence graph and employ GCNs with residual connection to effectively capture the long-term dependencies among tokens. We fuse the semantic and structural features and employ a classifier for prediction. To evaluate the effectiveness of our proposed method, we conduct extensive experiments on three public benchmark datasets. The experimental results demonstrate that CoGRVD outperforms all other baseline methods across all three datasets. Specifically, on the FFMPeg+Qemu, Reveal, and Big_vul datasets, its F1 scores reached 74.55%, 63.73%, and 90.23%, respectively, showing improvements of 7.61%, 13.16%, and 40.78% compared to the best baseline methods. SARV: Structure-Aware Retrieval and Multi-Agent Verification for Software Traceability Tongjia Ma (School of Information and Electrical Engineering, Hebei University of Engineering, Handan 056038, China; Center for Information Research, Academy of Military Sciences, Beijing 100142, China); Yuan Huang (School of Information and Electrical Engineering, Hebei University of Engineering, Handan 056038, China); Ziwei Lei, Chen Ling, and Yanfei Lv (Center for Information Research, Academy of Military Sciences, Beijing 100142, China); and Huihong He (China National Tendering Center of MACH. & ELEC. Equipment (Government Procurement Center of MIIT), No. 46 Enjizhuang, Haidian District, F Area, Beijing, 100142, P.R. China) Abstract Abstract In software engineering, establishing accurate traceability links between requirements and source code is essential for many development and maintenance tasks, yet it remains challenging in large and structurally complex software systems. Recent retrieval-augmented generation approaches based on large language models have introduced new opportunities for automated traceability link recovery. However, their effectiveness is often constrained by retrieval quality and the vulnerability of single-pass decision mechanisms to reasoning errors and hallucinations. This paper proposes SARV, a framework that explicitly decouples candidate link generation from link validation. In the retrieval stage, a structure-aware hybrid retrieval strategy is employed to identify potential links. In the validation stage, SARV introduces a multi-agent collaborative mechanism guided by auxiliary information, which applies a reflect-and-correct strategy to transform traceability decisions from a one-shot classification process into a verifiable and error-correctable validation process. Experimental results on multiple public datasets show that SARV achieves competitive and stable performance across diverse project settings. DSLA: End-to-End Feature Fusion for Superior Fault Localization of Software Qun Xiao (Institute of Information Engineering, CAS; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China); Ming Zhou (School of Cyber Science and Engineering, Nanjing University of Science and Technology); Jiaqian Peng (Institute of Information Engineering, CAS); Shouguo Yang (Zhongguancun Laboratory); and Zhanwei Song, Zhiqiang Shi, Zhi Li, and Limin Sun (Institute of Information Engineering, CAS) Abstract Abstract Current fault localization methods struggle to prioritize true faulty lines due to coverage ambiguities in dynamic analysis and missing semantic dependencies in static analysis. Existing feature-concatenation strategies fail to produce sufficiently discriminative suspiciousness scores for Top-N recall. To address this, we propose DSLA, which constructs a line-level property graph integrating dynamic coverage with static code semantics, enabling fine-grained code representation. DSLA performs deep feature fusion via a joint objective combining MSE, t-SNE constraints, and discriminative losses, explicitly enforcing geometric separability to prevent model collapse. On the Defects4J benchmark, DSLA improves Top-1 recall by 1.3× over state-of-the-art techniques, and on the C-language Codeflaws dataset, it achieves a 6.1× improvement, demonstrating its effectiveness in precise faulty line localization. DeepDebateCoder: Meta-Policy Modeling and Few-Shot Anchored Asymmetric Debate for Reliable Code Generation Yunpeng Wang, Yangzhi Han, Jiaying Gao, Wenrui Zhang, Jianfeng Wang, and Wei Zhang (Taiyuan University of Technology) Abstract Abstract Large language models (LLMs) are increasingly effective at generating code, but they still struggle on specification-intensive tasks where reliability is critical. Such tasks typically come with long problem statements, layered constraints, and many corner cases, and they are often judged by hidden tests. To address this gap, we introduce DeepDebateCoder, a closed-loop framework designed to improve robustness and structural consistency in LLM-generated code. DeepDebateCoder first builds an explicit Policy Graph that captures module boundaries, branching logic, state transitions, and I/O constraints. This graph then serves as a structured guide for code synthesis and refinement. To go beyond example-level correctness, it generates policy-aligned tests that target boundary conditions and exception paths, and validates them before using them for refinement. Finally, DeepDebateCoder uses execution feedback to drive an asymmetric debate between a strategy expert and a code expert. Instead of repeatedly patching surface-level errors, the debate focuses on revising the underlying decision path. Across HumanEval and APPS, DeepDebateCoder consistently improves pass@1 and robustness over competitive baselines. The gains are especially pronounced on APPS, with the largest improvements on long-specification and high-difficulty tasks, underscoring the value of meta-policy modeling and few-shot-anchored asymmetric debate for achieving highly reliable code generation. More broadly, our framework points to a new paradigm for reliability-oriented program synthesis. Monday Virtual Room 5 IJCNN Paper SS43 Computational Intelligence and Software Engineering III Session Chair: Hamed Jelodar (UNB), Yuchuan Chen (Changsha University of Science and Technology, School of Computer Science and Technology) MFVD: A Multimodal Feature-Based Vulnerability Detection Framework for Smart Contracts Tingxun Chen, Bo Yin, Yuan Zhu, and Yuchuan Chen (Changsha University of Science and Technology) Abstract Abstract Smart contract vulnerability detection (SCVD) has attracted increasing attention in both academia and industry. Pre-trained language models (PLMs) are trained on large-scale data, have shown strong generalization ability. However, existing PLMs-based SCVD methods are typically limited to processing inputs of 512 or 1024 tokens, or lack structural information, which degrades vulnerability detection performance. An important question arises: How effective are existing PLMs when applied to SCVD? In this paper, we introduce a new multimodal feature-based vulnerability detection framework named MFVD. We utilize segmentation encoding strategy and graph presentation to tackle the limitations above. MFVD enhances UniXcoder to process multimodal data—the source code sequence and the data flow graph—and enables binary classification. Extensive experiments demonstrate that MFVD outperforms state-of-the-art methods in recall, precision, and F1 scores, as well as in the stability of these metrics, achieving 96.14% F1 for reentrancy vulnerabilities and 96.71% F1 for timestamp dependency vulnerabilities. We also evaluate 11 PLMs on source code, with results indicating that UniXcoder outperforms the others in most scenarios. BAVul: Code Vulnerability Detection based on Bug Attention Chunyu Yang and Bo Zhao (Wuhan University, School of Cyber Science and Engineering) Abstract Abstract Machine learning is widely applied in software engineering tasks, particularly code vulnerability detection, owing to its ability to automatically identify patterns in data. However, existing machine-learning-based vulnerability detection methods exhibit several limitations. Reliance on a single feature type fails to capture comprehensive code semantics, leading to elevated false positive and false negative rates. Excessive dependence on fully automated feature learning overlooks potential inter-feature relationships, thereby impeding the model's capacity to recognize specific vulnerability patterns. Furthermore, the use of excessively high-dimensional features incurs substantial computational overhead during large-scale analysis, rendering real-time detection impractical. To address these shortcomings, we propose BAVul, a vulnerability detection system based on a bug attention mechanism. BAVul embeds source code using CodeBert language model to preserve vulnerability-relevant semantic information in the sequence. It then reconstructs the Control Flow Graph (CFG), Abstract Syntax Tree (AST), and Program Dependence Graph (PDG), and generates embeddings for these graphs via Graph Attention Network (GAT) to preserve structural information. During feature fusion, a bug attention mechanism is introduced to enable weighted integration of graph embeddings, thereby compensating for the limited modeling of vulnerability patterns in conventional graph-based approaches. In the detection phase, a refinement-based dimensionality reduction technique is combined with LightGBM to enhance both efficiency and accuracy. Compared with conventional machine-learning methods, BAVul achieves substantial improvements in F1-score and computational performance. Semantic and Temporal Feature Augmentation for Log-Based System Failure Prediction Yuan Tian, Shi Ying, and Junjun Li (Wuhan University) Abstract Abstract Early failure prediction from system logs is notoriously difficult because pre-failure signals are weak, heterogeneous, and easily overwhelmed by massive noisy logs. A key but under-exploited fact is that these early signals are structured across two levels: (i) rare semantic cues embedded in individual log messages, and (ii) temporal dynamics reflecting how the system state gradually drifts and occasionally transitions toward failure. Motivated by this observation, we propose a dual-level semantic and temporal feature augmentation approach for robust early failure prediction from noisy logs. At the semantic level, we introduce a random-walk-driven word reweighting mechanism that injects frequency-imbalance priors between normal and failure-related corpora, amplifying failure-indicative tokens without relying on log parsing or domain knowledge. At the temporal level, we design a complementary dynamics fusion predictor that couples an attention-based BiLSTM with multiple HMM components, where BiLSTM captures long-range non-linear evolution trends while HMMs model local Markovian state transitions associated with abnormal shocks. The two representations are then fused for accurate sequence classification. Extensive experiments on four public log datasets show that our method consistently outperforms strong failure prediction baselines and substantially surpasses representative log anomaly detectors, yielding higher F1 scores and improved recall under noisy conditions with practical inference overhead. These results demonstrate that jointly augmenting semantic saliency and temporal dynamics is crucial for reliable early failure prediction in real-world logging environments. LLM4CodeRE: Generative AI for Code Decompilation Analysis and Reverse Engineering Hamed Jelodar, Samita Bai, Tochukwu Emmanuel Nwankwo, Parisa Hamedi, Mohammad Meymani, Roozbeh Razavi-Far, and Ali A. Ghorbani (UNB) Abstract Abstract Code decompilation analysis is a fundamental yet challenging task in malware reverse engineering, particularly due to the pervasive use of sophisticated obfuscation techniques. Although recent large language models (LLMs) have shown promise in translating low-level representations into high-level source code, most existing approaches rely on generic code pretraining and lack adaptation to malicious software. We propose LLM4CodeRE, a domain-adaptive LLM framework for bidirectional code reverse engineering that supports both assembly-to-source decompilation and source-to-assembly translation within a unified model. To enable effective task adaptation, we introduce two complementary fine-tuning strategies: (i) a Multi-Adapter approach for task-specific syntactic and semantic alignment, and (ii) a Seq2Seq Unified approach using task-conditioned prefixes to enforce end-to-end generation constraints. Experimental results demonstrate that LLM4CodeRE outperforms existing decompilation tools and general-purpose code models, achieving robust bidirectional generalization. Monday Virtual Room 6 IJCNN Paper SS04 Tiny Machine Learning I Session Chair: Zonglin Yang (Guangdong Police College), Fei Ge (School of Computer Science, Central China Normal University, Wuhan, China) Soft Inductive Biases Accelerate In-Context Learning in Tiny Transformers Zonglin Yang (GuangDong Police Collage) Abstract Abstract Small Transformers trained from scratch on synthetic meta-learning corpora can learn to perform in-context learning (ICL), but doing so is notoriously slow and data-inefficient. We ask whether soft architectural priors-implemented as gated, learnable additions to attention logits can steer these tiny models toward ICL solutions faster while preserving flexibility. We augment attention heads with (i) a learnable relative-position prior and (ii) a content-matching prior, both controlled by learnable scalar gates regularised with weight decay so they can shrink toward zero once ICL circuitry forms. We evaluate two-layer Transformers on linear regression, associative recall, and algorithmic reasoning, comparing SOFT-BIAS against strong baselines including ALIBI and Learned Relative Positional Encodings (RPE). Our results show that SOFT-BIAS acts as an efficient catalyst: it accelerates convergence by approximately 2x compared to the VANILLA baseline. Crucially, in noisy regression settings where standard Learned RPE overfits, SOFT-BIAS remains robust due to its gating mechanism. Furthermore, on associative recall tasks, it achieves a superior trade-off between accuracy and training throughput compared to data-adaptive baselines. Analysis of learned gates confirms that the priors accelerate the emergence of ICL circuitry without constraining the model's final plasticity. Tiny Spiking Neural Network for On-Edge PPG-based Blood Pressure Estimation Rathore Shakti Singh, Francesco Carlucci, and Daniele Jahier Pagliari (Politecnico di Torino); Elisa Donati (UZH); and Benedetto Leto, Gianvito Urgese, Massimo Poncino, Enrico Macii, Vittorio Fra, and Alessio Burrello (Politecnico di Torino) Abstract Abstract Photoplethysmography (PPG)-based blood pressure (BP) estimation is a complex biosignal processing task, particularly challenging for resource-constrained wearable devices. Deep learning methods have achieved state-of-the-art performance on large public datasets but still require hundreds of kilobytes of memory for deployment, even after strong optimization. In this work, we show the application of the spiking neural networks (SNNs). Thanks to the inherent suitability of SNNs for temporal signals, our best quantized network achieves comparable or better performance on four benchmark datasets compared to state-of-the-art (SOTA) deep neural networks (DNNs), while strongly reducing the number of parameters by up to 13.8$\times$. We also demonstrate that our optimized SNN can be executed on wearable devices, deploying it on a low-power RISC-V-based microcontroller unit (MCU) and obtaining a memory footprint of 4.63 kB, a latency of 1.34 ms, and an energy consumption of 0.068 mJ. Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores Xiang Fu, Jixiang Ma, and Xinpeng Zhang (Nanchang Hangkong University); Peng Zhao (EEO Tech); Shuai Lu (Jiangxi University of Finance and Economics); and Xu Tony Liu (Amazon, University of Washington) Abstract Abstract Convolution is a principal computational bottleneck in deep neural networks, and its efficiency depends on tight integration between algorithms and GPU hardware. Existing GPU convolution methods suffer from large memory overhead, poor cache utilization, limited effectiveness across kernel sizes, or numerical instability. This work extends the im2win paradigm—a universal, memory-efficient convolution method with contiguous memory access for all kernel sizes—to run efficiently in full precision on CUDA cores and half precision on tensor cores. By introducing new kernel designs and optimizations such as zig-zag memory access and asynchronous data movement, im2win efficiently exploits hardware-accelerated half-precision matrix multiply-accumulate operations. Across twelve CNN benchmarks, im2win achieves up to 2.8$\times$ higher TFLOPS than its CUDA core implementation, 1.4$\times$ higher than cuDNN, and 6.4$\times$ higher than GEMM-based convolution with cuBLAS, while using as little as 53\% and 35\% of their memory, respectively. These results establish im2win as a unified, high-performance convolution framework for modern GPU architectures. ST-CA2: Lightweight CSI Input Refinement for Efficient and Tiny Wi-Fi Sensing Models Jiangyang Liu, Fei Ge, and Kang Gaoming (School of Computer Science, Central China Normal University, Wuhan, China) Abstract Abstract Wi-Fi sensing on resource-constrained platforms calls for efficient and tiny neural networks, yet CSI pipelines face a recurring trade-off between preserving the full multi-antenna CSI tensor and compressing the input by simple heuristics. We propose ST-CA2 (SpatioTemporal Channel–Antenna Attention), a lightweight input refinement module that performs subcarrier wise channel attention and antenna-pair-wise attention to output a compact C-by-T representation. Under a fixed downstream recognizer and in a door-opening person-identification setting, ST-CA2 achieves 99.20% OA with 6.208 GFLOPs and 2.006 GB peak GPU memory, compared with 97.20% OA with 53.232 GFLOPs and 8.601 GB for full-input. These results indicate a favorable accuracy–efficiency trade-off for compact CSI input refinement in the studied setting. We also provide analyses under signal-plus-noise assumptions and evaluate temporal statistical priors (GTA, variance, and RMS) for weight generation. Monday Virtual Room 7 IJCNN Paper SS04 Tiny Machine Learning II Session Chair: Chenkai Liao (Hunan University of Technology), jiaping Wang (East China Normal Univerisity) A Lightweight Face Image–Based Auxiliary Detection Model for Autism Spectrum Disorder Chenkai Liao and Wenqiu Zhu (Hunan University of Technology) Abstract Abstract Early diagnosis of Autism Spectrum Disorder (ASD) plays a crucial role in improving patients’ quality of life. In recent years, face image–based ASD detection has attracted increasing attention as an auxiliary diagnostic approach. However, existing lightweight models still show limitations in capturing fine-grained facial features. To address this problem, this paper proposes a lightweight face image–based auxiliary detection model for ASD, termed MN-ASD. First, MobileNetV4-S is selected as the baseline framework. To enhance the model’s ability to capture subtle facial details, the FReLU activation function is introduced, which strengthens spatial feature modeling. Second, to further improve the performance and accuracy of the model, the Coordinate Attention (CA) module is incorporated at the final stage. By jointly modeling positional and channel dependencies, the CA mechanism enables the model to better locate key facial regions and enhance global feature extraction. Finally, experimental results demonstrate that the proposed method achieves an accuracy of 93.33% on the Kaggle dataset, with precision, recall, and F1-score all outperforming other mainstream lightweight networks. These results demonstrate the effectiveness of the proposed MN-ASD model in capturing discriminative facial features, while highlighting its strong potential for deployment in resource constrained and TinyML-oriented environments. Insights into the Lottery Ticket Hypothesis and the Iterative Magnitude Pruning Algorithm Tausifa Jan Saleem, Ramanjit Ahuja, Surendra Prasad, and Brejesh Lall (Indian Institute of Technology Delhi) Abstract Abstract The Lottery Ticket Hypothesis for deep neural networks emphasizes the importance of initialization used to retrain the sparser networks obtained using the iterative magnitude pruning process. An explanation for why the specific initialization proposed by the lottery ticket hypothesis tends to work better in terms of generalization (and training) performance has been lacking. Moreover, the underlying principles in iterative magnitude pruning, like the pruning of smaller magnitude weights and the role of the iterative process, lack full understanding and explanation. In this work, we attempt to provide insights into these phenomena by empirically studying the volume/geometry and loss landscape characteristics of the solutions obtained at various stages of the iterative magnitude pruning process. Data-Augmented Quantization-Aware Knowledge Distillation Justin Kur, Jon Maravilla, and Kaiqi Zhao (Oakland University) Abstract Abstract Quantization-aware training (QAT) and Knowledge Distillation (KD) are combined to achieve competitive performance in creating low-bit deep learning models. Existing KD and QAT works focus on improving the accuracy of quantized models from the network output perspective by designing better KD loss functions or optimizing QAT’s forward and backward propagation. However, limited attention has been given to understanding the impact of input transformations, such as data augmentation (DA). The relationship between quantization-aware KD and DA remains unexplored. In this paper, we address the question: how to select a good DA in quantization-aware KD, especially for models with low precision? We propose a novel metric which evaluates DAs according to their capacity to maximize the Generalized Contextual Mutual Information--the information not directly related to an image's labels--while also ensuring the predictions for each class are close to the ground truth labels on average. The proposed method automatically ranks and selects DAs, requiring minimal training overhead, and it is compatible with any KD or QAT algorithm. Extensive evaluations demonstrate that selecting DA strategies using our metric significantly improves state-of-the-art QAT and KD works across various model architectures and datasets. Light Wings: More Efficient Speculative Decoding via Linear Attention Jiaping Wang (East China Normal University), Yifan Xu and Yifan Chen (Hong Kong Baptist University), and Jianwen Li (East China Normal University) Abstract Abstract Speculative decoding (SD) accelerates large language model (LLM) inference through introducing a fast draft model to propose candidate token sequences. However, conventional draft models employ softmax-based attention for capturing original transformers, incurring quadratic time and memory complexity during prefilling. Given the limited resources assigned to drafters, this design restricts their scalability in long-context scenarios. To this end, we propose a new drafter architecture built on linear attention, which achieves linear computational complexity. Specifically, our method maintains a fixed-size key-value cache during autoregressive generation, enabling efficient, scalable inference with minimal overhead. We evaluated the proposed approach on multiple model scales and benchmarks, demonstrating comparable token acceptance rates and downstream performance to standard attention mechanisms, while achieving an average 2.62 times speed-up. Monday Virtual Room 8 IJCNN Paper SS04 Tiny Machine Learning III Session Chair: Hatem Trigui (Hahn Schickard), BINHUA HUANG (University College Dublin) FedTinyProp: Adaptive Sparse Backpropagation for Efficient Federated Learning on Embedded Devices Hatem Trigui (Hahn Schickard); Marcus Rüb (ForestHub); Lilli Frison (Hahn Schickard, University Freiburg); Axel Sikora (Offenburg University of Applied Sciences); and Oliver Amft (Hahn Schickard, University Freiburg) Abstract Abstract We introduce FedTinyProp, a Federated Learning (FL) framework that adapts sparse backpropagation for on-device training. In our approach, clients (edge devices) dynamically sparsify gradients based on per-batch informativeness and skip updates when computation is unnecessary. Clients transmit sparsified weight deltas that are aggregated via a communication-aware Federated Averaging variant. LiteTSC: A Lightweight Neural Network for Time Series Classification on Edge Devices Tao Wang (Xiamen University) Abstract Abstract Human activity recognition (HAR) using wearable sensors has gained significant attention for healthcare monitoring and smart device applications. However, deploying deep learning models on resource-constrained edge devices remains challenging due to high computational costs and memory requirements. In this paper, we propose LiteTSC, an ultra-lightweight neural network for time series classification that achieves competitive accuracy with significantly reduced model complexity. Our approach integrates three key components: (1) multi-scale depthwise separable convolutions that capture temporal patterns at different granularities while reducing parameters, (2) a lightweight dual-pooling channel attention mechanism that enhances feature discrimination with minimal overhead, and (3) residual connections that compensate for the reduced representational capacity of lightweight operations. Extensive experiments on three benchmark datasets (UCI HAR, WISDM, and PAMAP2) demonstrate that LiteTSC achieves state-of-the-art accuracy (96.74% on UCI HAR) with only 20.8K parameters. Compared to TinyHAR, the ISWC 2022 Best Paper, LiteTSC achieves higher accuracy with 14.3× fewer parameters and 3.3× faster inference. The lightweight design makes LiteTSC suitable for deployment on resource-constrained devices. Onboard Implementation of Machine Learning for Autonomous Vortex Detection on Mars Anita Amofah, Daxell Wells, and Rajil Sajila (Fayetteville State University); Gary Doran (Jet Propulsion Laboratory, California Institute of Technology); and Sambit Bhattacharya (Fayetteville State University) Abstract Abstract Atmospheric vortices on Mars, such as dust devils, play a critical role in the planet’s climate, dust transport, and surface--atmosphere interactions. Capturing high-rate wind, pressure, and dust sensor data during vortex encounters provides valuable scientific insight; however, such events are rare and unpredictable. Continuous high-rate data collection is therefore impractical due to strict power and onboard data storage constraints, particularly for low-cost landers. Prior work has proposed using machine learning algorithms to detect approaching vortices from low-rate pressure measurements and autonomously trigger high-rate data acquisition during the most scientifically relevant portion of the event. In this work, we extend prior vortex detection studies by evaluating both classical machine learning models and a two-stage deep learning pipeline for early-warning detection under precision-critical deployment constraints. Specifically, we assess Random Forest and Extreme Gradient Boosting (XGBoost) classifiers trained on statistically engineered features, alongside a two-stage autoencoder-gated Temporal Convolutional Network (AE$\rightarrow$TCN) designed to filter nominal atmospheric behavior prior to classification. Model performance is evaluated using metrics based on precision, recall, and F1-score appropriate for rare-event detection. In addition, we examine the feasibility of onboard deployment by benchmarking inference latency and power consumption on a Qualcomm Snapdragon–class processor representative of spaceflight and edge-computing hardware. By jointly analyzing detection performance and computational efficiency, this study provides a comparative assessment of accuracy–resource trade-offs relevant to autonomous onboard triggering and demonstrates the potential of learning-enabled intelligent sensing to enhance Martian atmospheric science while respecting mission resource constraints. MoCrop: Training Free Motion Guided Cropping for Efficient Video Action Recognition BINHUA HUANG and Wendong Yao (University College Dublin), Shaowu Chen (Shenzhen Polytechnic University), Guoxin Wang and Qingyuan Wang (University College Dublin), and Soumyabrata Dev (Trinity College Dublin) Abstract Abstract Standard video action recognition models often process typically resized full frames, suffering from spatial redundancy and high computational costs. To address this, we introduce MoCrop, a motion-aware adaptive cropping module designed for efficient video action recognition in the compressed domain. Leveraging Motion Vectors (MVs) naturally available in H.264 video, MoCrop localizes motion-dense regions to produce adaptive crops at inference without requiring any training or parameter updates. Our lightweight pipeline synergizes three key components: Merge & Denoise (MD) for outlier filtering, Monte Carlo Sampling (MCS) for efficient importance sampling, and Motion Grid Search (MGS) for optimal region localization. This design allows MoCrop to serve as a versatile "plug-and-play" module for diverse backbones. Extensive experiments on UCF101 demonstrate that MoCrop serves as both an accelerator and an enhancer. With ResNet-50, it achieves a +3.5% boost in Top-1 accuracy at equivalent FLOPs (Attention Setting), or a +2.4% accuracy gain with 26.5% fewer FLOPs (Efficiency Setting). When applied to CoViAR, it improves accuracy to 89.2% or reduces computation by roughly 27% (from 11.6 to 8.5 GFLOPs). Consistent gains across MobileNet-V3, EfficientNet-B1, and Swin-B confirm its strong generality and suitability for real-time deployment. Our code and models are available at https://github.com/microa/MoCrop. Monday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V2 Session Chair: Xiangming Jiang (Xidian University) Unbounded Differential Evolution with Very Large Populations Xiuyi Guo (The University of Tokyo); Tomofumi Kitamura (The University of Tokyo, Meteorological Research Institute); and Alex Fukunaga (The University of Tokyo) Abstract Abstract Standard evolutionary search algorithms, including differential evolution (DE), use a relatively small, limited-size population that remains the same size or gradually shrinks during the search. It has long been believed that replacement (discarding some members of the population and inserting new candidate individuals in their place) is an important feature of evolutionary search. Recent work on Unbounded DE (UDE) has shown that non-increasing populations are not necessary, and that with an appropriate selection operator, a DE with a monotonically growing population can perform comparably to state-of-the-art adaptive DE. However, UDE has scalability issues, as the runtime complexity of each iteration increases with population size, and limited RAM capacity poses a limitation on the growth of the population. We propose extensions to UDE that overcome these limitations by (1) efficiently storing the population on external storage (SSD) and (2) reducing the computational complexity of tournament selection. We show experimentally that this allows UDE to scale to very large populations with 2e8 individuals. HADI: Multi-Objective Test Case Prioritization Framework Using NSGA-II Algorithm Abdullahi Hussein, Dennis Mugambi Kaburu, and Isaac Nyabisa Oteyo (Jomo Kenyatta University of Agriculture and Technology) Abstract Abstract Test case prioritization (TCP) addresses the challenge of efficiently ordering test cases to improve fault detection rates while managing execution time constraints in regression testing. This study presents HADI framework for multi-objective test case prioritization using Non-dominated Sorting Genetic Algorithm II (NSGA-II). The study evaluates the framework using the Defects4J benchmark. Our framework optimizes two competing objectives: (i) minimizing the execution time to fault detection (ETAF), and (ii) maximizing the Average Percentage of Faults Detected (APFD). The findings show that HADI achieves a mean APFD of 0.98 compared to 0.51 for original orderings, representing a 92% improvement. Statistical validation using Wilcoxon signed-rank tests confirms significance (p < 0.001) with large effect sizes (Vargha-Delaney A12 > 0.89). Convergence analysis shows that HADI typically reaches optimal solutions within 10-20 generations, with consistent performance across 30 independent runs per dataset. Index Terms—Test case prioritization, NSGA-II, multiobjective optimization, Defects4J, Pareto optimization. Adaptive Epsilon-Penalty Switching for Push-Pull Search from Vicinity to Feasibility Jiaping Hu, Jiachung Huang, and Wenji Li (Shantou University); Zhaojun Wang (Jinpeng Electronic Information Machine); Wulin Cai, Chenwen Ding, Hanyuan Zhang, and Biao Xu (Shantou University); and Zhun Fan (University of Electronic Science and Technology of China) Abstract Abstract The ϵ-constraint handling mechanism is a widely used method for dealing with constraints. However, when the value of ϵ is reduced to a very small magnitude, it becomes equivalent to the constrained dominance principle. Under these circumstances, maintaining the diversity of the algorithm becomes challenging. To overcome this challenge, we introduce AEP-PPS, an adaptive epsilon-penalty switching for push-pull search algorithm with the adaptive stage transition mechanism. AEP-PPS employs a two-stage search strategy. In the push stage, constraints are temporarily ignored to drive the population toward the UPF. Based on the symmetric cross-generational distance indicator, the algorithm then adaptively switches to the pull stage, in which an ϵ-constraint handling mechanism combined with a penalty function is adopted to progressively enhance feasibility pressure and guide the population toward the CPF. In addition, an archive updating strategy based on θ-domination is introduced to preserve population diversity. Experimental results on the LIR-CMOP and CF test suites demonstrate that AEP-PPS outperforms nine state-of-the-art algorithms in terms of IGD and HV metrics. These results confirm its effectiveness in handling complex infeasible regions while achieving a good balance between convergence and diversity. A Memetic-Driven Multi-objective Sparse Unmixing Method for Hyperspectral Image Zidong Wu, Xiangming Jiang, Xiaolong Fan, and Mingyang Zhang (Xidian University); Jianzhao Li (Xidian university); and Maoguo Gong Gong (Xidian University) Abstract Abstract This paper presents a novel memetic hyperspectral unmixing approach based on the MOEA/D framework, focusing on endmember identification using a spectral library. The population is optimized by alternately applying genetic and local search operators. The local search operator is designed to extract and utilize knowledge accumulated across iterations, particularly the endmember abundances of each individual. This enables it to guide individual optimization toward the direction dictated by cultural genes. Additionally, a voting mechanism integrates multiple Pareto-optimal solutions by selecting the most frequently occurring endmembers as the final set, thereby mitigating the instability inherent in relying on a single solution. Extensive experiments on simulated and real datasets demonstrate that incorporating cultural genes significantly enhances the evolutionary algorithm’s performance, yielding superior unmixing results. Monday Virtual Room 1 IJCNN Paper SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models I Session Chair: Di Wu (Hebei University of Engineering), Shuaichao Zhang (Southwest University of Science and Technology) Adolescent Psychological Support Dialogue Generation Model of Domain Data-Driven Di Wu, Tenghao Zhang, and Jiasen Wang (Hebei University of Engineering) Abstract Abstract To address the limited adaptability of existing psychological dialogue data in the field of adolescent mental health support, and the insufficient guidance ability of current models, the Adolescent Psychological Support Dialogue Generation Model of Domain Data-Driven (APS-DM) is proposed. By extracting multi-level adolescent psychological topics, using topic-related keywords to filter and integrate data corpus, a context-based role-playing prompt is designed according to the characteristics of adolescent psychological development stages. It is intended to guide GPT-4 in performing multi-turn semantic reconstruction, thereby constructing an adolescent psychological support dialogue dataset (AdoPsyDialogue). Besides, a structured thinking-guided prompt is designed to constrain the response generation scenarios. A parameter-efficient fine-tuning method is used to optimize the base model, enhancing its ability to provide adolescent psychological support responses and guidance. Experimental results on the QsnTest dataset demonstrate that the APS-DM model significantly outperforms models such as ChatGLM2-6B and ChatGPT-3.5 in both automatic and manual metrics. Patient-State–Conditioned Pharmacodynamics Modeling With Counterfactual Inference for ADR Risk in Peritoneal Dialysis Yudi Zhang (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications); Xinqiu Li (Renal Division, Department of Medicine, Peking University First Hospital; Institute of Nephrology, Peking University); Kai Niu (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications); Jie Dong (Renal Division, Department of Medicine, Peking University First Hospital; Institute of Nephrology, Peking University); and Zhiqiang He (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications) Abstract Abstract Patients with chronic kidney disease (CKD) undergoing peritoneal dialysis (PD) commonly experience multimorbidity and polypharmacy, making adverse drug reaction (ADR) risk highly patient-specific. Conventional trials and drug-centric computational models often lack sufficient patient context and longitudinal validity to capture context-specific safety signals in real-world PD follow-ups. We propose PS-PDNet, a patient-state–conditioned deep learning framework that integrates clinical indicators and longitudinal prescriptions to predict ADR risk in PD. By tailoring pharmacodynamics modeling to each patient’s clinical context and temporally valid exposure history, PS-PDNet supports individualized safety assessment. Coupled with counterfactual inference, it highlights high-risk drugs and combinations and summarizes them into clinically interpretable category-level exposure patterns, providing actionable cues for ADR surveillance and medication optimization in real-world PD care. SCFD: A Self-Calibrated Feature Denoising Framework for Robust Medical Image Classification Under Extreme Class Imbalance Guanqi Cheng, Mukesh Prasad, and Ali Braytee (University of Technology Sydney) Abstract Abstract Deep learning models for medical image classification often underperform when class imbalance is extreme, as supervised methods overfit the majority class, leading to biased feature representations. To address this, we propose Self-Calibrated Feature Denoising (SCFD), a framework that enhances the performance of self-supervised models in imbalanced settings. SCFD integrates a pretrained DINOv2 encoder for class-specific feature extraction, a lightweight autoencoder to denoise minority-class embeddings, and a latent-space regularization mechanism that leverages majority-class information for calibration. By combining these components, SCFD generates balanced, discriminative embeddings without modifying the backbone, enabling seamless integration into existing pipelines. Experiments on three disease diagnosis datasets, Lab-shanghai, Brain Tumor 4C, and Chest CT, demonstrate that SCFD consistently improves model robustness and generalization under extreme class imbalance, providing a reliable approach for effective medical image classification. A Multiple Instance Learning Framework for Breast Cancer Based on Pseudo-Bag Augmentation and Double-Layer Masking Shuaichao Zhang, Minxian Liu, Bo Su, and Zhengwei Du (Southwest University of Science and Technology) Abstract Abstract Breast cancer histopathological image analysis is crucial for diagnosis and treatment planning. Although multiple instance learning (MIL) methods based on whole slide images (WSIs) have shown great promise, existing approaches often suffer from two limitations: overfitting caused by insufficient training samples and a training bias dominated by easily classified instances, which leads to the neglect of hard examples. To address these issues, we propose a multiple instance learning framework for breast cancer based on pseudo-bag augmentation and double-layer masking (PADM-MIL). Specifically, PADM-MIL first expands the training set via pseudo-bag partitioning and mixing. It then integrates a hierarchical Transformer and introduces a teacher network to construct slide-level prototypes. Interpretable instance importance (importance score) is computed using the cosine similarity between instance representations and prototypes. These importance scores further guide instance-level masking to mine hard examples. Subsequently, the student network performs additional group-level masking at the sub-bag level on hard-example bags, enabling the model to stably focus on key tissue regions. Experiments on two public breast cancer datasets demonstrate that PADM-MIL consistently outperforms existing methods in terms of AUC, ACC, and F1-score. Results on the TCGA-LUNG lung cancer dataset further indicate that the proposed framework exhibits promising generalization. Monday Virtual Room 2 IJCNN Paper SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models II Session Chair: Pei Zhou (SiChuanUniversity), Risheng Xie (University of Science and Technology of China, School of Computer Science and Technology) Balance Accuracy and Efficiency: Segment 3D medical images with inter-slice context information guidance Risheng Xie, Shouhong Wan, Peiquan Jin, and Weiyi Zhen (University of Science and Technology of China, School of Computer Science and Technology) Abstract Abstract Medical image segmentation plays a crucial role in disease detection, diagnosis and treatment. However, with the rapid increase in three-dimensional (3D) image data size, existing medical segmentation methods struggle in balancing computational cost and model performance. This paper proposes a slice-based segmentation method on 3D images and several plug-and-play Slice Context Modules (SCMs) to enable 2D networks to perform 3D segmentation. With an autoregressive method to extract inter-slice context information, SCMs use image morphological operations to expand this information and predict the inter-slice morphological changes with previous inter-slice context information, and use a channel-wise fusion module to guide the segmentation. SCMs improve the model's ability to sense and predict target morphological changes, while balancing the accuracy and efficiency in model training process. This paper also designed a specific training method called random partitioning to further accelerate the training process. Experiments conducted on open datasets have shown that the proposed methods increased the performance of baseline networks by over 4.6% in Dice score, surpassing multiple state-of-the-art 3D segmentation networks. Furthermore, SCMs achieve 30% faster training speed and 40% less GPU memory usage during training, demonstrating their effectiveness in 3D medical segmentation tasks. Bidirectional Dynamic Transformer for Medical Image Segmentation Jing Tong, Ling Ma, Yanbiao Ji, Shaokai Wu, Yue Ding, and Hongtao Lu (Shanghai Jiao Tong University) Abstract Abstract Medical image segmentation is a challenging task, where U-shaped networks are widely used. In order to better model long-range dependencies, many recent works integrate Transformer or state-space model (SSM) into U-Net and claim superior performance. However, such methods either suffer from semantic redundancy and heavy computation in self-attention, or require unnatural scanning routes in SSM. To validate their limitations, we first conduct an empirical study to show that existing Transformer- or SSM-based methods may actually fail to outperform vanilla U-Net under fair comparison settings, i.e., with the same pre -trained encoder and proper ablations. Then, to mitigate the limitations of self-attention, we propose BIdirectional Dynamic TransFormer (BIDFormer) to model spatial relationships and semantic correlations separately. (i) spatial relationships: to efficiently model global context, we replace self-attention with Bidirectional Linear Attention (BLA) between image patches and learnable class embeddings; (ii) semantic correlations: to explicitly model the correlation among target classes, we utilize class embeddings as Dynamic Low-rank Classifiers (DLC). Extensive experiments on public datasets demonstrate that BIDFormer can be easily integrated into existing models, leading to remarkable performance improvements. Code is available at https://github.com/JingTongsh/UN-Segmentation. Hybrid Frequency--Spatial Attention and Arch-Aware Priors for Tooth Detection and FDI Numbering in Dental Images Xi Wu (College of Computer Science, Sichuan University, Chengdu 610065, China); Ruijie Huang (West China School/Hospital of Stomatology, Sichuan University, Chengdu 610065, China); Jiangping Zhu (College of Computer Science, Sichuan University, Chengdu 610065, China; School of Information Science and Technology, Tibet University, Lhasa 850000, China); Shiquan Min and Pei Zhou (College of Computer Science, Sichuan University, Chengdu 610065, China); and Zhongjian Wang (MoEntropy Science (Chengdu) Pharmaceutical Technology Co., Ltd., Chengdu 610094, China) Abstract Abstract Current oral imaging modalities, such as panoramic radiographs and intraoral photographs, suffer from challenges including noise, blurring, occlusions, and significant variations in pose and scale. Intraoral photographs further exhibit specular reflections and complex illumination. Moreover, tooth arrangement is governed by strict morphological and anatomical constraints. These factors collectively pose significant challenges to accurate tooth position detection. To address these challenges, we propose a novel network architecture. Specifically, we introduce a Hybrid Frequency--Spatial Attention (HFSA) capable of simultaneously capturing both frequency and spatial features from images, thereby enhancing the network's robustness against various noise conditions. Furthermore, we design a Differentiable Class-Uniqueness Gating (UniGate) Mechanism to impose unique indexing constraints on each tooth. Additionally, we develop an Arch-aware Query Graph Prior (QGP), which enforces morphological and anatomical constraints on the detection process. We validate the performance of our approach through experiments on a public DSLR intraoral dataset, a private smartphone in-the-wild intraoral photograph dataset, and the public DENTEX2023 benchmark. Our method achieves state-of-the-art performance across all three datasets with AP scores of 78.3%, 54.4%, and 54.1%, respectively. GQD-DETR: A Geometry-Aware and Quality-Adaptive Transformer for Robust Tooth Detection in Heterogeneous Dental Images Chen Zhang (College of Computer Science, Sichuan University, Chengdu 610065, China); Ruijie Huang (West China School/Hospital of Stomatology, Sichuan University, Chengdu 610065, China); Jiangping Zhu (College of Computer Science, Sichuan University, Chengdu 610065, China; School of Information Science and Technology, Tibet University, Lhasa 850000, China); Shiquan Min and Pei Zhou (College of Computer Science, Sichuan University, Chengdu 610065, China); and Zhongjian Wang (MoEntropy Science (Chengdu) Pharmaceutical Technology Co., Ltd., Chengdu 610094, China) Abstract Abstract Tooth position detection underpins automated analysis of oral images, yet existing methods often fail on low-quality radiographs with blur, overlapping teeth, and low contrast. We propose GQD-DETR, an encoder–decoder framework tailored for robust tooth localization under real-world imaging conditions. GQD-DETR comprises three key components: (1) a Geometry-Aware Attentive High-Resolution module that combines polar coordinate priors with a high-resolution attention pathway to capture dental-arch curvature and fine details; (2) a Quality-Adaptive Cross-Scale Interaction module that performs quality-aware feature filtering and cross-scale context aggregation to handle noise and illumination inhomogeneity; and (3) a Discriminative Anisotropic Decoding module with learnable anisotropic positional encoding and query channel attention, which stabilizes query refinement in crowded regions and improves discrimination between spatially adjacent, visually similar teeth. Experiments on our private dataset and the public DENTEX2023 benchmark show that GQD-DETR achieves 58.98 mAP and 54.67 mAP, outperforming strong baselines by 1.62 and 1.43 mAP, respectively, and effectively handling multi-source images including panoramic radiographs, smartphone photos, and DSLR images. Monday Virtual Room 3 IJCNN Paper SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models III Session Chair: 小雨 刘 (延边大学), yuxin wang (East China Normal University) Class-Conditional Center Alignment for Robust Vision--Language Contrastive Learning in Glaucoma Screening Xiaoyu Liu and Qi Wang (Yanbian University) Abstract Abstract Fine-tuningvision–languagemodels(VLMs)like CLIPformedicalscreeningfacestwocriticalchallenges:noisy supervisionfromweakly-alignedimage–textpairsandperfor-mancedisparitiesacrossdemographicsubgroups.Inglaucoma screening,wherefundusimagesarepairedwithcondensedclini-calnotes,theseissuesmanifestasheavy-tailedcontrastivelossdis-tributionsandunequalaccuracyacrossprotectedattributes.We proposeaunifedframeworkaddressingbothchallengesthrough threesynergisticcomponents:(1)TokenResidualFusion(TRF)strengthensvisualrepresentationsbyblendingintermediateand fnalViTtokenfeatures,(2)Student-tInfuenceReweighting(T-IRLS)stabilizesoptimizationbysuppressingextremeresiduals viarobuststatisticalweighting,and(3)Class-ConditionalCenter Alignment(CC-CA)reducessubgroupdisparitiesthroughen-tropicoptimaltransporttodynamicallyupdatedclasscenters.Ourmethodrequiresnoarchitecturalchangesatinference,addsminimaltrainingoverhead,anddemonstratesconsistent improvementsinbothoverallAUC(upto5.8%gain)andsub-groupfairnessacrossmultipleprotectedattributesinglaucoma screening. GeneMamba: Efficient and Effective Foundation Model on Single Cell Data Cong Qi, Hanzhang Fang, Siqi Jiang, Xun Song, Tianxing Hu, and Zhi Wei (New Jersey Institute of Technology) Abstract Abstract Transformer-based models have achieved strong performance in single-cell transcriptomics, yet scaling them to genome-wide gene sequences remains challenging due to the quadratic cost of self-attention. Recent state space models (SSMs) offer linear-time sequence modeling, but unidirectional designs may introduce ordering biases and overlook biological structure, potentially limiting robustness across datasets. To address these challenges, we propose GeneMamba, a bidirectional state space model tailored for single-cell RNA sequencing. GeneMamba integrates three key components: (1) Efficiency: BiMamba blocks enable bidirectional processing with O(L) complexity, providing substantial speedup over transformer baselines while maintaining linear scaling; (2) Biological integration: a pathway-aware contrastive objective incorporates curated pathway information to guide representation learning; (3) Scalability: large-scale pretraining on approximately 30 million curated cells (from about 50 million raw CELLXGENE profiles) supports transfer across multiple downstream tasks without task-specific architectures. Across benchmarks including multi-batch integration, cell type annotation, and gene rank reconstruction, GeneMamba demonstrates competitive performance relative to strong baselines such as scGPT and Geneformer, with improvements observed in several settings. Ablation studies further suggest that bidirectional modeling and pathway-aware pretraining contribute to these gains. Overall, GeneMamba provides an efficient and biologically informed framework for foundation modeling in single-cell transcriptomics. Low-Resource Diabetic Retinopathy Screening via AUM-ST with Vision Transformer Praneeth Rikka and S. M. Saiful Islam Badhon (University of North Texas); Mohammad Adibuzzaman (Oregon Health & Science University); Abu Saleh Mohammad Mosa (University of Alabama at Birmingham); and Ana D. Cleveland, Junhua Ding, and K. S. M. Tozammel Hossain (University of North Texas) Abstract Abstract Diabetic Retinopathy (DR) is a leading cause of preventable blindness globally, yet early diagnosis faces signif- icant challenges, including a lack of annotated data and limited access to ophthalmic expertise. While machine learning based screening systems (e.g., Convolutional Neural Networks) show promise for DR prediction using retinal image analysis, they typically require a large number of labeled images, an expensive and time-intensive demand that is unsuitable for resource- constrained settings. Semi-supervised learning (SSL) offers an alternative by harnessing abundant unlabeled images alongside limited labeled data. However, existing SSL methods for DR detection struggle with unreliable pseudo-labels. We propose a novel SSL framework that adapts area under margin self-training (AUM-ST), originally developed for natural language processing, to medical imaging for DR detection. Our approach addresses unreliable pseudo-labelling by leveraging training dynamics to assess label quality. We replace the text-based BERT encoder with BEiT (Bidirectional Encoder representation from Image Transformers) to process retinal fundus images while preserving AUM-ST’s core reliability assessment mechanism. Additionally, we implement saliency-preserving augmentation strategies that avoid aggressive transformations such as random rotations and crops, thereby retaining critical lesion features essential for accu- rate diagnosis. Experimental validation on a real-world dataset of 757 retinal fundus images demonstrates that the proposed method outperforms established baselines, including ResNet- 18, ResNet-50, DenseNet-121, MobileNet-v2, and EfficientNet-B0. Our approach achieves substantial improvements in accuracy and F1 scores when training with fewer than 50 labeled samples per class, an advancement over traditional semi-supervised methods. The results demonstrate the practical viability of the proposed method for medical imaging applications and its potential for DR screening in resource-limited environments, including rural healthcare facilities and developing regions. DETR-EHPose: An Edge-Enhanced and Heatmap-Guided Framework for DDH Ultrasound Landmark Detection zhuofan wan and simiao tao (East China Normal University); dandan zhang (International Peace Maternity and Child Health Hospital, School of Medicine, Shanghai Jiao Tong University); yuxin wang and qing zhang (East China Normal University); baoying ye (International Peace Maternity and Child Health Hospital, School of Medicine, Shanghai Jiao Tong University); and jiangtao wang and hailin pan (East China Normal University) Abstract Abstract Early intervention is critical for ultrasound screening of developmental dysplasia of the hip (DDH), yet speckle noise, weak textures, and blurred boundaries often hinder reliable detection and accurate localization of anatomical landmarks. Although structures such as the cartilage–bone interface and the femoral head contour provide strong edge cues, DETR-style detectors typically emphasize region-level semantics and make limited use of explicit boundary and structural priors. Moreover, Transformer decoders usually rely on iterative refinement, which slows convergence and can degrade performance on small or crowded targets.We propose a structure-aware DETR framework tailored for DDH ultrasound that jointly performs object detection and six keypoint localization. An Edge-aware Feature Propagation (EFP) module extracts multi-scale edge representations from shallow features and strengthens boundary sensitivity via convolutional edge fusion. In addition, we introduce a Heatmap-guided Keypoint Prior (HKP) branch and a Heatmap-guided Query Patch Embedding (HQPE). A coarse keypoint heatmap head on P3 converts sparse keypoint supervision into six-channel Gaussian heatmap regression, providing dense spatial constraints and improved robustness; the features are further enhanced through heatmap-aligned referencing. Experiments on our in-house DDH ultrasound dataset (1,255 training images and 300 test images) demonstrate consistent gains over DETR, DeiM, and strong YOLO baselines in terms of mAP@0.75 and PCK, with notably more stable localization under noise and boundary ambiguity. Monday Virtual Room 4 IJCNN Paper SS25 Integrating Large Language Models and Knowledge Graphs I Session Chair: Xinyue Fan (Qilu University of Technology), Fen Zhao (Nanjing Xiaozhuang University) Enhancing Conversational Question Answering through Reinforced Question Reformulation and Prompt Refinement in Children Application Fen Zhao (Nanjing Xiaozhuang University); Xinheng Wang (No); and Haifei Zhang, Jie Zhu, Lingling Zhang, Kexin Liu, Meiqi Shi, Siyi He, and Yi Wang (Nanjing Xiaozhuang University) Abstract Abstract While question answering (QA) technologies have progressed significantly, their deployment for primary-school children remains understudied. Children’s queries often exhibit semantic ambiguity, incomplete expressions, and unclear articulation, leading to unanswered questions or irrelevant responses. Leveraging external knowledge is key to addressing this common challenge of ambiguity. Consequently, effectively reformulating questions and refining prompts into precise SPARQL queries are crucial tasks. This paper presents a multi-task learning approach utilizing a text generation model for both question reformulation and prompt refinement. Additionally, to bridge the preference gap between the question reformulation module and the QA model, we introduce a training strategy that utilizes QA model feedback to further optimize the question reformulation processes. Evaluations across two datasets show our method outperforms baselines, achieving statistically significant gains of +5.96% to +12.70% in F1 and Acc scores. These consistent improvements confirm that unified question reformulation and prompt refinement, reinforced by downstream QA feedback, substantially advance child-oriented conversational QA. Enhancing Knowledge Graph Question Answering through Classified Path Retrieval and Deductive Reasoning Hao Wu, Xiangfeng Luo, Jianqi Gao, and Dian Huang (Shanghai University) Abstract Abstract Knowledge base question answering (KBQA) aims to retrieve precise answers from knowledge graphs using natural language queries. While large language models (LLMs) enhance KBQA with strong reasoning capacities, they still face challenges such as noisy retrieval paths and inherent difficulties in reasoning over graph structures. To address these issues, we propose a novel framework, Path Retrieval and Deductive Reasoning (PRDR), which reformulates path retrieval as a binary classification task to suppress irrelevant paths and facilitate accurate reasoning. Retrieved paths are converted into propositional forms, enabling LLMs to perform structured deductive inference. Our method achieves state-of-the-art results on the WebQuestionsSP and ComplexWebQuestions benchmarks, demonstrating robust performance in complex multi-hop question answering. CTS-CRL: Context Turn Selection for Conversational Rewriting with Joint Learning Yunpeng Zhang, Junyu Li, and Zhe Sun (Institute of Software, Chinese Academy of Science; University of Chinese Academy of Sciences); Yongji Wang (Integration Innovation Center, Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences); and Bei Guan (Institute of Software, Chinese Academy of Science; University of Chinese Academy of Sciences) Abstract Abstract Context Query Rewriting (CQR) is pivotal in conversational search for resolving context dependencies. However,existing methods typically employ a brute-force strategy that feeds the entire dialogue history into Large Language Models (LLMs). This indiscriminate approach introduces severe noise—particularly in multi-turn conversations with frequent topic shifts,which hinders both model robustness and training efficiency. To overcome these limitations, we propose CTS-CRL (Context Turn Selection for Conversational Rewriting with Joint Learning). Unlike previous methods, CTS-CRL incorporates a novel semantic-temporal context selection mechanism that dynamically filters irrelevant turns while preserving essential dependencies. Furthermore, we integrate this mechanism into a unified training framework combining Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). Experimental results on two benchmarks demonstrate the superiority of our approach. Notably, CTS-CRL establishes a new state-of-the-art on the TopiOCQA dataset, significantly outperforming baselines in handling complex topic transitions, while maintaining high efficiency with fewer parameters. Neuro-Symbolic Inductive Reasoning on Temporal Knowledge Graphs via LLM-Enhanced Rule Mining Xinyue Fan, Rui Yu, and Yinglong Wang (Qilu University of Technology) Abstract Abstract Inductive reasoning on Temporal Knowledge Graphs (TKGs) is pivotal for generalizing to entities unseen during training. However, existing symbolic methods heavily rely on historical interactions, rendering them ineffective under the extreme data sparsity of new entities (the "cold-start" problem). To address this, we propose Temporal Hierarchical Adaptive Inductive Logic (T-HAIL). First, we introduce an Open-to-Closed Knowledge Injection (OC-KI) mechanism, which leverages the parametric causal knowledge of Large Language Models (LLMs) to generate candidate rules, compensating for the lack of statistical signals in sparse graphs. To bridge the vocabulary gap, we employ a Semantic Vector Alignment module via Sentence-BERT to map candidates to the KG schema. These rules are refined through a rigorous data-grounding mechanism—validating candidates against observed facts—to eliminate hallucinations. To fully exploit these grounded rules under sparsity, we introduce a Hierarchical Adaptive Inference (HAI) mechanism that dynamically transitions between Exact Rule Matching, Semantic Similarity Borrowing, and Topology-based Fallback. Extensive experiments on ICEWS and GDELT benchmarks demonstrate that T-HAIL significantly outperforms state-of-the-art baselines in mitigating sparsity issues. Our code and datasets are available at: https://anonymous.4open.science/r/T-HAIL-B866. Monday Virtual Room 5 IJCNN Paper SS25 Integrating Large Language Models and Knowledge Graphs II Session Chair: 哲平 于 (天津师范大学), 钰琳 张 (延边大学) ISERA-KGC: Enhancing Knowledge Graph Completion with Interactive Semantic Enhancement and Representation Alignment Zheping Yu, Tongxuan Zhang, and Guiyun Zhang (Tianjin Normal University) Abstract Abstract Knowledge graph completion (KGC) aims to infer missing links by reasoning over triples. Recent work leverages Large Language Models (LLMs) to enhance graph understanding, yet three key challenges remain. First, LLMs often produce ambiguous entity descriptions lacking relational grounding. Second, relation semantics are context-dependent on the conditional entity, a factor frequently overlooked. Third, semantic embeddings from LLMs are misaligned with structural representations, limiting their effectiveness. In this paper, we propose ISERA-KGC (Entity–Relation Interactive Semantic Enhancement and Representation Alignment for Knowledge Graph Completion), a Language Model as a Service framework that leverages frozen LLMs to provide role-sensitive, context-aware semantic descriptions for KGC. Our method introduces two prompting strategies: Role-Aware Relation Prompting (RARP) for generating direction-specific relation semantics, and Relation-Anchored Entity Reasoning (RAER) for grounding ambiguous entities via relational context. Then the Contrastive Fusion and Alignment (CFA) module leverages contrastive learning to align and fuse these semantic and structural embeddings. Experiments on benchmark datasets FB15k-237 and WN18RR show that ISERA-KGC consistently improves mean reciprocal rank (MRR) by up to 2.4 points across multiple KGE models. These results demonstrate the effectiveness of our method in enhancing long-tail reasoning and bridging semantic-structural gaps in knowledge graph completion. DRAoG: A Dynamic Reasoning Adaptive on Graphs framework for Large Language Models Zili Zhou, Nuo Zhuang, and Chengyuan Xue (Qufu Normal University) Abstract Abstract Large Language Models (LLMs) have demonstrated powerful performance in knowledge-intensive reasoning tasks. However, when dealing with large-scale Knowledge Graphs (KGs), they still struggle with hallucinations, inefficient multi-hop reasoning, and sensitivity to noisy retrieval results. Existing graph-based reasoning frameworks typically rely on fixed-hop search strategies, discrete decision boundaries, and the blind linearization of subgraph inputs, which often lead to excessive contextual noise and unnecessary computational overhead. Retrieval, Reasoning, and Generation: A Universal Format Generation Method for Few-Shot Event Argument Extraction Pengfei Yin, Hao Li, Boxiang Hu, and Xixun Lin (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Abstract Abstract Event Argument Extraction (EAE) is essential for identifying event participants and their roles. However, in few-shot settings, existing generative approaches often depend on manually crafted, ontology-specific prompts and formats—an expensive and error-prone process that transfers poorly across domains and schemas. To mitigate these limitations, we propose FGEE, a universal JSON-based generation framework. FGEE tackles the data-scarcity bottleneck through three key innovations: (1) Unified I/O formats, which represent complex event structures with a standardized JSON schema and eliminate per-ontology prompt engineering; (2) Pseudo-output retrieval, which retrieves structurally similar pseudo-outputs to improve task understanding and mine salient event clues; and (3) Key sentence-based reasoning, which performs trigger-guided analysis to focus on informative evidence and deduce arguments from redundant contexts. We further introduce an N-Event-K-Shot evaluation framework to assess EAE under strict few-shot conditions. Experiments on ACE05, RAMS and WikiEvents datasets show that FGEE consistently outperforms the latest few-shot supervised models, yielding about a 9.6\% improvement in F1 across all benchmarks.These results show that FGEE substantially improves the robustness and generalization of generative EAE in low-resource environments. Implicit Relation Extraction with Large Language Models via Implicit Information Augmentation Yulin Zhang and Yahui Zhao (Yanbian University), Yiping Ren (Yanbian Vocational & Technical College), and Guozhe Jin and Rongyi Cui (Yanbian University) Abstract Abstract Relation Extraction (RE) is a core task in information extraction. Although existing approaches have achieved substantial progress in modeling explicit relations, their performance remains limited when confronted with implicit relations that require contextual reasoning or commonsense knowledge, particularly in complex scenarios involving multiple entities and multiple relations. To address this challenge, this paper proposes Implicit Relation Extraction with Large Language Models (IMP-RELLM). The proposed approach explicitly introduces implicit information generated by Large Language Models (LLMs) to enhance the model’s ability to capture deep semantic representations and latent relational patterns in text. Specifically, IMP-RELLM adopts an instruction fine-tuning paradigm that organizes task descriptions, predefined relation sets, implicit information, and the original context into a unified structured prompt, guiding the model to attend to semantically valid entity relations that are not explicitly expressed in the text during generation. To reduce training costs and improve scalability, parameter-efficient fine-tuning is employed to adapt multiple mainstream LLMs. Extensive experiments demonstrate that the proposed method substantially outperforms baseline models without implicit information on implicit relation extraction tasks, achieving consistent improvements in both F1 score and recall. The advantages are particularly pronounced in high relational complexity settings and pure implicit relation scenarios, highlighting the critical role of implicit information augmentation in enhancing the reasoning capability and relational coverage of LLMs. Our code is available at https://github.com/IIP408/IMP-RELLM. Monday Virtual Room 6 IJCNN Paper SS25 Integrating Large Language Models and Knowledge Graphs III Session Chair: Jinxin Liu (Tsinghua University), Zhongtian Bao (Nankai University) REPANA: Reasoning Path Navigated Program Induction for Transferable Reasoning over Heterogeneous Knowledge Bases Jinxin Liu (Tsinghua University); Shulin Cao (Tsinghua University, z.ai); Jiajie Zhang (Tsinghua University); Xiaoyin Che and Weichuan Liu (Siemens AG); Liangming Pan (Peking University); and Lei Hou and Juanzi Li (Tsinghua University) Abstract Abstract Program induction is a typical approach that helps Large Language Models (LLMs) in complex knowledge-intensive question answering over knowledge bases (KBs) to alleviate the hallucination of LLMs. However, accurate program induction requires extensive high-quality parallel data for a specific KB, which is scarce for low-resource KBs. Moreover, the heterogeneity of questions and KB schemas limits the transferability of models trained on a single dataset. To this end, we propose REPANA, a reasoning path navigated program induction framework that enables LLMs to reason over heterogeneous KBs. We decouple the program induction into perceiving the KB and mapping questions to program sketches. Accordingly, our framework consists of (1) an LLM-based navigator, which retrieves reasoning paths of the input question from the given KB; (2) and a KB-agnostic parser trained on multiple heterogeneous datasets, which takes the retrieved paths and the question as input and generates the corresponding program. Experiments show that REPANA exhibits strong generalization and transferability. It can directly perform inference on datasets not seen during training, outperforming other state-of-the-art low-resource methods, even approaching the performance of supervised methods. KG-Hopper: Empowering Compact Open LLMs with Knowledge Graph Reasoning via Reinforcement Learning Shuai Wang and Yinan Yu (Chalmers University of Technology) Abstract Abstract Large Language Models (LLMs) demonstrate impressive natural language capabilities but often struggle with knowledge-intensive reasoning tasks. Knowledge Base Question Answering (KBQA), which leverages structured Knowledge Graphs (KGs) exemplifies this challenge due to the need for accurate multi-hop reasoning. Existing approaches typically perform sequential reasoning steps guided by predefined pipelines, restricting flexibility and causing error cascades due to isolated reasoning at each step. To address these limitations, we propose KG-Hopper, a novel Reinforcement Learning (RL) framework that empowers compact open LLMs with the ability to perform integrated multi-hop KG reasoning within a single inference round. Rather than reasoning step-by-step, we train a Reasoning LLM that embeds the entire KG traversal and decision process into a unified “thinking” stage, enabling global reasoning over cross-step dependencies and dynamic path exploration with backtracking. Experimental results on eight KG reasoning benchmarks show that KG-Hopper, based on a 7B-parameter LLM, consistently outperforms larger multi-step systems (up to 70B) and achieves competitive performance with proprietary models such as GPT-3.5-Turbo and GPT-4o-mini, while remaining compact, open, and data-efficient. Enhancing Factual Consistency in Cross-Lingual Dialogue Summarization via Self-Guidance Prompting Zhongtian Bao, Wenjian Ding, and Yao Zhang (Nankai University); Jun Wang (Ludong University); Zhe Sun (Juntendo University); Andrzej Cichocki (Systems Research Institute of Polish Academy of Sciences, and UMK, Poland); and Zhenglu Yang (Nankai University) Abstract Abstract Recent breakthroughs in large language models have made it feasible to effectively summarize cross-lingual dialogue information, proving essential for the global communication context. However, existing methodologies encounter difficulties maintaining factual consistency across multiple dialogue exchanges and lack clear explanations of the summarization process. This paper presents a novel factually consistent prompting technology with large language models to address these challenges in cross-lingual dialogue summarization. First, we propose a factual replacement mechanism to enhance information analysis by incorporating noise information into summarization candidates. We adopt a self-guidance framework to enforce factual consistency, enhancing information flow tracking in cross-lingual hybrid dialogue scenarios with the assistance of GPT-based models. Furthermore, we introduce a view-aware chain-of-thought-driven architecture to improve the interpretability and transparency of the cross-lingual dialogue summarization process. Comprehensive experimental evaluations on cross-lingual summarization tasks, spanning English, French, Spanish, Russian, Chinese, and Arabic, and hybrid cross-lingual tasks substantiate that the proposed model achieves superior performance relative to state-of-the-art baselines. Hierarchical Code Abstraction via Unified GraphRAG on Multi-Grained Graphs Zhongtian Bao and Wenjian Ding (Nankai University); Jun Wang (Ludong University); Yao Zhang and Zhenglu Yang (Nankai University); Zhe Sun (Juntendo University); and Andrzej Cichocki (Systems Research Institute of Polish Academy of Sciences, and UMK, Poland) Abstract Abstract Code summarization plays a crucial role in helping developers quickly grasp the purpose and behavior of source code. However, the majority of existing approaches concentrate solely on the function level, overlooking the richer contextual signals and hierarchical relationships that naturally exist within large software repositories. To address these shortcomings, we present a new hierarchical code summarization framework built upon GraphRAG. Our key observation is that the community structures revealed by GraphRAG inherently align with the layered organization of real-world software systems—for instance, functions forming classes, and classes composing files. Leveraging this insight, we construct a multi-level code summary graph and dynamically retrieve the most relevant subgraph to supply a large language model with structured and contextually coherent information. This design alleviates the context fragmentation commonly seen in traditional retrieval-augmented generation pipelines. Comprehensive experiments demonstrate that our framework markedly improves summary quality, especially when applied to complex or large-scale codebases. Monday Virtual Room 7 IJCNN Paper SS39 Computational Audio Intelligence for Perception & Representation I Session Chair: Ravindrakumar Purohit (Dhirubhai Ambani University), Li Xiang (DongHua University) BeatMamba: A Drum-Separated Beat Tracking Method Using Mamba Xiaopeng Zhao, Xiang Li, and Xiaoliang He (Donghua University) Abstract Abstract In this paper, we propose BeatMamba, a drum-separated beat tracking framework built on the Mamba state space architecture. Beat tracking in polyphonic music is challenged by complex mixtures and high computational costs. To address these issues, We first apply music source separation to isolate the drum stem, preserving salient percussive cues while suppressing interference from vocals and harmony. This representation improves rhythmic clarity and reduces ambiguity in mix audios. For sequence modeling, we employ selective state space modules to capture long-range temporal dependencies with linear-time and memory complexity, avoiding the quadratic cost of self-attention. Compared to Transformer baseline models, BeatMamba significantly improved computational efficiency and reduced model size, making it more suitable for low-latency and real-time applications. Experimental results on benchmark datasets demonstrate that BeatMamba reduces parameter counts while maintaining or even surpassing the accuracy of heavy models. SpecSlice-ViT: Preserving Spectral Integrity via Slicing Transformers for Underwater Acoustic Target Recognition Ruiting Sun, Zhangjie Cai, Zhenhong Liao, Guanwen Zhang, and Wei Zhou (Northwestern Polytechnical University) Abstract Abstract To address the challenges of non-stationary signals and the disruption of spectral continuity by standard patch-based Transformers in underwater acoustic target recognition, we propose SpecSlice-ViT. This novel architecture preserves spectral integrity via an acoustic-aware Feature Slicing strategy that slices completely along the frequency dimension to maintain harmonic structure. We introduce a dual-view fusion mechanism, combining fine-grained Short-Time Fourier Transform (STFT) details with robust Mel-Frequency Cepstral Coefficients (MFCC) envelopes through cross-attention and adaptive gating. Additionally, a Hierarchical Encoder cascades window attention with token compression to efficiently model multi-scale temporal dependencies. Experimental results on the DeepShip dataset demonstrate that SpecSlice-ViT outperforms existing Convolutional Neural Networks and Transformer baselines, achieving competitive accuracy and efficient inference speeds, indicating potential for future deployment on underwater platforms. RUCL: Integrating Regularization and Unsupervised Contrastive Learning for Automatic Speech Recognition Ze Ting Li, Jiang Ping Zhu, Pei Zhou, and Zheng Zhong Zhu (SiChuanUniversity) Abstract Abstract Automatic Speech Recognition (ASR) aims to tran scribe speech signals into textual sequences, serving as a funda mental technology in human-computer interaction. However, end to-end ASR systems often suffer from inconsistent predictions across different stochastic passes and insufficient representation learning capability. To mitigate these issues, we propose RUCL (Regularization and Unsupervised Contrastive Learning), a novel framework that hierarchically integrates consistency regulariza tion and unsupervised contrastive learning.Specifically, at the feature level, we employ unsupervised contrastive learning to enhance the discriminability of speech representations, thereby strengthening the model’s ability to capture robust acoustic features. At the output level, we apply consistency regularization to minimize the divergence between two CTC distributions generated under different dropout masks, promoting prediction stability. Extensive experiments on AISHELL-1 and LibriSpeech datasets demonstrate that RUCL achieves superior performance, with a CER of 4.63% on AISHELL-1 and WERs of 4.44%/9.96% on LibriSpeech test-clean/test-other, significantly outperforming strong baselines. SSMamba-VC: Learning Linear-Time Sequence Modeling with Selective State Spaces for Voice Conversion. Ravindrakumar Purohit, Kaustubh Wade, and Hemant A. Patil (Dhirubhai Ambani University Gandhinagar) Abstract Abstract Voice Conversion (VC) modifies the speaker identity of speech while preserving its linguistic content. It has transformed the domains of voice dubbing, voice cloning, and assistive technologies by improving the naturalness of the synthesized speech. Recent advances in generative models, particularly diffusion-based frameworks, have significantly improved perceptual quality but come at a high computational cost and face challenges, including degraded prosody and reduced speech intelligibility. These primarily arise from attention-based sequence modeling, which scales quadratically. In this work, we propose SSMamba-VC, a selective state space (SS)-based hierarchical discussion framework that replaces attention mechanism with linear-time sequence modeling. Compared to the existing state-of-the-art VC systems, SSMamba-VC achieves notable improvements across \textit{subjective}, \textit{objective}, and \textit{quantitative} metrics. In particular, the proposed approach (DiffHier-VC + Ours) achieves the NISQA-MOS of 3.90 (+0.02) and a UTMOS of 2.92 (+0.16), outperforming existing diffusion baselines in perceptual stability. The proposed approach reduces $F_{0}$ RMSE to 42.44, improving prosodic accuracy, while also improving intelligibility, with WER reduced to 18.14 and CER to 8.73 compared to DiffHier-VC. Furthermore, with 98\% fewer computational resources, SSMamba-VC achieves a $\sim$15.6\% improvement in real-time factor (RTF), demonstrating that linear-time SS modeling enables faster inference without degrading the quality of the synthesized VC. SSMamba-VC further reduces the real-time factor to achieve high-fidelity VC for practical deployment. The demo utterances are publicly available at: \textit{\url{https://iamshreeji-copy2.github.io/SSMamba-VC/}} Monday Virtual Room 8 IJCNN Paper SS39 Computational Audio Intelligence for Perception & Representation II Session Chair: Junbin Zhang (Peking University), Hui Zhang (Southwest University Of Science And Technology) AudioControl: Efficient Audio-to-Image Generation via Global Latent Modulation and Audio–Visual Alignment Junbin Zhang (Peking University), Tao Fang (Macau Millennium College), and Peng Jin (Peking university) Abstract Abstract Recent advances in diffusion and flow-matching models have enabled high-fidelity text-to-image synthesis. However, conditioning these models on audio remains challenging, as token-level fusion through cross-attention increases the attention length and introduces substantial computational overhead. We present AudioControl, an efficient audio-to-image conditioning module for flow-matching generators that adapts global affine conditioning for audio guidance with minimal overhead. AudioControl extracts a compact global audio embedding with a pre-trained Audio Spectrogram Transformer (AST) and maps it through a lightweight projection to obtain modulation vectors. These vectors are fused additively with pooled text embeddings to form a unified global condition, which is then injected into Transformer blocks via adaptive layer normalization. By modulating feature statistics instead of concatenating audio tokens, AudioControl preserves the original attention sequence length and avoids the quadratic scaling associated with longer contexts, while still steering the full visual latent space. Experiments on VGGSound and GreatestHits demonstrate improved audio-visual alignment with competitive image quality, while substantially reducing the parameter footprint and inference latency relative to token-based conditioning strategies. AnchorAlign++:用于联合抒情音符转录和强健歌唱评估的计数感知和锚驱动的神经框架 魏 胡 and 洪峰 高 (北京化工大学) Abstract Abstract Automatic singing assessment relies on aligning sung notes with a reference score; however, achieving robust alignment remains challenging for beginners due to pitch instability and phrasing errors. We propose AnchorAlign++, an end-to-end framework that explicitly models the number of notes aligned to each lyric token as a structural variable and converts this dis- crete prediction into actionable temporal constraints through an anchor-driven inference strategy. The model consists of a shared acoustic encoder, a non-autoregressive note decoder for frame- level pitch and onset estimation, and an autoregressive lyric decoder that jointly predicts lyric tokens and their corresponding note counts. To further enhance robustness, we introduce a unidi- rectional acoustic–symbol fusion mechanism that injects frame- level pitch and onset cues into lyric decoding, thereby stabilizing long-tail note count prediction. We further propose a count– onset consistency loss, which couples discrete note counts with continuous onset probabilities to reduce boundary drift. During inference, the predicted note counts guide the selection of onset anchors, forming sharp lyric–note boundaries without manual correction. Experiments on multiple singing datasets demonstrate that AnchorAlign++ consistently improves transcription accuracy and alignment robustness, while maintaining practical compu- tational efficiency for deployment, providing reliable automatic singing assessment for both novice and professional singers. LM-AdvVC: Localized Multi-Granularity Targeted Adversarial Attacks on Voice Conversion Xinning Song (Tongji University), Rui Wang (IFLYTEK Research), and Liwei Zheng and Zhihua Wei (Tongji University) Abstract Abstract Despite recent progress in generating natural and personalized speech, voice conversion (VC) systems that aim to modify paralinguistic features of speech while preserving linguistic content remain vulnerable to degraded inputs, where even minor adversarial perturbations cause severe performance drops. Prior studies mainly focused on perturbations that affect speaker identity in classic encoder-decoder models (e.g., AdaIN-VC, AutoVC), while the impact of linguistic-targeted perturbations, especially on modern neural audio codecs, remains unexplored. To address this gap, we propose LM-AdvVC, a two-stage localized adversarial attack framework that applies speaker-guided perturbations directly to content embeddings. Firstly, a critical-segment local ization with phoneme- and word-level analysis adaptively identifies and perturbs vulnerable regions. Secondly, the waveform perturbations are refined using CMA-ES optimization under a teacher forcing cross-entropy objective, yielding transferable adversarial attacks across gray-box scenarios. Experimental results demonstrate that LM-AdvVC successfully manipulates linguistic content under both white and gray-box settings and reveal the content vulnerability of VC systems. The codes are available at https://github.com/windforestfiremountain/LM-AdvVC. Temporal Dynamics Perception for Audio-Visual Question Answering Mingxiang Wen, Yixin Wang, Xujian Zhao, Hongyou Chen, Yin Long, and Bo Li (Southwest University of Science and Technology); Peiquan Jin (University of Science and Technology of China); and Chunming Yang and Hui Zhang (Southwest University of Science and Technology) Abstract Abstract The Audio-Visual Question Answering (AVQA) task aims to mine temporal dynamic information in videos and perform comprehensive spatio-temporal reasoning to answer given questions accurately. Although pre-trained models have achieved significant success in audio-visual question-answering tasks, they still face several challenges. On the one hand, the length of long-term videos makes it difficult for image-based and audio-based pre-trained models to capture temporal dynamic information effectively. On the other hand, the textual semantic gap between the question and the original text of the text-based pre-trained models limits the prediction results. To address the above challenges, we propose a Temporal Dynamics Perception Model (TDPM) to capture temporal dynamic dependencies in videos through a visual temporal alignment module and an audio temporal alignment module. Those temporal alignment modules are constructed via a language-guided auto-regressive task and can combine historical clues of videos and language guidance to predict the future state of videos. Also, TDPM mitigates the semantic gap between the original text of the question and the pre-trained models through a textual semantic alignment module. Experimental results on the MUSIC-AVQA dataset demonstrate the effectiveness of our proposed model. It outperforms all the compared models in terms of average accuracy, with a maximum improvement of 17.02% and an average improvement of 10.86%. Monday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V3 Session Chair: Tam Nguyen (Ho Chi Minh City University of Technology) Integrating Priority Structures into NSGA-II for Interval Data-based Multi-Objective Nonlinear Fixed-Cost Transportation Problem Cristian Pop, Adrian Petrovan, and Petrica Pop Sitar (Technical University of Cluj-Napoca) Abstract Abstract In this paper, we propose a novel approach for solving the interval data-based multi-objective non-linear fixed-cost transportation problem with different means of transport by integrating priority structures into the framework of the NSGA-II. We analyzed different ways of calculating the crowding distance corresponding to a solution of the investigated problem and proposed new genetic operators specifically designed to handle different means of transport. The computational experiments on the existing instances from the literature show that our approach is suitable and efficient for solving the considered multi-objective non-linear fixed-cost transportation problem. Multiform Surrogate-Assisted Evolutionary Algorithm for Expensive Optimization Problems with Mixed Variables Shenghao Wu (South China Agricultural University) and Xiaoming Xue (China University of Petroleum (East China)) Abstract Abstract Expensive optimization problems with mixed contin-uous and categorical variables (EOPMVs) arise in engineering and scientific design, yet remain difficult because evaluations are costly and the two variable types interact in non-smooth ways. To address this challenge, we propose multiform transfer ant colony optimization (MFT-ACO), a surrogate-assisted framework grounded in the multiform transfer optimization paradigm. MFT-ACO reformulates the target problem into three complementary auxiliary formulations by exploiting the distinct modeling strengths of different surrogates for heterogeneous variables: an RBF-based full-space formulation for smooth interpolation, an LSBT-based combinatorial formulation for categorical discontinuities, and a category-conditioned local refinement formulation for continuous improvement in promising regions. These forms co-evolve through an iterative multiform selection mechanism that enables additional surrogate-guided refinement before expensive evaluations are performed. Experimental results on benchmark EOPMVs show that MFT-ACO is competitive under a limited evaluation budget, suggesting that multiform reformulation is a viable strategy for expensive mixed-variable optimization. An LLM-evolved adaptive operator control mechanism for evolutionary Sudoku solving Tam Nguyen (Ho Chi Minh City University of Technology) Abstract Abstract Evolutionary algorithms have demonstrated strong performance on combinatorial problems such as Sudoku; however, their effectiveness heavily depends on handcrafted local search operators and fixed control heuristics. Designing adaptive operator control mechanisms that generalize across problem instances remains a challenging and largely manual process. This paper proposes an LLM-evolved adaptive operator control framework for Sudoku solving, in which a large language model (LLM) is integrated into the evolutionary loop to evolve state-aware local search control heuristics. Rather than directly constructing solutions or performing search, the LLM generates symbolic heuristic policies that regulate swap selection and mutation pressure based on runtime indicators of the search process. These heuristics are encoded as lightweight control rules and embedded into a local search-based genetic algorithm (LSGA). To ensure robustness and mitigate hallucination effects, all LLM-generated heuristics are subjected to an empirical verification procedure across multiple Sudoku instances, and only validated heuristics are retained through an evolutionary selection mechanism. Experimental results on standard Sudoku benchmarks demonstrate that the proposed method achieves competitive performance compared to state-of-the-art approaches and consistently improves convergence speed relative to the original LSGA. Multitask Node Combinatorial Optimization in Complex Networks via Mixture-of-Experts Xiaojie Yang and Zhijie Cao (Shenzhen University, College of Computer Science and Software Engineering); Yeming Yang (Shenzhen University, College of Electronics and Information Engineering); Lingjie Li (Shenzhen Technology University, School of Artificial Intelligence); and Lijia Ma (Shenzhen University, College of Computer Science and Software Engineering) Abstract Abstract The node combinatorial optimization (NCO) problem in complex networks tries to select a set of critical nodes to achieve objectives such as influence maximization, robustness enhancement, and edge coverage, etc. Owing to its significant applications in social network analysis, infrastructure protection, and system design, NCO has attracted extensive research attention. Recently, evolutionary deep reinforcement learning (EDRL) emerges as a promising approach by reformulating discrete NCO into continuous parameter optimization of deep Q networks. However, existing EDRL studies primarily focus on single-task scenarios, while how to efficiently solve multiple NCO problems simultaneously remains a significant challenge. To address this issue, we introduce a multi-gate mixture-of-experts (MMoE)-based multifactorial EDRL algorithm (called MMoE-EDRL) that employs a MMoE-based deep Q network (called MDQN) to model the multitask NCO problem as a weight optimization task. Specifically, the MDQN employs shared expert layers to learn general node representations across tasks and equips each task with an independent gating network to explicitly capture task similarities and differences. Subsequently, this MDQN is evolved by integrating EDRL with the multifactorial evolutionary algorithm (MFEA) for effective multitask learning (MTL). Experimental results across six real-world networks demonstrate the efficiency of MMoE-EDRL over state-of-the-art methods in addressing the NCO problems. This work provides an efficient solution for multitask node combinatorial optimization in complex networks. Monday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V4 Session Chair: Ishara Hewa Pathiranage (Adelaide University) A Feature and Similarity Based Grouping Algorithm for Efficient Time Series Storage Yao Wang, Jiong Zhao, Jinhai Pan, Zhenkuo Kang, and Qicai Zhou (Tongji University) Abstract Abstract For time series databases, single-column storage and single-column-group storage are common storage schemes. The former stores each time series independently but results in redundant storage of timestamps. The latter stores a batch of time series by sharing a time column but requires extra disk space for recording null values. This paper aims to design an algorithm that finds a compromise storage scheme between the two. We divide a batch of time series into multiple column-groups to minimize the overall space cost. We first model the time series grouping problem as an optimal set partitioning problem. Subsequently, we propose a two-phase algorithm to solve this problem. In Phase 1, we propose a similarity metric based on the time overlap degree between time series, and we employ a clustering algorithm to generate initial solutions. In Phase 2, based on the genetic algorithm (GA), we propose genetic operators derived from the features of time series to improve the quality of solutions obtained in Phase 1. Finally, we validate the effectiveness of the proposed algorithm through experiments. On the Use of Evolutionary Optimization for the Dynamic Chance Constrained Open-Pit Mine Scheduling Problem Ishara Hewa Pathiranage and Aneta Neumann (Adelaide University) Abstract Abstract Open-pit mine scheduling is a complex real-world optimization problem that involves uncertain economic values and dynamically changing resource capacities. Evolutionary algorithms are particularly effective, as they can easily adapt to such environments. However, uncertainty and dynamic changes are often studied in isolation in real-world problems. In this paper, we study a dynamic chance-constrained open-pit mine scheduling problem where block economic values are stochastic and mining and processing capacities vary over time. We adopt a bi-objective evolutionary formulation that simultaneously maximizes expected discounted profit and minimizes its standard deviation. To address dynamic changes, we propose a diversity-based change response mechanism that repairs a subset of infeasible solutions and introduces additional feasible solutions whenever a change is detected. We evaluate the effectiveness of this mechanism across four multi-objective evolutionary algorithms and compare it with a baseline re-evaluation–based change-response strategy. Experimental results demonstrate that the proposed approach consistently outperforms the baseline methods across different uncertainty levels and change frequencies. Closed-Loop Success-History Driven Reconstructed Differential Evolution for Constrained Multi-Objective Optimization Chennuo Jin and Bingchao Shi (Hirosaki University, Faculty of Science and Technology); Sichen Tao (Tohoku University, Cyberscience Center); and Yifei Yang (Hirosaki University, Faculty of Science and Technology) Abstract Abstract Constrained multi-objective optimization (CMOP) is challenging because the search must balance feasibility maintenance, convergence toward the constrained Pareto-optimal set, and diversity preservation under complex constraints and limited evaluation budgets. This paper proposes Closed-Loop Success-History Driven Reconstructed Differential Evolution(CLSHRDE), which is built upon the Reconstructed Differential Evolution (RDE) baseline under the Institute of Electrical and Electronics Engineers (IEEE) Congress on Evolutionary Computation (CEC) 2025 CMOP benchmark setting, where RDE has demonstrated strong overall performance. On this basis, CLSHRDE introduces an explicit two-operator search structure by integrating an extra stochastic mutation rule with the original order-based mutation. To enable robust stage wise resource allocation between the two operators, CLSHRDE develops a closed-loop scheduling mechanism driven by an exponentially weighted success-history estimate, providing noise attenuated feedback for online operator regulation across search stages. Extensive experiments on the Scalable Decision space Constraints (SDC) benchmark suite and real-world constrained multi-objective problems (RWMOPs) show that, under unified parameter settings, CLSHRDE achieves competitive or improved performance in terms of Inverted Generational Distance (IGD), Hypervolume (HV), and feasible success rate (FSR), with improvements statistically supported by one-sided Wilcoxon rank sum tests. Monday Virtual Room 1 IJCNN Paper SS12 XSTASys: Explainability and Security in Trustworthy Artificial Intelligence Systems I Session Chair: Dakai Zhai (Tsinghua University), Xi Zhong (Institute of Computing Technology,Chinese Academy of Sciences; University of Chinese Academy of Sciences) DCAdapt: A Dual-Cognition Driven Framework for Adaptive Hate Speech Detection Xi Zhong, Cheng Ouyang, Yan Guo, Yuanhai Xue, and Xinran Liu (Institute of Computing Technology,Chinese Academy of Sciences) Abstract Abstract Hate speech detection faces significant challenges due to inconsistent criteria across datasets, rooted in inherent subjectivity and annotation bias. This paper argues that the core of this task lies in deep adaptation to specific annotator stances within particular scenarios rather than merely fitting data distributions across dataset. Inspired by the Dual-Process Theory in human cognitive science, we propose DCAdapt, a dual-cognition driven framework for adaptive hate speech detection that achieves high adaptability by simulating the synergy between intuitive association and logical calibration. The framework decomposes the detection task into two collaborative stage: decoupled dual-knowledge base construction and dual-pathway cognitive reasoning. In the construction stage, a mutual-indexing mechanism decouples objective knowledge from subjective precedents to achieve synergistic modeling of general knowledge and annotation criteria. In the reasoning stage, the framework emulates the decision sequence from affective association to rational judgment: an implicit pathway leverages a general knowledge base to uncover latent contextual clues, while an explicit pathway employs a specialized knowledge base to enforce semantic constraints and stance alignment based on precedents for informed adjudication. This dual-pathway coordination enables DCAdapt to guide Large Language Models (LLMs) to precisely identify complex decision boundaries and align with target-scenario annotation stances. Experimental results demonstrate that DCAdapt exhibits superior effectiveness, architectural flexibility, and cross-scenario transferability, particularly in detecting complex and implicit hate speech. Explainable Recommendation with Topic-enhanced Personalized Prompt Learning Rui Tang (Communication University of China) Abstract Abstract Explainable recommendation systems aim to provide both recommendation results and corresponding explanations to enhance user trust and engagement. Existing post-hoc methods leverage pre-trained language models to generate textual explanations. However, these works hardly consider the latent topical structures in user reviews, resulting in generic and repetitive explanations with limited personalization. To address this issue, this paper introduces TPLER (Topic-enhanced Prompt Learning for Explainable Recommendation), a novel model that integrates topical preferences into prompt learning to generate more personalized and higher-quality explanations. Our approach employs a topic-enhanced feature extraction strategy based on Latent Dirichlet Allocation (LDA) to explicitly model user and item topical preferences from review texts. These topic-aware features are fed into a pre-trained transformer model via continuous and discrete prompts, guiding the generation process. Continuous prompts encode rating information to capture personalized user-item interactions, while discrete prompts are obtained from topic-enhanced features. A two-stage multi-task training strategy is proposed to jointly optimize the rating prediction and explanation generation tasks. Experiments on three benchmark datasets show that TPLER achieves an average improvement of 7.93% across text quality metrics (BLEU-1/4, ROUGE-1/2/L) and 21.63% for personalization metrics, verifying the effectiveness of our proposed TPLER model. Architecture-Level Backdoors in Malware Detection Models for Trustworthy Artificial Intelligence Ziyi Yu, Xiaobing Xiong, Ju Yang, and Fei Kang (Key Laboratory of Cyberspace Security, Ministry of Education, China; Information Engineering University, Zhengzhou 450001, China) Abstract Abstract Deep learning--based static malware detection models face severe supply chain security threats. Existing defense strategies against backdoored static malware detection models are largely built upon a core assumption: that fine-tuning contaminated models or retraining them on clean data can effectively remove backdoors. However, our study reveals critical limitations of this widely trusted security assumption. To expose this defense blind spot, we are the first to propose a Model Architecture Backdoor attack in the domain of static malware detection.Considering the strict format and semantic constraints of binary files, we design a universal architecture-level backdoor scheme based on a specific trigger mechanism. Unlike traditional data poisoning or weight manipulation attacks, the proposed attack introduces malicious structural components that are independent of model weights, enabling the backdoor to survive and persist even after full retraining in a clean environment. Extensive experiments conducted on the SOREL-20M and MOFIT benchmarks demonstrate that mainstream architectures, including CNNs, LSTMs, and Transformers, remain vulnerable to this architecture-level attack despite rigorous retraining procedures. Our findings expose the inherent limitations of defenses that rely solely on weight verification and highlight the necessity of incorporating architecture auditing and structure-aware defenses into the model supply chain to enhance the security and trustworthiness of AI systems in real-world deployment scenarios. Robust Multi-Bit Watermarking for LLM-Generated Text via Semantic-Aware Bucket Encoding Dakai Zhai and Kun Hu (Tsinghua University), Hong Zou (South China University of Technology), Yangdongjie Pan (Nanjing Forestry University), and Qianhui Zhu and Xingjun Wang (Tsinghua University) Abstract Abstract The rapid development of Large Language Models (LLMs) has raised concerns about the authenticity of LLM-generated text, and existing multi-bit watermarking schemes have low robustness to cropping, limited capacity, and high false positives. We propose a robust framework based on clustered multi-bucket coding. First, for token embedding, we use frequency-balanced adaptive k-means for unsupervised clustering, which achieves the goal of improving the robustness of the framework. Then, in each cluster, we use multiple buckets to subdivide again to form a combination of red and green buckets, which in turn improves the embedding capacity. Finally, we employ two-stage detection with statistical validation and combinatorial decoding to achieve a low false positive rate. Experiments on GPT-2 and OPT-1.3B show that our framework outperforms state-of-the-art methods in terms of robustness ($92.65\%$ accuracy after $50\%$ cropping), capacity ($96.33\%$ at 20 bits), and false positives ($\le 10^{-6}$), providing an effective solution for traceable LLM-generated text. Monday Virtual Room 2 IJCNN Paper SS12 XSTASys: Explainability and Security in Trustworthy Artificial Intelligence Systems II Session Chair: Guangyu Gong (Shandong University), Song Xu (University of Science and Technology of China) PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification Guangyu Gong (School of Cyber Science and Technology, Shandong University) and Zizhuang Deng (School of Cyber Science and Technology, Shandong University; Suzhou Research Institute of Shandong University) Abstract Abstract Large Language Model (LLM) agents are increasingly integrated into critical systems, leveraging external tools to interact with the real world. However, this capability exposes them to Indirect Prompt Injection (IPI), where attackers embed malicious instructions into retrieved content to manipulate the agent into executing unauthorized or unintended actions. Existing defenses predominantly focus on the pre-processing stage, neglecting the monitoring of the model's actual behavior. In this paper, we propose PlanGuard, a training-free defense framework based on the principle of Context Isolation. Unlike prior methods, PlanGuard introduces an isolated Planner that generates a reference set of valid actions derived solely from user instructions. In addition, we design a Hierarchical Verification Mechanism that first enforces strict hard constraints to block unauthorized tool invocations, and subsequently employs an Intent Verifier to validate whether parameter deviations are benign formatting variances or malicious hijacking. Experiments on the InjecAgent benchmark demonstrate that PlanGuard effectively neutralizes these attacks, reducing the Attack Success Rate (ASR) from 72.8% to 0%, while maintaining an acceptable False Positive Rate of 1.49%. Furthermore, our method is model-agnostic and highly compatible. Cross-Domain Vulnerability Detection using LLMs: Knowledge Transfer from Contemporary Software to AI/ML-Specific Software Miltiadis Siavvas (Centre for Research and Technology Hellas, Aristotle University of Thessaloniki) and Ilias Kalouptsoglou, Dionysios Kehagias, and Dimitrios Tzovaras (Centre for Research and Technology Hellas) Abstract Abstract The increasing integration of artificial intelligence (AI), and particularly machine learning (ML), into software-intensive systems raises security concerns for AI/ML-specific code. Although AI-based vulnerability detection models (VDMs) achieve accurate prediction on contemporary software, their cross-domain effectiveness on AI/ML-specific software remains underexplored due to limited labeled data. To fill this gap, in this paper we aim to examine whether VDMs trained on contemporary software can accurately detect vulnerabilities in AI/ML-specific software as well, and to assess how domain adaptation can affect their performance. We construct a dataset of vulnerable and non-vulnerable Python functions split into contemporary and AI/ML-specific subsets, fine-tune CodeBERT and CodeGPT on the contemporary subset, and evaluate them on the AI/ML subset. We then perform gradual adaptation by adding 5%, 10%, 15%, 20%, 25%, 40%, and 50% target- domain samples. We also analyze representation drift with UMAP and attention drift via self-attention-based explainability. Results suggest that AI-based VDMs built on contemporary software cannot be used effectively on AI/ML-specific software without proper domain adaptation, indicating that AI/ML-specific soft- ware exhibits unique characteristics; however, even a small- moderate amount of domain-specific data can enable satisfactory cross-domain detection accuracy SpAdvGAN: Learning Imperceptible Adversarial Perturbations via Sparse Semantic Attention Jiajie Xing (Inner Mongolia University) and Delong Yang (Baotou Vocational & Technical College。) Abstract Abstract Adversarial attacks pose serious security risks to deep neural networks (DNNs) by inducing erroneous predictions. While gradient-based and GAN-based attack methods have achieved high success rates, they often produce perturbations that are diffusely distributed across images, resulting in poor imperceptibility. This limitation restricts their applicability in scenarios such as privacy protection for face recognition systems. To address this issue, we propose SpAdvGAN, a semantic-aware adversarial example generation framework that integrates sparse attention with generative adversarial networks. SpAdvGAN introduces a sparse semantic attention extractor that explicitly constrains perturbations to a limited set of semantically important regions, preventing unnecessary noise diffusion and preserving visual fidelity. This design enables precise regulation of perturbation granularity while maintaining strong attack effectiveness. Extensive experiments on CIFAR-10 and ImageNet datasets demonstrate that SpAdvGAN substantially improves the imperceptibility of adversarial examples, while achieving comparable or higher attack success rates than state-of-the-art gradient-based attacks and the GAN-based baseline LpAdvGAN. RationAnomaly: Log Anomaly Detection with Rationality via Chain-of-Thought and Reinforcement Learning Song Xu (University of Science and Technology of China); Yilun Liu, Minggui He, Mingchen Dai, Ziang Chen, Chunguang Zhao, Jingzhou Du, Shimin Tao, Weibin Meng, and Mengyao Piao (Huawei); Shenglin Zhang and Yongqian Sun (Nankai University); and Boxing Chen and Daimeng Wei (Huawei) Abstract Abstract Logs serve as the foundational audit trails for monitoring the integrity and security of modern software systems. Automated log anomaly detection is crucial for ensuring the reliability of modern software infrastructures. However, existing approaches encounter significant limitations. Traditional deep learning models frequently lack interpretability and struggle with generalization, whereas Large Language Models are prone to hallucinations, introducing new security risks through factual fabrication. To address these challenges, we propose RationAnomaly, a novel framework designed to enhance log anomaly detection by synergizing Chain-of-Thought fine-tuning with reinforcement learning. Our approach initially instills expert-level reasoning patterns using supervised fine-tuning guided by Chain-of-Thought data, grounded in a high-quality dataset corrected through a rigorous expert-driven review process. Subsequently, we introduce a reinforcement learning phase employing a multi-faceted reward function to optimize the model. This function simultaneously evaluates format adherence, answer accuracy, and logical consistency, thereby effectively mitigating hallucinations. Experimental results demonstrate that RationAnomaly outperforms state-of-the-art baselines by achieving superior F1-scores on key benchmarks while providing transparent and step-by-step analytical outputs. Monday Virtual Room 3 IJCNN Paper SS29 Graph-Based Solutions for Explainable and Efficient AI I Session Chair: Enguang Zuo (Tsinghua University, School of Intelligence Science and Technology), yanqin luo (yunnan university) MRV-GCN: Multi-Relational Semantic Graph Learning with Explainable Subgraph Fusion for Smart Contract Vulnerability Detection Yanqin Luo and Gehao Lu (Yunnan University) Abstract Abstract Smart contracts are immutable once deployed, making vulnerabilities prone to severe and irreversible losses. Existing deep-learning--based detectors often operate as black boxes and fail to capture vulnerability-relevant semantics underlying real attack mechanisms. We propose MRV-GCN (Multi-Relational Vulnerability-aware Graph Convolutional Network), a graph-based framework that integrates explicit expert knowledge into deep representation learning for smart contracts. MRV-GCN constructs a multi-relational semantic graph (MRSG) that explicitly models AST structures together with control- and data-flow dependencies, and introduces hierarchical vulnerability semantic features (HVSF) to inject expert-defined cues---such as external calls, state updates, unsafe arithmetic, and delegatecall chains---directly into node representations. An interpretable hybrid readout (R-Hybrid) module further fuses semantic subgraph evidence with non-semantic context to produce robust predictions. Experiments on real-world datasets demonstrate that MRV-GCN achieves strong performance across multiple vulnerability types, while the learned fusion coefficient α provides a lightweight semantics-aware signal for interpretation by quantifying the influence of expert-defined semantic evidence in each prediction. MEHGB: Meta-Path Pruned and Explainable Heterogeneous Graph Backdoor for Stealthy NIDS Evasion Zhonghang Sui, Fei Kang, Yuntian Zhao, and Yan Guang (Key Laboratory of Cyberspace Security, Ministry of Education, China; Information Engineering University, Zhengzhou 450001, China) Abstract Abstract Heterogeneous Graph Neural Networks (HGNNs) have become essential for encrypted traffic analysis, yet existing backdoor attacks fail to exploit their vulnerabilities due to a fundamental mismatch with homogeneous assumptions. Current methods rely on additive assumptions that fail to capture cooperative interactions for precise localization, suffer from signal dilution caused by massive structural redundancy, and generate semantically invalid triggers that violate rigid network protocols. SEGCN-AIKC: Trustworthy Herbal Prescription Recommendation via Knowledge-Guided Verification and Correction Quan Gan, Jiaqing Shang, and Lulu Wei (Jiangsu Ocean University); Chen Li (Zhejiang University School of Medicine); and Chuanxia Liu and Li Jia (Jiangsu Ocean University) Abstract Abstract Traditional Chinese Medicine (TCM) herbal prescription recommendation involves selecting appropriate herb combinations from patient symptoms under symbolic constraints derived from TCM theory. While most data-driven approaches achieve strong predictive accuracy, they typically encode domain knowledge only implicitly and lack explicit mechanisms to verify and correct knowledge-inconsistent predictions, which limits their reliability for clinical decision support. In this paper, we propose SEGCN-AIKC, a Semantic-Enhanced Graph Convolutional Network with Adaptive Inconsistency-aware Knowledge Correction, which formulates TCM herbal recommendation as a constrained hypothesis inference problem. Specifically, a semantic-enhanced multi-view graph convolutional network captures heterogeneous symptom–herb relations and constructs a syndrome-level hypothesis, which is then verified and minimally corrected by AIKC under symbolic TCM constraints. By integrating neural perception with constraint-guided verification and repair, SEGCN-AIKC bridges data-driven learning with symbolic reasoning, yielding recommendations that are both accurate and theoretically consistent. Experiments on a public TCM prescription dataset demonstrate consistent improvements over baselines, while ablation and case studies further validate interpretability and clinical plausibility. Adaptive Context Evolution for Zero-Fine-Tuning Generalist Graph Anomaly Detection Haoran Dan (Xinjiang University, School of Computer Science and Technology); Chen Chen and Ruishuang Sun (Xinjiang University, School of Software); Yinhong Li (Xinjiang University, School of Computer Science and Technology); Wei He (Xinjiang Jiuding Testing Technology Co.LTD); Ziwei Yan and Xiaoyi Lv (Xinjiang University, School of Software); ChangYong Liu (Xinjiang Academy of Agricultural and Reclamation Sciences); and Enguang Zuo (Xinjiang University, School of Intelligence Science and Technology) Abstract Abstract Graph Anomaly Detection (GAD) aims to identify anomalous nodes whose behavioral patterns or structural relationships deviate from the majority, a task that has attracted significant research attention. However, constructing a universal GAD framework capable of adapting to diverse real-world applications faces the following technical bottlenecks: on one hand, existing models struggle to achieve a balance between capturing complex long-range dependencies and maintaining high computational efficiency; on the other hand, mainstream methods are constrained by "generalization silos," lacking the cross-domain capability to handle data distribution shifts. To alleviate these issues, we propose AceGAD, a universal graph anomaly detection method. Specifically, AceGAD consists of two core components: first, to address the dilemma of reconciling long-range dependencies with computational efficiency, we designed the Linear Graph State-Space Encoder (LG-SSE), which reformulates graph propagation as a continuous linear state evolution process, achieving precise memory of global context while maintaining linear time complexity; second, to break through the limitations of generalization silos, we proposed the Context-Aware Metric Projector (CAM-Pro), which enables the model to adapt to new environments using only a few normal samples by constructing an adaptive metric space, achieving Zero-Fine-Tuning inference on new graphs. Extensive experiments conducted on eight public benchmark datasets across different domains demonstrate that AceGAD achieves superior performance, efficiency, and universality. Monday Virtual Room 4 IJCNN Paper SS29 Graph-Based Solutions for Explainable and Efficient AI II Session Chair: Nannan Hu (Shandong Normal University), YUNQI HAN (Universiti Putra Malaysia) BDSAM: Fusing Graph Attention Networks and Adapters to Address Domain Imbalance in Cross-Domain Segmentation Yuefeng Zhao, Siting Zhou, Pengfei Sun, Junjie Wang, Haoyu Wang, and Nannan Hu (Shandong Normal University) Abstract Abstract Cross-domain image segmentation plays a crucial role in the field of image segmentation, as it helps to improve the generalization ability of models across different datasets. However, existing models overemphasize domain-invariant features that are not specific to a particular domain, such as cells in medical images or buildings in remote sensing images. Consequently, they neglect the importance of domain-specific features, such as object contours or boundaries.Therefore, we propose Balanced Domain SAM(BDSAM), a model designed to balance domain-specific features while enhancing domain-invariant features.Specifically, we introduce an Iterative Optimization Selection Module (IOSM). By leveraging query features to continuously filter for more relevant pixels within the support features, we automatically select these pixels as point prompts. Furthermore, we propose a Domain-specific Feature Fusion Module (DFFM). By extracting and fusing features from different levels, we enhance the focus on domain-specific features.Extensive experiments on four cross-domain datasets show that our model has better average accuracy on 1-Shot and 5-Shot settings than the state-of-the-art approaches respectively . H$^2$-GCN: Towards Complex Higher-order Interactions for Session-based Recommendation Qi Zhang, Wei Zhou, and Huayi Shen (Chongqing University) Abstract Abstract Session-based recommendation (SBR) aims to predict the next item based on anonymous user sessions. However, existing studies still face challenges below: \textbf{(1)} Inadequate capture of both pairwise and high-order relationships in graph relations. \textbf{(2)} Spurious edges in the global graph lead to redundant global relations extraction. \textbf{(3)} Complex coupling of different user intents caused by noise and dynamics within sessions. To address these issues, we propose a novel \underline{H}igher-order \underline{H}ybrid Graph Convolutional Network (H$^2$-GCN) for SBR, which effectively captures multi-level relationships and refines diverse user intents, enabling effective recommendations. Specifically, we first present a higher-order graph based on simplicial complexes (SCs) to learn the multilevel complex dependencies of items and sessions separately. Moreover, two intent expert modules are designed to independently learn user intents within sessions: the short-term intent expert module is proposed to block long-term dependencies and activate true short-term intents from user noise behaviors, while the long-term intent expert module employs a dynamic-static graph collaborative propagation to filter redundant global relations and derive true long-term intents. Eventually, we guide the intents obtained from the two expert modules to form the final user intents. Extensive experiments show that H$^2$-GCN outperforms state-of-the-art methods in three real-world datasets. Attribution-Aware Log Anomaly Detection via Index-Preserving Multi-View Fusion Yuan Tian and Shi Ying (Wuhan University) Abstract Abstract Log anomaly detection is essential to the reliability of large-scale software systems. In real-world deployment, anomaly alarms should not only be accurate but also reveal when anomalies occur, where they arise, and which components and events are most responsible. However, most existing methods model logs from a single view and lose fine-grained traceability during representation learning, which limits their value for actionable diagnosis. To address this issue, we propose IP-Log, an attribution-aware framework for log anomaly detection based on index-preserving multi-view fusion. IP-Log models system behavior from three complementary views, namely component, event, and time, and structures their interactions according to two principles. One is level-confined information flow, which prevents premature entanglement across views. The other is index-preserving computation, which maintains the correspondence between latent representations and original log entities throughout the model. As a result, anomaly evidence can be decomposed into fine-grained contributions over components, events, and time segments, and the model can further identify top-ranked evidence triplets for diagnosis and troubleshooting. Experiments on public log benchmarks show that IP-Log achieves strong anomaly detection performance while also providing interpretable and index-aligned evidence attribution. PIRGL: Prior and Implicit Reliability Graph Learning for Social Recommendation Yunqi Han, Jidong Wang, and Yunjie Guo (Universiti Putra Malaysia); Yifan Chen and Zhiyuan Chen (University of Nottingham Malaysia); and Dario Landa-Silva (University of Nottingham) Abstract Abstract Social recommendation plays a critical role in modern recommender systems. Graph-based methods effectively model relational signals to improve performance. However, real-world interaction data contains abundant noise, while social links remain typically sparse and unreliable. These limitations significantly hinder the robustness of current models. Existing research rarely explores implicit social relations induced by user preferences. To address these challenges, Prior and Implicit Reliability Graph Learning is proposed. This unified framework explicitly models reliability across interaction, relation, and fusion levels. Specifically, the model integrates a self-paced interaction denoising strategy with a hybrid social graph containing preference-based neighbors. A reliability-aware fusion gate further balances these heterogeneous user representations. Experiments on three real-world benchmarks demonstrate consistent improvements in Recall and NDCG. These findings establish reliability-oriented graph learning as a principled direction for advancing robust social recommendation. Monday Virtual Room 5 IJCNN Paper SS32 Deep Neural Networks and Generative AI for Multi-Agent Smart Vehicle Perceptron, Learning, Automation and Optimization I Session Chair: Haoqian Song (Institute of Automation, Chinese Academy of Sciences, Beijing, China; Pengcheng Laboratory, Shenzhen, China), Ya Zhang (Southeast University) SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving Kangyu Wu, Peng Cui, Guoxi Chen, and Ya Zhang (Southeast University) Abstract Abstract Ensuring both safety and efficiency in decision-making of autonomous driving systems remains a fundamental challenge. Traditional Deep Reinforcement Learning (DRL) suffers from unsafe random exploration and slow convergence, while Large Language Models (LLMs) demonstrate inherent latency in real-time inference operations. To address these limitations, this paper proposes SARAD, a novel safety-aware hybrid framework that synergizes LLMs and DRL for autonomous driving. SARAD substitutes DRL's random exploration with Retrieval-Augmented Generation (RAG)-enhanced, LLM-guided decisions sourced from a dynamic expert knowledge repository. An attention discriminator that integrates LLM’s prior knowledge into DRL policy optimization is proposed. A collision predictor module, which is fine-tuned with historical collision data, is further designed to guarantee the safety of the vehicle. Extensive experiments show that SARAD achieves significant performance improvements in Highway-Env simulator, validating the effectiveness of the model in autonomous driving. DW-YOLO: A Robust Object Detection Framework for Adverse Weather and Domain Shifts Shuo Cai, Shuai Chen, Yuanzhi Tang, Wei Shuai, Zeyang Deng, and Hao Xiong (Changsha University of Science and Technology) Abstract Abstract Under adverse environmental conditions such as fog, rain, snow, and low illumination, object detection models often suffer significant performance degradation in real-world deployments due to domain shifts and visual feature degradation. To address this problem, this paper adopts YOLOv8 as the baseline detector and proposes a robust framework, termed DWYOLO, which is designed to learn domain-invariant object representations in complex environments.At the core of the framework is the proposed Dual-Context Attention Module (DCAM). This module integrates short-distance attention and long-distance attention with multi-scale contextual fusion, enabling simultaneous modeling of global contextual relationships and local fine-grained details. By explicitly leveraging complementary contextual information, DCAM substantially improves the robustness of the characteristics against weather-induced domain drift. Furthermore, an adaptive optimization strategy, Weighted Intersection over Union (WIoU), is incorporated to emphasize challenging and underrepresented samples during training. This strategy improves generalization across diverse environmental conditions while preserving an end-to-end training pipeline and without incurring significant computational overhead. Extensive experiments on multiple public benchmark datasets demonstrate that DW-YOLO consistently outperforms baseline detectors under adverse weather conditions, yielding notable improvements in both detection accuracy and robustness. These findings indicate that DW-YOLO provides an effective solution to object detection under domain shift and shows strong potential for deployment in real-world intelligent perception systems. A Point Cloud Completion Network Via The Latent Space-driven Two-stage Noise Synthesis And Restoration Strategy Xiaofei Qin, Anluo Yi, Shiwei Tao, Jie Zhang, and Xuedian Zhang (China/University of Shanghai for Science and Technology) Abstract Abstract Raw point cloud data collected in real-world scenarios often encounter complex interferences such as sensor noise and uneven density, making high-fidelity point cloud completion tasks extremely challenging. Current mainstream point cloud completion methods generally suffer from insufficient noise resistance, struggling to meet practical application requirements. Moreover, publicly available datasets for noisy point cloud completion are relatively scarce. We proposes a point cloud completion network via the latent space-driven two-stage noise synthesis and restoration strategy, constructing a novel unified framework. Within this framework, noise synthesis and restoration are achieved through gradient guidance and constrained projection mechanisms under the constraint of hyperspherical distribution in the feature-level latent space. Experimental results on public datasets demonstrate that our method significantly improves both noise robustness and completion accuracy compared to the current state-of-the-art (SOTA) techniques, offering a new solution for noisy point cloud completion tasks. A Unified Framework for Occlusion Reasoning and De-occlusion in Open Scene Understanding Haoqian Song (Institute of Automation, Chinese Academy of Sciences, Beijing, China; Pengcheng Laboratory, Shenzhen, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China); Long Cheng (Institute of Automation, Chinese Academy of Sciences, Beijing, China; School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China); and Xuchen Liu, Junjie Wen, and Jinqiang Cui (Pengcheng Laboratory, Shenzhen, China) Abstract Abstract The semantic understanding of occlusion relationships is essential for applications such as autonomous driving, robot navigation, and object grasping, yet existing scene understanding methods often lack this capability. Recognizing this limitation, this paper proposes a unified framework of occlusion reasoning and de-occlusion for open scene understanding. This framework consists of three modules: 1) Visible-region-based scene understanding: the framework leverages RAM-BLIP-Grounded-SAM to automatically generate segmentation masks and semantic information for the visible regions of objects. 2) Occlusion reasoning: an improved partial completion network (PCNet) with an occlusion completion prior module is proposed to ensure robust zero-shot generalization in open scenes through a self-supervised manner, and achieves occlusion ordering with semantic information. 3) Automatic de-occlusion: analyzing occlusion ordering results, the framework enables automatically amodal segmentation and completion based on pix2gestalt, eliminating the need for manual prompts. The results on COCOA dataset show that improved PCNet achieves an occlusion ordering accuracy of 91.534%, representing a 4.357% improvement. Moreover, the unified framework for occlusion reasoning and de-occlusion demonstrates outstanding performance in open scenes. Monday Virtual Room 6 IJCNN Paper SS32 Deep Neural Networks and Generative AI for Multi-Agent Smart Vehicle Perceptron, Learning, Automation and Optimization II Session Chair: Huangnan Zheng (Zhejiang University), Xianchang Wang (Shenyang Aerospace University) AeroGraph: Window-Level Large-Graph Association and Uncertainty-Guided Fusion for Multi-Object Tracking in UAV Videos Xianchang Wang and Qi Li (Shenyang Aerospace University) Abstract Abstract variations, and non-stationary motion caused by platform jitter and detection noise. Most online tracking-by-detection pipelines mainly rely on frame-wise matching with limited temporal evidence, where unreliable motion priors and noisy observations can easily trigger identity drift and trajectory fragmentation. To tackle these issues, we propose AeroGraph, a window-level largegraph association framework for online UAV MOT. Specifically, we formulate data association as inference on a unified largegraph within a sliding window, which jointly models intra-frame structural context and inter-frame correspondences to aggregate multi-frame cues. Moreover, an uncertainty-guided motion fusion strategy is designed to adaptively balance motion consistency and graph affinity according to motion reliability, preventing error propagation under jitter and detection noise. In addition, a trajectory memory with window-level consistency learning is introduced to stabilize long-term identity representations and mitigate feature drift during online updates. Extensive experiments on representative UAV tracking benchmarks demonstrate that AeroGraph consistently improves association robustness and achieves competitive performance under challenging aerial conditions. BC-Attack: Smooth Adversarial Trajectory Attack Based on Bézier Curves Shuai Li, Libing Wu, Lijuan Huo, Zhuangzhuang Zhang, and Yi Yu (Wuhan University) Abstract Abstract Trajectory prediction is a critical component for autonomous vehicles (AVs) to execute safe planning. While extensive work has been done on adversarial attacks against deep learning-based trajectory prediction models, most approaches generate adversarial trajectories by adding noise to trajectory points. This approach often results in unsmooth, or even jagged, trajectories, which suffer from poor stealthiness and are impractical for real-world implementation. To address this limitation, we propose a novel adversarial attack method based on Bézier Curves, named BC-Attack (Bézier Curves Attack). BC-Attack constructs adversarial trajectories by searching for control point positions. It leverages the continuous curvature of Bézier curves to ensure the generated adversarial trajectory naturally possesses geometric smoothness, while applying both soft and hard constraints to limit the offset error, velocity, and acceleration of the adversarial trajectory, improving its feasibility in the real world. Experiments on the Apolloscape and nuScenes datasets show that BC-Attack improves the trajectory prediction error by an average of 210\% compared to normal prediction. Visualization of the trajectories also demonstrates that the generated adversarial trajectories are more stealthy and smoother than other attack methods. ProSpectNet: A Prototype-Guided Spectral-Temporal Framework for Efficient Traffic Forecasting Ziang Ma and Wenhao Zhang (Civil Aviation University of China) Abstract Abstract Accurate traffic flow forecasting is a cornerstone of Intelligent Transportation Systems (ITS).However, existing methods struggle to strike a balance between performance and efficiency: high-performance models (e.g., Transformers) achieve superior accuracy yet suffer from quadratic computational complexity O(N^2), limiting their scalability for real-time deployment; conversely, efficient lightweight models often neglect high-frequency details or global spatial information, consequently compromising prediction accuracy. To address this dilemma, we propose ProSpectNet, a novel framework that integrates spectral analysis with prototype learning. Specifically, we introduce: (1) A Multi-Frequency Decomposition(MFD) module, which explicitly disentangles traffic series into trend and detail components via gated fusion to prevent frequency aliasing; and (2) A Latent Prototype Router (LPR), which captures global semantic dependencies by routing information through learnable prototypes with linear complexity (O(NK)). Extensive experiments on four real-world traffic datasets (PEMS03, PEMS04, PEMS07, and PEMS08) demonstrate that ProSpectNet consistently outperforms state-of-the-art baselines in terms of long-term accuracy and efficiency.It achieves competitive prediction accuracy comparable to state-of-the-art Transformers (e.g., STAEformer) while outper-forming recent efficient models (e.g., STG-Mamba). Crucially,ProSpectNet reduces inference latency by approximately 14× and memory footprint by 88% compared to high-capacity baseline models, making it an ideal solution for large-scale, real-time ITS applications, which serves as a critical macro-level prior for downstream multi-agent autonomous driving and generative AI-based traffic simulations. HMF-Mapnet: Hybrid Memory Fusion for Accurate and Robust Vectorized Map Perception Huangnan Zheng, Ruihang Li, Heng Chen, Kehan Wang, Wangliang Guo, and Zhijie Pan (Zhejiang University) Abstract Abstract The accurate and efficient Vectorized high-definition (HD) Map perception is a fundamental prerequisite for reliable planning and decision-making in autonomous driving. To enhance the perception of complex road geometries and temporal dynamics, dominant paradigms either leverage explicit map priors or employ sequential modeling. Nevertheless, these two methods suffer from inherent weaknesses: the former is highly sensitive to prior quality, while the latter is prone to error accumulation, both of which ultimately limit the model's accuracy and robustness. To this end, we propose HMF-MapNet, a robust framework featuring a parallel dual-path architecture that achieves an optimal trade-off between efficiency and performance. Specifically, the geometric path utilizes a historical map bank that interfaces directly with the BEV feature map to bolster its geometric reasoning. Meanwhile, the temporal path employs a lightweight recurrent mechanism that iteratively decodes static map data over time, robustly separating static and dynamic elements.HMF-MapNet not only significantly reduces memory usage but also achieves a remarkable +6.3 mAP gain over SOTA methods under extreme rotation noise , proving its reliability for real-world deployment. Extensive experiments on nuScenes and Argoverse2 show that our method is remarkably superior to other existing methods. Monday Virtual Room 7 IJCNN Paper SS38 Deep Learning in Computational Biology and Biomedicine: from Biomedical Data to Drug Discovery I Session Chair: Zhe Liu (Jiangnan University), Luojian Xie (East China Normal University) LKSleepNet: An explainable lightweight model for sleep staging based on large kernel convolution Zhe Liu and Jinlong Yang (Jiangnan University) Abstract Abstract Automatic sleep staging is essential for diagnosing and treatment of sleep disorders. However, most convolutional based methods rely on empirically designed architectures and overlook the intrinsic time-frequency structure of EEG signals, resulting in elevated computational complexity, limited generalization and insufficient clinical interpretability that restrict both edge deployment and clinical adoption. Addressing these challenges, we propose an explainable and lightweight network LKSleepNet that integrates physiological priors into its architecture. The Fourier Transform Convolution (FTConv) and Large Kernel Convolution (LKConv) jointly capture global time-frequency patterns by combining transfer-domain filtering with large receptive field modeling. The lightweight 2D temporal folding module (LKTimes) provides efficient contextual representation by folding temporal information into a compact 2D feature map. The Prior-Guided Spectral Attention (PGSA) module actively extracts stage-discriminative spectral patterns by generating learnable attention masks. The spectral constraint loss regularizes feature learning to ensure physiological consistency with AASM guidelines, effectively supervising the model to focus on valid frequency bands. Furthermore, we employ Grad-CAM and Conditional Random Fields (CRF) to visualize backbone features, revealing the contributions of different frequency bands to each sleep stage. Experiments on three public datasets Sleep-EDF-20, Sleep-EDF-78, and SHHS demonstrates that LKSleepNet outperforms state-of-the-art methods, achieving macro F1-scores of 78.5%, 75.6%, and 79.5%, and overall accuracies of 85.8%, 83.2%, and 86.9%, respectively. Experimental results demonstrate the capable of LKSleepNet for robust and interpretable sleep staging on edge devices. Learning Class-Consistent Metacell Prototypes via Quantization for Few-Shot Single-Cell Annotation Yiwei Li (College of Software Engineering, Sichuan University); Yongqing Zhang (School of Computer Science, Chengdu University of Information Technology); Zixuan Wang (College of Electronics and Information Engineering, Sichuan University); Junlong Cheng (College of Computer Science, Sichuan University); Tianhao Li (School of Computer Science, Chengdu University of Information Technology); and Min Zhu (College of Computer Science, Sichuan University) Abstract Abstract Single-cell annotation plays a critical role in developmental biology, disease research, and precision medicine. However, most existing approaches rely on large-scale labeled datasets for supervised training and exhibit limited generalization in few-shot settings due to the scarcity of rare cell samples, sequencing-induced technical noise, pronounced intra-class representation bias, and weak inter-class separability. To address these challenges, we propose COMET, a few-shot single-cell annotation method based on class-consistent single-cell quantization. By introducing class-consistency constraints into the quantization process, COMET maintains mutually decoupled and structured latent code domains for each cell type, discretizes continuous single-cell representations in latent space, and constructs high-purity metacell prototypes through representation purification. This process effectively compresses redundant intra-class variation while amplifying key discriminative features, endowing metacells with stronger transferability and prior representational capacity under few-shot learning scenarios. Extensive experiments on eight real-world single-cell transcriptomic datasets demonstrate that COMET significantly outperforms non-quantized few-shot annotation baselines under limited supervision, particularly in rare cell recognition tasks. Moreover, the purified metacell representations exhibit superior structural consistency in clustering evaluations compared with raw single-cell representations, validating the robustness and reliability of COMET in complex single-cell analysis settings. Code is available in https://github.com/lilyway18/COMET. Physicochemically Informed Dual-Conditioned Generative Model of T-Cell Receptor Variable Regions for Cellular Therapy Jiahao Ma (The University of Hong Kong), Hongzong Li (HKUST), yefan Hu (Bayvaxbio Inc.), and Jiandong Huang (The University of Hong Kong) Abstract Abstract Physicochemically informed biological sequence generation has the potential to accelerate computer-aided cellular therapy, yet current models fail to jointly ensure novelty, diversity, and biophysical plausibility when designing variable regions of T-cell receptors (TCRs). We present PhysicoGPTCR, a large generative protein Transformer that is dual-conditioned on peptide and HLA context and trained to autoregressively synthesise TCR sequences while embedding residue-level physicochemical descriptors. The model is optimised on curated TCR-peptide-HLA triples with a maximum-likelihood objective and compared against ANN, GPTCR, LSTM, and VAE baselines. Across multiple neoantigen benchmarks, PhysicoGPTCR substantially improves edit-distance, similarity, and longest-common-subsequence scores, while populating a broader region of sequence space. Blind in-silico docking and structural modelling further reveal a higher proportion of binding-competent clones than the strongest baseline, validating the benefit of explicit context conditioning and physicochemical awareness. Experimental results demonstrate that dual-conditioned, physics-grounded generative modelling enables end-to-end design of functional TCR candidates, reducing the discovery timeline from months to minutes without sacrificing wet-lab verifiability. RetroRecs: Enhancing single-step Retrosynthesis via Reaction Context Sequence Luojian Xie, Zehui Wang, Zixian Cheng, Dongliang Chen, Huibin Wang, Liang Dou, and Ying Qian (East China Normal University) Abstract Abstract Retrosynthesis identifies reactants to synthesize a target molecule, which is essential for drug discovery. Graph-edit-based methods transform products into reactants via molecular edits. While existing approaches reflect chemical transformations, they often model molecular structures and edit content separately. To bridge this gap, we propose RetroRecs, a graph-edit-based retrosynthesis model that unifies this encoding. RetroRecs builds a structured Reaction Context Sequence that unifies the dynamic encoding of intermediate molecules with the contextual embedding of each molecular edit. This design supports stepwise and correlation-aware retrosynthetic prediction. This linkage is established through two modules in each edit prediction step: the Reaction Context Reasoner (RCR), which analyzes relationships among reaction contexts (describing the "what, where, and when" of past edits), and the Molecular Environment Perceiver (MEP), which refines this contextual reasoning by incorporating the current synthon's structural features. Additionally, we propose a contrastive learning objective, guided by reaction types, to enhance the model's ability to differentiate between various edit process groups. These designs allow the model to make edit decisions that are sequentially coherent and thus chemically valid. Experiments on USPTO-50K show RetroRecs achieving 92.3% and 70.2% in top-10 exact match and round-trip accuracy, setting a new state-of-the-art for semi-template retrosynthesis models. Monday Virtual Room 8 IJCNN Paper SS38 Deep Learning in Computational Biology and Biomedicine: from Biomedical Data to Drug Discovery II Session Chair: Pengwei Hu (Chinese Academy of Sciences, University of Chinese Academy of Sciences), Zhiyuan Chen (Beijing University of Technology; Institute of Automation, Chinese Academy of Sciences) Primate Spatio-Temporal Relational Network for Unified Behavior Quantification in Cage Environments Zhiyuan Chen (Beijing University of Technology; Institute of Automation, Chinese Academy of Sciences); Shuangqingyue Zhang (Northeastern University; Institute of Automation, Chinese Academy of Sciences); Xueyin Li (State Grid Gansu Electric Power Company); Zongli Jiang (Beijing University of Technology); and Xibo Ma (Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Abstract Abstract Quantifying macaque position, identity, and action is essential for neuroscience and pharmacology studies, yet remains challenging in cage environments due to frequent occlusions and rapid limb movements. This paper introduces the Primate Spatio-Temporal Relational Network (PSTRN), a framework designed to extract robust behavioral information from short video clips. PSTRN combines a hierarchical spatial encoder with a cross-frame relational attention module that applies temporal attention across all frames, adaptively aggregating complementary visual information to resolve ambiguities caused by occlusion or motion blur. A hierarchical 3D convolutional fusion then integrates multi-scale spatio-temporal features to jointly capture fine-grained limb motions and global body dynamics, enabling unified prediction of bounding boxes, identities, and actions. Experiments show that PSTRN achieves state-of-the-art performance, with 99.5% mAP50 and 85.1% F1 score on the single-macaque benchmark, and 92.6% mAP50 and 82.0% F1 score on the multi-macaque benchmark. These results demonstrate that PSTRN effectively captures macaque dynamics in realistic cage environments and offers a reliable solution for comprehensive behavior quantification. DM-VMUNet: Dual-Context Perception and Multi-Scale Feature Alignment in Vision Mamba for Medical Image Segmentation Tian Zhou (XinJiang University) and Yehang Li and Hui Zhao (Xinjiang University) Abstract Abstract Accurate segmentation of colorectal polyps and skin lesions is critical for clinical diagnosis. However, morphological variability and blurred lesion boundaries pose significant challenges to automated segmentation. Although the Mamba architecture excels in long-range modeling, its inherent 1D scanning mechanism often compromises 2D local geometric fidelity. Furthermore, conventional skip connections fail to adequately address semantic misalignment across different feature levels. To overcome these limitations, we propose DM-VMUNet, a hybrid architecture designed to integrate global context with local precision. Specifically, we introduce a Dual-Context Large-Small Convolution (DCLSConv) module. By leveraging multi-scale large-kernel perception and dynamic small-kernel aggregation, DCLSConv is designed to decouple receptive field expansion from computational cost, facilitating adaptive sharpening of lesion boundaries.Additionally, we propose a Multi-Scale Feature Alignment (MFA) module to semantically align encoder-decoder features prior to fusion. Extensive experiments on the Kvasir-SEG, ISIC 2017, and ISIC 2018 datasets demonstrate that DM-VMUNet achieves Dice coefficients of 90.75%, 89.87%, and 90.25%, respectively.These results validate the effectiveness of our approach, particularly its robust capability in delineating complex lesion boundaries compared to representative methods. Texture-Refined Probabilistic Prototype Learning for Semi-Supervised Medical Image Segmentation Liangjie wang, Ning Jiang, Xiaoqi Kuang, and Wanli Dong (Southwest University of Science and Technology) Abstract Abstract Given the scarcity of expert annotations, semi-supervised learning has become standard practice for medical image segmentation. The scarcity of expert annotations in medical imaging necessitates semi-supervised learning approaches. In semi-supervised medical image segmentation, existing methods typically impose point-wise consistency regularization or contrastive learning on individual voxels. However, such voxel-centric paradigm inherently lacks global structural constraints over class distributions, thus failing to disambiguate ambiguous boundaries with intra-class heterogeneity. While prototype learning provides distribution-level guidance, its semi-supervised application remains constrained by the reliability challenge of learning prototypes from noisy pseudo-labels. To this end, we present TR-PC, a framework that performs feature purification and semantic decoupling in probabilistic projection space. It first introduces a Texture-Refined Probabilistic Head (TR-Head) to suppress feature noise in 3D convolutions via local contextual modeling and dimension-balanced projection, thereby mapping raw voxels into high-fidelity multivariate Gaussian distributions. Afterwards, leveraging predictive uncertainty as a dynamic discriminator, it explicitly decouples each class center into cohesive core prototypes and dispersed boundary prototypes, and enforces divide-and-conquer likelihood matching. Therefore, the model is encouraged to infer and delineate highly challenging ambiguous regions via these decoupled core-boundary distributional conditions. TR-PC achieves state-of-the-art performance on three authoritative SSMIS benchmarks (LA, ACDC, and Pancreas-CT), with particularly significant improvements in boundary metrics (HD95). FCNDA: A Fuzzy Consensus Network for ncRNA–Drug Association Prediction Yujie Qi (Xinjiang Technical Institute of Physics and Chemistry, Xinjiang University); Xi Zhou (Xinjiang Technical Institute of Physics and Chemistry, University of Chinese Academy of Sciences); Xiaobo Zhu (Xinjiang Technical Institute of Physics and Chemistry); and Lun Hu and Pengwei Hu (Xinjiang Technical Institute of Physics and Chemistry, University of Chinese Academy of Sciences) Abstract Abstract Non-coding RNAs (ncRNAs) have been increasingly recognized as key regulators in diverse biological processes, and accumulating evidence suggests that interactions between ncRNAs and drugs play a critical role in disease regulation and therapeutic outcomes. Systematic identification of potential ncRNA–drug associations is therefore essential for understanding disease mechanisms and supporting drug discovery. Recent computational methods commonly integrate multiple similarity views of ncRNAs and drugs to improve prediction accuracy; however, heterogeneous biological similarity information often varies in reliability and exhibits inherent uncertainty, posing challenges for effective multi-view integration. To address this issue, we propose FCNDA, a novel ncRNA–drug association prediction model based on fuzzy consensus learning. FCNDA performs association prediction independently under each combination of ncRNA and drug similarity views and interprets view-specific prediction scores as fuzzy membership degrees, enabling explicit modeling of uncertainty and partial biological evidence. An adaptive fuzzy consensus coordination mechanism is then employed to aggregate multiple fuzzy memberships into a unified global prediction, where informative view combinations are emphasized while moderate discrepancies across views are tolerated. Experimental results on benchmark datasets demonstrate that FCNDA achieves competitive predictive performance and exhibits strong robustness to noisy and heterogeneous similarity information. Moreover, the learned consensus weights provide interpretable insights into the relative contributions of different biological similarity views. Monday Virtual Room 1 IJCNN Paper SS10 Trustworthy and Explainable Federated Learning: Towards Security and Privacy Future I Session Chair: Sen Yu (Yunnan University), Hongzhan Ma (Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Model Drift via Neural Collapse and Negative Label Knowledge Distillation Sen Yu (Yunnan University), Ruizhi Pu (Southeast University), Shaojie Zhan (Texas Tech University), Xiuting Weng (Yunnan University), Yuanhang Yao (Southeast University), and Lixing Yu (Yunnan University) Abstract Abstract Federated learning (FL) is severely impacted by model drift under data heterogeneity, manifesting as simultaneous feature representation drift and classifier drift across local clients—a dual challenge that collectively impairs global model aggregation performance. Current methodologies predominantly address a singular aspect of this challenge. In pursuit of a comprehensive resolution, we introduce a novel framework grounded in neural collapse theory. Our framework aims at achieving thorough drift mitigation by decoupling features from classifiers. We replace cross-entropy loss with hyperspherical uniformity gap loss in FL frameworks to mitigate classifier bias through data-agnostic means. Additionally, we reconstruct latent negative label knowledge distributions using knowledge distillation to ensure geometric consistency in feature representations across diverse clients. Extensive experiments validate the effectiveness of the proposed method in heterogeneous data. FedDetox: Robust Federated SLM Alignment via On-Device Data Sanitization Shunan Zhu, Jiawei Chen, Yonghao Yu, and Hideya Ochiai (The University of Tokyo) Abstract Abstract As high quality public data becomes scarce, Federated Learning (FL) provides a vital pathway to leverage valuable private user data while preserving privacy. However, real-world client data often contains toxic or unsafe information. This leads to a critical issue we define as unintended data poisoning, which can severely damage the safety alignment of global models during federated alignment. To address this, we propose FedDetox, a robust framework tailored for Small Language Models (SLMs) on resource-constrained edge devices. We first employ knowledge distillation to transfer sophisticated safety alignment capabilities from large scale safety aligned teacher models into light weight student classifiers suitable for resource constrained edge devices. Specifically, during federated learning for human preference alignment, the edge client identifies unsafe samples at the source and replaces them with refusal templates, effectively transforming potential poisons into positive safety signals. Experiments demonstrate that our approach preserves model safety at a level comparable to centralized baselines without compromising general utility. FedPAL: Adaptive Head Aggregation for Personalized Federated Learning Guided by Generative Feature Distributions Heng Zhang, Ping Zhang, Defeng Wang, Wenhui Huang, and Cankun Yue (Henan University of Science and Technology) Abstract Abstract Federated learning, as an emerging distributed machine learning paradigm, offers a key advantage in effectively mitigating data silos while enhancing data privacy protection. However, heterogeneity in client data distributions often leads to insufficient generalization of the global model, thereby weakening personalized adaptation across clients. To address this issue, we propose a novel personalized federated learning method, FedPAL. On the server side, a class-conditional generator is constructed and trained using a classification loss and a supervised contrastive loss to generate feature distributions corresponding to class labels, so that the feature distributions exhibit intra-class compactness and inter-class separability. After downloading the generator, the client leverages the generated feature distributions to learn fusion weights between the local model head parameters and the global model head parameters. This process enables adaptive aggregation, which balances the absorption of global knowledge with the preservation of critical local information. As a result, an effective local model is initialized for each client at every communication round. To evaluate the effectiveness of FedPAL, extensive experiments are conducted on four benchmark datasets in the computer vision domain. The results demonstrate that the proposed method significantly outperforms ten state-of-the-art federated learning baseline methods. TrustedMix: Mixup with TEE for improving the data distribution heterogeneity of federated learning Hongzhan Ma (Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences); Muyan Shen (University of Chinese Academy of Sciences); and Yu Qin (Institute of Software, Chinese Academy of Sciences) Abstract Abstract Data heterogeneity (Non-IID) remains a critical challenge in federated learning (FL), often leading to model divergence and performance degradation. We propose TrustedMix, a novel framework that leverages Trusted Execution Environments (TEEs) to facilitate secure, privacy-preserving data sharing for distribution rectification. In TrustedMix, clients offload encrypted subsets of local data to a server-side TEE enclave, which computes the global label distribution and performs cross-client sample mixing within its hardware-isolated boundary. These synthesized samples are redistributed to clients to augment local training, effectively mitigating the impact of data skewness. Evaluated on CIFAR-10 and Tiny-ImageNet benchmarks, TrustedMix consistently outperforms state-of-the-art methods including FedAvg, FedProx, and FedMix. Specifically, in extreme non-IID scenarios ($\alpha=0.1$), the test accuracy of TrustedMix increases to up to 89.0\%, significantly surpassing the performance of FedMix while maintaining robust data confidentiality and minimal communication overhead. Monday Virtual Room 2 IJCNN Paper SS10 Trustworthy and Explainable Federated Learning: Towards Security and Privacy Future II Session Chair: 周 荣博 (Xinjiang University), Manel MILI (Faculty of Sciences of Monastir, Medical Technology and Imaging Laboratory LTIM-LR12ES06) FedPAF: Federated Personalized Alignment Framework Rongbo Zhou, Jiao Tian, Liejun Wang, and Yihao Yin (xinjiang university) Abstract Abstract Federated learning under client heterogeneity faces non-IID data distributions and client-specific feature statistics that degrade performance. While prior personalized federated learning methods partially mitigate statistical mismatch through local adaptation or normalization decoupling, they struggle to jointly model client-dependent discriminative cues and maintain cross-client semantic consistency. We propose FedPAF, a personalized federated framework that improves performance by jointly addressing normalization heterogeneity, feature heterogeneity, and semantic drift. FedPAF adopts a personalization strategy with normalization decoupling, leveraging Per-FedAvg-style meta-learning for adaptation while maintaining FedBN-style local batch normalization, enabling transferable representation learning while preserving client-specific statistics. It further highlights client-dependent discriminative cues through a lightweight multi-spectral channel reweighting module with negligible overhead, and aligns class semantics via global prototype aggregation to reduce semantic drift and enhance class-wise separability. Extensive experiments on Digits-Five, PACS, and DomainNet demonstrate consistent improvements over strong baselines in both accuracy and Macro-F1 under heterogeneous client settings. Split to Spot: Multi-Reference Byzantine Detection via Gradient Splitting in Federated Learning Zihan Yu, Zukang Ai, Yilin Su, Dan Yu, Yongle Chen, and Jianhua Wang (Taiyuan University of Technology) Abstract Abstract Federated learning under non-independent and identically distributed (non-IID) data settings faces severe Byzantine security challenges. To mitigate the performance degradation of robust AGgregation Rules(AGRs) caused by gradient heterogeneity and the curse of dimensionality, a GrAdient Splitting (GAS) mechanism has been proposed. However, existing GAS method typically relies on a single reference when identifying Byzantine clients, making their defense performance sensitive to reference selection. This sensitivity arises because different attack types often require different reference choices to achieve satisfactory robustness. Motivated by this observation, we propose GAS-MR, namely Gradient Splitting with Multi-Reference Consistency-Based Byzantine Detection, which leverages multiple references to quantify the deviation-based inconsistency of Byzantine clients after gradient splitting, thereby enabling robust Byzantine client identification. Experimental results demonstrate that the proposed method consistently outperforms GAS under various nonIID data distributions and five representative Byzantine attack scenarios. Even in cases where performance is comparable, our approach exhibits superior robustness and stability, thereby significantly reducing the dependence on reference selection and prior knowledge of attack strategies Code is available at https://github.com/SKLIIS-AIS/GAS-MR.git. RBA: Representation Backdoor Attack on Personalized Federated Learning Hefeng Zhou and Dingxuan Zhang (Shanghai Jiao Tong University), Zhihong Su (Ant Group), Honghong Zeng and Jiong Lou (Shanghai Jiao Tong University), Wugedele Bao (Hohhot Minzu College), and Chentao Wu and Jie Li (Shanghai Jiao Tong University) Abstract Abstract Personalized federated learning (FL) tackles the challenge of data heterogeneity by training personalized models instead of a single global model. Backdoor attacks pose a fatal threat to FL systems. There is an urgent need to investigate on backdoor attacks under personalized FL. Existing representation-based personalized FL frameworks share a global extractor, which creates a large attack surface for backdoor attacks. This work explores previously unknown backdoor risks in FedPer and FedRep through data poisoning and parameter scaling. Experimental results show that the attack success rate (ASR) degrades dramatically under different pFL frameworks. We further propose a backdoor attack leveraging the shared representation extractor to inject backdoors more effectively. Experimental results demonstrate that our method improves the ASR by up to 72% while maintaining the ASR after stopping the attack for 200 rounds. Code is available at https://github.com/RezinChow/RBA. Weakly Supervised ViT-Conditioned Text Explanations in Glioma Histology Manel Mili (Faculty of Sciences of Monastir, University of Monastir, Tunisia; Laboratory of Technology and Medical Imaging (LR12ES06), Faculty of Medicine, University of Monastir, Tunisia); Abderrahman Ben Abdeljelil (Sfax National Engineering School, University of Sfax, Tunisia; Laboratory of Technology and Medical Imaging (LR12ES06), Faculty of Medicine, University of Monastir, Tunisia); Asma Ben Abdallah (Higher Institute of Computer Science and Mathematics of Monastir, University of Monastir, Tunisia; Laboratory of Technology and Medical Imaging (LR12ES06), Faculty of Medicine, University of Monastir, Tunisia); and Mohamed Hedi Bedoui (Faculty of Medicine of Monastir, University of Monastir, Tunisia; Laboratory of Technology and Medical Imaging (LR12ES06), Faculty of Medicine, University of Monastir, Tunisia) Abstract Abstract Predicting the promoter methylation of O$^{6}$-methylguanine--DNA--methyltransferase (MGMT) from hematoxylin–eosin (H\&E)–stained glioma biopsies using deep learning (DL) remains challenging due to limited interpretability. In this work, we propose a weakly supervised framework that jointly performs classification and generates textual explanations grounded in visual evidence. We employ a lightweight convolutional neural network (CNN) to predict MGMT status from histology images. Building on these predictions, we introduce a Vision Transformer (ViT)–conditioned language decoder guided by SmoothGrad-CAM saliency maps to produce class-aware textual explanations. Patch-level predictions and saliency maps guide the generation of concise rationales aligned with visually relevant tissue regions. We further assess explanation faithfulness using perturbation-based insertion and deletion analyses. On a glioblastoma (GBM) cohort, the proposed model achieves a patient-level F1-score of 0.94, while the generated explanations remain consistent with both saliency maps and model predictions, demonstrating that class-conditioned textual rationales can provide a faithful and interpretable interface for histology-based MGMT prediction. Monday Virtual Room 3 IJCNN Paper SS19 Advances in Trustworthy XAI: Novel Methodologies, Benchmarking, and Diverse Data Modality Contexts I Session Chair: Keyuan Wang (Northwest A&F University ), Liangzhou Qu (Shenzhen University) GLM-Based: CAM Based on Gradient Local Mean for Enhanced Explainability in Visual Transformer Keyuan Wang and Yixin Ren (Northwest A&F University), Lingyuan Kong (Harbin Institute of Technology), and Zhiyi Zhang (Northwest A&F University) Abstract Abstract The rapid advancement of Visual Transformers (ViT) has created an urgent need for model interpretability, however, existing gradient-based explanation methods often suffer from significant visual noise caused by local gradient fluctuations. To address this challenge, this paper proposes the Gradient Local Mean-based (GLM) method, a novel framework designed to enhance the transparency of ViT models. The GLM method employs a stochastic sampling strategy by introducing Gaussian noise into the input image and aggregating the resulting Class Activation Maps (CAMs)—derived from gradient and attention matrices—to effectively mitigate noise instability. Extensive experiments on the ImageNet ILSVRC 2012 and CUB-200-2011 datasets demonstrate the method's superiority. On the fine-grained CUB-200-2011 dataset, the GLM method achieves a Top-1 Localization Intersection over Union (IoU) of 0.5958, significantly outperforming the state-of-the-art AGCAM method (0.4235) by 40.7%, while attaining a Pixel Accuracy of 0.8405. Furthermore, in pixel perturbation tests on ImageNet, the proposed method yields a superior Area Between Perturbation Curves (ABPC) score of 0.3739, statistically outperforming leading alternatives (p<0.05). These results confirm that the GLM method provides robust and precise visual explanations, facilitating reliable decision-making in practical applications. SwinBERT-Fake: Boosting Multimodal Fake News Detection via Hierarchical Visual Modeling and Explainable AI Yihao Zhao, Yitian Lu, Runyang Yu, Zhan Shi, and Yan Tu (Wuhan University of Technology) Abstract Abstract Multimodal fake news detection faces two critical challenges: capturing fine-grained visual artifacts (e.g., splicing traces) and bridging the semantic gap between textual and visual modalities. Existing state-of-the-art approaches typically rely on standard Vision Transformers (ViT) or augment semantic information via automated image captioning. However, we argue that ViT's global attention mechanism often overlooks local discriminative details, while generated captions introduce semantic hallucinations and significant computational latency. To address these limitations, we propose SwinBERT-Fake, a novel framework that integrates Swin Transformer with BERT to prioritize hierarchical visual modeling over semantic augmentation. By leveraging Swin’s shifted window mechanism, our model effectively localizes subtle visual inconsistencies that standard ViT misses. Extensive experiments on the Fakeddit dataset demonstrate that SwinBERT-Fake achieves a state-of-the-art accuracy of 88.09%, significantly outperforming ViT-based baselines. Furthermore, our rigorous ablation studies reveal that semantic augmentation via captioning is redundant, offering no statistical performance gain while increasing inference latency by approximately 4.5 times. Finally, we introduce a comprehensive Explainable AI (XAI) module—incorporating Eigen-CAM, Integrated Gradients, and SHAP—to provide transparent visual, textual, and modal-level evidence for detection. Our findings suggest that optimizing the granularity of visual backbones is a more effective and efficient pathway for misinformation detection than relying on external generative models. Reliable Multimodal Skin Lesion Analysis using Dermoscopic Images and Clinical Metadata Sonia Bouzidi, Imen Jdey, and Fadoua Drira (University of Sfax) Abstract Abstract Reliable decision-making in medical applications has become increasingly crucial, particularly in skin lesion analysis, where misdiagnosis can lead to severe health consequences. While recent advances in computer vision have achieved high performance, relying solely on image-based models often overlooks essential clinical context. To address this limitation, we propose a novel multimodal framework that integrates dermoscopic image analysis with structured clinical metadata to improve both classification performance and decision reliability. The proposed approach combines an optimized Vision Transformer with a Deep Fuzzy Neural Network (ViT-DFNN) for image-based representation learning and a TabNet model for tabular data, enabling effective fusion of heterogeneous information through a decision-level weighted strategy. Experimental results on the HAM10000 dataset demonstrate the effectiveness of the proposed framework, achieving an accuracy of 93.15%, AUC of 96.1%, precision of 92.95%, sensitivity of 93.30%, specificity of 98.85%, and G-mean of 96%. These results confirm that integrating clinical metadata significantly enhances predictive performance and robustness compared to unimodal approaches. To ensure transparency, the framework incorporates Explainable Artificial Intelligence (XAI) techniques, where Score-CAM highlights important image regions and Global Feature Importance (GFI) identifies influential clinical variables. Overall, the proposed method provides a reliable and interpretable solution for AI-assisted skin lesion diagnosis. FINS is your favor: A Task-Driven Data Evaluation Framework for Financial Sentiment Analysis Liangzhou Qu (Shenzhen University); Di Han (Guangdong University of Finance); Junjie Mao (Shenzhen University); and Qixian Li, Zikun Guo, and Weiquan Fan (Guangdong University of Finance) Abstract Abstract Currently, datasets for Financial Sentiment Analysis (FSA) often contain rich multimodal information, yet the task-oriented evaluation metrics for processing this sentiment remain singular. Consequently, performance comparisons across different datasets are challenging, which hinders researchers from selecting appropriate resources for specific tasks. To address this problem, we propose FINS (Financial Integration Nexus of Standards), a data evaluation framework that systematically summarizes existing studies and abstracts the main application tasks of FSA into five core dimensions: cycle, risk, causality, entity and wording. On this basis, we establish a unified quantitative evaluation system for data. FINS combines large language models with financial toolchains to achieve an automated evaluation process, and presents the adaptability of data across different task dimensions through multi-dimensional quantitative score maps. Empirical results and real financial scenario experiment on several representative data sources show that FINS can effectively capture structural differences, which provides systematic support for researchers in data selection and modeling practices for specific tasks. Monday Virtual Room 4 IJCNN Paper SS19 Advances in Trustworthy XAI: Novel Methodologies, Benchmarking, and Diverse Data Modality Contexts II Session Chair: Mufti Mahmud (King Fahd University of Petroleum and Minerals), Yue Guo (南京航空航天大学) Cross-Model Interpretability of Alzheimer’s Disease Classification Using ADNI Dataset Ghada Abdulsalam and Moataz Aly Kamaleldin Ahmed (King Fahd University of Petroleum & Minerals) and Mufti Mahmud (King Fahd University of Petroleum and Minerals) Abstract Abstract Alzheimer’s Disease (AD) staging from tabular clinical and neuropsychological data remains highly relevant for real-world decision support, yet model interpretability is hindered by multicollinearity among overlapping cognitive and functional measures. Using baseline Alzheimer’s Disease Neuroimaging Initiative (ADNI) data, we study one-vs-rest classification of Cognitively Normal (CN), Mild Cognitive Impairment (MCI), and AD with linear and tree-based models: Elastic Net Logistic Regression (LR), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). Beyond predictive performance, we provide a structured interpretability workflow that integrates Top-K global importance, cross-model feature agreement, correlation analysis of consensus predictors, and SHapley Additive exPlanations (SHAP). Results show a stable, clinically coherent core of predictors across models, with stage-specific shifts from global cognition and memory signals in CN to combined cognitive–functional impairment in MCI, and to severity and activities-of-daily-living measures in AD. Cross-model agreement highlights clinical dementia rating and functional measures as robust AD indicators. SHAP summaries further demonstrate consistent directional effects for key measures (e.g., higher CDRSB increasing the likelihood of AD and higher MMSE decreasing it). Overall, the proposed framework improves interpretability by contextualizing feature importance with cross-model robustness and correlation structure, supporting more reliable identification of AD-relevant predictors from tabular data. WCI-Net: A Wavelet-guided Cross-domain Interaction Network for Underwater Image Enhancement Yupeng Ma and Qiuling Yang (Hainan University, School of Computer Science and Technology) Abstract Abstract Underwater images suffer from color distortion and detail loss due to wavelength-dependent absorption and scattering. Existing deep learning-based enhancement methods primarily operate in the pixel domain or rely on a single frequency-domain transform, limiting their ability to jointly preserve fine-grained details and model global dependencies. To address this gap, we propose WCI-Net, a wavelet-guided cross-domain interaction network for underwater image enhancement. WCI-Net decomposes features into multi-scale wavelet subbands to capture local structural details, while a cross-domain interaction module integrates wavelet-domain representations with Fourier-domain global refinement. A dual-domain loss further regularizes learning in both spatial and frequency domains. Extensive experiments on real-world benchmarks demonstrate that WCI-Net achieves state-of-the-art enhancement performance with low computational cost. Strat-LLM: Stratified Strategy Alignment for LLM-based Stock Trading with Real-time Multi-Source Signals Wenliang Huang (Zhejiang University of Technology) and Zengyi Yu (East China Normal University) Abstract Abstract Large Language Models (LLMs) are evolving into autonomous trading agents, yet existing benchmarks often overlook the interplay between architectural reasoning and strategy consistency. We propose Strat-LLM, a framework grounded in Stratified Strategy Alignment. Operating in a live-forward setting throughout 2025, it integrates heterogeneous data including sequential prices, real-time news, and annual reports to eliminate look-ahead bias. Extensive stress tests on A-share and U.S. markets reveal: (1) reasoning-heavy models achieve peak utility in Free Mode via internal logic, whereas standard models require Strict Mode as a vital risk anchor; (2) alignment utility is regime-dependent, with Free and Guided modes capturing momentum in uptrending markets, while Strict Mode mitigates drawdowns in downtrends; (3) mid-scale models (35B) show optimal fidelity under strict constraints, whereas ultra-large models (122B) suffer an alignment tax under rigid rules but gain a performance premium in Guided Mode; (4) standard LLMs often fall into a high win-rate trap, optimizing for small gains at the expense of total returns, which can only be mitigated through deep reasoning or strict external guardrails. Project details are available at https://Strat-LLM.github.io. RACER: Risk-Aware Conformal Expert Routing for Risk-Constrained Efficient Inference in Cascaded Language Models and Mixture-Of-Experts Architectures Yue Guo, Aoyu Li, Qingkang Tang, and Huiping Ma (Nanjing University of Aeronautics and Astronautics) Abstract Abstract Mixture-of-experts (MoE) language models promise favorable capacity–compute trade-offs, yet their routing decisions remain brittle: a single misrouted token can trigger compounding hallucinations or costly fallback heuristics, undermining reliability in real deployments. We propose RACER, a post-hoc, riskcontrolled routing layer that turns a black-box MoE (or any cascaded language system) into a selective generator: the system either (i) answers with a cheap expert path or (ii) escalates to a stronger expert (or retrieval-augmented path) when its own signals indicate elevated risk. RACERis built around conformal risk control and “learn-then-test” (LTT) calibration, yielding ffnitesample guarantees on selective risk under mild exchangeability assumptions. We evaluate RACERon two real tasks—TruthfulQA factuality and Natural Questions open-domain QA—using both a Mixtral-8x7B MoE system and a LLaMA-2-7B→70B cascade. Across settings, RACERsatisffes user-speciffed risk constraints while achieving 26–45% compute savings over always-large baselines at matched error rates, outperforming six competitive baselines including entropy thresholding, temperature scaling, and self-consistency routing. Ablations conffrm robustness to calibration set size (stable with n ≥ 150), distribution shift (topic/time splits), and score function choice. Code will be released upon acceptance. Monday Virtual Room 5 IJCNN Paper SS37 AI in Healthcare: Harnessing Emerging, Generative and Agentic Technologies for Responsible Innovation Session Chair: Matheus Becali Rocha (Universidade Federal do Espírito Santo, Nature Inspired Computing Laboratory), Qishen Chen (Shanghai University) De-biasing multi-question learning in medical visual question answering via semantic-aware natural direct effect subtraction Qishen Chen, Wenxuan He, and Huahu Xu (Shanghai University) Abstract Abstract Medical Visual Question Answering (MedVQA) supports clinical decision-making by answering diagnostic questions based on medical images. Multi-Question Learning (MQL) has been adopted to jointly model multiple questions per image, improving performance by leveraging inter-question dependencies. However, MQL also amplifies language bias, where models rely on textual patterns rather than visual content, by introducing a new shortcut: reasoning from correlated questions alone. Using only questions, MQL can exhibit severe language bias, yet achieves 84.35\% accuracy on the SLAKE dataset, nearly 30\% higher than traditional single-question learning methods, highlighting the risk of spurious shortcuts in reasoning. Further experiments conducted on the proposed SLAKE-CP-Co dataset, designed to isolate the language bias introduced by co-occurring questions, provide evidence for the presence of multi-question language bias in the MQL. Thus, this paper proposes a novel causal framework for MedVQA under MQL. The proposed method explicitly identifies two types of language bias: the conventional Single-Question Language Bias and the newly introduced Multi-Question Language Bias. Both are mitigated by subtracting their Natural Direct Effects from the total effect, enforcing greater reliance on visual information. Additionally, this paper introduces a Semantic-Aware Bias Subtraction module that dynamically adjusts the subtraction strength based on question semantics, preserving beneficial language cues for domain-knowledge questions. Experiments on SLAKE, VQA-RAD, SLAKE-CP and SLAKE-CP-Co demonstrate that the proposed method outperforms state-of-the-art biased and de-biased models in both accuracy and robustness. The proposed approach offers a principled, effective strategy to balance multi-question reasoning and language bias mitigation, advancing the state of the art in MedVQA. RALSTM: Autoregressive Multi-Phase Contrast-Enhanced CT Synthesis via Spatio-Temporal Priors Jiajian Xie and Wenfeng Xu (South China Normal University, School of Computer Science); Cong Lai and Cheng Liu (Sun Yat-sen Memorial Hospital, Sun Yat-sen University; Department of Urology); Gansen Zhao (South China Normal University, School of Computer Science); and Kewei Xu (Sun Yat-sen Memorial Hospital, Sun Yat-sen University; Department of Urology) Abstract Abstract To extend the benefits of dynamic lesion characterization to patients with contrast agent contraindications, synthesizing multi-phase contrast-enhanced computed tomography (CECT) from non-contrast CT (NCCT) has emerged as a vital tool for personalized radiotherapy planning. However, existing methods often struggle to fully utilize the anatomical context priors of NCCT and the temporal dependencies between different phases. This results in anatomical distortions across different phases and a significant increase in training iterations and parameter count, limiting their practical application. To overcome these challenges, this paper introduces a novel framework named RALSTM for the autoregressive and sequential synthesis of multi-phase CECT from abdominal NCCT. The framework leverages a Residual ConvLSTM module to model encoded features and capture the temporal transitions across phases. Additionally, a Region-Aware Gating Unit is used to localize contrast-enhancing regions and impose spatial constraints on the decoded features. Experimental results on both internal and external datasets demonstrate that this method achieves superior metrics and delivers outstanding synthesis quality. RKANet: Reconstruction-Consistent KAN–Mamba Framework for Medical Image Segmentation Wei Wang, Junxiao Lv, and Xin Wang (Changsha University of Science and Technology) Abstract Abstract Medical image segmentation remains challenging due to complex anatomy, ambiguous boundaries, and elongated structures, where structural inconsistency is frequently observed during feature reconstruction. In many encoder–decoder architectures, local upsampling operations are insufficient to preserve geometric continuity and long-range dependencies, resulting in fragmented predictions and boundary artifacts. To address this issue, a reconstruction-consistent segmentation framework, termed RKANet, is presented, in which decoding is explicitly strengthened as a structured reconstruction process. A cooperative modeling module is designed to jointly enforce local geometric refinement and global dependency propagation. In addition, a Boundary-Guided Skip Branch is incorporated to enhance skip feature fusion with boundary-aware shallow features. Extensive experiments on five public datasets demonstrate that RKANet achieves consistent performance improvements over representative methods, yielding a 3.00% Dice gain on BUSI and stable improvements of 0.25% – 0.55% on other benchmarks, with an average Dice score of 88.22%. Enhancing Diagnostic Accuracy for Urinary Tract Disease through Explainable SHAP-Guided Feature Selection and Classification Filipe Ferreira de Oliveira, Matheus Becali Rocha, and Renato A. Krohling (Federal University of Espírito Santo, Nature Inspired Computing Laboratory) Abstract Abstract This paper proposes an explainable machine learning framework to support the diagnosis of urinary tract diseases, particularly bladder cancer, by employing SHAP (SHapley Additive exPlanations)-based feature selection to enhance model transparency and predictive performance. Six binary classification scenarios were designed to differentiate bladder cancer from other urological and oncological conditions. The models were implemented using XGBoost, LightGBM, and CatBoost algorithms, with hyperparameter optimization conducted via Optuna and class imbalance addressed through the SMOTE technique. Predictive variables were selected according to their importance values derived from SHAP analyses. The results demonstrated that SHAP-based feature selection efficiently reduced data dimensionality while maintaining or even improving performance metrics such as balanced accuracy, precision, and specificity. These findings underscore the potential of explainable artificial intelligence in medical diagnostics. The proposed methodology contributes to the development of transparent, reliable, and efficient clinical decision support systems, helping to optimize the screening and early diagnosis of urinary tract diseases, particularly bladder cancer. Monday Virtual Room 6 IJCNN Paper SS01 Privacy-Preserving Machine and Deep Learning Session Chair: Wen Yan (Heilongjiang university), Weigang Wu (Sun Yat-sen University) Joint Optimization of Adaptive Gradient Compression and Aggregation for Efficient Asynchronous Federated Learning Yingwei Hou, Danyang Xiao, Linlin You, and Weigang Wu (Sun Yat-sen University) Abstract Abstract Asynchronous Federated Learning (AFL) enables distributed model training without client synchronization, offering enhanced scalability in heterogeneous and bandwidth-limited environments. However, in many AFL approaches, the use of gradient compression is coupled with fixed aggregation intervals and compression rates, which often leads to suboptimal performance caused by excessive model staleness and communication cost. We propose FedJACA, a novel AFL framework that jointly and adaptively configures the aggregation interval and gradient compression rate, so as to maximize gradient utility while minimizing staleness and communication cost. We model their intertwined impact and formulate a joint optimization problem under training time and bandwidth constraints. And we design a customized solution algorithm based on the Trust Region Newton method that yields high-quality solutions for achieving a balanced trade-off between efficiency and cost. Extensive experiments on four benchmarks show that FedJACA improves accuracy, reduces training time by 40%, and lowers communication cost by 35%, outperforming state-of-the-art baselines. Noise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models Guo Li and Weihong Chen (South China University of Technology) and Yongfu Fan (University of Electronic Science and Technology of China) Abstract Abstract Diffusion models have demonstrated powerful performance in generating high-quality images. A typical example is text-to-image generator like Stable Diffusion. However, their widespread use also poses potential privacy risks. A key concern is membership inference attacks, which attempt to determine whether a particular data sample was used in the model training process. Existing membership inference attacks against diffusion models either directly exploit sample loss differences or rely on image-level reconstruction differences. Both approaches commonly ignore the consistency characteristics of noise prediction during the diffusion process, resulting in either low inference accuracy or high computational costs. To address these shortcomings, we propose a membership inference method based on noise aggregation analysis, and introduce a single-step, low-intensity noise injection diffusion strategy to amplify differences between member and non-member samples. Our proposed approach substantially reduces model query requirements while delivering more efficient and accurate membership inference. SMRAM: Boosting Class-specific Face Privacy Protection with Stochastic Mask Regularized Adaptive Momentum jing han, yuanbo li, cong hu, and xiaojun wu (Jiangnan University) Abstract Abstract To ensure secure face privacy protection, existing class-specific adversarial methods have shown substantial performance improvement by tailoring perturbations to individual identities. However, due to high local similarity in same-identity images, these methods focus perturbations on specific regions. Though some studies use generic features, they still fail to adequately balance generic and fine-grained features, yielding limited privacy protection improvements. To address these issues, we innovatively propose Boosting Class-specific Face Privacy Protection with Stochastic Mask Regularized Adaptive Momentum (SMRAM) for privacy protection. This approach introduces Stochastic Mask Regularized Adaptive Momentum, which dynamically integrates fine-grained features and global consistency features through Adaptive Momentum (AM). Additionally, Stochastic Mask Regularization (SMR) increases the diversity of the perturbation distribution, thereby capturing richer generic features. When fine-grained features dominate, AM refines the generation direction by incorporating generic features and stabilizes the generation process through momentum accumulation, thereby improving the success rate of privacy protection. Furthermore, SMR forces the model to learn more generic features, enhancing the generalization of the perturbations. Extensive experiments conducted on two benchmark datasets demonstrate that the proposed SMRAM outperforms the state-of-the-art approaches, achieving a higher success rate in privacy protection for same-identity images under black-box recognition models. Dual Adaptive Augmentation with Difficulty-Aware Guidance for Semi-Supervised Semantic Segmentation Wen Yan, Jinghua Zhu, and Heran Xi (Heilongjiang university) Abstract Abstract Semi-supervised semantic segmentation has made significant advances, yet current methods often overlook variations in sample difficulty and rely on limited data augmentation strategies, which constrains both accuracy and the exploration of diverse perturbation spaces. This paper proposes a Dual Adaptive Augmentation framework. First, we introduce a quantitative hardness evaluation mechanism based on the Intersection over Union (IoU) of segmentation predictions, which assesses the learning difficulty of each unlabeled sample. Guided by the derived hardness coefficient, we dynamically adjust the mixing ratio of weak and strong augmentations, aligning perturbation intensity with the model's evolving learning state. Second, we expand the perturbation space by applying two independent strong augmentation strategies to unlabeled images, thereby improving generalization. Finally, we introduce additional random perturbations, such as Dropout2D, between the encoder and decoder layers to enhance feature-level diversity and robustness. Extensive experiments demonstrate that our model consistently outperforms existing methods across different data partition settings. Monday Virtual Room 7 IJCNN Paper SS03 Physics-Informed Neural Networks: Advancements and Applications Session Chair: Xiaohui Jia (North University of China), Junqi Qu (Florida State University) COMPOL: A Scalable Neural Operator Framework for Multi-Physics Simulations Junqi Qu (Florida State University), Tao Wang (The University of Texas at Austin), Yushun Dong (Florida State University), Hewei Tang (the University of Texas at Austin), and Shibo Li (Florida State University) Abstract Abstract Multi-physics simulations play an essential role in accurately modeling complex interactions across diverse scientific and engineering domains. Although neural operators, particularly the Fourier Neural Operator (FNO), have significantly improved computational efficiency, they often fail to capture the intricate correlations inherent in coupled physical processes. To address this limitation, we introduce COMPOL, a novel coupled multi-physics operator learning framework. COMPOL extends conventional operator architectures by incorporating a sophisticated attention-based aggregation mechanism that effectively models interdependencies among interacting physical processes within latent feature spaces. Our approach is architectureagnostic and seamlessly integrates into various neural operator frameworks involving latent space transformations. Extensive experiments on diverse benchmarks, including biological reactiondiffusion systems, pattern-forming chemical reactions, multiphase geological flows, and thermo-hydro-mechanical processes, demonstrate that COMPOL consistently achieves superior predictive accuracy compared to state-of-the-art methods. Our code and data are available at https://github.com/AriaQJ/COMPOL. GAF-Flow: An Auditory-Perceptual Flow Model for Enhanced Speech Synthesis Rao Deng (Beijing University of Posts and Telecommunications, Shanghai Key Laboratory of Forensic Medicine and Key Laboratory of Forensic Science); Weike You and Linna Zhou (Beijing University of Posts and Telecommunications); and Hong Guo (Shanghai Key Laboratory of Forensic Medicine and Key Laboratory of Forensic Science) Abstract Abstract This paper introduces GAF-Flow, a novel auditory-perceptual flow model for high-fidelity neural Text-To-Speech (TTS) synthesis. To address the limitations of “machine-centric” acoustic modeling, which often overlooks the non-linear frequency selectivity of human hearing, we propose three key innovations. First, a Gammatone-Mel Hybrid Filterbank is designed to mimic cochlear resolution and enrich high-frequency spectral details. Second, a Channel-wise Adaptive Normalization module harmonizes the heterogeneous statistical distributions between filter domains. Third, an auditory-prior supervision scheme provides targeted perceptual guidance during training. Experimental results on LJSpeech and VCTK datasets demonstrate that GAF-Flow significantly outperforms state-of-the-art baselines in both subjective (MOS, SMOS) and objective metrics (MCD, PESQ), successfully restoring fine spectral details and reducing artifacts while maintaining efficient inference. Our work bridges the gap between physiological auditory mechanisms and generative speech modeling, providing a perceptually-enhanced paradigm for high-quality synthesis. TPRformer: A Sea Surface Temperature Prediction Method Based on Temporal Periodic Feature Mining and Reconstruction Transformer Jiabao Zhang and Qian Li (the College of Meteorology and Oceanography in National University of Defense Technology, the High Impact Weather Key Labora- tory of CMA); Bing Sui (the Institute of Meteorological Sciences of Hunan Province, the Dongting Lake National climatological Observatory); and Yangweng Wang and Sheng Li (the College of Meteorology and Oceanography in National University of Defense Technology, the High Impact Weather Key Labora- tory of CMA) Abstract Abstract Sea Surface Temperature (SST) is a key factor in ocean-atmosphere interactions, which exerts significant influence on climate, ecology and maritime activities. Existing methods are limited to the modeling and optimization of long-term temporal dependence, and fail to adequately capture the complex temporal patterns in SST sequences, which are formed by the combined effects of long-term periodic trends and short-term non-stationary variations, resulting in suboptimal performance for long-term SST prediction tasks. To address this issue, we propose a temporal periodic feature mining and reconstruction transformer (TPRformer). In TPRformer, the encoder employs a temporal decomposition block to separate the historical sequence into periodic and non-stationary components, with a focus on mining long-term temporal dependencies within the periodic component to construct periodic features. The decoder performs the temporal decomposition on the SST sequence, employs temporal cross-attention to dynamically align historical periodic features, and ultimately achieves SST prediction by reconstructing the two decomposed components. Experimental results on multiple SST datasets demonstrate that TPRformer is superior to state-of-the-art baseline methods across multiple forecasting horizons. Geometry-Aware Learnable Gabor Networks for High-Fidelity Unsigned Distance Field Reconstruction Xiaohui Jia, Yuan Zhang, Caiqin Jia, Xiaowen Yang, and Min Pang (North University of China) Abstract Abstract Unsigned Distance Functions (UDFs) provide flexible representations for non-watertight geometries but struggle with a critical spectral mismatch. Since 3D surfaces exhibit non-stationary frequency distributions, demanding high frequencies for sharp edges but low frequencies for flat regions, traditional global encodings are inherently inefficient. They typically enforce excessive bandwidth globally to resolve local details, inevitably inducing computational redundancy and ripple artifacts in smooth areas. To address this, we propose a geometry-aware Learnable Gabor Network. First, leveraging the time-frequency locality of Gabor wavelets, our network adaptively modulates its spectral bandwidth, concentrating representational capacity on geometric details while preventing noise in flat regions. Second, to mitigate the gradient ambiguity of UDFs, we introduce a Dual-Query Geometric Consistency mechanism. By imposing local differential constraints, we regularize the gradient field direction without relying on ground-truth normal supervision. Experiments demonstrate that our approach improves upon state-of-the-art methods, achieving a superior balance between reconstruction fidelity and efficiency. Monday Virtual Room 8 IJCNN Paper SS21 Novel Networks in Human-Machine Collaboration: Paradigms, Methods, and Applications Session Chair: Haojie Luo (Fudan University), Lian Guo (Huazhong Agricultural University) Agile Imitation: Efficient Robot Skill Learning via Iterative Human Feedback and Data Aggregation Chenjun Liu, Haojie Luo, and Wei Li (Fudan University) Abstract Abstract Interactive Imitation Learning (IIL) is a promising paradigm for the online optimization of robotic skills, enabling agents to learn from sparse expert demonstrations. However, in few-shot scenarios, effectively combining human feedback with data aggregation to mitigate distributional shift remains a significant challenge. While previous works explored this integration, systematic research into internal mechanisms—such as dynamic feedback weighting, delay compensation, and lightweight trajectory resampling—remains insufficient under extreme data constraints. This paper proposes the Agile Iterative Aggregation (AIA) framework, which systematically optimizes these core iterative feedback and data processing mechanisms using a human-in-the-loop approach. We evaluate AIA on six simulated manipulation tasks and further validate it on a real-world platform. The results show that our optimized processing and aggregation strategies significantly enhance policy performance and robustness. Compared to baselines, AIA achieves superior success rates and finer sub-skill acquisition, offering a cost-effective solution for high-efficiency robotic skill learning. ComGrasp: Component-Aware Whole-Body Grasp Motion Generation via Autoregressive Model Lin Zhou, Binghui Zuo, and Yangang Wang (Southeast University) Abstract Abstract Generating long, realistic, and semantically aligned full-body human--object interactions from natural language remains challenging due to complex hand--object contacts, long-range temporal dependencies, and multi-part coordination. Existing methods often struggle to jointly model body motion, hand grasping, and object dynamics. To address these challenges, we propose a language-guided framework for full-body grasp motion generation with explicit contact awareness and part-wise coordination. We introduce a unified representation that encodes SMPL-X body motion, object motion, contact labels, and relative spatial features, and decompose interaction sequences into four semantic components that are discretized using component-specific VQ-VAEs. A masked autoregressive Transformer with a part coordination mechanism is then employed to generate coherent motion tokens conditioned on text. Experiments on the GRAB dataset demonstrate that our method generates longer and more coordinated interaction sequences, outperforming existing approaches. The code will be made publicly available upon acceptance. DualMixFormer: A Dual-Mixing Transformer for Long-Term Multivariate Time Series Forecasting Ruiqi Liu and Dingju Zhu (South China Normal University) Abstract Abstract Long-term multivariate time series forecasting (LTSF) demands accurate long-horizon prediction while modeling cross-variable dependencies. Most Transformer-based methods focus attention along the temporal axis, underutilizing inter-variable interactions and suffering from error accumulation over long horizons. We propose DualMixFormer, a variable-centric framework that treats each variable as a token and applies self-attention across variables. To enhance token representations amid heterogeneous periodicity, we introduce MEVA, a multi-embedding module that injects channel identity and phase embeddings to guide cross-variable attention explicitly. For robustness, DualMixFormer employs a dual-branch head: a lightweight MLP branch provides a low-variance continuation bias for smooth components, while a MEVA-conditioned Transformer branch models nonlinear residuals. Predictions are fused via learnable channel-wise weights, adaptively combining smooth and residual components. RevIN is integrated to address instance-wise distribution shifts. Experiments on seven benchmarks show competitive performance across horizons. Ablations and visualizations confirm the model’s effectiveness, interpretability, and robustness. A Siamese Network with Difference-Aware Residual Attention for Fine-Grained Cattle Face Recognition Lian Guo, Sheng Li, Guoyuan Zhou, Jiawei Li, and Guoliang Li (Huazhong Agricultural University) Abstract Abstract Non-contact cattle face recognition is the prerequisite for precise livestock management. However, modern intensive farming systems predominantly employ frozen semen technology for reproduction, which leads to highly similar genetic characteristics among individuals (particularly half-siblings sharing the same paternal lineage), presenting a significant challenge for individual identification. To address this issue, this paper proposes the DiffSiam network, which incorporates a Difference-Aware Residual Attention Module (DA-RAM). This approach models the recognition task as a fine-grained similarity measurement for image pairs. Specifically, the Swin Transformer V2 is employed as the core architecture for deep feature extraction. Subsequently, the DA-RAM module is utilized to explicitly compute feature difference maps, which guide the network to focus on highly distinctive local discrepancies, such as facial texture and contours, facilitating identity determination based on pairwise similarity scores. Additionally, the P-K sampling strategy combined with a weighted binary cross-entropy loss is employed to optimize the distribution of positive and negative samples. Experimental results on a dataset comprising 198 categories and 27,118 genetically similar cattle faces demonstrate that the proposed method achieved 98.09% Top-1 accuracy and 98.12% F1 score, outperforming existing mainstream approaches. The ablation experiments further confirmed that the DA-RAM module delivered a 15.04% performance improvement, which demonstrates its efficacy for fine-grained recognition. Monday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V5 Session Chair: Sunith Bandaru (University of Skövde) Embedded 2-opt for LNS: A Process-Aware Window Mechanism Approach Donghan Wu and Jialong Shi (Xi'an Jiaotong University) Abstract Abstract Large Neighborhood Search (LNS) and 2-opt local search are key algorithms for routing problems like TSP and CVRP. Standard hybrid frameworks usually separate these steps, using LNS for global exploration and 2-opt for later optimization. This order often leads to fault propagation, where suboptimal edges from repair are only fixed after costly global scans. To address this limitation, we investigate a \textit{Process-Aware Window Mechanism} for 2-opt based LNS. Unlike other granular search methods that use fixed geometric metrics, our method uses the dynamic information from the repair operator. We set an optimization 'window' around the insertion index of new nodes. This change makes 2-opt a proactive, built-in correction tool instead of just a post-processing step. We test six window selection strategies, from static to adaptive. Our experimental analysis shows that this frequent, low-complexity micro-optimization can improve convergence speed and result stability compared to traditional global search methods. A Non-Reductionist, Homeodynamic Simulator Of Ancient-Medicine, Inspired-Artificial Immune Systems For Emergent Intelligence Analysis Joshika Singh (Independent Research) Abstract Abstract Abstract— This work presents a mathematically grounded, non-reductionist framework for simulating emergent dynamics in a canonical 10-dimensional ancient medicine-artificial immune inspired system . The system is realized as a stochastic dynamical network with state-dependent drift, adaptive control operators, and hysteresis-based memory, enabling trajectory-level exploration of homeodynamic plasticity. Interventions are modeled as smooth, time-gated operators deforming the state-space geometry without imposing fixed-point constraints. Simulations demonstrate separable pathological regimes, regime-specific stability, and adaptive recovery under structured interventions, illustrating the system’s intrinsic self-regulation, non-commutativity, and memory effects. This framework unifies emergent intelligence analysis, trajectory-based disease modeling, and intervention dynamics in a fully differentiable, reproducible computational pipeline, providing a platform for exploring complex adaptive systems beyond traditional reductionist approaches A Dual-Layer Clustering Guided Genetic Algorithm for Multiobjective Multitype Satellite Observation Scheduling Yong-Chao Chen, Xiao-Fang Liu, Zhi-Hui Zhan, and Jun Zhang (Nankai University) Abstract Abstract Earth observation satellites have been widely applied and become indispensable in various domains. The collaboration of optical, synthetic aperture radar, and electromagnetic satellites enables diverse mission requirements through complementary capabilities. However, scheduling observation tasks across multiple satellite types to satisfy diverse objectives still remains a critical challenge, especially on large-scale instances. This paper focuses on two optimization objectives, i.e., maximizing the total profit of scheduled tasks and maximizing the average profit efficiency across satellites. To solve this problem, a dual-layer clustering guided genetic algorithm is proposed to obtain high-quality and well-distributed solution sets. The algorithm decouples the scheduling process into two layers, i.e., task sequence construction and time window assignment, and adopts dual sequence encoding. The two layers dynamically switch based on the performance improvement to generate new task sequences or new time-window assignments. At both of task and time-window layers, clustering is performed to guide the crossover and mutation for improving profits and resource utilization based on task characteristics and historical selection frequencies. In addition, a time-window allocation strategy is adopted in both layers to improve solution quality by utilizing historical information and current resource states. Experimental results on fifteen instances with up to 1300 tasks demonstrate that the proposed algorithm outperforms state-of-the-art algorithms in terms of convergence and diversity. Multi-Objective Optimization and Decision Support for Conformal Cooling in Injection Molding António Gaspar-Cunha, João Melo, Tomás Marques, and António Pontes (IPC-Institute for Polymers and Composites, University of Minho) and Sunith Bandaru (School of Engineering Science, University of Skövde) Abstract Abstract The design of conformal cooling channels (CCC) in injection molding is a complex, multi-objective engineering task in which geometric, thermal, and mechanical constraints interact in nonlinear ways. Although evolutionary multi-objective optimization (EMO) algorithms, such as NSGA-II, have been successfully applied to this domain, the interpretation of large Pareto-optimal sets often relies on manual visual inspection, which limits knowledge transfer and decision reproducibility. This paper examines how interactive knowledge discovery, facilitated by the Mimer platform, can enable the systematic interpretation of EMO results in the injection molding process. Four bi-objective CCC optimization problems were solved using high-fidelity simulations accelerated by an Artificial Neural Network surrogate model based on Moldex3D data. Mimer was then used to analyze the resulting solution sets through linked visualizations, region-of-interest-based preference articulation, and flexible pattern mining for rule extraction. The generated knowledge was compared against the independent assessment of an expert in mold design. Results show that Mimer consistently complements expert insights, identifies the most influential geometric variables, and uncovers additional non-obvious relations across runs, including stable structural rules governing channel diameter, distance to the cavity surface, and gate-dependent performance. These findings demonstrate that Mimer can act as a semi-automatic analyst in EMO-based CCC design, increasing transparency, supporting informed decision-making, and improving the reusability of optimization knowledge in industrial injection molding. Tuesday Virtual Room 1 IJCNN Paper Graph Neural Networks I Session Chair: Xiao Yue (Oakland University), Qian Tao (South China University of Technology) Diversity-Aware Forward Prototype Learning of Graph Neural Networks for Node Classification Enze Zhang and Qian Tao (South China University of Technology) Abstract Abstract Graph neural networks (GNNs) have achieved great success in many applications. Most existing GNNs rely on backpropagation for end-to-end training. Recently, forward-only learning methods have been considered as promising alternatives, but directly applying them to GNNs faces challenges in efficiency and training stability. In this paper, we propose the Prototype-based Forward algorithm for Graph Neural Networks (ProF-GNN), a forward-style GNN learning framework that introduces class-aware virtual nodes to explicitly model class prototypes. By aligning node representations with their corresponding prototypes at each layer, ProF-GNN enables direct and efficient prediction without exhaustive label-wise inference or the generation of multiple negative samples. Furthermore, we incorporate a diversity-aware objective to prevent prototype collapse, ensuring stable and discriminative layer-wise learning. Extensive experiments on real-world datasets demonstrate the effectiveness of the proposed learning framework for GNNs for node classification. Graph Learning with Consistency and Diversity against Label Noise Jie Zhang (Shandong University of Science and Technology) Abstract Abstract Graph Neural Networks (GNNs) have shown remarkable effectiveness in learning graph representations across diverse applications. However, in many real-world scenarios, graph-structured data often contains noisy labels, which substantially compromise the performance of GNNs. This degradation occurs because noisy labels weaken effective supervision during training, causing GNN-based models to memorize mislabeled samples and consequently impair their performance on downstream tasks. Existing approaches for handling label noise in images and text typically rely on large-scale labeled datasets and assume data independence, making them unsuitable for addressing the label sparsity and structural dependencies inherent in graphs. To overcome these challenges, we propose a multi-network learning paradigm in which models supervise each other, thereby reducing the bias accumulated when learning from noisy samples. We introduce a novel multi-network framework, Joint Training with Consistency and Diversity (JoCaD), which is designed to maximize prediction consistency across networks while maintaining sufficient diversity in their representation learning. Specifically, we incorporate graph augmentation through link addition to enrich the aggregated information, and employ a consistency loss to ensure alignment between the node representations learned by both networks. Extensive experiments on four real-world datasets demonstrate the effectiveness of the proposed JoCaD in addressing the node classification task under varying levels of label noise. Fusion-Rewired One Class Graph Autoencoder Marcos Gôlo (University of São Paulo), Edgard Marx (Byondis), João Gama (University of Porto), and Ricardo Marcacini (University of São Paulo) Abstract Abstract Heterogeneous graphs provide a representation for real-world systems that involve multiple entity types and relations. In many of these applications, the practical goal is to detect one class of interest while treating the remaining instances as outliers, which naturally leads to one-class learning on graph neural networks (GNNs). However, most one-class GNN approaches either rely on unsupervised methods, decouple representation learning from the one-class objective in two-step pipelines, or focus on homogeneous graphs and do not explicitly leverage heterogeneity in generic settings. We propose FOLGA (Fusion-rewired One-cLass Graph Autoencoder), an end-to-end method for one-class node classification on heterogeneous graphs. FOLGA enriches heterogeneous structure through complementary rewiring strategies, fuses information from multiple rewired graph views via a fusion GNN encoder, and combines graph reconstruction with a state-of-the-art sphere loss to learn interpretable 3D embeddings. We also show that FOLGA naturally extends to dynamic data streams for evolving dynamic heterogeneous graphs in online learning. Experiments with 12 representative baselines in three datasets demonstrate that FOLGA outperforms them, while providing real-time interpretability of the learning dynamics and competitive results. GIN-Graph: A Generative Interpretation Network for Model-Level Explanation of Graph Neural Networks Xiao Yue and Guangzhi Qu (Oakland University) Abstract Abstract One significant challenge of exploiting Graph neural networks (GNNs) in real-life scenarios is that they are treated as black boxes, therefore leading to the requirement of interpretability. To address this, model-level interpretation methods have been developed to explain what patterns maximize probability of predicting to a certain class. However, existing model-level interpretation methods pose several limitations such as generating invalid explanation graphs and lacking reliability. In this paper, we propose a new Generative Interpretation Network for Model-Level Explanation of Graph Neural Networks (GIN-Graph), to generate reliable and high-quality model-level explanation graphs. The implicit and likelihood-free generative adversarial networks are exploited to construct the explanation graphs which are similar to original graphs, meanwhile maximizing the prediction probability for a certain class by adopting a novel objective function for generator with dynamic loss weight scheme. Experimental results demonstrate that GIN-Graph consistently generates high-quality model-level explanation graphs with high stability and reliability across diverse graph datasets. Tuesday Virtual Room 2 IJCNN Paper Graph Neural Networks II Session Chair: Ningyun Chen (BNBU, Zuse School ELIZA), Zhongming Mei (Donghua university) UNGIA: Unnoticeable Injection Attack on Graph Neural Networks Zixuan Wang (HKUST-GZ, BNBU); Ningyun Chen (BNBU, Zuse School ELIZA); Donglong Chen (BNBU); and Leo Yu Zhang (Griffith University) Abstract Abstract Graph neural networks excel in classification and prediction tasks but are highly susceptible to adversarial attacks. Among attack strategies, graph injection attacks are more practical than graph modification attacks as they avoid altering the original graph structure. However, existing GIA methods often sacrifice stealth for effectiveness, making fake nodes easily detectable and reducing their impact. To address this, we propose UNGIA, a steganographic GIA method. UNGIA employs clustering to identify anchor nodes and an adaptive generator to learn the feature distribution of original nodes, enabling the creation of indistinguishable fake nodes. Potential links between nodes are established using a mask learning mechanism to ensure high homogeneity. A two-layer optimization framework balances stealth and attack effectiveness. Experiments on Cora, OGB-arxiv, and Reddit datasets demonstrate that UNGIA is covert, transferable, and outperforms state-of-the-art methods in both effectiveness and adaptability. Cross-Stage Reward-Aware Node Injection Attacks against Graph Neural Networks for Fraud Detection Weiqi Kang, Linghao Ying, and Li Han (East China Normal University) Abstract Abstract Graph Neural Networks (GNNs) have become an effective tool for financial fraud detection. However, beyond continuously evolving fraud patterns in real transaction systems, GNN-based fraud detectors are increasingly exposed to adversar- ial attacks, where attackers deliberately manipulate transaction graphs to evade detection. In realistic financial scenarios, ex- isting GNN attack methods are often ineffective, due to strict constraints such as limited access to model parameters and training data, dynamic transaction networks, extreme sparsity of fraud behaviors and severe class imbalance. To address these challenges in financial scenarios, we propose Cross-stage Reward- aware Node Injection (CRNI), a black-box node injection attack for realistic financial scenarios. Under the condition that only partial graph features, limited local structure information, and model outputs can be accessed, CRNI formulates node injection as a cross-stage optimization problem with stage-specific learning signals. To enhance the concealment and effectiveness of the attack, CRNI jointly optimizes feature generation and structural connection strategies through reward modeling and topological consistency constraints, while incorporating stable training and sample reuse mechanisms to alleviate the problems of strategy fluctuations and data imbalance. Experimental results show that CRNI achieves better performance than the state-of-the- art methods on multiple real financial fraud datasets and GNN models. LPSA: Leaf-Prior Structural Attack on Graph Neural Networks at Scale Yu Ran, Chenbo Ma, Shiqi Liu, Zihan Chen, Yi Pan, Xin Zhang, Maoyi Xiong, and Wentao Zhao (National University of Defense Technology) Abstract Abstract Adversarial attacks on Graph Neural Networks (GNNs) have exposed critical robustness vulnerabilities in graph-based learning. However, existing attack methods face a fundamental trade-off between effectiveness and scalability: exhaustive greedy attacks achieve strong effectiveness on small graphs but become computationally intractable at scale, while scalable optimization-based attacks suffer from slow convergence, requiring excessive computational overhead to achieve effective perturbations. To bridge this gap, we propose Leaf-Prior Structural Attack (LPSA), a simple yet efficient and effective structural attack for large-scale graphs. LPSA is built upon two key insights. First, structural asymmetry in graphs makes leaf nodes highly effective adversarial injection points, enabling strong influence on the target node's prediction with minimal structural modifications. Leveraging this, LPSA restricts the perturbation search space to a compact leaf-prior subspace, dramatically reducing the computational overhead of structural attacks. Second, inductive biases and gradient misalignment in surrogate models degrade attack transferability. To address this issue, we introduce a manifold-smoothed surrogate fine-tuning strategy that regularizes the surrogate’s gradient field to align with the intrinsic data distribution, thereby obtaining more transferable adversarial gradients. By combining leaf-prior search with single-step gradient-based perturbation generation, LPSA eliminates costly iterative optimization while maintaining strong attack effectiveness. Extensive experiments on four benchmark datasets, including the million-scale MAG dataset, demonstrate that LPSA not only achieves competitive or superior attack success rates but also reduces runtime by over 80\% compared to state-of-the-art baselines, particularly on large-scale graphs. The source code is available at https://github.com/rannyu/LPSA. LDF-GNN: Alleviating Structural Misguidance in Graph Neural Networks via Local Dynamic Fusion Filter Zhongming Mei (Donghua University), Rui Duan (Guangzhou University), and Wenzhuo Fan and Mingjian Guang (Donghua University) Abstract Abstract Graph neural networks (GNNs) have emerged as a fundamental framework for graph representation learning. A major research direction in this field is spectral graph filters, which process different frequency components of graph signals through polynomial approximations. However, most existing studies primarily focus on optimizing the design of the graph filters while ignoring the interaction between input signals and the filtering process and applying the same filtering mechanism for all regions of the graph. This uniform approach may lead to the structural misguidance problem, making the filter susceptible to regional variations or heterophilic connections. To address this issue, we propose LDF-GNN, a spectral GNN based on the \textbf{l}ocal \textbf{d}ynamic \textbf{f}usion (LDF) filter. Specifically, we first generate localized graph signals using a random walk-based method to ensure adaptive filtering across different regions. Then we employ the \textbf{d}ynamic \textbf{f}usion (DF) filter which dynamically adjusts its filtering mechanism in response to the frequency characteristics of each localized signal. By dynamically integrating input signals with the filtering process, LDF-GNN can capture both local discrepancies and structural heterophily to alleviate the structural misguidance problem. Experimental results demonstrate that LDF-GNN achieves state-of-the-art performance, and ablation studies further validate the importance of its two dynamic properties in mitigating structural misguidance. Tuesday Virtual Room 3 IJCNN Paper Image Restoration and Enhancement I Session Chair: Zhicheng Qian (Fuyang Normal University), Jinao Li (Qilu University of Technology) DCAAFusion: A Diffusion Model-Based Cross-Attention Adaptive Fusion Network for Nighttime Infrared and Visible Image Fusion Jinao Li, Jinyong Chen, and Kening Cui (Qilu University of Technology) Abstract Abstract Infrared and visible image fusion seeks to combine complementary information from heterogeneous sensors into a unified composite image, thus providing a more comprehensive scene representation. However, fusion in nighttime low-light degradation scenarios remains a significant challenge. Low illumination weakens the brightness and details of visible images, leading to the loss of critical information for effective fusion. Furthermore, existing fusion approaches lack adaptive modeling mechanisms that fully leverage the complementary information from both modalities, often resulting in blurred salient features and the loss of fine textures in the fused image. To address these challenges, we propose a Diffusion Model-Based Cross-Attention Adaptive Fusion Network (DCAAFusion). Specifically, we employ the denoising network of a pre-trained conditional diffusion model to extract visible diffusion features, and that of a pre-trained diffusion model for infrared diffusion features. Subsequently, we design a Degradation Embedding Modulation Module (DEMM), which modulates the visible diffusion features using the degradation embedding extracted by an introduced degradation-aware encoder, compensating for the loss of brightness and details caused by low light at the feature level. Finally, we propose a Cross-Attention Adaptive Fusion Module (CAAFM), which employs multiple attention mechanisms to adaptively fuse infrared structural information and visible textural details, fully exploiting complementary information from both modalities. Extensive experiments demonstrate that DCAAFusion outperforms current state-of-the-art methods, achieving superior performance in both visual quality and quantitative metrics. CDL-FusionNet: A Two-Stage Illumination-Invariant and Luminance-Guided Framework for Infrared and Visible Image Fusion Wang Linyang, Wang Lei, and Wang Hengyang (Heilongjiang University) Abstract Abstract Infrared and visible image fusion (IVIF) plays a pivotal role in enhancing scene perception under challenging conditions such as nighttime surveillance and autonomous driving. However, existing deep learning-based approaches often suffer from significant performance degradation in low-light environments. Specifically, direct fusion methods tend to introduce noise and artifacts from visible images, while the prevalent decompose-then-fuse paradigm struggles with imperfect decoupling of illumination and reflectance information. This frequently results in inadequate exploitation of illumination cues, leading to fused images with degraded textures and attenuated target features. To address these limitations, this paper proposes a two-stage fusion framework termed CDL-FusionNet. In the first stage, we develop an illumination-invariant decoupling network incorporating a self-supervised mechanism that enforces output consistency under varying lighting conditions. This enables the extraction of a physically consistent reflectance component and effectively suppresses interference from noise and artifacts. In the second stage, a brightness-guided complementary-enhanced cross-attention module is introduced. This module leverages the illumination map from the first stage as a pixel-level prior to dynamically modulate fusion weights, thereby preserving fine visible details in well-lit regions while enhancing infrared thermal features in dark areas. Extensive qualitative and quantitative evaluations on multiple public datasets demonstrate the superiority of CDL-FusionNet over state-of-the-art methods, particularly in low-light conditions. The proposed framework yields fused images with superior visual quality, higher contrast, and enhanced detail preservation, validating its robustness and effectiveness for dual-modal information integration. Unified Text-Guided Degradation-Aware Fusion of Infrared and Visible Images Yueqiang Zhao, Yijun Lin, and Xiongxin Tang (Institute of Software, Chinese Academy of Sciences; Chinese Academy of Sciences) Abstract Abstract Existing image fusion methods often perform poorly in real-world scenarios with complex degradations due to a lack of understanding of diverse degradation characteristics and the absence of relevant datasets. To address this issue, we propose a unified text-guided degradation-aware fusion framework that introduces textual prompts to adaptively mitigate multiple degradations. A two-stage training strategy is adopted, together with a Cross-Prompt Attention Aggregation Module and a JointGate module, to enable cross-degradation knowledge sharing. In addition, we construct an infrared-visible fusion dataset that covers various representative degradation scenarios, providing a solid benchmark for evaluating fusion performance under complex conditions and facilitating future research on degradation-aware image fusion. Extensive experiments and ablation studies demonstrate that the proposed method achieves superior performance in both single and compound degradation cases, showing stronger generalization capability. Our code will be available at https://github.com/zhaoyueqiang/UTDFusion.git LiteFusion: A Lightweight Multimodal Feature Fusion Model for Adverse Driving Conditions Zhicheng Qian, Hongzhi Zhou, and Nina Ling (FuYang Normal University) Abstract Abstract With the rapid advancement of intelligent driving, object detection technology faces robustness challenges in complex environments, particularly the loss of RGB image information caused by low-light and adverse weather conditions. To address this, we propose the LiteFusion model—a novel lightweight multimodal object detection framework based on mid-term feature fusion. It effectively combines visible and infrared modalities to enhance detection robustness in complex driving scenarios. First, the model introduces the lightweight LFNet architecture, employing a dual-branch heterogeneous and hierarchical gated fusion strategy to achieve refined integration of visible and infrared features. To further improve localization accuracy, the WIoU v3 loss is integrated into the detection head. Its dynamic non-monotonic focusing mechanism emphasizes high-quality samples, enhancing bounding box precision and convergence speed, thereby boosting detection robustness. Finally, the hybrid convolution module C3k2\underline{ }SCConv is introduced. By embedding spatial and channel reconstruction convolutions into the bottleneck structure, it enables adaptive re-calibration of features across both spatial and channel dimensions, mutually enhancing convolution efficiency and fusion performance. Experimental results demonstrate that LiteFusion achieves lightweight detection while improving performance, validating its practical value in intelligent driving scenarios. Tuesday Virtual Room 4 IJCNN Paper Image Restoration and Enhancement II Session Chair: Chuancheng Fu (Wuhan University of Science and Technology), Xuanchao Lin (Shanghai University) SCI-Net: Semantic-Guided Color and Intensity Decoupling Network for Low-Light Image Enhancement Chuancheng Fu, Zhaokun He, and Li Chen (Wuhan University of Science and Technology) Abstract Abstract Low-light image enhancement (LLIE) is vital for vision system reliability. However, most existing methods remain "semantically agnostic" pixel-wise mappings that ignore scene content, leading to blurred edges and color deviations. While foundation models like SAM offer powerful semantic extraction, the low-light-induced domain shifts impair the reliability of captured priors. To address this, we propose the Semantic-Guided Color and Intensity Decoupling Network (SCI-Net), focusing on two key challenges: (1) acquiring reliable semantic priors from degraded images, and (2) efficiently incorporating these features into image restoration process. For the first challenge, we design a Residual Fourier Guidance Module (RFGM) that generates high-quality pseudo-enhanced images via phase structural compensation and amplitude refinement, ensuring reliable inputs for semantic extraction. For the second challenge, we introduce the Cross-Modal Guided Attention (CMGA) module. By employing a cross-modal interaction with semantic features as the Key, CMGA precisely injects guidance into the decoupled HVI branches (HV and I), effectively leveraging the inherent color-intensity separation of the HVI color space. Extensive experiments on eight datasets demonstrate that SCI-Net outperforms existing state-of-the-art methods in contrast enhancement, detail recovery, and semantic consistency. Learning to Enhance in One Step: Average Residual Flow for Efficient Low-light Image Enhancement Weixuan Huang, Zhenghan Zhang, and Yiyang Li (Tongji University) Abstract Abstract While diffusion and flow models exhibit strong generative ability for low-light image enhancement, they commonly suffer from trajectory drift accumulation: color bias builds up along curved sampling trajectories from Gaussian noise, and iterative sampling further amplifies this drift, causing color distortion and slow inference that hinder real-time perception in nighttime autonomous driving. Experiments show that the recursive updates in conventional flow models introduce redundant computation and, more critically, break the direct mapping from low-light to normal-light domains. To address these problems, we propose the Average Residual Flow (ARF), which redefines the generative process as a direct mapping in the residual space. By averaging multi-step residual fields, ARF eliminates the Markov chain dependency and compresses the original multi-step iteration into a one-step transformation. To better approximate average residuals, we design a lightweight Conditioned Illumination-Temporal Enhancement Network (CITE-Net), which explicitly models the joint representation of temporal evolution and illumination disparity. Experiments show that the joint design of ARF and CITE-Net achieves state-of-the-art performance across six benchmarks from general and real-world driving datasets, delivering 2–5× speedup over comparable methods, validating its effectiveness and practicality. PhyIC-Net: A Robust Unsupervised Framework for Underwater Image Restoration via Intrinsic Consistency Hongxun Gao (Jiangnan University); Gaoe Qin (Shanghai Jiao Tong University); and Tao Zhang (Jiangnan University, Central South University) Abstract Abstract Underwater image restoration remains a challenging task due to complex optical degradation and the scarcity of paired training data. Existing learning-based methods often rely on paired training data or simplified physical models, which limits their robustness and generalization in real-world scenarios. To address these issues, this paper proposes PhyIC-Net, a robust unsupervised framework that effectively bridges physical modeling with deep learning. To mitigate the overfitting issue common in unsupervised training, we design a Spatial-Channel Decoupled Network (SCD-Net) as the backbone, which explicitly disentangles spatial blurring from color attenuation using depthwise separable convolutions. Furthermore, to constrain the ill-posed solution space of the physical model inversion, we introduce a Physics-Aware Perturbation (PAP) strategy. This strategy simulates diverse water turbidity conditions to enforce an intrinsic consistency constraint, compelling the network to learn invariant scene representations without relying on ground truth data. Finally, a Multi-Prior Guided (MPG) loss is incorporated to regulate color balance and structural smoothness. Extensive experiments on multiple underwater image datasets demonstrate that PhyIC-Net consistently outperforms existing unsupervised and weakly supervised methods in terms of quantitative metrics and visual quality, exhibiting strong generalization across diverse underwater conditions. BiMIN: Bilateral Modality Imagination Network for Sentiment Analysis under Modality Uncertainty Xuanchao Lin and Junjie Peng (Shanghai University) Abstract Abstract Human communication relies on the synergy of textual, acoustic, and visual cues to convey sentiment. While Multimodal Sentiment Analysis (MSA) has made significant strides in modeling these interactions, most existing models are predicated on the idealized assumption of complete data availability. In real-world scenarios, however, one or more modalities may be entirely absent, causing standard models to suffer severe performance degradation. To address this challenge, we propose the Bilateral Modality Imagination Network (BiMIN), a unified framework designed for robust sentiment analysis under modality uncertainty. BiMIN overcomes the limitations of existing approaches through two core innovations. First, we design a novel Bilateral Auto-Encoder featuring parallel Core and Detail branches. This structure effectively disentangles global semantics from fine-grained details, ensuring robust feature preservation. Second, to uniformly handle diverse absence scenarios, we incorporate an Absence-Gating Mechanism within a cascaded architecture designed for robust iterative imputation. This mechanism dynamically modulates the integration of original and imagined features based on real-time modality availability, preventing the hallucination of noise. Experimental results on the CMU-MOSI, CMU-MOSEI, and CH-SIMS datasets demonstrate that BiMIN achieves superior stability and the highest average accuracy. Specifically, it delivers remarkable performance gains, exceeding baselines by 3.36% on MOSEI, 1.04% on MOSI, and 4.97% on SIMS, while maintaining competitive proficiency even in full-modality settings. Tuesday Virtual Room 5 IJCNN Paper Image Segmentation and Dense Prediction I Session Chair: ziyang Tong (Wuhan University of Technology), Xiaohong Jia (Lanzhou Jiaotong University) A Semantic-Guided Multimodal Image Segmentation Framework with Graph-Structured Reasoning ziyang Tong and yixin Su (Wuhan University of Technology) Abstract Abstract Image segmentation driven by complex natural language instructions is a significant challenge in computer vision, with wide applications in scenarios such as street-scene analysis and autonomous driving. Existing methods often fall short in handling semantic ambiguity, modeling inter-object relationships, and achieving fine-grained segmentation. To address this, we propose a novel image segmentation framework that fuses multimodal semantic guidance with graph-structured reasoning. Our method first leverages the LLaMA3-V multimodal large language model to jointly encode the image and language instruction, generating a task-relevant semantic guidance vector. Subsequently, a cascade of GroundingDINO and SAM generates semantically-aware candidate masks, whose visual features are extracted using DINOv2.The core innovation is a graph-based mask selection module that models candidates as nodes and leverages a graph convolutional network to perform collaborative reasoning based on spatial and semantic relationships, significantly improving the differentiation of similar objects. To reduce training costs, we employ a parameter-efficient modular fine-tuning strategy, inserting a QLoRA structure into LLaMA 3-V and a lightweight Adapter into the selection module, with tunable parameters accounting for less than 1% of the total. Experiments on street-scene benchmarks confirm our method's superior performance in comprehending complex instructions, achieving fine-grained segmentation accuracy, and maintaining high computational efficiency. Two-stage Gradient Vectorization via SAM Guidance and Differentiable Rendering Anshu Hu, Guixiang Nie, and Qing Xie (Wuhan University of Technology, Engineering Research Center of Intelligent Service Technology for Digital Publishing); Yanchun Ma (Wuhan Vocational College of Software and Engineering, Hubei Engineering Research Center for Intelligent Detection and Identification of Complex Parts); and Jiachen Li and Jinyu Xu (Wuhan University of Technology, Engineering Research Center of Intelligent Service Technology for Digital Publishing) Abstract Abstract Image vectorization aims to convert raster images into editable and resolution-independent vector representations. While recent learning-based methods based on differentiable rendering achieve high automation, most rely on solid-color primitives and struggle to model smooth color gradients efficiently. Existing gradient-aware approaches improve gradient fitting but often suffer from limited structural robustness and high computational cost on complex images. MEI: Mutual-Enhanced Integration Between Pretrained SAM and Lightweight Model for Medical Image Segmentation Zhiyan Wang and Changjian Wang (College of Computer Science and Technology, National University of Defense Technology); Zhongshun Tang (Zhujiang Hospital of Southern Medical University); Maolin Luo (College of Computer Science and Technology, National University of Defense Technology); Daozhong Lei (Hunan College of Information); and Qian Deng (the 921st Hospital of the People's Liberation Army) Abstract Abstract The success of pretrained large models has provided new and important technical approaches for medical image segmentation (MedISeg). However, pretrained models also have two significant shortcomings: (i) high-cost fine-tuning is usually required when they are applied to specific tasks; (ii) it is difficult to incorporate new attributes not used during the pretraining process. To address these issues, we propose the Mutual-Enhanced Integration framework Between Pretrained SAM and Lightweight Model for MedISeg, named MEI. Firstly, Adaptive Prompt Enhancement (APE) module is proposed to optimize SAM-Med by enhancing prompt quality based on the inconsistency between SAM-Med and the lightweight model, thereby introducing new features to the pretrained SAM and improving its specificity on specific task datasets. Secondly, Ensemble Structure Boundary Enhancement (ESBE) module is designed to optimize the lightweight model by providing more accurate boundary information via the fusion of prediction results from both the pretrained SAM and the lightweight model, thereby enhancing the segmentation accuracy of the lightweight model. Extensive experiments on four public 2D/3D segmentation datasets demonstrate that our model achieves the state-of-the-art performance relative to a strong baseline. KAN-Based Superpixel Segmentation with Boundary Constraint and Semantic Guidance Xiaohong Jia, Fuhai Wang, Tong Tong, Long Ma, and Guanghui Yan (Lanzhou Jiaotong University) Abstract Abstract Superpixel segmentation serves as a critical computational primitive for downstream vision tasks by clustering perceptually similar pixels into compact regions. Existing methods based on Fully Convolutional Networks often grapple with an inherent trade-off between boundary adherence and semantic consistency, particularly when handling complex scenarios. To resolve this inherent trade-off, we present the KAN-Based Superpixel Network (KBSNet), a tailored framework constructed to enforce geometric boundary constraints and simultaneously establish robust deep semantic representations. Structurally, we construct a hybrid encoder that pioneers the integration of Kolmogorov-Arnold Networks into superpixel generation. To refine feature representation, we design a Geometric Boundary Constraint Module (GBCM) to explicitly model physical contours. We introduce a Semantic-Guided Local Attention (SGLA) mechanism, which leverages deep features as ``semantic agents'' to suppress texture noise via top-down modulation. Extensive experiments on four benchmark datasets demonstrate that KBSNet outperforms state-of-the-art methods across ASA, BR-BP, UE, and CO metrics. Tuesday Virtual Room 6 IJCNN Paper Knowledge Graphs and Question Answering I Session Chair: JunWei Yang (East China Normal University), Yi Zhou (Southwest University of Science and Technology) DSS: Dynamic Sparse Selection for Multi-modal Knowledge Graph Completion Yi Zhou and Qingming Zhang (Southwest University of Science and Technology) Abstract Abstract Multi-modal Knowledge Graph Completion (MMKGC) aims to uncover hidden knowledge by fusing multi-modal and structured information. Existing fusion methods use attention, weighting, or mixture-of-experts strategies. However, they treat all feature segments equally at the fine-grained level. This equal treatment prevents effective filtering of image background noise and textual redundancy. Consequently, discriminative features are weakened. To address this, this paper introduces a sparse mechanism into the model, proposing a Dynamic Sparse Selection (DSS) model. Its core is the Sparse-Global Fusion (S-GF) method, which utilizes Sparse Prompt Attention (SPA) and Global Dense Attention (GDA) in parallel. The SPA branch uses a lightweight controller to learn attention masks. It focuses on the most discriminative local segments. Meanwhile, the GDA branch captures global context to preserve semantic coherence. The two are adaptively fused via a gating mechanism. Additionally, we propose Sparse Gated Contrastive Learning (SGCL). DSS proposes the Sparse Gated Selection Mechanism (SGSM) to dynamically construct tailored anchor representations for entities. Optimization is performed using the Multi-view Focused Adaptive Contrastive Loss (MFAC Loss) with progressive focusing capability, thereby enhancing both the discriminative power and alignment robustness of entity representations. Experiments on DB15K, MKG-W, and MKG-Y show that DSS improves the MRR metric by 6.48%, 2.75%, and 0.97%, respectively, compared to existing MMKGC models, demonstrating the effectiveness of the introduced sparse mechanism in focusing on critical information and suppressing multi-modal noise. SAMEF-DHNN: Sample-Aware Multimodal Expert Fusion and Dynamic Hard Negatives Network for Multimodal Knowledge Graph Completion Yi Zhou and Qingming Zhang (Southwest University of Science and Technology) Abstract Abstract Multimodal Knowledge Graph Completion (MMKGC) leverages multimodal data such as images and texts to predict missing triples. Two major challenges are static multimodal fusion strategies overlook differences in modality reliability across samples, and contrastive learning balancing negative sample quality with computational efficiency. To address these issues, we propose a Sample-Aware Multimodal Expert Fusion and Dynamic Hard Negatives Network (SAMEF-DHNN). First, to address the limitation of static fusion, we propose the Sample-Aware Multimodal Expert Fusion Block. Inspired by Mixture of Experts, it treats each modality as an expert and employs a novel gating network to dynamically assign adaptive fusion weights for each input entity sample, thereby enhancing fusion robustness against variable modality quality. Second, we introduce the Dynamic Hard Negative Sampling Contrastive Learning. Instead of static or pre-generated negatives, Its core is a real-time mining strategy that, for each anchor, selects only the top-K most semantically similar samples in the current batch as hard negatives. This improves discriminative ability while maintaining efficiency. Experiments on DB15K and MKG-W demonstrate that SAMEF-DHNN outperforms 16 existing baselines across multiple metrics, validating its effectiveness. Query-Aware Hierarchical Contrastive Path Learning for Inductive Temporal Knowledge Graph Reasoning Lin Jie Shi, Zhi Xin Shi, Xiao Yu Kang, Rui He, and Shi Hao Zhao (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Abstract Abstract Extrapolation reasoning over Temporal Knowledge Graphs is crucial for forecasting future facts involving emerging entities in intelligent systems. Existing inductive path-based approaches are often limited by insufficient contextual adaptability and weak global logical consistency, mainly due to static attention mechanisms and local edge-level supervision. To address these limitations, we propose the Query-Aware Hierarchical Contrastive Path (QHCP) framework, a novel architecture that combines dynamic query-aware attention with hierarchical contrastive supervision for temporal knowledge graph reasoning. Specifically, QHCP dynamically prioritizes relevant evidence at each reasoning step by incorporating evolving path contexts, while hierarchical contrastive learning jointly regularizes message-level semantics and path-level logic to improve both local coherence and global consistency of reasoning chains. Extensive experiments on four real-world TKG datasets demonstrate that QHCP achieves significant improvements over state-of-the-art baselines, with gains of up to 2.62% in MRR on the WIKI dataset and retains over 96% of its performance when transferred to a dataset containing entirely unseen entities, validating its superior inductive generalization capability. Knowledge-based Graph Network with Semantic-Neighbor Items for Government Service Recommendation Junwei Yang, Zilong Wu, Yiming Yin, and Hongyan Mao (East China Normal University) Abstract Abstract Government service recommendation suffers from sparse user–item interactions. This data sparsity hinders the precise learning of embedding representation, leading to limited performance. Therefore, we construct a government service knowledge graph, which enriches the description of service items by introducing multi-dimensional entities. To enhance recommendation performance by semantically exploiting the auxiliary knowledge information, we propose a Relation-aware contrastive Knowledge-based graph network with Semantic-Neighbor Items (RKSNI). Firstly, semantic-neighbor items feature is proposed to directly capture valuable inter-item association. It guides the relation-attention graph convolution network in generating semantic-neighbor representation, while also serving as interpretable signal for contrastive learning. Secondly, holistic-structural representation is constructed through propagation over the collaborative knowledge graph, thereby capturing global information. To effectively integrate the two complementary representations, discriminative contrastive learning is designed to enrich the representation with interpretable difference, while alignment contrastive learning is applied to ensure fine-grained consistency between the two representations. Experiments on a real e-government service dataset and two common public datasets demonstrate that RKSNI outperforms state-of-the-art methods, verifying its effectiveness and generalization. These improvements arise from its focus on reducing noise in hidden relationship modeling and improving the semantic interpretation of contrastive learning signals. Tuesday Virtual Room 7 IJCNN Paper LLM Adaptation and Fine-Tuning I Session Chair: Yichen Liu (CASIA), Wentao Hu (University of Science and Technology of China) OWM-LoRA: Projecting LoRA Gradients onto Orthogonal Subspaces Yichen Liu, Liangxuan Guo, Yuming Dai, Pengcheng Pan, Canyang Canyang, Ziheng Li, and Shan Yu (CASIA) Abstract Abstract Large language models (LLMs) have achieved state‑of‑the‑art results across a wide range of NLP tasks, yet they remain prone to catastrophic forgetting when fine‑tuned sequentially on multiple tasks. In this paper, we introduce Orthogonal Weight Modification Low‑Rank Adaptation (OWM‑LoRA), a novel continual learning method that integrates low‑rank parameter updates with orthogonal gradient projection. By restricting Low-Rank Adaptation (LoRA) updates to the subspace orthogonal to the gradient directions of previously learned tasks, OWM‑LoRA enables the model to acquire new knowledge without perturbing knowledges associated with earlier tasks. Our method retains the parameter efficiency and flexibility of LoRA. We conducted extensive experiments on standard continual‑learning benchmarks using various LLMs and demonstrated that OWM‑LoRA consistently outperforms existing approaches, significant gains in end‑task performance. Furthermore, OWM‑LoRA scales gracefully to models with billions of parameters and requires neither storage of historical gradients nor retention of past data, thereby preserving data privacy and computational practicality. These results establish OWM‑LoRA as a simple and efficient approach for enabling robust continual learning in LLMs. Robust-LoRA: Confidence-Aware Fine-Tuning with Knowledge Distillation Hao Wu, Xiangfeng Luo, and Jianqi Gao (Shanghai University) Abstract Abstract Low-Rank Adaptation (LoRA) efficiently adapts pre-trained models to downstream tasks by freezing the original parameters and introducing low-rank decomposable trainable matrices. Although this approach significantly reduces computational costs, it suffers from limited performance gains and poor generalization on downstream tasks. To address this challenge, inspired by knowledge distillation that mitigates overfitting risks through soft labels conveying inter-class relationships, we propose Robust-LoRA, a robust low-rank adaptation framework that integrates knowledge distillation with margin-based constraints. Specifically, during fine-tuning, Robust-LoRA leverages knowledge distillation to enable the trainable LoRA components to retain the model's general representations, while simultaneously employing a margin loss function to widen the probability gap between correct and incorrect predictions, thereby stabilizing the decision boundary. Experimental results across multiple pre-trained backbone models and benchmark datasets demonstrate that Robust-LoRA consistently outperforms conventional low-rank adaptation methods, and exhibits superior generalization capabilities in cross-task transfer and few-shot learning scenarios. The Intrinsic Low-Rank Geometry of LLM Alignment: Universality and Scaling Laws of the Gradient Changguo Fang (Anhui Agricultural University; Intelligence Institute of Machine, Hefei Institutes of Physical Science, Chinese Academy of Sciences); Wang Xi (University of Science and Technology of China; Intelligence Institute of Machine, Hefei Institutes of Physical Science, Chinese Academy of Sciences); Quan Shi (Changzhou University); Yucheng Fang (Anhui University); and Zenghui Ding (Intelligence Institute of Machine, Hefei Institutes of Physical Science, Chinese Academy of Sciences) Abstract Abstract Parameter-Efficient Fine-Tuning (PEFT) methods such as LoRA achieve strong alignment quality while updating only a tiny fraction of parameters, yet the geometry of the underlying alignment updates is still not well understood. This paper studies the unconstrained full-parameter alignment gradients and tests the hypothesis that their effective dimension is intrinsically small. Concretely, for each weight matrix gradient G, we quantify its effective rank by the stable rank r_s(G) = ||G||²_F / ||G||²_2, which is robust to small singular-value noise. We propose a unified perspective (information bottleneck, feature learning beyond lazy regimes, and a task-manifold view) that predicts low rs for alignment gradients, and we empirically validate this LowRank Gradient Hypothesis via large-scale spectral measurements across multiple mainstream LLM families (e.g., Llama, Qwen,Mistral) and scales (7B–70B). Across models and tasks, we consistently observe that alignment gradients admit accurate lowrank approximations and that their stable rank decreases with model size, following an approximate scaling law r_s ∝ N^(-γ) with γ ≈ 0.13. These results provide a first-principles account of why low-rank adaptation can be a high-fidelity proxy for full finetuning, and they offer practical guidance for choosing LoRA rank from measurable gradient spectra. GLAD: A Gating LoRA-Adapter Approach for Efficient Transfer Learning in Recommender Systems Wentao Hu and Hong Xie (University of Science and Technology of China) Abstract Abstract Leveraging large foundation models for modality-based transferable recommender systems (TransRec)—which learn directly from raw item features like text and images—is a promising paradigm. However, adapting these models via full fine-tuning is resource-intensive and often leads to catastrophic forgetting and overfitting, eroding their valuable pre-trained representations. To achieve parameter-efficient and generalizable TransRec, we propose GLAD (Gating LoRA-Adapter) that adaptively re-weights the influence of Low-Rank Adaptation (LoRA) and Adapter by employing a simple gating module, designed as a learnable controller that generates context-dependent scores from the input representation. This approach transforms the originally static, independent components into a cooperative ensemble, thereby forming an adaptive hybrid strategy, enabling more fine-grained control over the adaptation process. We conduct an extensive evaluation on two textual and one visual recommendation tasks. To assess robustness to data scale, we further evaluate our method on progressively smaller training subsets. While tuning only a small fraction of the parameters compared to full fine-tuning, our method outperforms all baselines on the textual tasks. On the visual task, our approach substantially narrows the performance gap to full fine-tuning, achieving nearly comparable outcomes. Tuesday Virtual Room 8 IJCNN Paper LLM Adaptation and Fine-Tuning II Session Chair: Md Zarif Hossain (Florida Atlantic University), 杉杉 陈 (沈阳航空航天大学) Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models Md Zarif Hossain and Ahmed Imteaj (Florida Atlantic University) Abstract Abstract Vision–Language Models (VLMs) rely heavily on pretrained vision encoders to support downstream tasks such as image captioning, visual question answering, and zero-shot classification. Despite their strong performance, these encoders remain highly vulnerable to imperceptible adversarial perturbations, which can severely degrade both robustness and semantic quality in multimodal reasoning. In this work, we introduce Sim-CLIP, an unsupervised adversarial fine-tuning framework that enhances the robustness of the CLIP vision encoder while preserving overall semantic representations. Sim-CLIP adopts a Siamese training architecture with a cosine similarity objective and a symmetric stop-gradient mechanism to enforce semantic alignment between clean and adversarial views. This design avoids large-batch contrastive learning and additional momentum encoders, enabling robust training with low computational overhead. We evaluate Sim-CLIP across multiple Vision–Language Models and tasks under both targeted and untargeted adversarial attacks. Experimental results demonstrate that Sim-CLIP consistently outperforms state-of-the-art robust CLIP variants, achieving stronger adversarial robustness while maintaining or improving semantic fidelity. These findings highlight the limitations of existing adversarial defenses and establish Sim-CLIP as an effective and scalable solution for robust vision–language representation learning. Pre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders Qingtao Xia, Jiahua Bao, Siyao Cheng, and Jie Liu (Harbin Institute of Technology) Abstract Abstract We propose a pre-fine-tuning probing method for Parameter-Efficient Fine-Tuning (PEFT) layer selection, aiming to obtain more stable and higher gains with fewer trainable parameters when adapting large vision--language models (VLMs). Unlike the common practice of applying LoRA and other adapters to all layers at once---where layer selection often relies on heuristic rules---we focus on the vision encoder and directly evaluate the "adaptability'' of each Transformer layer. Specifically, we characterize each layer from two perspectives: (i) the statistical properties of its Q/K/V projection weights (e.g., norms and condition numbers); (ii) robustness under controlled parameter perturbations. We then systematically compare these indicators with the downstream performance gains brought by applying PEFT to a single layer. Across experiments covering seven benchmarks and five PEFT variants, we observe a consistent correlation: layers (or matrices) with larger weight norms and higher condition numbers are usually more robust to perturbations and are more likely to yield larger fine-tuning gains. These results show that distribution-statistics analysis and perturbation tests before fine-tuning can provide practical signals for adaptation-layer selection, thereby maintaining or improving performance while reducing trainable parameters. Bound the Risk, Gate the Bias: A Mechanistic-Driven Sim-to-Real Transfer Method for Rebust Yield Prediction Tiansong Wu, Yuling Fan, and Ning Li (College of Informatics, Huazhong Agricultural University) and Dong Lin (Liaoning Technical University) Abstract Abstract Precision fertilization is the core of sustainable agricultural development and relies heavily on high-precision yield prediction. Crop model-based yield prediction is crucial for sim-to-real transfer learning. However, under conditions of data scarcity, the extrapolation ability is generally limited, risks are difficult to quantify, and prediction accuracy is low. Therefore, this paper proposes a bias analysis method based on `perfect calibration' to achieve sim-to-real transfer learning, aiming to obtain robust and reliable yield prediction. First, a crop mechanistic model is calibrated on real data to generate a large amount of simulated data as ideal yield observation data (pseudo-observation data). And the `perfect error (PE)' is calculated under controlled conditions. Subsequently, the PE is used as a mechanistic reference lower bound (MRLB) for extrapolation error through a zero-error test and statistical analysis for prediction risk assessment. Secondly, the bias decomposition for PE is proposed, and a gated debiasing procedure is designed to obtain a low-bias dataset for pre-training the neural network. Fine-tuning is then performed on a small amount of real data to achieve mechanistic-fused transfer learning for yield prediction. Experimental results show that, under conditions of extreme data scarcity, the proposed method significantly outperforms multiple baseline methods and provides robust yield confidence intervals. This method provides a new approach for the uncertainty analysis and trustworthy prediction in agriculture under extreme data scarcity by deeply integrating mechanistic modeling with machine learning, and can be extended to the calibration and evaluation of mechanistic models under similar data scarcity conditions. Fault-Controlled Routing: Proactive Failure Avoidance for Cost-Effective and Robust LLM Reasoning Chunlong Fan and Shanshan Chen (Shenyang Aerospace University) Abstract Abstract Frameworks for complex reasoning tasks with Large Language Models (LLMs) are often bottlenecked by computa- tional inefficiency, largely due to repeated exploration of invalid solution paths—a cost we term the “failure tax.” To address this, we propose Fault-Controlled Routing (FCR), a framework that shifts the paradigm from reactive error correction to proactive failure avoidance. FCR combines a Risk Gate mechanism with Stratified Collaborative Representation (SCR) to enable real- time risk assessment of reasoning trajectories. Upon detecting high-recurrence failure patterns, it dynamically escalates to more robust strategies only when necessary, thereby minimizing redundant computation while maintaining strong performance. Extensive experiments on major coding benchmarks show that FCR outperforms leading baselines in accuracy and achieves substantially lower token consumption. These results demonstrate FCR’s superior trade-off between performance and inference efficiency. Tuesday Virtual Room 1 IJCNN Paper Image Segmentation and Dense Prediction V Session Chair: Shaoguo Cui (Chongqing Normal University), guoqiang zheng (西南科技大学) A Morphology-Aware Deformable Hybrid CNN-Mamba Network for Nuclei Instance Segmentation and Classification Saiyi Ma, Shaoguo Cui, Guofen Wang, and Yanhe Gong (Chongqing Normal University) Abstract Abstract Accurate nuclei instance segmentation and classification in pathological images are crucial for quantitative cancer diagnosis. However, nuclei in histopathological images are often densely distributed and exhibit irregular morphologies, which makes precise nuclei instance segmentation and classification highly challenging. CNN-based methods are limited by local receptive fields and fail to global contextual dependencies, while Transformer-based approaches incur high computational costs on high-resolution images. Although Mamba state-space models can efficiently model global contextual information, they remain insufficient in capturing fine-grained local nuclei morphological variations. To address this, this paper proposes a morphology-aware deformable hybrid CNN-Mamba network, MDH-CMambaNet, for nuclei instance segmentation and classification. The network employs a CNN encoder to extract multi-scale features and introduces DeformHybridMamba (DHMamba) block during decoding as the core modeling unit, which adopts a dual-branch parallel structure to jointly model local morphological variations of nuclei and global contextual dependencies. Moreover, an adaptive feature fusion mechanism (AFM) is incorporated at skip connections to dynamically fuse encoder and decoder features via channel and spatial weighting, enhancing critical nuclei regions while suppressing background tissue interference. Experimental results demonstrate that the proposed method outperforms existing approaches in overall performance, particularly in scenarios with irregular nuclei morphology and dense distribution, validating the effectiveness and superiority of the proposed framework. HCF-UNet: Selective Skip Fusion and Calibrated Decoding for Coronary Vessel Segmentation xiu yang, li li, xiqianlong yuan, and guoqiang zheng (西南科技大学) Abstract Abstract Coronary vessel segmentation is essential for vascular modeling and quantitative analysis, yet remains challenging in coronary angiography due to elongated morphology, large scale variation, and heavy background interference. Many U-Net-based methods employ plain upsampling and skip-connection concatenation, which often fails to reconcile multi-scale semantics with high-resolution details, resulting in fragmented thin branches and inaccurate boundaries. We propose an efficient encoder-decoder network, termed HCF-UNet, to improve fine-structure reconstruction through redesigned skip connections and an enhanced decoder. Specifically, we introduce a Vessel-Dimension-Aware Selective Integration (V-DASI) module that combines channel-wise selection with spatial attention to adaptively fuse cross-level features and reduce redundancy. Moreover, we design a Residual-Channel-Spatial (RCS) decoder, where stacked residual blocks strengthen reconstruction and dual attention progressively calibrates vessel responses, enhancing sensitivity to tiny branches and boundary regions. Experiments on the public XCAD dataset demonstrate that HCF-UNet consistently outperforms representative baselines, yielding clearer delineation of fine vessels and more coherent vascular connectivity. These results indicate that selective fusion and attention-enhanced decoding offer an effective yet efficient solution for coronary angiography vessel segmentation. SGEA-Net: Stimulus-Guided Gating with Multi-Scale Edge-Enhanced Aggregation for Retinal Vessel Segmentation Zhiyuan Liu and Ziying Lu (School of Artificial Intelligence and Computer Science, Nantong University); Yifan Jiang (State Key Laboratory of Internet of Things for Smart City, University of Macau); and Shu Jiang (School of Artificial Intelligence and Computer Science, Nantong University) Abstract Abstract Retinal vessel segmentation plays a critical role in computer-aided screening and diagnosis of ocular diseases such as diabetic retinopathy. However, existing methods often suffer from discontinuous vessel boundaries, insufficient multi-scale semantic fusion, and weak channel-wise detail preservation, especially for thin capillaries. To address these challenges, we propose SGEA-Net, a novel network that integrates Multi-scale Edge-Enhanced Aggregation with Stimulus-Guided Gating for retinal vessel segmentation. Specifically, an Edge Attention Enhancement (EAE) block is designed to strengthen boundary continuity by exploiting dual-scale gradient cues while suppressing background noise. A Multi-scale Global Aggregation (MGA) block aligns hierarchical features into a unified resolution and enables adaptive global–local fusion across scales. In addition, a Stimulus-guided Adaptive Gated Bottleneck (SAGB) block dynamically recalibrates channel responses, improving the balance between global context and local vessel peaks. Experiments on three public benchmarks (DRIVE, CHASEDB1, and STARE) demonstrate that SGEA-Net consistently outperforms representative baseline methods, while identifying more continuous and detailed vessel structures. This work provides an effective and extensible solution for robust retinal vessel segmentation in fundus image analysis. S²A-Net: A Semantic-Structural Aware Network for Nuclei Instance Segmentation and Classification Can Jiang, Shaoguo Cui, and Guofen Wang (Chongqing Normal University) Abstract Abstract Nuclear instance segmentation and classification are fundamental for quantifying the tumor microenvironment. However, strictly distinguishing adherent nuclei while identifying their fine-grained types remains challenging. Standard singleencoder architectures struggle to accommodate the conflicting feature requirements of fine-grained localization and high-level classification. Conversely, existing multi-branch designs tend to over-decouple these sub-tasks, breaking the intrinsic consistency between an instance’s boundary and its semantic category. To address this, we propose S²A-Net, featuring an asymmetric dualstream architecture to explicitly disentangle task-specific features. We introduce the Dynamic Morpho-Semantic Rectifier(D-MSR) to inject attribute-level textual knowledge from medical VisionLanguage Models into the visual pipeline. Acting as a semantic mediator, this module utilizes statistical priors to dynamically calibrate features for confusing cell subtypes without requiring explicit attribute annotations. Furthermore, a Directional Interaction Attention(DIA) module integrates direction-field constraints to explicitly disentangle adherent boundaries. Extensive experiments on the CoNSeP dataset and related benchmarks demonstrate that S²A-Net achieves state-of-the-art performance and significantly enhances recognition accuracy for challenging cell categories. Tuesday Virtual Room 2 IJCNN Paper Image Segmentation and Dense Prediction VI Session Chair: Jin-Chun Piao (Yanbian University), Xiaopeng Liu (Shandong University of Science and Technology) ULSNet: Underwater-Aware Preference–Value Attention with a Boundary-Refined Head for Underwater Instance Segmentation Zhi-Dong Liang, De Li, and Jin-Chun Piao (Yanbian University) Abstract Abstract Underwater instance segmentation is critical for marine robotics and ecological monitoring. However, it remains challenging due to color cast, low contrast, scattering haze, and sensor noise, as well as frequent occlusions, ambiguous boundaries, and severe instance merging. To address the trade-off between accuracy and efficiency in existing methods, we propose ULSNet, an efficient two-stage underwater instance segmentation framework built upon an LSNet+FPN feature extractor. ULSNet introduces three lightweight, task-oriented modules to strengthen RoI representations and mask decoding under underwater degradations. Specifically, Underwater-Aware Feature Gating (UWAF) recalibrates RoI features via complementary spatial and channel gating guided by hetero-scale pyramid cues. Preference–Value Attention (PVA) further enhances instance evidence by modulating a shared value branch with content-adaptive preference cues from two lightweight heads. Finally, the UQBR-Head performs boundary-refined mask decoding with an explicit boundary branch to mitigate boundary uncertainty and instance merging. Extensive experiments on UIIS and USIS10K show that ULSNet improves the accuracy–efficiency trade-off, with consistent gains in mAP and AP75 under challenging underwater conditions. LW-UVENet:An Efficient and Lightweight Underwater Video Enhancement Network Xinnan Fan, Qi Sun, and Yuanxue Xin (Hohai University, College of Information Science and Engineering) and Pengfei Shi (Hohai University, College of Artificial Intelligence and Automation) Abstract Abstract Underwater Video Enhancement (UVE) aims to alleviate color distortion, low contrast, haze-like degradation, and motion blur caused by light absorption and scattering in water. Existing end-to-end UVE models, such as UVENet, already demonstrate the benefit of temporal modeling, but their computational cost and limited deblurring capability still restrict deployment on resource-constrained underwater robotic terminals. To address this practical bottleneck, we present LW-UVENet, a deployment-oriented lightweight redesign of UVENet for underwater video enhancement. The proposed network combines a Lightweight ConvNeXt-based Encoder, an Underwater Deblurring Module (UDM), an Enhanced Feature Alignment and Aggregation Module (Enhanced FAAM), a Lightweight Decoder, and a Lightweight Global Recovery Module (LW-GRM). Rather than introducing a completely new alignment paradigm, our goal is to integrate low-cost yet task-effective components to improve the quality–efficiency trade-off in underwater video enhancement. Experiments on the synthetic SUVE dataset and the real-world MVK dataset show that LW-UVENet maintains competitive and in some cases superior enhancement quality than the original UVENet, while reducing single-frame inference time from 0.241s to 0.111s and peak memory usage from 19.25G to 10.26G. These results indicate that LW-UVENet makes underwater video enhancement more practical as a front-end module for resource-limited underwater robots and other embodied visual systems. SAM2-MAS: Boundary-Aware Frequency-Adaptive Framework for Marine Animal Segmentation Ziyi Bao, Jichao Jiao, Ning Li, Yingchao Zeng, Yingjian Zhang, Peiyuan Zhao, Yifan Li, Chunyang Wu, Tianxiang Zhang, and Jiajie Huang (Beijing University of Posts and Telecommunications) Abstract Abstract Marine animal segmentation (MAS) in real underwater environments remains challenging due to optical degradation (scattering, attenuation, blur) and target--background ambiguity caused by camouflage and low contrast. We repurpose the SAM2 image encoder as a strong generic backbone, but its geometry priors alone can struggle to disambiguate category-level semantics and boundaries under severe degradations. SFPNet-UIE: Spatial-Frequency Perception Network for Underwater Image Enhancement Xiaopeng Liu and Hui Xin (Shandong University of Science and Technology), Cong Liu (Nanjing University of Science and Technology), Long Chen (University College London), and Junyu Dong (Ocean University of China) Abstract Abstract Underwater Image Enhancement (UIE) aims to restore the visual quality of degraded underwater images and is essential for various underwater applications. While most deep learning approaches focus solely on spatial domain restoration, overlooking the valuable information in the frequency domain. To address this limitation, we propose SFPNet-UIE, a Spatial-Frequency Perception Network for Underwater Image Enhancement. SFPNet-UIE incorporates multiple Dynamic Fusion Blocks (DFBs), each containing a Spatial-Frequency Perception Module (SFPM) and Multi-Scale Receptive Field Module (MSRFM). The SFPM employs a dual-path structure, where the spatial branch captures local details and the frequency branch models global features. The MSRFM applies multi-kernel receptive field aggregation to fuse information from both domains while preserving fine details. Experiments on four real-world underwater datasets demonstrate that SFPNet-UIE delivers state-of-the-art (SOTA) or competitive performance across multiple benchmarks. Code is available at https://github.com/xnogjlnla/SFPNet-UIE. Tuesday Virtual Room 3 IJCNN Paper Image Segmentation and Dense Prediction VII Session Chair: zhiyu xiao (Beijing Information Science and Technology University), Yingqi Liang (Guangxi University) FSECrossNet: Frequency–Spatial Enhancement with Cross-Modal Attention for Remote Sensing Semantic Segmentation Yingqi Liang (Guangxi University) Abstract Abstract Multi-modal remote sensing semantic segmentation benefits from complementary cues such as RGB appearance and DSM elevation, yet effectively integrating fine boundary details with global context remains challenging. Moreover, many fusion pipelines rely solely on spatial-domain processing and employ heavy cross-modal attention, leading to limited efficiency. We propose FSECrossNet, a Transformer-style encoder-decoder that performs staged fusion by coupling frequency-spatial enhancement with lightweight cross-modal reasoning. Specifically, Gated Feature Fusion (GFF) builds balanced multi-scale skip features by adaptively weighting RGB and DSM cues. At the deepest stage, Dual Frequency-Spatial Fusion (DFSF) strengthens discriminative textures via DCT-based frequency enhancement and improves local structures using multi-scale depthwise convolutions. Finally, Cross-Modal Attention Fusion (CMAF) reuses shared projections across self- and cross-attention to enable efficient intra-/inter-modal interaction with reduced parameters. Experiments on ISPRS Vaihingen and Potsdam demonstrate consistent improvements over strong baselines, yielding higher segmentation accuracy and clearer boundary delineation in qualitative comparisons under a favorable accuracy-efficiency trade-off. CFEFNet: Cross-Modal Feature Enhancement and Fusion Network for RGB-T Semantic Segmentation Ziyang Zhang and Lei Sun (Sun Yat-sen University) Abstract Abstract Semantic segmentation based on RGB images and thermal infrared (TIR) images can enhance the robustness of scene parsing and improve segmentation performance. However, existing RGB-thermal (RGB-T) segmentation methods still inadequately exploit the complementary information between modalities and exhibit limited performance in segmenting targets of diverse scales and shapes in complex scenes. Moreover, cross-attention mechanism incurs considerable computational overhead when modeling global dependencies. To address these issues, we propose a novel Cross-Modal Feature Enhancement and Fusion Network (CFEFNet). Specifically, we design a Bidirectional Feature Enhancement (BFE) module to bridge the gap between two modalities. Then, a Frequency-Domain Cross-Attention (FDCA) module is proposed to efficiently promote cross-modal information exchange. Furthermore, we introduce a Spatially-Aware Feature Integration (SAFI) module to capture multi-scale information and anisotropic spatial context. Experimental results on two public datasets MFNet and PST900 show that our proposed CFEFNet achieves state-of-the-art performance. The code is available at https://github.com/ziyangbest/CFEFNet. IDC-YOLO: Efficient Visible-Infrared Object Detection via Illumination-Guided Modulation and Difference-Aware Interaction Xiao Ma, De Li, and Jin-Chun Piao (Yanbian University) Abstract Abstract Visible-infrared imagery provides complementary benefits that reduce the limitations of single-modality object detection, particularly under challenging illumination conditions. However, effectively exploiting cross-modal information while maintaining computational efficiency remains difficult for detector-oriented architectures. In this paper, we propose IDC-YOLO, a dual-stream visible-infrared object detection framework built upon a YOLO-style architecture. The framework processes visible and infrared inputs in parallel and introduces two modules to enhance cross-modal interaction. An Illumination-Guided Modulation (IGM) module implicitly estimates global scene illumination and adaptively regulates modality responses, improving robustness across varying lighting conditions. Meanwhile, a Difference-Aware Interaction Module (DAIM) explicitly models cross-modal structural discrepancies and enhances complementary texture and shape cues via residual difference-based interaction. To facilitate efficient feature fusion, a lightweight Channel Switching and Spatial Attention (CSSA) module is incorporated at the detector feature level with limited computational overhead. Experiments on the M3FD and FLIR benchmarks demonstrate that IDC-YOLO achieves a 5% improvement in mAP50 on the M3FD dataset and a 3.8% improvement on the FLIR dataset, compared to the baseline, while maintaining a favorable accuracy-efficiency trade-off. Further evaluations on an embedded platform verify its real-time inference capability. Furthermore, consistent performance gains across different YOLO backbones indicate good generalization capability of the proposed framework. CFSFusion: A Cross-Domain Interaction and Frequency-Spatial Collaborative Network for Infrared and Visible Image Fusion Zhiyu Xiao and Qiang Tong (Beijing Information Science and Technology University), Yuli Chen (Beijing University of Posts and Telecommunications), and Xiulei Liu and Siyu Zhu (Beijing Information Science and Technology University) Abstract Abstract Fusing infrared and visible images is a key technique for intelligent perception in complex environments, such as nighttime surveillance and autonomous driving. Existing methods often combine frequency and spatial features without sufficient interaction, resulting in limited structure–saliency integration, degraded detail fidelity, and less prominent object representation. Moreover, high-frequency details crucial for textures and contours are commonly enhanced in a non-selective manner, which can introduce noise and blurred edges. To address these issues, we propose CFSFusion, a collaborative framework that integrates frequency-informed and spatial features. Specifically, we design a Cross-Domain Interaction Module (CDIM) that enables structure–saliency guided interaction between frequency-informed representations and spatial features. In addition, a structure-aware high-frequency enhancement mechanism (StructHF) is introduced to selectively emphasize key details, while an improved Multi-Scale Enhancement Channel Attention (MECA) further reinforces salient structures and suppresses noise. Extensive experiments on public datasets demonstrate that our method achieves high-quality fusion across multiple metrics and shows strong robustness in downstream object detection tasks. We release the code1 to facilitate future research. Tuesday Virtual Room 4 IJCNN Paper Image Segmentation and Dense Prediction VIII Session Chair: Chaoli Wang (上海理工大学), Yan Wan (DongHua University) Dual-Domain Feature Enhancement and Boundary-Guided Multi-Branch Attention Network for Polyp Segmentation in colonoscopy images Sisi Chen, Chaoli Wang, and Zhanquan Sun (上海理工大学) Abstract Abstract In recent years, deep learning methods have made significant progress in polyp image segmentation tasks. However, effectively modeling the complementary relationship between spatial and frequency domain features within a unified framework and accurately describing the complex and blurred polyp boundaries remains a challenging task. Most existing methods focus on spatial domain feature learning, and even when frequency domain information is introduced, the differences in spectral energy distribution across different feature layers are often neglected. Meanwhile, cross-layer feature fusion and region-specific modeling are insufficient, and the high-level semantic information is not effectively used to guide low-level spatial features, lacking differentiated modeling mechanisms for foreground, background, and boundary regions. To address these issues, we propose a Dual-domain Feature Enhancement and Boundary-guided Multi-branch Attention Network (DBMA-Net) for polyp image segmentation. The network first introduces a Multi-scale Context and Attention Fusion Module (MCAF) to enhance the multi-scale structural and texture representation of shallow features. Next, a Dual-domain Feature Enhancement Module (DFE) is designed to strengthen the expression of deep features by collaboratively modeling spatial semantics and frequency domain global texture information. In addition, a Boundary-guided Multi-branch Attention Module (BMAM) is proposed, which adaptively focuses on foreground, background, and boundary regions using high-level prediction results, thereby improving the ability to describe complex boundaries. Experimental results on five publicly available polyp segmentation datasets demonstrate that the proposed DBMA-Net outperforms existing methods across various evaluation metrics. MTFNet: Multimodal Fusion with Topographic Priors for Landslide Semantic Segmentation Zilan Ning, Juan Luo, Kexuan Feng, Ying Qiao, and Anping Liu (HUNAN UNIVERSITY) Abstract Abstract Landslide segmentation from remote sensing images is important for risk mapping and disaster response, but it remains difficult in complex scenes with unclear boundaries and confusing background. Traditional methods that rely only on RGB images often produce missed regions and false detections, since landslides can look similar to bare soil and roads. To address this, we propose Multimodal Topographic Fusion Network (MTFNet), a three branch framework that combines a frozen SAM ViT-L spatial encoder with lightweight ATL adaptation, a frequency based texture encoder that strengthens high frequency boundary details, and a Digital Elevation Model (DEM) based topographic encoder that provides geometry information such as slope change and surface shape. The multi level features from the three branches are fused with the Global Cross-attention Context Gating (GCCG) to build global interaction across branches and position wise gating, and then refined with the Cross scale Dual Residual Mixing (CDRM) to mix stable semantics with fine local details. The proposed method is evaluated on the Bijie and Landslide4Sense datasets and performs better than baseline models, with overall improvements in key metrics such as mIoU, Precision, and F1-score. These results demonstrate the effectiveness of the proposed multimodal fusion design for robust landslide segmentation in complex scenes. DINOv2-UNet: Deployment-Friendly Colonoscopy Polyp Segmentation with Self-Supervised ViT Priors Chengfei Cai, Jia Yu, Zehao Liu, Lixing Tan, Zenan Lu, and Zhenyu Song (Jiangsu Key Laboratory of Intelligent Drug Screening and Repositioning, College of Information Engineering, Taizhou University) Abstract Abstract Accurate polyp segmentation is crucial for colorectal cancer screening but remains challenging due to low tissue contrast, imaging artefacts, and cross-domain variability. We propose DINOv2-UNet, a deployment-friendly framework that transfers large self-supervised vision transformers to colonoscopic segmentation through a representation-first design. The model couples a DINOv2 ViT-B/14 encoder with a lightweight U-shaped decoder, avoiding complex fusion modules to preserve efficiency. To ensure stable adaptation on label-scarce medical datasets, we introduce a principled fine-tuning recipe that involves partial optimisation of deep transformer blocks, differential learning rates, hybrid BCE–Dice supervision, and cosine scheduling with warm-up. Experiments on four public datasets, including Kvasir-SEG, CVC-ClinicDB, CVC-ColonDB, and ETIS-LaribPolypDB, achieve Dice scores of 0.927, 0.948, 0.925, and 0.931, respectively, consistently outperforming CNN- and transformer-based baselines. Despite the high-capacity backbone, the model runs at 53.9 FPS on an RTX 4050 GPU and over 23 FPS on a GTX 1080 Ti. These results demonstrate that well-adapted self-supervised priors can replace heavy decoders, enabling accurate and real-time clinical deployment. ColorControlNet: A Diffusion Model-Based Dual-Branch Framework for Image Colorization Yan Wan, Xiaojing Wen, and Li Yao (Donghua University) and Ru Zhou (Department of General Surgery, RuiJin Hospital LuWan Branch, Shanghai Jiaotong University School of Medicine,Shanghai, China) Abstract Abstract Image colorization is an essential task in image generation, aimed at producing realistic color results for grayscale images to enhance visual performance. Despite significant progress in deep learning-based methods for natural images, challenges such as uneven color distribution, color bleeding, and structural instability persist in scenes with limited semantic information. To address these issues, ColorControlNet, a dual-branch image colorization framework based on diffusion models, is proposed in this paper. The model takes grayscale images, palette images, and text descriptions as multimodal inputs. The gating mechanism adaptively combines structural features and color priors, enabling precise feature fusion. To efficiently extract structural and color control features, a lightweight encoding network is designed, integrating multi-scale context fusion and axial self-attention mechanisms to enhance local texture representation and consistency modeling across regions. Additionally, an inference-stage enhancement strategy is proposed to suppress local color bleeding and enhance details without increasing training costs. Experimental results demonstrate that the proposed method outperforms existing methods on both a custom flat design dataset and the DomainNet-painting dataset. Tuesday Virtual Room 5 IJCNN Paper Information Extraction and Text Mining I Session Chair: Boyan Xu (Guangdong University of Technology), Zhiyang Yu (Shandong Normal University) PRAF: A Part-of-speech-Ruled Adaptive Framework for Aspect Sentiment Triplet Extraction Zhiyang Yu and Guangjin Wang (Shandong Normal University); Fuyong Xu (Nanjing University of Aeronautics and Astronautics); and Zhipeng Wang, Ru Wang, and Peiyu Liu (Shandong Normal University) Abstract Abstract Aspect Sentiment Triplet Extraction (ASTE), as a fine-grained sentiment analysis task, extracts triplets containing aspect terms, opinion terms, and corresponding sentiment polarities. Recently, most approaches focus to investigating the effect of additional linguistic knowledge on triplet extraction, including dependent syntactic relations and semantic roles. However, these approaches lack in-depth analysis of inherent characteristics and rules of part-of-speech, ignoring potential impact of certain non-target part-of-speech. Therefore, this paper proposes a Part-of-speech-Ruled Adaptive Framework (PRAF) for ASTE task. Specifically, we design a POS-Powered Network (P-PN) to predict potential part-of-speech distributions, which overcomes the limitations of previous methods in exploring part-of-speech rules. Additionally, in order to filter irrelevant information we introduce a Rules-aware Attention Filtering Layer (RAFiL) to adaptively adjust the contribution of part-of-speech features. Experiments on four benchmark datasets conclusively demonstrate the superiority of our approach. Learning Contrastive Interaction Relations for Aspect Sentiment Triplet Extraction Shiman Zhao (Peking University) and Jianyuan Ni (Juniata College) Abstract Abstract Aspect Sentiment Triplet Extraction (ASTE) aims to extract triplets consisting of an aspect, an opinion, and their corresponding sentiment polarity. A key challenge in ASTE is the presence of numerous Overlapping Triplets (OT), which increase the task’s complexity. We categorize these OTs into three types based on their complexity: Single-OT, Multi-OT, and Conflicted-OT. While most existing approaches primarily focus on Single-OT, Multi-OT and Conflicted-OT remain relatively underexplored. To address this gap, we propose a novel ASTE method that effectively decodes complex relations between aspects and opinions to perform Multi-OT and Conflicted-OT. Specifically, we introduce a pointer network to locate aspect and opinion indexes, and then propose a Contrastive Sentiment-enhanced Mechanism (CSM) to enhance sentiment interaction. CSM integrates a Markov Graph Convolutional Network to capture syntactic dependencies for context enhancement and constructs contrastive pairs to enrich sentiment knowledge, improving performance on OTs. Extend experiments demonstrate that our method achieves significant performance, especially in Multi-OT and Conflicted-OT. Syntax-Enhanced Network with Gated Relational Attention for Aspect Sentiment Triplet Extraction Bowen Wang, Xiaoxu Zhu, and Peifeng Li (Soochow University) Abstract Abstract Aspect Sentiment Triplet Extraction (ASTE) has emerged as a critical task in sentiment analysis, specifically aimed at identifying aspect terms, opinion terms, and their corresponding sentiment polarities within textual data. While recent span-based tagging models have achieved promising results, they typically rely on coarse-grained contextual representations, either without explicitly modeling syntactic structure or without mechanisms to selectively control its influence, which limits their ability to effectively distinguish task-relevant information from noise during triplet extraction. To address these limitations, we propose a Syntax-Enhanced Network with Gated Relational Attention (SENGRA), which augments span-based tagging through the selective integration of syntactic structures. SENGRA employs a Gated Relational Graph Attention Network (G-RGAT) to encode typed dependency relations, in which an input-dependent gating mechanism modulates the graph-aggregated token representations. This design effectively reduces noise from irrelevant syntactic information while improving the model’s nonlinear expressive capacity. Furthermore, we design a syntax-guided inference strategy to accurately align aspect and opinion terms within sentiment spans. Extensive experiments on four widely adopted ASTE benchmarks show that SENGRA consistently surpasses state-of-the-art methods, confirming its efficacy and robustness in complex sentiment extraction scenarios. OER: Opinion Externalization via Large Language Model Reasoning for Implicit Sentiment Analysis Bingfeng Chen and Yongqi Luo (Guangdong University of Technology), Shaobin Shi (gdut), Yuguang Yan and Ruichu Cai (Guangdong University of Technology), Zhifeng Hao (Shantou University), and Boyan Xu (Guangdong University of Technology) Abstract Abstract Implicit Sentiment Analysis (ISA) requires models to infer sentiment polarity without explicitly given opinions. To accomplish this, models must either capture opinions from the surrounding context or leverage implicit clues for sentiment inference. Existing large language model (LLM)-based approaches decompose ISA through multi-step reasoning but suffer from hallucination and susceptibility to contextual interference. To address these challenges, we propose the \underline{O}pinion \underline{E}xternalization via LLM \underline{R}easoning (OER) framework, which externalizes implicit opinions to enhance sentiment inference. OER consists of two key modules: (1) an \emph{Explicit Opinion Extraction} module, which guides LLMs in extracting opinions while reducing hallucination through self-correcting in-context learning and a self-debate mechanism; and (2) an \emph{Implicit Sentiment Clue Extraction} module, which retrieves sentiment-enhancing information to infer implicit opinions while mitigating contextual interference. Experimental results on two ISA benchmark datasets demonstrate that OER consistently outperforms existing methods in both overall performance and implicit sentiment accuracy. Tuesday Virtual Room 6 IJCNN Paper Information Extraction and Text Mining II Session Chair: Weixin Zuo (Shanghai University of Electric Power, Faculty of Artificial Intelligence), yifan huo (Zhejiang Sci-Tech University) EPNER: An External Prototype Augmented Framework for Few-Shot Named Entity Recognition Weixin Zuo, Xilong Wang, and Nuoxian Liu (Shanghai University of Electric Power, Faculty of Artificial Intelligence) Abstract Abstract Few-Shot Named Entity Recognition (NER) aims to identify novel entities relying on scarce annotations. However, the traditional two-stage span-based model is still constrained by two major bottlenecks: the increase of false-positive spans and the instability of prototypes derived from sparse support sets, which exacerbates the indistinguishability of fine-grained categories. We propose EPNER to tackle these issues by augmenting prototypes with external knowledge. In the detection stage, we introduce an Entity-Non-Entity Contrastive Learning mechanism to enhance representational distinctiveness, complemented by a learnable width bias to suppress spurious candidate spans. In the classification stage, we construct a high-quality external knowledge prototype library by large language models (LLMs) and frequent entities from external corpora. They are used as stable prior knowledge. These priors are integrated with support prototypes via a Dynamic Attention Fusion mechanism. Furthermore, we have also designed a prototype separation loss function to widen the inter-class distance between different prototype categories. A large number of experiments have shown that our proposed method significantly outperforms state-of-the-art approaches. Synergizing Sparse Structures and Hyperbolic Geometry for Nested Named Entity Recognition Chunyue Lu, Xilong Wang, and Weixin Zuo (Shanghai University of Electric Power, Faculty of Artificial Intelligence) Abstract Abstract Nested Named Entity Recognition (NER) involves complex layered entities but is often hindered by boundary ambiguity and intricate dependencies. Current span-based paradigms are frequently limited as they overlook the non-Euclidean nature and discrete structural correlations of nested entities. To address this, we propose a Structure-Geometry Synergistic Framework that integrates sparse structural graphs with hyperbolic mapping mechanisms. Specifically, we employ a structural enhancement layer to construct sparse graphs for dependency learning, and a hyperbolic mapping layer to project representations into the Poincaré ball to explicitly capture hierarchy. Additionally, a calibration strategy guided by boundary priors is introduced to refine decision boundaries. Comprehensive evaluations show that our framework outperforms state-of-the-art (SOTA) baselines. Dual-Perspective Modeling for Chinese NER with Selective SSM and Semantic Difference Convolution Jianquan Ouyang and Xinxin Li (Xiangtan University) Abstract Abstract Chinese Named Entity Recognition (Chinese NER) faces challenges in boundary identification due to the absence of explicit word delimiters. While lexicon-based methods improve performance, their high maintenance cost limits domain applicability. Moreover, models often overemphasize high-frequency boundary characters, leading to biased predictions and reduced generalization. To address these issues, we propose a dual-perspective modeling framework that integrates the Selective State Space Model (Selective SSM) and Semantic Difference Convolution (SDC). First, biaffine attention transforms the input sentence into a character-pair grid matrix, capturing inter-character relationships. The Selective SSM then captures long-range dependencies from a global perspective, providing rich contextual information. Concurrently, the SDC enhances local context understanding by modeling subtle semantic variations between adjacent characters, mitigating boundary bias. Ultimately, the fusion of features from both perspectives enables the decoding of entities and categories. Experimental results demonstrate that our method significantly outperforms or matches state-of-the-art approaches on four datasets: Resume, MSRA, Weibo, and OntoNotes 4.0, thereby validating the effectiveness of the proposed dual-perspective paradigm. MVCCLF: Gradient Surgery-Driven Multi-View Contrastive Learning for Multimodal NER Kai Zhuo, Yifan Huo, Junhong Zheng, and Lili He (Zhejiang Sci-Tech University) Abstract Abstract Multimodal Named Entity Recognition (MNER) significantly improves entity recognition performance by integrating textual and visual information from social media posts. However, existing methods still face challenges such as imprecise text-image semantic alignment, visual noise interference, and multi-task optimization conflicts, resulting in insufficient fine-grained cross-modal interaction and unstable training. To address these issues, this paper proposes a Gradient Surgery-Driven Multi-View Cross-Modal Contrastive Learning Framework (MVCCLF). The framework enables fine-grained cross-modal alignment via a coarse-to-fine multi-view contrastive learning mechanism, while introducing dynamic weight scheduling and a dual-layer gradient surgery mechanism to effectively mitigate gradient conflicts between the primary NER task and auxiliary contrastive learning. Experimental results on the Twitter15 and Twitter17 datasets demonstrate that MVCCLF achieves F1 scores of 75.55\% and 87.31\%, respectively, demonstrating competitive performance on both datasets. Ablation studies and sensitivity analysis further validate the effectiveness of each component and the robustness of the framework, providing an efficient fine-grained fusion solution for MNER in social media scenarios. Tuesday Virtual Room 7 IJCNN Paper Information Extraction and Text Mining III Session Chair: Shiao Meng (Tsinghua University), Yuan Gao (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) I2DRE: Intra- and Inter-Pair Interactions based Document-Level Relation Extraction with Binary Hill Loss Shiao Meng and Lijie Wen (Tsinghua University) Abstract Abstract Document-level relation extraction (DocRE) aims to identify semantic relations between entity pairs within a document and plays a crucial role in various knowledge-centric applications. Though have achieved certain progress, existing studies still expose two critical vulnerabilities: 1) For a given entity, they typically apply heuristic pooling methods to aggregate all its mention representations into a globally unique entity representation, ignoring that the importance of each mention varies when the entity is paired with different entities. This may lead to wrong predictions due to the disturbance of irrelevant mentions. 2) Most approaches use the features of entity pair itself for relation prediction, without fully utilizing the information from other entity pairs in the document. This limits the prediction of relation instances that rely on logical reasoning. To tackle these problems, we propose a novel intra- and inter-pair interactions based DocRE model (I2DRE). We first introduce an interaction mechanism within each entity pair. By allowing the two entities in each pair to guide each other's aggregation of mention representations, we derive pair-aware entity representations which can focus on key mentions for the entity pair. We further devise a gated biased attention module to selectively propagate information across entity pairs. This inter-pair interaction enables the model to effectively capture global reasoning clues. Moreover, we design a binary hill loss to tackle the missing relation problem prevalent in real-world DocRE datasets, where numerous relation instances are missing in annotations. The loss re-weights the gradients for negative class in traditional binary cross-entropy loss as hill-shaped to alleviate the adverse effect of missing relations. Experimental results on two benchmark datasets demonstrate that our method consistently outperforms existing methods. Fine-Grained Prototype Adaptation for Few-Shot Document-Level Relation Extraction Hongxia Jin, Dazhuang Wang, and Xiaowang Zhang (Tianjin University) Abstract Abstract Few-shot document-level relation extraction (FSDLRE) is often hindered by prototype dilution, where critical relational evidence is obscured by irrelevant document-level context and semantic overlap. To address this, we propose a framework that enhances prototype discriminability through three key innovations: (1) A Query-driven Evidence Adaptation (QE) module that dynamically aligns support instances with query-specific evidence to filter contextual noise; (2) A Relation-guided Feature Attention (RF) module that leverages relational semantic anchors to dynamically recalibrate feature importance, thereby enhancing the discriminative power for overlapping relations; and (3) A Task-specific NOTA Representation (TN) that adaptively refines decision boundaries for relation-absent instances. Experiments on FREDo and ReFREDo benchmarks show our model achieves state-of-the-art results, outperforming competitive baselines. Visualization and case studies further confirm our model's robustness in capturing intra-class consistency and inter-class distinction in complex document contexts. SABRE: Semantic Anchored Biaffine Relation Extraction for Chinese Metaphors Mingchao Liu and Shuo Wang (Hebei Key Laboratory of Machine Learning and Computational Intelligence, College of Mathematics and Information Science, Hebei University) Abstract Abstract Metaphorical Relation Extraction (MRE) is a fundamental task that bridges linguistic expressions and conceptual mappings by identifying directed connections between Target and Source spans. Current state-of-the-art methods typically rely on a generate-then-classify paradigm, where relations are predicted only for pre-filtered candidate spans. This cascaded structure introduces irreversible error propagation, resulting in a severe “Recall Bottleneck” in which implicit metaphors are discarded at early stages. Context-Construction-Based Prompt Injection in Document-Centric LLM Decision Systems Yuan Gao, JinWen He, Yue Zhao, and Kai Chen (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Abstract Abstract Large Language Models (LLMs) are increasingly deployed in document-centric workflows such as resume screening and academic peer review, where uploaded files are automatically parsed and converted into model input. This processing pipeline introduces new security risks: hidden content embedded in documents may be extracted by parsers and incorporated into the LLM context without being visible to human users. Tuesday Virtual Room 8 IJCNN Paper Information Extraction and Text Mining IV Session Chair: Hao Zhang (School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China), Jipeng Guo (Beijing University of Chemical Technology) Hierarchical Contrastive Graph Attention Network for Aspect-Based Sentiment Analysis Hao Zhang and Baiyou Qiao (School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China) Abstract Abstract Aspect-based sentiment analysis aims to identify the sentiment polarity of specific aspect terms in sentences. Existing graph neural network methods based on dependency trees face two limitations: processing dependency structures at fixed scales without recognizing the distinct roles of syntactic relationships at different granularities, and lacking self-supervised regularization to enhance representation learning. To address these issues, we propose Hierarchical Contrastive Graph Attention Network (HCGAT). The model captures syntactic information at three granularities: word-level graph attention (1-hop) for lexical modifications, phrase-level graph attention (2-hop) for phrasal composition, and sentence-level multi-head attention for global semantics. We design cross-granularity consistency contrastive learning with intra-granularity consistency and inter-granularity discriminability constraints to enhance hierarchical fusion effectiveness. Experiments on three benchmark datasets demonstrate that HCGAT achieves state-of-the-art performance, with particularly significant improvements on Laptop14 which has fewer training samples. A Frequency-Channel Enhanced Dynamic GCN for Aspect-Based Sentiment Analysis Shaokun Liu, Mieradilijiang Maimaiti, and Wushouer Silamu (Department of Computer Science, Xinjiang University) Abstract Abstract Aspect-based Sentiment Analysis (ABSA) is one of the an essential down-stream tasks in natural language process- ing. Generally, most of current methods for ABSA mainly focus on exploiting the Graph-based approaches which rely heavily on dependency trees, yet they often suffer from parsing noise and the rigidity of static graph structures. To address these drawbacks, we propose a Frequency-Channel Enhanced Dynamic Graph Convolutional Network (FCD-GCN). To contrast with the previ- ously presented static models, the our proposed approach FCD- GCN introduces a dynamic graph mechanism that iteratively updates edge weights across layers, allowing the topology to adap- tively capture evolving aspect-context associations. Our method is composed of three steps, firstly, constructing a layer-wise dynamic graph to refine topology; secondly, enhancing features via frequency-channel attention; and finally, aggregating these representations for sentiment prediction. Besides, we incorporate a dual-attention strategy to improve the representation quality, a frequency-domain module utilizes Fourier transforms to model long-distance and multi-scale dependencies, while a channel aware self-attention module recalibrate the feature subspaces to suppress redundancy. We conduct vast experiments for the verifications of our presented method. Extensive experiments on benchmark datasets demonstrate that FCD-GCN consistently outperform strong baselines, validating the effectiveness of com- bining dynamic graph refinement with frequency and channel level modeling. The code is publicly available. DDD: Image Description-Guided Dual-Chain Enhancement Model Based on Dynamic Fusion Minghua Luan and Jun Lu (Heilongjiang University, Jiaxiang Industrial Technology Research Institute) Abstract Abstract Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to extract aspect terms and perform sentiment classification from text-image pairs. Existing methods suffer from two limitations: Firstly, the differences in information density of intra-modal features are overlook. Secondly, the semantic correlation and collaboration mechanisms between modalities are insufficient. This paper proposes an Image Description Guided Dual Chain Enhancement Model Based on Dynamic Fusion (DDD), which synergistically tackles the above problems from two dimensions: intra-modal enhancement and inter-modal collaboration. Specifically, the Image Description Guided module (IDG) leverages image descriptions generated by GPT-3.5 to weight and guide text features, making up for the weak semantic correlation between text and images. The Dual Chained Enhancement (DCE) Module adopts a dual-chain architecture to filter and enhance features within each modality. The Dynamic Fusion module (DFM) dynamically optimize modal collaboration based on inter-modal correlations. Experiments on two benchmark datasets demonstrate that DDD achieves performance improvements compared with existing methods. Multi-order Filter Fusion and Triple Contrastive Learning for Attribute Graph Clustering Tianyou Yu, Kailin Zhang, and Jianlin Zheng (Beijing University of Chemical Technology); Tianruo Liu (Duke University); and Tianxiang Zhao, Haopeng Yang, and Jipeng Guo (Beijing University of Chemical Technology) Abstract Abstract Deep contrastive graph clustering has attracted growing attention due to its powerful self-supervised representation learning paradigm on attributed graphs. However, two challenges still hinder further improvements. First, most existing methods heavily rely on the observed adjacency in neighboring information propagation, which fails to capture attribute-induced global semantics and becomes fragile under sparse, incomplete, or noisy topology. Second, most existing methods employ a single contrastive strategy, resulting in shallow self-supervised representation learning. To this end, this paper proposes Multi-order Filtering Fusion (MFF) and Triple Contrastive Learning (TCL) for Attribute Graph Clustering (MT-AGC), a semantic feature-enhanced topology augmentation with comprehensive contrastive learning. Specifically, MFF constructs an attribute-similarity-induced global semantic graph and fuses it with the original topology, then performs multi-order low-pass filtering with self-attentive fusion to obtain discriminative multi-scale embedding representation. Further, TCL jointly optimizes representation contrastive consistency, structural consistency, and pseudo-label guided semantic matching, achieving thorough self-supervised training. More importantly, a two-stage strategy introduces high-confidence pseudo-labels to form a closed and reliable feedback loop in crucial semantic contrastive learning. Extensive experiments on six benchmark datasets demonstrate the effectiveness of MT-AGC, and ablation results and visualizations further verify the contribution of each component. The code could be available at https://github.com/Sky-Right/MT-AGC. Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V6 Session Chair: yilin fang (Wuhan University of Technology) Decoupling Numerical and Structural Parameters: An Empirical Study on Adaptive Genetic Algorithms via Deep Reinforcement Learning for the Large-Scale TSP Hongyu Wang, Yuhan Jing, Yibing Shi, Enjin Zhou, Haotian Zhang, and Jialong Shi (Xi’an Jiaotong University) Abstract Abstract Proper parameter configuration is a prerequisite for the success of Evolutionary Algorithms (EAs). While various adaptive strategies have been proposed, it remains an open question whether all control dimensions contribute equally to algorithmic scalability. To investigate this, we categorize control variables into numerical parameters (e.g., crossover and mutation rates) and structural parameters (e.g., population size and operator switching), hypothesizing that they play distinct roles. This paper presents an empirical study utilizing a dual-level Deep Reinforcement Learning (DRL) framework to decouple and analyze the impact of these two dimensions on the Traveling Salesman Problem (TSP). We employ a Recurrent PPO agent to dynamically regulate these parameters, treating the DRL model as a probe to reveal evolutionary dynamics. Experimental results confirm the effectiveness of this approach: the learned policies outperform static baselines, reducing the optimality gap by approximately 45% on the largest tested instance (rl5915). Building on this validated framework, our ablation analysis reveals a fundamental insight: while numerical tuning offers local refinement, structural plasticity is the decisive factor in preventing stagnation and facilitating escape from local optima. These findings suggest that future automated algorithm design should prioritize dynamic structural reconfiguration over fine-grained probability adjustment. To facilitate reproducibility, the source code is available at https://github.com/StarDream1314/DRLGA-TSP Adaptive Single-Photon 3D Imaging via Recurrent Uncertainty Estimation Rui Fan, rui lai, and Juntao Guan (Xidian University) Abstract Abstract Active single-photon sensors can timestamp the returning photons with picosecond accuracy, enabling high-resolution time-resolved 3D imaging. However, they typically employ fixed acquisition cycles in histogram-building, leading to significant power inefficiency when scene conditions enable reliable depth-estimation with fewer cycles. To tackle this, we present the first learning-based framework for adaptive acquisition through recurrent uncertainty estimation. In specific, our approach establishes a probabilistic foundation connecting photon accumulation to depth reliability via Poisson-Skellam statistics, then employs a lightweight recurrent neural network (1.8K parameters) to predict real-time uncertainty from histogram dynamics. We also contribute the ADP-SPAD dataset with over 10,000 uncertainty-labeled accumulation sequences spanning diverse imaging conditions. Experimental results demonstrate 67.3\% average cycle reduction while maintaining 98.2\% depth accuracy, with high signal-to-background scenarios requiring only 11\% of traditional cycles. Our compact architecture is amenable to on-chip integration, establishing pathways toward intelligent, power-efficient single-photon imaging systems that adapt acquisition strategies based on scene-dependent uncertainty estimates. TKG-DG-DMOEA: Temporal Knowledge Graph-Guided Domain Generalization for Dynamic Multi-Objective Disassembly Line Balancing Yilin Fang and Ziyan Fang (Wuhan University of Technology) and Kang Lin (Zhongfu Shenying Carbon Fiber Co.,Ltd.) Abstract Abstract Dynamic disassembly line balancing (D-DLB) in remanufacturing faces time-varying optimization landscapes as End-of-Life product quality fluctuates across batches. Existing disassembly task AND/OR graph (D-TAOG)-based approaches inadequately capture structural relationships among products, operations, robots, and workstations, limiting knowledge transfer. We propose TKG-DG-DMOEA, a temporal knowledge graph-guided domain generalization framework with three innovations: (1) a unified temporal knowledge graph (TKG) integrating product structures, resource capabilities, and dynamics; (2) a graph-conditioned solution generative model using multi-relational GNNs to generate disassembly plans; and (3) a meta-learning strategy discovering environment-invariant patterns from historical data. Experiments on 36 D-DLB instances against five baselines show TKG-DG-DMOEA achieves best MIGD on 14/18 small-scale cases, fastest runtime on 4/7 benchmarks, and consistently low NIGD (near-zero negative transfer) across dissimilarity levels (70% comparisons significant at p<0.05). Dynamic Clustering-Based Dynamic Multiobjective Optimization for Disassembly Line Balancing Yilin Fang and Yizhe Zhang (Wuhan University of Technology) and Xin Gu (Hubei Standardization and Quality Institution) Abstract Abstract Disassembly line balancing plays a key role in improving the efficiency and sustainability of remanufacturing systems. In practice, however, the conditions of end-of-life (EOL) products, task processing times and resource availability often change over time, which naturally leads to a dynamic disassembly line balancing problem (D-DLBP). This problem is inherently a dynamic multiobjective optimization problem (DMOP), where the Pareto optimal set (POS) and the Pareto front (PF) vary with the environment. Existing dynamic multiobjective evolutionary algorithms (DMOEAs) can react to environmental changes by preserving diversity, using multiple populations, memorizing historical solutions, or predicting future environments. Nevertheless, most of them focus primarily on the quality of the historical solutions and ignore the structural distribution and density characteristics of the historical POS samples. As a result, they may suffer from redundancy, negative transfer and insufficient adaptability when the environment changes significantly. To address this issue, this paper proposes a Dynamic Clustering-Based Dynamic Multiobjective Evolutionary Algorithm (DC-DMOEA) for dynamic disassembly line balancing. The algorithm collects POS samples from all historical environments and applies HDBSCAN, a hierarchical density-based clustering method, to extract representative cluster exemplars in the objective space. For each cluster, we define a shape descriptor that captures its mean, spread and density; based on these descriptors, we construct an environment-shape similarity measure between the new environment and existing clusters. The initial population in the new environment is then generated via weighted sampling of cluster exemplars according to their similarity. In this way, DC-DMOEA can produce environment-adaptive initial populations that simultaneously exploit stable historical knowledge and avoid overfitting to environment-specific outliers. The proposed framework is particularly suitable for discrete, constrained and combinatorial DMOPs such as D-DLBP, where the number of historical POS samples is large and their distributions are highly irregular. Tuesday Virtual Room 1 IJCNN Paper Knowledge Graphs and Question Answering II Session Chair: Jibing Wu (National University of Defense Technology), Nurmemet Yolwas (Xinjiang University) Few-Shot Temporal Knowledge Graph Completion via Frequency-Based Relation Encoding and Conditional Diffusion Lai Xu, Jibing Wu, Yanmin Li, Liwei Qian, Hang Zhang, and Lihua Liu (National University of Defense Technology) Abstract Abstract Few-shot temporal knowledge graph completion (FTKGC) has emerged as an important task for predicting missing facts in temporal knowledge graphs (TKGs) with limited training instances. However, existing methods face two major challenges: (1) They mainly focus on capturing the temporal evolution of entities, overlooking the heterogeneous patterns inherent in relations, which are characterized by a combination of short-term fluctuation and long-term stability; (2) The data sparsity problem in few-shot settings leads to inadequately optimized entity representations, consequently undermining the reliability of prediction on query quadruples. To this end, we propose frequency-based relation encoding and conditional diffusion (FRECD) for FTKGC. Specifically, FRECD presents a novel relation encoding strategy, which integrates the fast Fourier transform (FFT) to decompose relation embeddings and performs further modulation with temporal information. Regarding insufficiently informative entity representations, FRECD leverages a conditional denoising network built upon the diffusion transformer (DiT) to perform generative feature enhancement. Comprehensive experiments are carried out on three benchmark datasets to demonstrate the superiority of our model. The code is available at https://github.com/caker2077/FRECD. Multi-tendency negative samples for multi-modal knowledge graph completion based on diffusion 啸尘 曾 and 晖 赵 (Xinjiang University) Abstract Abstract Multimodal Knowledge Graph Completion (MMKGC) seeks to automatically uncover unseen true knowledge from multimodal knowledge graphs by jointly modeling entity triplets and multimodal data. However, real-world MMKGs face two key challenges: modality diversityand modality missingness.To tackle these challenges, this paper proposes a Multimodal Knowledge Graph Completion framework based on Relation-Guided Adaptive Fusion and Different Modality Tendency Negative Sampling (MTNS), featuring two core innovations: 1) adaptive fusion of any modalities via dynamic weight allocation; 2) a diffusion model-driven negative sampling strategy that generates diverse negative samples ), effectively mitigating modality missingness issues.Empirical results against SOTASOTA baselines confirm MTNS's superiority. Dual Quaternion Decoupling and Enhanced Temporal Embeddings for Temporal Knowledge Graph Completion Yue Zhao, Wenbin Zhang, Jiazheng Guo, Tianyi Xu, Mei Yu, and Mankun Zhao (Tianjin University) Abstract Abstract Temporal Knowledge Graph Completion models are proposed to predict facts at specific times. However, existing models treat temporal information as complementary to the semantic information of entity (relation) and ignore independence between them, which leads to the loss of details and accuracy of temporal information. SPaR: Instance-Conditioned Routing for Multimodal Knowledge Graph Completion Xue Ma, Nurmemet Yolwas, Lixu Sun, Yizhen Wu, and Ailing Xiong (Xinjiang University) Abstract Abstract Predicting missing links in Multimodal Knowledge Graphs (MMKGs) requires balancing structural patterns with visual and tex- tual evidence. However, most existing methods apply the same fusion strategy across all queries of a given relation, overlooking that individ- ual instances may favor entirely different modality mixes based on their local neighborhoods. Meanwhile, sparse expert models—though compu- tationally appealing—often degenerate as routers fixate on a small expert subset. This paper introduces SPaR-MoE, which conditions expert se- lection on instance-specific signals: we aggregate k-hop neighbors around the head entity via attention and combine this with relation-level pat- tern descriptors to form a rich routing key. Two auxiliary mechanisms complement the core design: RouteReg penalizes uneven expert load and low routing entropy to discourage collapse, while RobustScore ran- domly masks modalities during training so the model degrades gracefully when inputs are incomplete. Evaluations on MKG-W, MKG-Y, DB15K, and KVC16K show SPaR-MoE outperforms existing approaches, lifting average MRR by 2.45 points over the strongest baseline. Tuesday Virtual Room 2 IJCNN Paper Knowledge Graphs and Question Answering III Session Chair: Xiangfeng Luo (Shanghai University, School of Computer Engineering and Science), Zheng Lin (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) SubTabQA: Benchmarking Subtable Extraction in LLM-based Table Question Answering Hanwen Zhang and Jia Wang (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences); Peng Fu (Institute of Information Engineering, Chinese Academy of Sciences); Chuanyu Qin (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences); Zheng Lin (Institute of Information Engineering, Chinese Academy of Sciences); Chao Xu (State Grid Jiangsu Electric Power Co., Ltd); and Weiping Wang (Institute of Information Engineering, Chinese Academy of Sciences) Abstract Abstract Table Question Answering (TQA) remains challenging for Large Language Models (LLMs) due to the abundance of noisy information in tables. A prevalent solution is the Decompose-Then-Reason (DTR) paradigm, which first extracts a relevant subtable before reasoning to produce an answer. As a critical prerequisite in DTR, SubTable Extraction (STE) not only simplifies downstream reasoning but also offers a lens into model behavior. However, there is currently a lack of dedicated benchmarks for fine-grained evaluation of STE. To address this, we introduce SubTabQA, the first benchmark dedicated to assessing STE across both flat and hierarchical tables. We further develop TableFocus, a specialized STE model trained via Supervised Fine-Tuning (SFT) and Reinforcement Learning with Verifiable Rewards (RLVR). Comprehensive experiments on SubTabQA reveal that strong TQA performance does not necessarily correlate with robust STE capabilities, highlighting the need for explicit evaluation of the decomposition stage. Notably, TableFocus achieves STE performance comparable to advanced large-scale and closed-source LLMs, demonstrating the effectiveness of our data and training strategy. RoGTR: Robust Table Question Answering via Graph-based Textual Reasoning Xiaoke Guo and wen Zhang (Zhejiang University) Abstract Abstract Large language models (LLMs) have advanced table-based question answering (Table QA), but current approaches face a critical trade-off. Symbolic reasoning methods, such as addressing specific questions via SQL generation, offer precision but are brittle, often failing on structurally complex or irregular tables. Conversely, purely End-to-End textual reasoning handles flexibility but struggles with context overload, poor structural comprehension, and computational inaccuracies—key limitations identified in complex QA scenarios. To bridge this gap, we introduce Structured Textual Reasoning, a new paradigm that integrates the precision of symbolic methods with the flexibility of textual inference. We present its first implementation, RoGTR (Robust Graph-based Textual Reasoning), a framework designed to overcome these limitations. RoGTR begins by converting diverse tables into a unified, triple-based textual representation. This structured format is then processed by LLM using a dual-channel reasoning strategy. This representation allows the model to effectively filter irrelevant information and infer latent semantic relationships, thereby addressing the core challenges of Table QA without resorting to fragile, hard-coded parsers or code generation. Extensive experiments demonstrate that RoGTR achieves state-of-the-art performance on complex table benchmarks while maintaining robust accuracy on simple tables. Beyond Answer Symbols: Discovering Key Layers of VLMs for Visual Processing in Multiple Choice Question Answering Zhiling Jin (ZhejiangUniversity) and Rongge Mao (University of Science and Technology of China) Abstract Abstract While multiple-choice question answering (MCQA) has emerged as a standard paradigm for evaluating Vision-Language Models (VLMs), the internal mechanisms through which these models process visual information during decision-making remain largely opaque. Unlike text-only LLMs where critical layers for answer prediction have been established, it is unknown whether VLMs harbor specialized neural circuits dedicated to visual information encoding, or how such visual processing is hierarchically organized across transformer layers. In this study, we employ activation patching and vocabulary projection techniques to systematically investigate visual token processing in state-of-the-art VLMs. Our experiments reveal the existence of visual processing key layers—distinct, localized early-to-middle transformer layers where visual information undergoes intensive encoding and integration prior to downstream reasoning. Furthermore, these critical layers exhibit task-specific functional specialization: recognition tasks engage earlier layers (layers 14--16), while grounding and counting tasks recruit progressively deeper layers (layers 15--20). These findings delineate a hierarchical visual-to-linguistic processing architecture in VLMs, wherein distinct visual competencies activate discrete neural circuits at specific architectural depths. Bridging the Selectivity Gap in Zero-Shot Visual Question Answering via Intrinsic Answer Verification xuan tang, Xiangfeng Luo, and Liyan Ma (Shanghai University) Abstract Abstract Zero-shot visual question answering (VQA) has recently benefited from large language models that utilize generated captions as intermediate reasoning contexts.We observe that while high Answer Hit Rates (AHR) in captions are a prerequisite for visually grounded reasoning, they do not guarantee reliable final predictions. Specifically, our approach integrates two key components to address this limitation. First, we introduce Question-Guided Context Refinement (QGCR) to generate semantically aligned captions, designed to enrich the extraction of question-critical visual details. Second, to ensure the correct answer is selected, we design an intrinsic answer verification mechanism based on the internal representation consistency of a frozen vision-language encoder.Extensive experiments on GQA and OK-VQA demonstrate that our framework consistently improves final accuracy, showing the effectiveness of intrinsic verification in strictly training-free reasoning scenarios. Tuesday Virtual Room 3 IJCNN Paper Knowledge Graphs and Question Answering IV Session Chair: Mian Wu (School of Software, Beihang University), wen Zhang (Zhejiang University) MedLTRAG: Co-Augmentation of Knowledge Graphs and Large Language Models for Long-tail Medical Question Answering Mian Wu and Xu Wang (School of Software, Beihang University); Jie Sun (Zhongguancun Laboratory); and Bing Wei (School of Cyberspace Security, Hainan University) Abstract Abstract Despite the remarkable success of large language models (LLMs), their ability to effectively incorporate long-tail domain knowledge remains a fundamental challenge. In specialized domains such as medicine, reliable question answering requires structured and precise knowledge that is often poorly captured by purely parametric representations. We observe that the key limitation of existing LLM-based approaches to long-tail medical QA lies in a misalignment between how specialized knowledge is represented and how it is accessed during inference. To address this limitation, we propose MedLTRAG, a framework that tightly integrates LLMs with knowledge graphs (KGs) to enable grounded medical reasoning. MedLTRAG follows an offline-to-online paradigm: domain-specific KGs are first constructed from medical literature to encode specialized expertise, and are then selectively retrieved at inference time to support reasoning through a multi-stage beam search mechanism. By externalizing long-tail knowledge into structured representations and explicitly incorporating them into the inference process, MedLTRAG reduces the brittleness of purely parametric reasoning. Extensive experiments on multiple disease-centric medical QA benchmarks demonstrate that MedLTRAG consistently outperforms state-of-the-art methods. Beyond empirical improvements, this work outlines a generalizable paradigm for integrating structured knowledge with LLMs, offering a principled pathway toward robust reasoning in knowledge-intensive, long-tail domains. LLM-SPM:Semantic-Driven Temporal Knowledge Graph Forecasting via Large Language Models and Pattern Mining Jun Zhu (University of Electronic Science and Technology of China); Tianxiang He (Chengdu Union Big Data Tech Inc.); and Yan Fu, Junlin Zhou, and Duanbing Chen (University of Electronic Science and Technology of China, Chengdu Union Big Data Tech Inc.) Abstract Abstract Temporal Knowledge Graph (TKG) forecasting is a critical task in the field of Knowledge Graphs (KGs), aiming to predict events that may occur in the future. However, only structured information or timestamps are considered in the process of reasoning, and the semantics of entities and relations are ignored in the current models. This semantic information includes potential evolutionary laws, which are beneficial for reasoning for more accurate results. To address this limitation, we propose a novel framework driven by Large Language Models (LLMs) for enhancing TKG forecasting by Semantic Pattern Mining, named LLM-SPM. By leveraging the natural language understanding and zero-shot learning capabilities of LLMs, entity type information in TKGs is automatically extracted. Next, based on the inferred entity types and relations, semantic pattern triples are constructed. Finally, historical facts in the TKG are reweighted or filtered through a comprehensive scoring mechanism that integrates multiple factors, including temporal decay, pattern relevance, historical recurrence, and an entity specificity measure, enabling the model to focus on facts that are more relevant and semantically coherent. Through this approach, deep semantic knowledge is effectively incorporated into the prediction process, substantially improving reasoning accuracy in various scenarios. Experiments on benchmark datasets demonstrate that the LLM-SPM framework outperforms both previous methods and standard LLMs. Enhancing GNN-RAG with Adaptive Subgraph Optimization and Bidirectional Feedback for Robust Knowledge Graph Question Answering bo ma (Peking University); LuYao Liu (China University of Political Science and Law); and Xin Zhang, Din Li, and Hang Li (Peking University) Abstract Abstract Knowledge graph question answering (KGQA) has emerged as a critical task that combines structured knowledge from knowledge graphs with natural language understanding. Recent advances in GNN-RAG have demonstrated the effectiveness of combining graph neural networks (GNNs) for retrieval with large language models (LLMs) for reasoning. However, existing GNN-RAG frameworks suffer from several limitations: static subgraph extraction leads to incomplete answer coverage (only 79.3% in complex datasets), shortest path retrieval may miss critical reasoning paths, unidirectional information flow prevents model interaction, and retrieval augmentation introduces redundancy. To address these issues, we propose an enhanced GNN-RAG framework with four key innovations: (1) dynamic subgraph optimization with entity linking correction, (2) multi-strategy path filtering with cross-modal completion, (3) bidirectional feedback mechanism between GNN and LLM, and (4) intelligent fusion with domain-adaptive retrieval augmentation. Experimental results on WebQSP and CWQ benchmarks show that our approach achieves state-of-the-art performance, with subgraph answer coverage improved from 79.3% to 92.7%, and F1 scores increased by 4.5-4.6% over the baseline GNN-RAG+RA. SMuRT: A Structured Multi-Stage Rule-guided LLM Framework for Triple Set Prediction Yuan Yuan, Juan Li, Songze Li, Huajun Chen, and Wen Zhang (Zhejiang University) Abstract Abstract Triple Set Prediction (TSP) aims to discover missing facts in Knowledge Graphs (KGs) from scratch, which is more challenging and realistic than conventional KG Completion (KGC) tasks (e.g., link prediction) that rely on partially known triple elements. While conventional embedding and rule-based methods lack semantic flexibility and require frequent retraining, Large Language Models (LLMs) offer a transformative paradigm for TSP with their advanced semantic reasoning capabilities, training-free adaptability, and ability to harness external knowledge beyond graph facts. However, direct application of LLMs to TSP is hindered by unguided discovery and factual hallucinations. To address these challenges, our key insight is to bridge the semantic-symbolic gap by transforming LLMs from unconstrained generators into rule-guided reasoners. We propose SMuRT, a Structured Multi-Stage Rule-guided LLM framework that implements a ``plan-retrieve-validate'' pipeline. SMuRT synergizes a Hybrid Rule Curation Engine, Structured Task Planning, Retrieval-augmented Prediction, and Dual-stage Validation to enforce logical grounding. This architecture allows LLMs to capture latent semantic patterns while remaining anchored in verifiable graph evidence, effectively mitigating hallucinations. Experiments on Wiki79k and CFamily benchmarks demonstrate that SMuRT achieves state-of-the-art performance for LLM-based TSP, effectively merging generative reasoning with specialized structural inference. Tuesday Virtual Room 4 IJCNN Paper Knowledge Graphs and Question Answering V Session Chair: Jiebin Huang (South China University of Technology), yining liu (shandong university) StrucNID: LLM-Guided Structural Prior for Disentangled New Intent Discovery Jiebin Huang, Jindian Su, and Ruiqi Wang (South China University of Technology) and Xiaobin Ye and Dandan Ma (Guangdong Unicomm) Abstract Abstract New Intent Discovery (NID) seeks to recognize both known and novel user intents from unlabeled utterances, serving as a fundamental capability for dialogue systems operating in dynamic environments. Existing methods confront two critical challenges, notably Semantic Entanglement, characterized by monolithic embeddings conflating intent with stylistic variations to cause intra-class scatter, and Structural Deficiency, stemming from the lack of topological guidance in LLM-Enhanced approaches that leaves boundary samples ambiguous. To address these issues, we propose StrucNID, a framework explicitly designed for structural disentanglement. We leverage LLMs to construct fine-grained Semantic Mini-Clusters as a binary adjacency graph, providing reliable topological constraints rather than simple data augmentation. We then enforce strict orthogonality between intent and instance subspaces via dual-stream projection, mathematically separating class-invariant semantics from stylistic noise. Finally, we design a Spatio-Temporal voting mechanism that fuses temporal momentum with spatial consensus to rectify boundary ambiguities. Extensive experiments on three benchmarks demonstrate that StrucNID achieves consistent and significant performance improvements across diverse datasets, validating the efficacy of our structural disentanglement paradigm. Topology-Enhanced Experts for Continuous Intent Trigger Modeling in Disentangled SLU Chuanwang Xiong and Hongwei Ge (Jiangnan University) Abstract Abstract Multi-intent Spoken Language Understanding (SLU) is critical for deciphering complex user utterances in real-world dialogue systems. However, current methods are constrained by two fundamental limitations: the entanglement of multi-granularity semantics leading to feature collapse, and the indiscriminate treatment of tokens which ignores the structural locality of intent triggers. To address these challenges, we propose TECT (\textbf{T}opology-enhanced \textbf{E}xperts for \textbf{C}ontinuous \textbf{T}riggers), a unified framework that injects explicit inductive biases for granular disentanglement and structural awareness. Our approach introduces two key mechanisms: 1) \textbf{Orthogonal Multi-view Experts (OME)}, which integrates utterance-, chunk-, and token-level semantics. Crucially, OME imposes a rigorous pairwise orthogonality constraint to compel experts to capture complementary, non-redundant representations within the semantic manifold; and 2) \textbf{Topology-Aware Trigger Selection (TATS)}, which redefines the selection process to identify coherent semantic spans. TATS incorporates a relative positional advantage bias to leverage topological structure and employs adaptive sparsity to dynamically modulate the focus area based on intent complexity. Extensive experiments on the MixATIS and MixSNIPS benchmarks demonstrate that TECT establishes new state-of-the-art performance, highlighting the superior efficiency and robustness of our structural design in capturing multi-granularity linguistic patterns. A Pointwise Method for Search Result Diversification via Intent Matching Chunyang Li (Xidian University), Zhichao Huang (JD Intelligent Cities Research), and Xubo Qin (Renmin University of China) Abstract Abstract Search result diversification aims to mitigate query ambiguity by presenting documents that collectively satisfy multiple user intents. However, many supervised diversification models focus on listwise interaction modeling while implicitly assuming that the initial candidates already provide high-quality intent signals; in practice, weak intent matching introduces noise that limits the effectiveness of even sophisticated diverse rankers. We propose Intent Matching Enhanced Diversification (IMEDiv), a decoupled two-stage framework that explicitly separates intent matching from diverse re-ranking. In the first stage, we fine-tune an instruction-following, decoder-only LLM as a pointwise intent matcher to score each document by its intent coverage given a query and its subtopics, producing an intent-enhanced ranking list. In the second stage, a lightweight subtopic-aware ranker reduces redundancy on top of these intent-enhanced candidates. Importantly, this stage can be replaced by existing diverse ranking models, making IMEDiv a plug-and-play relevance module. Experiments on two public datasets demonstrate that strengthening intent matching alone is highly effective IMEDiv (Intent) improves $\alpha$-nDCG@20 from 0.452 to 0.467 over DSSA), and combining it with diverse re-ranking further improves diversification (IMEDiv (Intent+Div) reaches $\alpha$-nDCG@20 0.476). Detailed analyses show that the gains primarily come from improving intent coverage at top ranks, highlighting the overlooked but foundational role of intent matching in diversification. Ambiguity-Aware Multi-Round Text-Guided Image Editing via Intent Knowledge Graph Compilation Yining Liu (Shandong University) Abstract Abstract Text-driven image editing with diffusion models has achieved significant progress, and most existing methods assume that user instructions are clear and directly actionable. However, in real interactions, users often provide vague or underspecified requests (e.g., “make it more cinematic” or the polysemous word “bank”). To address this problem, we propose DiaKG-Edit, an ambiguity-aware multi-round dialogue image editing framework with two components: Ambiguity-Aware Modeling (AAM) and Graph-to-Token Compilation (G2T). In AAM, we maintain an Editable Knowledge Graph (EKG) as a persistent representation of editing intent, which explicitly records entities, attributes, relations, and their strengths and is incrementally updated across interaction rounds. For each incoming instruction, the system determines whether it is clear or ambiguous, applying conservative updates in clear cases and resolving ambiguity via dialogue-driven refinement, thereby progressively updating the EKG while preserving previously confirmed edits. Based on the stabilized intent representation, G2T maps incremental differences between successive EKG states to executable token-level control strategies for diffusion-based editors, enabling deterministic and stable multi-round execution. Experiments show that DiaKG-Edit outperforms HRV under ambiguous instructions, improving CLIP instruction alignment from 0.77 to 0.79 (+2.6%) and reducing LPIPS from 0.152 to 0.147 (-3.3%). Tuesday Virtual Room 5 IJCNN Paper Knowledge Graphs and Question Answering VI Session Chair: Ziqiong Liu (Southern University of Science and Technology), Yang Liu (Shanxi University) Structure Matters: Principled Dual-View Augmentation for Unsupervised Graph OOD Detection Wenping Zheng, Yuanyuan Guo, Yuzhi Wang, and Yang Liu (Shanxi University) Abstract Abstract Unsupervised Graph Out-of-Distribution (OOD) detection is pivotal for ensuring model reliability in open-world scenarios, particularly absent label supervision. Existing approaches predominantly rely on random perturbations, such as feature masking and edge dropout, to construct augmented views. However, these heuristic strategies often disrupt semantic consistency—inadvertently introducing noise rather than informative variance—and fail to capture the intricate interplay between homophilic and heterophilic topologies. To address these limitations, we propose HGOOD, a novel framework featuring a principled dual-view augmentation strategy via hierarchical structural disentanglement. Specifically, HGOOD utilizes an invariant subgraph extractor to decompose the graph into distribution-stable invariant and noise-sensitive environmental subgraphs. A learnable edge separator then refines the invariant structure into homogeneous and heterogeneous views, which are strategically perturbed by the environmental subgraph to generate complementary, semantically faithful augmentations. Furthermore, we integrate a hierarchical contrastive learning scheme with differentiable Softmax-based clustering, enabling the adaptive optimization of group-level representations against dynamic feature distributions. Extensive experiments on multiple graph classification benchmarks demonstrate that HGOOD consistently outperforms state-of-the-art baselines, validating its effectiveness in robust OOD detection. DICE: A Dynamic Subgraph and Interaction–Type Collaborative Enhancement Framework for One-shot Subgraph Reasoning Zhiyong Wang, Di Ying, and Hui Zhao (Xinjiang University) Abstract Abstract Large-scale knowledge graph completion (KGC) must balance inference efficiency and prediction reliability. Although one-shot subgraph reasoning greatly reduces computation, existing methods often use fixed-budget heuristic sampling that cannot adapt to query-specific structural distributions,causing either insufficient evidence coverage or redundant noise.Moreover, ranking candidates mainly by subgraph propagation may yield false positives that are structurally reachable but semantically or type-inconsistent, degrading top-rank quality.We propose a dynamic subgraph and interaction–type collaborative enhancement framework for one-shot subgraph reasoning(DICE). DICE determines subgraph size via a cumulative Personalized PageRank (PPR) mass threshold with a minimum size constraint and key-node retention. On top of a replaceable subgraph reasoning backbone, DICE adds an Entity Interaction Module (EIM) for relation-conditioned interaction patterns and a Type-Aware Constraint (TAC) for tail-type preferences, with a gating mechanism to fuse evidence adaptively. Experiments on WN18RR,NELL-995, and YAGO3-10 show that DICE outperforms strong baselines; ablations verify complementary gains; and subgraph/backbone tests demonstrate robustness and transferability. MTAM: Multi-Level Trustworthiness Assessment via Neurosymbolic Synergy for Knowledge Graph Error Detection Jinzhu Wang, Yongchao Gao, and Heng Qian (Qilu University of Technology (Shandong Academy of Sciences), Shandong Computer Science Center (National Supercomputer Center in Jinan)) and Yan Wang (Shandong Standard Institute of Emerging Technologies and Innovations) Abstract Abstract Automated knowledge graph (KG) construction inevitably introduces noise, resulting in inaccurate or logically inconsistent triples. Existing error detection methods typically rely on either structural embeddings or static constraints, often failing to capture the complex interplay between graph topology and ontological logic. To address this limitation, this paper proposes a Multi-level Trustworthiness Assessment Model (MTAM) via neurosymbolic synergy. Unlike static fusion approaches, our framework dynamically integrates three complementary dimensions: (1) Entity Association Strength, which filters low-confidence links based on local connectivity patterns; (2) Hierarchical Path Reasoning, which validates logical plausibility over multi-hop networks using an adaptive attention mechanism; and (3) Type-Structure Consistency, a novel metric that employs dual hypergraphs to align neural structural embeddings with symbolic type constraints. This dual-hypergraph design effectively detects sophisticated semantic anomalies where structural existence contradicts entity types. Extensive experiments on FB15K and YAGO43K demonstrate that MTAM significantly outperforms state-of-the-art baselines in both error detection precision and ranking correlation, validating its effectiveness in identifying diverse noise patterns. RASA: Reliability-Aware Selective Adaptation via Hybrid Reliability Threshold Ziqiong Liu, Yijia Wang, Junyang Ji, Yushun Tang, and Zhihai He (Southern University of Science and Technology) Abstract Abstract Test-time adaptation (TTA) has emerged as a highly promising paradigm for adapting pre-trained models to target domains. However, in open-world scenarios, standard TTA often suffers from noise overfitting due to indiscriminate entropy minimization, especially when faced with severe semantic shifts and out-of-distribution (OOD) samples. To address these challenges, we propose RASA, a Reliability-Aware Selective Adaptation framework that employs a hybrid reliability threshold for selective test-time adaptation. The core of RASA is a novel History-Aware Thresholding (HAT) module, which integrates hybrid prior information by combining theoretical class priors with running historical entropy statistics, and dynamically adjusts the reliability criteria as the adaptation process progresses. Furthermore, we introduce a dual-stream Transformer architecture that explicitly decouples semantic classification and OOD inference into different feature subspaces, significantly enhancing sensitivity to outliers while maintaining accuracy on in-distribution samples. Extensive experiments demonstrate that RASA consistently outperforms existing state-of-the-art methods, exhibiting superior stability and robustness in open-world environments with severe data perturbations and a high proportion of OOD samples. Tuesday Virtual Room 6 IJCNN Paper LLM Adaptation and Fine-Tuning V Session Chair: Wenge Rong (School of Computer Science and Engineering, Beihang University, China; Engineering Research Center of Integration and Application of Digital Learning Technology Ministry of Education,China), baojun tian (Inner Mongolia University of Technology, Inner Mongolia Key Laboratory of Intelligent Perception and System Engineering) HDAF-LLM: Hierarchical Decoupled Adaptive Fusion with Large Language Models for Multi-Domain Recommendation LiChang zhao (Inner Mongolia University of Technology); baojun tian (Inner Mongolia University of Technology, Inner Mongolia Key Laboratory of Intelligent Perception and System Engineering); and pengyu chen (Inner Mongolia University of Technology) Abstract Abstract Multi-domain recommendation (MDR) leverages shared user interests across domains to improve personalized recommendations. However, existing methods face two critical challenges: (1) cross-domain preference interference, where naive fusion causes negative transfer, and (2) semantic-structural misalignment between LLM-based features and graph-based collaborative patterns. Prior approaches address these separately through implicit regularization or static integration, lacking uni- fied solutions with explicit guarantees. We propose HDAF-LLM (Hierarchical Decoupled Adaptive Fusion with Large Language Models), a novel framework that systematically tackles both challenges. First, we introduce explicit architectural decoupling through separate Global-HGCN and Domain-HGCN modules with isolated parameter sets, ensuring zero cross-domain gradient interference. Second, we design layer-wise contrastive learning with augmentation-based hard negatives to progressively align LLM semantics with graph embeddings. Third, we develop a two-stage adaptive fusion mechanism with density-aware weights that dynamically balances domain-specific and global informa- tion. Extensive experiments on six Amazon sub-domain datasets demonstrate that HDAF-LLM achieves statistically significant improvements of 1.15%–6.78% over state-of-the-art baselines (paired t-test, p < 0.05), with particularly notable gains of 7.8%– 19.1% in cold-start scenarios. D-Con: Domain-Partitioned Contrastive Learning Framework for Mitigating Domain Interference in Multi-Domain Text Mining Jiarui Zhang, Mingzhe Lu, and Yifan Deng (Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China) Abstract Abstract Large-scale text mining mixes heterogeneous domains, where semantically similar yet domain-conflicting samples distort the embedding space and hurt generalization. Standard contrastive learning lacks explicit mechanisms to separate domain semantics. We present D-Con, a two-stage intermediate pre-training framework with Domain-Partitioned Hard Negative Mining (D-PHNM), which retrieves cross-domain hard negatives via partitioned FAISS and optimizes a tailored contrastive loss to form domain-disentangled subspaces. During downstream adaptation, parameter-efficient fine-tuning (e.g., LoRA) can leverage these subspaces without test-time domain labels. On multi-domain machine translation and sentiment analysis, D-Con improves in-domain accuracy and reduces forgetting versus strong baselines in our settings, offering a scalable solution for heterogeneous data mining. SLoRA: Shared Low-Rank Adaptation for Parameter-Efficient Fine-Tuning of Large Language Models Qingjun Zhang, Ji Feng, and Ruisheng Ran (College of Computer and Information Science, Chongqing Normal University) Abstract Abstract Low-Rank Adaptation serves as the predominant framework for parameter-efficient fine-tuning of large language models. Despite its effectiveness, standard implementations introduce considerable storage and parameter requirements that may be prohibitive in resource-constrained environments. This work introduces Shared Low-Rank Adaptation, an approach designed to facilitate global parameter sharing across homogeneous modules situated in distinct layers. The methodology involves the construction of a shared weight pool dedicated to modules of identical types. Through the application of selection and splice operators, the framework synthesizes layer-specific matrix pairs characterized by diverse effective ranks derived from the common pool. This architectural design substantially diminishes the total parameter budget while maintaining robust adaptation performance. Comprehensive empirical evaluations indicate that the proposed method achieves a more favorable balance between parameter economy and fine-tuning efficacy than previous variants. Leveraging Fine-Tuned Large Language Models for Personalized Feedback in Undergraduate Object-Oriented Programming Education Zuchen Li (School of Computer Science and Engineering, Beihang University, China); Ji Wu (School of Computer Science and Engineering, Beihang University, China; Engineering Research Center of Integration and Application of Digital Learning Technology, Ministry of Education,China); Qitong Liu (School of Cyber Science and Technology, Beihang University,China); Qing Sun (School of Computer Science and Engineering, Beihang University, China; Engineering Research Center of Integration and Application of Digital Learning Technology, Ministry of Education,China); Jing Li (Engineering Research Center of Integration and Application of Digital Learning Technology, Ministry of Education, China; The Open University of China); and Wenge Rong (School of Computer Science and Engineering, Beihang University, China; Engineering Research Center of Integration and Application of Digital Learning Technology, Ministry of Education,China) Abstract Abstract In undergraduate software engineering education, architectural design courses are pivotal for fostering higher-order engineering skills through complex, iterative projects. However, providing effective feedback in these courses remains challenges: traditional static analysis verifies functional correctness but cannot evaluate architectural quality, while manual review is unscalable for large student cohorts. To address these challenges, this paper proposes a generalizable LLM-based feedback framework, named Helper, which integrates three core modules: course-adaptable data preparation; lightweight fine-tuning for general models; and context-aware recommendation generation. This paper also presents a dedicated fine-tuning dataset tailored to the Object-Oriented Design (OOD) course. The Helper's specific instantiation of the OOD course is named OO-Helper. Experimental validation demonstrates that OO-Helper outperforms general foundation models in modification accuracy and design compliance. Qualitative feedback from students confirms that it can act as an effective pedagogical scaffold. The modular framework enables seamless cross-disciplinary migration with minor adjustments, offering broad practical value for CS education. Tuesday Virtual Room 7 IJCNN Paper LLM Adaptation and Fine-Tuning VI Session Chair: Yixian Kong (Beijing University of Posts and Telecommunications), Shayok Chakraborty (Florida State University) An Analysis of Active Learning Algorithms using Real-World Crowd-sourced Text Annotations Varun Totakura, Ankita Singh, Yushun Dong, and Shayok Chakraborty (Florida State University) Abstract Abstract Active learning algorithms automatically identify the most informative samples from large amounts of unlabeled data and tremendously reduce human annotation effort in inducing a machine learning model. In a conventional active learning setup, the labeling oracles are assumed to be infallible, that is, they always provide correct answers (in terms of class labels) to the queried unlabeled instances, which cannot be guaranteed in real-world applications. To this end, a body of research has focused on the development of active learning algorithms in the presence of imperfect / noisy oracles. Existing research on active learning with noisy oracles typically simulate the oracles using machine learning models; however, real-world situations are much more challenging, and using ML models to simulate the annotation patterns may not appropriately capture the nuances of real-world annotation challenges. In this research, we first collect annotations of text samples (from 3 benchmark text classification datasets) from crowd-sourced workers through a crowd-sourcing platform. We then conduct extensive empirical studies of 8 commonly used active learning techniques (in conjunction with deep neural networks) using the obtained annotations. Our analyses sheds light on the performance of these techniques under real-world challenges, where annotators can provide incorrect labels, and can also refuse to provide labels. We hope this research will provide valuable insights that will be useful for the deployment of deep active learning systems in real-world applications. The obtained annotations can be accessed at https://github.com/varuntotakura/al_rcta/. Mitigating Prompt Bias: Representativeness-Aware Active Learning for VLM Prompt Tuning Hanlei Li, Peng Wang, Menglin Yang, and Jun Zhou (Southwest University) Abstract Abstract Prompt learning is the leading paradigm for adapting Vision-Language Models (VLMs), yet its performance hinges on high-quality supervision. While Active Learning (AL) can mitigate data scarcity, we identify a fundamental misalignment in traditional uncertainty-based strategies, termed Prompt Bias. Specifically, prompt learning seeks global class centroids, whereas conventional AL prioritizes decision-boundary outliers, causing prompts to degenerate into descriptors of anomalies and leading to feature space drift. To address this, we propose Distribution-Aware Representative Active Learning (DARAL), which shifts the selection focus from difficulty mining to representativeness. DARAL integrates a rarity-weighted metric with Stratified Pseudo-Label Sampling to explicitly target intra-class diversity. By partitioning the unlabeled pool into category-specific strata, our approach ensures the learned prompt robustly covers both typical prototypes and long-tail instances. Extensive experiments demonstrate that DARAL consistently outperforms state-of-the-art methods, achieving significant gains particularly under extremely low-budget constraints. Adaptive Regional Feature Fusion for Brain WSI Iron Concentration Grading in Neurodegeneration Xiaomin Liu (Jilin University); Bo Yu (Xiaomi Corporation); Daniëlle Toen, Louise van der Weerd, and Jouke Dijkstra (Leiden University Medical Center); and Yunke Zhang, Jiuman Song, and Hechang Chen (Jilin University) Abstract Abstract Iron accumulation in cerebral gray matter is a critical research focus in neurodegenerative diseases. Whole-slide images (WSIs) offer new opportunities for iron grading, yet manual WSI scoring remains subjective and prone to bias. Developing objective deep learning models for iron concentration grading faces two key challenges. First, natural tissue heterogeneity causes stained white matter to appear darker than gray matter, leading to the misclassification of normal white matter as iron-rich gray matter. Second, non-uniform iron deposition across cortical layers makes it difficult to capture global distribution patterns beyond sparse local cues, often biasing predictions toward a few high response patches. To address these limitations, we propose an Adaptive Regional Feature Fusion (ARFF) framework. ARFF employs a discriminative region mask extraction module to isolate gray matter and suppress structural and background noise. It then leverages cluster-based adaptive feature sampling and fusion weights learned via proximal policy optimization (PPO) to construct a comprehensive WSI-level representation. Experiments show that ARFF outperforms state-of-the-art methods, delivering average AUC gains of 9% on the Frontotemporal Dementia dataset and 16.4% on the Alzheimer’s Disease dataset, thereby helping reduce subjectivity in neuropathological assessment. Code: https://github.com/xiaominLiu2001/ARFF. FD-MIL: Frequency-Guided Differential Attention for Whole-Slide Image Classification Yixian Kong (Beijing University of Posts and Telecommunications) Abstract Abstract Multiple instance learning (MIL) has gained increasing adoption in the classification of histopathology whole slide images (WSIs). However, unlike supervised learning, MIL relies solely on bag-level labels, introducing interference from challenging negative samples, such as those morphologically similar to positive cancer tissues or affected by tissue staining variations. These interferences are difficult to characterize in the visual domain but can be effectively separated in the frequency domain. To address this challenge, this paper proposes a framework that leverages global frequency-domain information to extract critical features, uses inter-instance difference in frequency domain to realign the distribution distance between positive and negative samples, thereby mitigating the interference caused by sample preparation variations, and simultaneously fuses frequency-domain features with visual spatial features to collectively enhance the classification capability of the model. Experiments conducted on the TGA-BRCA, TCGA-NSCLC, and Camelyon16 datasets demonstrate substantial performance improvements, validating that the frequency-spatial fusion approach effectively addresses MIL challenges in WSI analysis by leveraging spectral feature separability. Tuesday Virtual Room 8 IJCNN Paper LLM Adaptation and Fine-Tuning VII Session Chair: Jianwei Xu (Sichuan University), Yicheng Pan (Hangzhou Institute for Advanced Study,UCAS; University of Chinese Academy of Sciences) Beyond Paper-to-Paper: Structured Profiling and Rubric Scoring for Paper-Reviewer Matching Yicheng Pan (Hangzhou Institute for Advanced Study,UCAS; University of Chinese Academy of Sciences); Zhiyuan Ning (Westlake University); Ludi Wang (Computer Network Information Center, Chinese Academy of Sciences; University of Chinese Academy of Sciences); and Yi Du (Hangzhou Institute for Advanced Study,UCAS; Computer Network Information Center, Chinese Academy of Sciences) Abstract Abstract As conference submission volumes continue to grow, accurately recommending suitable reviewers has become a critical challenge. Most existing methods follow a ``Paper-to-Paper'' matching paradigm, implicitly representing a reviewer by their publication history. However, effective reviewer matching requires capturing multi-dimensional expertise, and textual similarity to past papers alone is often insufficient. To address this gap, we propose P2R, a training-free framework that shifts from implicit paper-to-paper matching to explicit profile-based matching. P2R uses general-purpose LLMs to construct structured profiles for both submissions and reviewers, disentangling them into Topics, Methodologies, and Applications. Building on these profiles, P2R adopts a coarse-to-fine pipeline to balance efficiency and depth. It first performs hybrid retrieval that combines semantic and aspect-level signals to form a high-recall candidate pool, and then applies an LLM-based committee to evaluate candidates under strict rubrics, integrating both multi-dimensional expert views and a holistic Area Chair perspective. Experiments on NeurIPS, SIGIR, and SciRepEval show that P2R consistently outperforms state-of-the-art baselines. Ablation studies further verify the necessity of each component. Overall, P2R highlights the value of explicit, structured expertise modeling and offers practical guidance for applying LLMs to reviewer matching. SWIS: An Efficient Structured Pruning Framework for LLMs with Spectral Geometric Alignment Lan Xu and Ping Li (Changsha University of Science and Technology) Abstract Abstract Network pruning is a promising approach to alleviate the computational burden of deploying and serving large language models (LLMs). Training-free paradigms are particularly appealing in this context due to the prohibitive cost of retraining. However, most existing training-free methods implicitly rely on magnitude-based criteria that effectively ignore the anisotropic geometry of LLM representations, overlooking the directional structure of information flow and struggling to adapt to modern architectures. In this paper, we propose SWIS (Spectral-Weighted Importance Scoring), a training-free structured pruning framework that offers a geometry-aware perspective on LLM compression. SWIS employs Eigenvalue Decomposition (EVD) on activation covariance matrices to identify geometric directions governing dominant information flow. By integrating weight magnitudes with geometric alignment to high-energy principal subspaces, SWIS quantifies true contribution of channels, distinguishing critical signals from orthogonal noise. Furthermore, to optimize layer-wise sparsity allocation, SWIS-Pruner uses Feature Angular Distance to quantify intensity of geometric transformations across layers. Extensive evaluations on language benchmarks demonstrate that, without retraining, SWIS outperforms state-of-the-art baselines, including LLM-Pruner and structured extensions of Wanda. Budgeted Cascaded Routing for LLM-Based Structured Extraction from Noisy Conversations Tianzhe Ning and Guohua Liu (Donghua University) Abstract Abstract Structured extraction from noisy conversational transcripts is challenging due to fragmented cross-turn evidence and evolving taxonomies. Although large language models (LLMs) achieve strong accuracy, exhaustive inference is prohibitively costly, while lightweight rules or retrieval-based methods lack robustness. We propose a budget-constrained cascaded inference framework that constructs context blocks, screens (block, slot) candidates through dense semantic matching and lexical gating, and invokes the LLM only for routed candidates, followed by slot normalization and order-level aggregation. The framework also supports an optional uncertainty-based deferral mechanism. We release a de-identified benchmark with standardized splits, dual-granularity evaluation, and strong baselines. Under matched calls/order budgets, our cascade consistently outperforms competitive baselines while substantially reducing LLM usage. MGLS: Malicious User Detection on Movie Review Platforms with Multi-Channel Graphs and LLM-assisted Semantic Guidance Jianwei Xu and Haizhou Wang (Sichuan University) Abstract Abstract As online movie rating platforms play an increasingly important role in users' viewing decisions and movie marketing, rating manipulation and spam flooding driven by commercial or adversarial motives have become increasingly severe, making it urgent to identify malicious users in the complex user–movie–review interaction network. However, existing studies in the movie review domain focus mainly on individual reviews and have not yet addressed malicious user detection at the user level. To fill this gap, we construct a user-level malicious user dataset based on data from the Douban movie platform and propose a malicious user detection model that uses multi-channel graphs and LLM-assisted semantic guidance (MGLS). Specifically, we first build user representations by integrating BERT-based review embeddings with sentiment intensity and off-topic probability. From four perspectives—sentiment, content, temporal patterns, and behavior—we construct multi-channel user relation graphs and use multiple GCNs to learn channel-specific representations. Finally, we design an LLM-assisted semantic guidance module and a dual-stream attention fusion mechanism that adaptively fuses GCN-learned channel weights with movie-level semantic channel weights produced by the LLM, enabling collaborative modeling of multi-channel information and robust malicious user detection. Experimental results show that our model is more effective than state-of-the-art baselines with an F1-score of 87.16% for malicious user detection in movie reviews and provides a useful reference for future work on malicious user detection in movie review platforms. Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V7 Session Chair: Zhenan He (Sichuan University) A System-Level Multi-Objective Framework for Solar Tracking with Chaotic ABC MPPT Justin London and Tingjun Lei (University of North Dakota), Lin Gong (Prairie View A&M University), Guoming Li (University of Georgia), and Chaomin Luo (Mississippi State University) Abstract Abstract Solar tracking systems operating in dynamic and partially shaded environments require not only effective maximum power point tracking (MPPT) but also system-level trade-offs among energy yield, tracking accuracy, and control effort. Most existing approaches address MPPT or mechanical tracking optimization in isolation, lacking an integrated formulation that captures these competing objectives within a unified framework. In this paper, solar tracking is formulated as a system-level multi-objective optimization problem, and a hierarchical optimization framework is proposed in which a Chaotic Artificial Bee Colony (CABC)-based MPPT controller is integrated as an inner-loop power optimization mechanism. The outer-loop optimization explicitly accounts for multiple competing objectives, including power maximization, tracking stability, and control performance, subject to physical and operational constraints of a multi-axis solar tracking system. The proposed framework is evaluated through simulation and experimental studies. Results show that the CABC-based MPPT improves convergence and reduces steady-state oscillations under dynamic irradiance and partial shading conditions, while system-level multi-objective optimization demonstrates Pareto trade-offs across competing performance metrics. Geometric Modeling and Global Optimization for Plant Line Identification in Agricultural Orthomosaics Willian Moreira Antunes, Bruno Moraes Rocha, Emilia Alves, Juliana Felix, and Fabrizzio Soares (Universidade Federal de Goiás) Abstract Abstract Accurate detection of planting lines in high- resolution orthomosaics is a prerequisite for many precision agriculture applications, yet existing methods struggle with field- specific challenges such as vegetation gaps, segmentation artifacts, and partial occlusions. We propose a hybrid pipeline that first segments crop rows by exploiting the HSV color space and a lightweight K-Nearest Neighbors (KNN) classifier, followed by morphological refinement to suppress noise.The resulting binary mask is then interpreted as a set of parallel lines whose orienta- tion, spacing, and offset are jointly optimized in a global paramet- ric framework. A compound cost function—penalizing deviations from the mask, penalizing large residuals, and encouraging uniform spacing—is minimized using Differential Evolution, a population-based optimization algorithm that is robust to local minima. Because the method relies only on generic color cues and a simple line model, it does not require large annotated datasets. Experiments on a diverse set of crop fields show that our ap- proach achieves line coverage exceeding 95%, with high accuracy and improved performance over segmentation-only approaches in challenging scenarios. These results demonstrate that the proposed technique is well suited for autonomous navigation, gap detection, and other downstream tasks in precision agriculture. Index Terms—Differential evolution, Combinatorial optimiza- tion, Real-world applications, Intelligent Agriculture. Predicting When LLMs Enhance Evolutionary Optimization: A Landscape-Aware Framework Dongxin Guo (The University of Hong Kong), Jikun Wu (Brain Investing Limited), and Siu Ming Yiu (The University of Hong Kong) Abstract Abstract Large Language Models (LLMs) have emerged as promising components in evolutionary optimization, yet understanding when and why LLMs improve optimization performance remains a fundamental open question. This paper establishes a landscape-aware prediction framework connecting fitness landscape properties to LLM-assisted optimization effectiveness through the lens of search bias alignment. We introduce the concept of LLM-induced search bias and formalize its alignment with problem landscape structure through a novel landscape-operator compatibility function. Building upon classical autocorrelation analysis and fitness-distance correlation theory, we establish three analytical results: (1) necessary landscape conditions for positive LLM contribution with theoretical justification, (2) a convergence characterization under verifiable drift conditions, and (3) an information-theoretic bound on maximum optimization acceleration. Comprehensive multi-dimensional validation across d in {10, 30, 50, 100} and two benchmark suites (BBOB and CEC 2022) demonstrates that our prediction framework achieves 93.1% overall accuracy (Pearson r = 0.91, p < 0.001) across 144 benchmark configurations, validated against six state-of-the-art baselines including CMA-ES, L-SHADE, jSO, and NL-SHADE-RSP. Multi-LLM validation across GPT-4-turbo, GPT-3.5-turbo, and Claude-3-Sonnet confirms framework generalization. This work provides both theoretical understanding and practical deployment guidelines for LLM-assisted evolutionary optimization. Evolutionary Transformer Architecture Search via LLM-Guided Coarse-to-Fine Transformations Lei Liu, Gary Yen, Chenyang Liu, and Zhenan He (Sichuan University) Abstract Abstract Practical Transformer architecture search requires navigating a highly discrete and strongly coupled design space under limited computational budgets, while achieving an effective trade-off between model performance and complexity. However, in the absence of explicit search directionality, existing methods often generate a large proportion of infeasible candidates under resource constraints, thereby substantially reducing search efficiency. To address this issue, we propose CF-TAS, an LLM-guided coarse-to-fine framework for constrained Transformer architecture search. Built upon a multi-objective evolutionary backbone, CF-TAS employs a fine-tuned LLM to generate direction-aware architecture edit scripts, rather than directly synthesizing complete architectures. Specifically, Directional Architecture Transformation (DAT) generates edit scripts for a small set of seed architectures during initialization to improve the feasibility of the initial population, and is further applied to representative parents, including knee and boundary solutions, during iterative search to prioritize exploration of critical regions on the Pareto front. To handle infeasible candidates, Directional Local Refinement (DLR) performs lightweight local corrections that reduce constraint violation while largely preserving the original architecture. The LLM-guided branch operates in parallel with conventional evolutionary search, and all candidates are evaluated under a unified constraint-handling and environmental selection mechanism. Experiments on ImageNet-1K under two parameter budgets show that CF-TAS achieves performance comparable to, or better than, representative state-of-the-art baselines. These results demonstrate the effectiveness of combining evolutionary exploration with LLM-guided directional transformations for constrained Transformer architecture search. Tuesday Virtual Room 1 IJCNN Paper LLM Adaptation and Fine-Tuning VIII Session Chair: zhichao zhang (Northeast Electric Power University), Lindong Wang (Shanghai Jiao Tong University) CREF-LLM: Contrastive Representation Encoding with Frozen Language Models for Few-Shot Wind Power Forecasting Jingdong Wang and Zhichao Zhang (School of Computer Science Northeast Electric Power University), Lina Zhou (Beijing Institute of Computer Technology and Application), and Fanqi Meng (School of Computer Science Northeast Electric Power University) Abstract Abstract Accurate wind power forecasting is critical for power system scheduling and electricity market operations. However, newly deployed wind farms face significant challenges due to limited historical data, as conventional forecasting methods typically require substantial training samples and tend to overfit in few-shot scenarios. Pre-trained large language models (LLMs) offer a promising solution—frozen LLM backbones can leverage pre-trained sequence modeling capabilities while avoiding overfitting to limited data. Yet, existing LLM-based methods primarily focus on general-purpose adaptation strategies without designs specific to wind power forecasting: they lack mechanisms for learning robust temporal representations from limited data, and fail to address the unique characteristics of wind power such as cross-variable dependencies and inherent periodicity. To address these limitations, we propose CREF-LLM, a framework that integrates a contrastive temporal encoder with a frozen language model backbone. Contrastive pre-training enables learning robust temporal representations without requiring external datasets, while multivariate joint encoding captures cross-variable dependencies and explicit temporal features model periodic patterns. Experiments on two real-world wind farms with different operational characteristics demonstrate that, with only 30 days of training data, our method achieves consistent superiority over state-of-the-art methods across all forecasting horizons. CuraLight: Debate-Guided Data Curation for LLM-Centered Traffic Signal Control Qing Guo, Xinhang Li, Junyu Chen, Zheng Guo, Shengzhe Xu, Lin Zhang, and Lei Li (Beijing University of Posts and Telecommunications) Abstract Abstract As a core component of intelligent transportation systems (ITS), traffic signal control (TSC) mitigates congestion, reduces emissions, and improves travel-time reliability. Recent TSC methods based on rules, reinforcement learning (RL), and large language models (LLMs) have improved adaptivity across urban networks. Yet key limitations persist: rule-based methods and RL methods provide limited interpretability of timing actions; LLM-based methods face challenges in selecting high-dimensional timings at heterogeneous intersections; and because LLMs are typically trained with far less task-aligned interactive data than RL, intersection-specific operational knowledge remains limited. This paper presents CuraLight, an LLM-centered framework in which an RL agent assists the fine-tuning of an LLM agent. RL exploration gathers traffic states and actions, enhanced prompting constructs imitation pairs for fine-tuning the base model, timing decisions are structured to elicit interpretable preferences, and a multi-LLM ensemble deliberation system conducts adversarial debates that defend each RL-filtered phase action and consolidates a consensus outcome for RL-assisted fine-tuning. Evaluation in SUMO across heterogeneous networks from Jinan (17 intersections), Hangzhou (19), and Yizhuang (177) shows improvements over recent state of the art: ATT by about 5.34%, AQL by about 5.14%, and AWT by about 7.02%. Code and configurations are available at https://anonymous.4open.science/r/CuralightCode-6437/. FLOOD: Fine-tuning LLM with Offline Optimal Oracle for Online Decision Making Tao Qiu (Laboratory for Big Data and Decision, National University of Defense Technology); Kaiming Xiao (Wuxi Research Institute of Applied Technologies, Tsinghua University); and Hang Zhang, Zhihao Chen, Haoyu Yang, Xuan Li, and Mao Wang (Laboratory for Big Data and Decision, National University of Defense Technology) Abstract Abstract Online decision refers to making sequential choices without future information, which is widely applied in real world scenarios such as dynamic resource allocation, scheduling and online advertising. Motivated by the remarkable reasoning and planning capabilities of Large Language Model (LLM), researchers have applied these models to such online problems. However, existing LLM-based approaches primarily rely on parameter-frozen LLM, which often exhibit weak performance in complex, non-stationary and long-horizon environment. Although fine-tuning LLM for online decision making is a promising direction, this remains challenging due to the lack of supervision labels and effective training mechanisms for online scenarios. To address this, we propose the FLOOD framework (Fine-tuning LLM for Online decision via Oracle Distillation), which distills the capabilities of an offline optimal oracle into an online decision LLM. Specifically, our framework leverages an offline optimal oracle (with full information access) to generate optimal decision trajectories. These trajectories are then distilled into an online LLM via fine-tuning. Crucially, to adapt to the characteristics of online decision making, we design Decision Period Aligned Gradient Accumulation (DPAGA) strategy where gradients are accumulated to match the granularity of decision periods. We train FLOOD solely on synthetically generated oracle trajectories, and evaluate its performance on both synthetic data and a real world sequential emergency dataset. These experiments demonstrate that FLOOD generalizes effectively, achieving competitive performance compared to existing baselines. Further experiment confirms that the DPAGA strategy is essential for the distillation process. DiffMoE: Differential-Encoded MoE Predictor for Variable-Length Flight Trajectory Prediction Lindong Wang, Hongya Tuo, and Zhongliang Jing (Shanghai Jiao Tong University) Abstract Abstract Flight trajectory prediction (FTP) is a essential task for future air traffic management (ATM), enabling conflict detection and traffic flow control. Unlike fixed-window sliding methods that require sequence truncation or padding, variable-length prediction preserves complete historical information from takeoff to landing, reducing information loss and improving accuracy. However, variable-length prediction must balance accuracy across both short sequences (takeoff phase) and long sequences (landing phase). Additionally, dynamic transitions among flight modes, including climbing, turning, cruising, and descending, further complicate the prediction task. To address these challenges, we propose DiffMoE, a framework incorporating differential normalization and encoding to convert absolute coordinates into relative displacements, effectively handling scale variations in variable-length sequences. We further introduce an MoE regression head with Top-K routing for multi-expert soft-threshold learning, enabling adaptive handling of flight mode transitions. Experiments on a real-world ADS-B dataset demonstrate that DiffMoE outperforms all comparison methods and achieves distance errors as low as 154 meters. Ablation studies further confirm the effectiveness of both differential encoding and MoE modules. Tuesday Virtual Room 2 IJCNN Paper LLM Agents and Tool Use V Session Chair: Yang Han (Institute of Software, Chinese Academy of Sciences; Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences), Yue Liu (Shanghai University) ConfigSwarm: A Multi-Agent Framework on Versioned Kernel Configuration Graphs Yang Han (Institute of Software, Chinese Academy of Sciences; Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences) and Kaichun Yao and Libo Zhang (Institute of Software, Chinese Academy of Sciences) Abstract Abstract Linux kernel configuration is critical for workload-specific performance, security, and resource footprint, but it remains difficult to automate. Any automated approach must generate configuration edits that satisfy executable Kconfig constraints and remain robust under rapid kernel evolution, where options and constraint logic drift across releases. Existing LLM-based approaches often produce non-executable edits, treat each release as an isolated snapshot, and fail to retain execution feedback as reusable knowledge. We propose ConfigSwarm, a multi-agent framework for kernel configuration QA and tuning under version drift. ConfigSwarm grounds agent reasoning in the Versioned Kernel Configuration Knowledge Graph (VKConfig-KG), which combines executable Kconfig constraints with a lightweight semantic taxonomy for intent grounding. To enable efficient cross-version migration, ConfigSwarm adopts a Backbone-and-Diff factorization that reuses an anchor backbone and localizes maintenance to sparse deltas. Four specialized agents (Interaction, Parsing, Alignment, and Validation) coordinate via typed artifacts and a validation-in-the-loop policy with staged checks from static feasibility to compile/boot and workload measurements, while writing structured outcomes back to VKConfig-KG for iterative refinement. Experiments show that ConfigSwarm improves QA reliability and closed-loop tuning validity, reduces cross-version adaptation cost by over 86%, and raises build success rate to 73.3%. ICMOS: Incremental Concept Mining for OS Kernel Configuration via LLMs Agentic Reasoning Yang Han (Institute of Software, Chinese Academy of Sciences; Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences) and Kaichun Yao and Libo Zhang (Institute of Software, Chinese Academy of Sciences) Abstract Abstract Linux kernel configuration is critical for system performance, security, and adaptability. However, its vast configuration space, comprising over 17,000 configuration options, renders manual tuning both time-consuming and prone to errors. Existing methods largely rely on static heuristics or limited semantic rules, which struggle to capture complex configuration dependencies or adapt across diverse workloads. We introduce ICMOS, a framework that integrates large language models (LLMs) with a heterogeneous knowledge graph of kernel configuration concepts (OSKC-KG). By grounding LLM reasoning in structured semantics, ICMOS supports context-aware mining of configuration concepts and agentic concept evolution in response to new requirements and kernel updates. We evaluate ICMOS on configuration QA tasks and diverse real-world workloads, including databases, web servers, in-memory caches, and system benchmarks. ICMOS consistently outperforms LLM-only baselines, delivering higher accuracy, faster optimization, and robust system performance. Notably, it halves optimization time, reduces tail latency by 58.1%, and more than doubles configuration success rates. These results demonstrate that ICMOS provides a scalable and reliable framework for grounding LLM reasoning in structured semantics, thereby advancing kernel configuration understanding and optimization. RE-MCDF: Closed-Loop Multi-Expert LLM Reasoning for Knowledge-Grounded Clinical Diagnosis Shaowei Shen, Xiaohong Yang, and Jie Yang (Xiamen University); Lianfen Huang (the Key Laboratory of Intelligent Manufacturing Equipment and Industrial Internet Technology, Fujian Provincial Universities / School of Information Science and Technology, Xiamen University Tan Kah Kee College / the Department of Informatics and Communication Engineering, Xiamen University, Fujian , China); Yongcai Zhang (Xiamen University); Yang Zou (Tongji University); and Seyyedali Hosseinalipour (Department of Electrical Engineering,University at Buffalo-SUNY) Abstract Abstract Electronic medical records (EMRs), particularly in neurology, are inherently heterogeneous, sparse, and noisy, which poses significant challenges for large language models (LLMs) in clinical diagnosis. In such settings, single-agent systems are vulnerable to self-reinforcing errors, as their predictions lack independent validation and can drift toward spurious conclusions. Although recent multi-agent frameworks attempt to mitigate this issue through collaborative reasoning, their interactions are often shallow and loosely structured, failing to reflect the rigorous, evidence-driven processes used by clinical experts. More fundamentally, existing approaches largely ignore the rich logical dependencies among diseases, such as mutual exclusivity, pathological compatibility, and diagnostic confusion. This limitation prevents them from ruling out clinically implausible hypotheses, even when sufficient evidence is available. To overcome these, we propose RE-MCDF, a relation-enhanced multi-expert clinical diagnosis framework. RE-MCDF introduces a generation–verification–revision closed-loop architecture that integrates three complementary components: (i) a primary expert that generates candidate diagnoses and supporting evidence, (ii) a laboratory expert that dynamically prioritizes heterogeneous clinical indicators, and (iii) a multi-relation awareness and evaluation expert group that explicitly enforces inter-disease logical constraints. Guided by a medical knowledge graph (MKG), the first two experts adaptively reweight EMR evidence, while the expert group validates and corrects candidate diagnoses to ensure logical consistency. Extensive experiments on the neurology subset of CMEMR (NEEMRs) and on our curated dataset (XMEMRs) demonstrate that RE-MCDF consistently outperforms state-of-the-art baselines in complex diagnostic scenarios (the source code and datasets are publicly available at https://github.com/shenshaowei/RE-MCDF). A Closed-Loop Framework for Material Text Annotation Correction and Validation Based on Large Language Models Yue Liu, Zi Liu, Wei Zuo, Zhengwei Yang, Jie Wu, and Tianlin Dai (School of Computer Engineering and Science, Shanghai University) and Siqi Shi (State Key Laboratory of Materials for Advanced Nuclear Energy & School of Materials Science and Engineering, Shanghai University) Abstract Abstract High-quality and standardized annotations are a prerequisite for extracting reliable structure–activity relationships in materials science. Despite rapid progress in extraction models, their performance is fundamentally bottlenecked by annotation noise and inconsistency inherent in domain-specific corpora. While Large Language Models (LLMs) offer a potential solution for automated annotation, they typically fail to generate high-fidelity data in specialized vertical domains due to severe hallucinations and the lack of explicit boundary constraints. To break this quality-efficiency trade-off, we propose CoMETAM (Collaborative Model and Expert-knowledge-constrained Text Annotation for Materials), a closed-loop automated data accuracy governance framework designed to refine noisy annotations into high-quality datasets via multi-model collaboration. CoMETAM first introduces a generic insight metric, operationalized by high-precision extraction models, to precisely identify low-quality datasets that degrade downstream learning. Subsequently, it employs a knowledge-guided structured prompting strategy to inject domain knowledge as rigorous constraints into the inference process, thereby effectively suppressing LLM hallucinations. Finally, a consistency verification module ensures label alignment, closing the loop for automated data hygiene. Experiments on four diverse materials datasets show that CoMETAM consistently boosts downstream extraction performance, delivering F1 gains of up to 13.98%. These results indicate that CoMETAM provides a robust basis for constructing reliable domain knowledge bases. Tuesday Virtual Room 3 IJCNN Paper LLM Agents and Tool Use VI Session Chair: Yue Liu (Shanghai University), Zhilin Zhang (Chongqing Institute of Green and Intelligent Technology,Chinese Academy of Sciences; Chongqing School,University of Chinese Academy Sciences) Task-Tree A-Star: A Hierarchical Framework for Long-Horizon Task Planning with Large Language Models Zihao Li, Zhilin Zhang, and Yun Lu (Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences; Chongqing School, University of Chinese Academy of Sciences) and Guang Li (Chongqing Institute of Green and Intelligent Technology, Chinese Academy of Sciences) Abstract Abstract While Large Language Models show promise in agentic planning, their reliability drops sharply in long-horizon scenarios. Existing methods often face a trade-off between robustness and planning granularity. Local replanning often misses root errors buried in earlier steps, causing repetitive failures, whereas global replanning wastes valid progress by starting over. Additionally, LLM-generated plans often lack physical grounding, leading to actions that violate environmental constraints. We propose Task-Tree A-Star, a hierarchical framework that decouples high-level task decomposition from low-level action execution. By structuring tasks into a tree, TTA* enables backtracking to specific decision points rather than restarting, thereby preserving completed subtasks. Inside each node, a localized A* search generates executable actions without the combinatorial overhead of global search. Experiments in a complex simulation environment demonstrate that TTA* achieves higher success rates than baselines, particularly in unseen scenarios. CPGR-Mem: Cross-Platform Graph-based Routing Memory for Multimodal Intelligent Agents Jiale Kong and yuqi Zhao (Central China Normal University, computer science department) Abstract Abstract Robust long-horizon execution in multimodal agents is severely hampered by inefficient cross-platform experience transfer. Existing memory mechanisms struggle to balance platform adaptability with the structured representation required for reliable reasoning. To address this, we introduce CPGR-Mem, a Cross-Platform Graph-based Routing Memory framework. CPGR-Mem organizes knowledge into a novel tri-graph architecture representing high-level strategies, platform-aware routing, and fine-grained execution traces. By employing a learnable routing mechanism and bidirectional retrieval, it dynamically constructs transferable task paths across heterogeneous environments. Evaluated on the KGCE benchmark across three leading multimodal backbones, CPGR-Mem significantly outperforms state-of-the-art baselines. Most notably, it achieves a 91.3\% task completion ratio while reducing the backtracking rate by over 40\%, offering a highly scalable, model-agnostic solution for cross-platform automation. A Hierarchical Memory-Based Exploration Mechanism for Language Model Creative Agents Changhong He, Zhen Bi, Jungang Lou, Zhenlin Hu, Zihao Xue, Qing Shen, Zhengfang Liu, and Kang Zhao (Huzhou University) Abstract Abstract With the rapid advancement of Large Language Models (LLMs), autonomous agents capable of integrating perception, decision-making, and action have become essential for addressing complex, long-horizon tasks. However, existing agent frameworks still face limitations in creative reasoning, primarily due to an over-reliance on foundational model capabilities, inefficient memory management, and superficial tool invocation that neglects cross-scenario semantic logic. These constraints hinder agents' generalization and innovative problem-solving abilities in unconventional tasks. To address these challenges, we propose HMemAgent (Hierarchical Memory Agent), a novel framework centered on a three-level hierarchical memory mechanism. This architecture comprises buffer memory for capturing real-time environmental interactions, reflective memory for distilling errors and clues into structured knowledge through introspection, and associative memory for inspiring context-aware creative decisions via semantic links. Together, these levels form a closed-loop ``recording–refining–associating" pipeline, significantly enhancing experience reuse efficiency and the exploration of tool semantics. Experimental results demonstrate that HMemAgent markedly strengthens creative reasoning and cross-scenario adaptation in dynamic, complex environments while maintaining robust decision stability. LoG-ADFormer: A Battery Long-Horizon Capacity Forecasting Method Yue Liu, Jiawei Ni, and Wei Zuo (School of Computer Engineering and Science, Shanghai University); Xinyue Jin (Materials Genome Institute, Shanghai University); and Siqi Shi (State Key Laboratory of Materials for Advanced Nuclear Energy & School of Materials Science and Engineering, Shanghai University) Abstract Abstract Long-horizon forecasting for lithium-ion batteries is essential for practical health management. Yet reliable multi-horizon prediction remains challenging under non-stationary operation, rest-induced recovery, and nonlinear aging. We propose LoG-ADFormer, a local-to-global forecasting framework that improves long-horizon stability by mitigating level drift. It introduces an Abs-Delta dual-view prediction design to jointly model the absolute capacity level and incremental evolution, and uses horizon-aware fusion to adaptively balance the two views across future steps. This design enhances robustness to non-stationary operating conditions, rest-induced capacity recovery, and nonlinear aging dynamics. We conduct experiments on four public battery degradation datasets, spanning different cell chemistries and temperature conditions, to evaluate the model under multiple prediction horizons and lookback lengths. LoG-ADFormer achieves the best performance in 75% of long-horizon forecasting settings, with an overall average MAE of 0.3468. These results suggest that LoG-ADFormer offers a robust and practical solution for long-horizon capacity prognosis, enabling more reliable maintenance planning and risk-aware battery operation. Tuesday Virtual Room 4 IJCNN Paper LLM Agents and Tool Use VII Session Chair: Bijia Liu (Independent Researcher), Nirali Sanghvi (Indian Institute of Technology, Roorkee) FinRS: A Risk-Sensitive Trading Framework for Real Financial Markets Bijia Liu (Independent Researcher) and Ronghao Dang (Alibaba DAMO Academy) Abstract Abstract Large language models (LLMs) have demonstrated strong reasoning capabilities and are increasingly applied to financial trading. However, existing LLM-based trading agents typically address risk in a mechanistic or reactive manner, relying on post-hoc risk-adjusted metrics, heuristic rules, or prompt-level adjustments, without embedding genuine risk awareness into their decision-making process. As a result, these agents lack an internalized understanding of downside exposure and often behave myopically under market volatility. To address this gap, we propose FinRS, a risk-sensitive trading framework that integrates hierarchical market analysis, a dual-decision agent architecture, and multi-timescale reflection to explicitly align trading actions with both return objectives and downside risk constraints. Extensive experiments demonstrate that FinRS achieves superior returns and improved drawdown stability compared to state-of-the-art baselines, representing a meaningful step toward more robust and risk-aware LLM-driven trading systems. Risk-Aware Decision-Theoretic Routing for Reliable Use of Large Language Models Lujin Zhao, Sujuan Qin, and Yijie Shi (School of Cyberspace Security, Beijing University of Posts and Telecommunications) Abstract Abstract Routing queries to large language models (LLMs) is a knowledge-driven decision-making problem under uncertainty. Most existing routing approaches optimize expected performance and implicitly adopt a uniform failure cost assumption, which does not hold in safety-critical settings. We formulate instance-level LLM routing from a decision-theoretic perspective and cast it as a cost-aware optimization problem targeting tail-risk control. We introduce lightweight and interpretable risk abstractions that integrate performance prediction, predictive uncertainty, and extrapolation uncertainty, enabling principled risk estimation prior to model invocation. Based on these abstractions, we develop a one-shot decision-theoretic routing framework that balances reliability and inference cost without static thresholds or multi-model cascading. Experiments across reasoning-intensive, professional, and safety-critical benchmarks demonstrate that the proposed approach substantially improves reliability on high-risk and out-of-distribution queries while maintaining competitive inference efficiency. By providing interpretable decision-relevant signals, our framework supports robust and explainable model selection for mission-critical LLM deployment. EEVEE: Ensemble Expertise from Visual-Temporal-Symbolic Embeddings Enhancing Financial Time Series Forecasting Yue Yuan (Key Laboratory of Interdisciplinary Research of Computation and Economics ,Shanghai University of Finance and Economics; Advanced Micro Devices, Inc); Yahao He (Advanced Micro Devices, Inc); and Songqiao Han (Key Laboratory of Interdisciplinary Research of Computation and Economics ,Shanghai University of Finance and Economics) Abstract Abstract While human traders have decoded market movements through candlestick patterns for centuries, modern neural networks often discard this visual wisdom, processing markets purely as numbers. We propose EEVEE, a framework that enables machines to leverage the same multi-perspective reasoning that expert traders use by unifying three complementary knowledge perspectives. First, general visual knowledge: we distill professional K-line chart analysis from commercial Language Models into vision-language models, enabling machines to recognize patterns like head-and-shoulders. Second, general temporal knowledge: we discover that synergistically combining DMMV-A and DMMV-S decomposition methods reduces substantially prediction error on financial forecasting. Third, finance-specific knowledge: we integrate Kronos foundation model with hierarchical processing of its coarse-fine tokenization to capture market-specific trading dynamics. Through cross-modal attention fusion, EEVEE achieves 64.9 percent directional accuracy with 0.88 percent MAPE on stock forecasting, outperforming eight state-of-the-art baselines. Experiments across stocks and cryptocurrencies demonstrate that ensemble expertise from visual-temporal-symbolic perspectives provides robust and interpretable market predictions. Neuro-Symbolic Risk Semantics for Safe Multi-Agent Coverage in Hazardous Environments Nirali Sanghvi and Rajdeep Niyogi (Indian Institute of Technology, Roorkee) Abstract Abstract Ensuring safety in Multi-Agent Reinforcement Learning (MARL) is critical for real-world deployment, but remains brittle under epistemic mismatch, scalar penalties obscure failure mechanisms, and fixed symbolic priors can become miscalibrated under distribution shift. We introduce FMEA-Net, a neuro-symbolic risk semantics module for safe multi-agent coverage that makes engineering-style Failure Mode and Effects Analysis (FMEA) differentiable and adaptive. A neural cause recognizer maps each agent’s local observation to hazard causes, which is propagated through learnable cause-to-mode and modeto-effect tables to obtain interpretable mode/effect beliefs and a severity-weighted risk score. These outputs form a compact safety context that conditions decentralized actors, enabling agents to distinguish why a region is hazardous (e.g., slip vs. collision), rather than merely how much and to learn risk-aware behaviors without collapsing coverage. Across hazardous multi-agent coverage benchmarks, the proposed method improves safe coverage and reduces safety violations and overall risk relative to baselines that rely on unstructured scalar penalties, with consistent gains across unseen maps and risk-model mismatch. Tuesday Virtual Room 5 IJCNN Paper LLM Agents and Tool Use VIII Session Chair: Weizhi Kong (Harbin Institute of Technology, Shenzhen), Yinong Chi (Qilu University of Technology; Shandong Provincial Key Laboratory of Industrial Network and Information System Security , Shandong Fundamental Research Center for Computer Science) Efficient Multi-Turn LLM Serving with Contextual Similarity and Resource-Aware Scheduling Weizhi Kong and Weizhe Zhang (Harbin Institute of Technology, Shenzhen, China) and Yiming Wang (China University of Geosciences, Beijing) Abstract Abstract Large Language Models (LLMs) have achieved tremendous success. However, their autoregressive nature results in highly uncertain execution times for user requests, making it difficult to prioritize requests and often causing Head-of-Line (HOL) blocking in inference systems. Existing approaches frequently optimize scheduling by building lightweight predictors for model response lengths. Yet, these predictors have limited context input capacity, and in multi-turn dialogue scenarios, naively concatenating the entire context may lead to truncation and loss of critical information. In this work, we experimentally reveal a correlation between the context of multi-turn dialogues and the lengths of their responses. By extracting key features from the dialogue context, we construct a multi-turn dialogue predictor that achieves more accurate length predictions. Leveraging these predictions, we simulate the request generation process to effectively reduce preemptions caused by insufficient GPU memory. Comparative experiments on a state-of-the-art LLM inference system show that our method reduces the Time to First Token (TTFT) by 47.98\% and the average per-token generation latency by 76.56\%. MinorEval: A Behavior-driven Framework for Multi-Turn Safety Evaluation of LLMs for Minors Qunpeng Lei, Hongwei Hu, Hanghui Yang, Zhiwen Zhang, and Zijiao Zhang (Zhengzhou University) Abstract Abstract As Large Language Models (LLMs) increasingly permeate K-12 education, ensuring their safety for minor users is paramount. However, existing benchmarks primarily focus on single-turn content filtering, overlooking dynamic risks stemming from interaction patterns specific to minors. To this end, we propose MinorEval, a psychologically-grounded automated framework. We distinguish between Child and Adolescent personas and leverage four behavioral templates derived from psychological theories to simulate realistic interactions, conducting a total of 800 dynamic multi-turn simulation experiments across five representative models. Results indicate that, excluding Claude, multi-turn interactions increased the Final Success Rate (FSR) by over 10 times compared to single-turn baselines; model defense capabilities exhibit significant stratification, where Claude demonstrates robust intent recognition, GPT-5 displays defense resilience, while Mistral faces near-total collapse. Generalized Linear Mixed Models (GLMMs) analysis further reveals a ``Fragile Shield'' phenomenon: although the Child persona triggers significant defense enhancement, this intrinsic alignment remains insufficient to withstand the impact of strategies such as Role Playing. These findings highlight the failure of current defenses against cumulative contextual risks and emphasize the urgency of establishing context-aware guardrails tailored for children. Dynamic Semantic Adaptation for Efficient Large Language Model Inference Guanghui Han (Qufu Normal University), Yidan Yan and Lei Wei (China Agricultural University), Da Gao (Qufu Normal University), and Shougang Feng (Shandong Branch of China Mobile Group) Abstract Abstract Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of reasoning and generation tasks, but their inference-stage inefficiencies remain a significant bottleneck for scalable deployment, particularly in latency-sensitive or resource-constrained environments. To solve this, this paper introduces a unified and adaptive inference optimization framework. It addresses three critical limitations in existing methods. These limitations include inter-layer semantic redundancy, static attention computation, and rigid precision scheduling. First, we propose a cross-layer context reuse mechanism with dynamic drift correction, enabling conditional bypass of redundant Transformer layers while preserving semantic fidelity. Second, we introduce a predictive attention path pruning strategy guided by token-level contextual complexity, allowing the model to selectively allocate attention computation based on entropy and dependency metrics. Third, we propose a self‑evolving precision scheduler that adjusts bit‑width based on token‑level confidence. An online feedback mechanism further refines the precision policy during inference. Experiments on AlpacaEval, HotpotQA, and Vicuna QA demonstrate that our method achieves up to 2.43× decoding speedup and 41.7% memory savings, while maintaining comparable output quality to state-of-the-art baselines including Flash-Decoding, Radix Attention, and DEFT-Flatten. These results highlight the potential of fine-grained semantic adaptivity and runtime feedback as core principles for efficient and intelligent LLM inference. DEC-MAC: A Device-Edge-Cloud Dynamic Resource-Aware Security Task Scheduling Framework Based on Multi-Agent Collaboration Yinong Chi, Xiaohui Han, Guangqi Liu, and Lin Chen (Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center, Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Industrial Network and Information System Security , Shandong Fundamental Research Center for Computer Science) Abstract Abstract Mobile terminals face challenges in deploying complex security detection mechanisms due to resource constraints. Although the device-edge-cloud collaborative architecture can mitigate some of these limitations, its adaptability in dynamic environments remains inadequate. To address this issue, this paper formulates the security task scheduling problem in device-edge-cloud environments as a Partially Observable Markov Decision Process (POMDP). A reward function is constructed based on security task latency and detection accuracy, transforming the problem into one of maximizing the expected cumulative discounted reward. Subsequently, we propose a scheduling algorithm based on multi-agent reinforcement learning, which employs Multi-Agent Proximal Policy Optimization (MAPPO) to achieve cooperative training. To enhance the state estimation capability of the value network, a value decoupling mechanism is designed. It uses a Siamese network to extract robust feature representations from high-dimensional environmental states, while integrating a hierarchical perceptual contrastive learning mechanism and a reward-guided adaptive weighting mechanism to improve the discrimination of key decision-making states. Simulation results demonstrate that, compared to existing state-of-the-art methods, the proposed framework achieves a better balance between task processing latency and detection accuracy under dynamic workloads and resource variations. Tuesday Virtual Room 6 IJCNN Paper LLM Evaluation and Benchmarking III Session Chair: Ponnurangam Kumaraguru (IIIT Hyderabad), Xichen Lin (Nanjing University) LAWRE: Evaluating Legal Case Retrieval for Practical Legal Reasoning Xichen Lin, Yi Yang, Shuxian Xie, Yi Feng, and Chuanyi Li (Nanjing University) Abstract Abstract Legal Case Retrieval (LCR) involves retrieving legally relevant cases from a candidate corpus given a query case. While state-of-the-art (SOTA) LCR models have achieved high recall scores on public datasets (e.g., LeCaRD), such datasets fail to assess key legal reasoning skills needed in practice (e.g., distinguishing confused cases, identifying legal elements). To bridge this gap, we introduce LAWRE, a comprehensive benchmark for evaluating LCR systems, facilitating an understanding of their reasoning processes and offering detailed diagnostic feedback. Testing eight SOTA models on LAWRE exposes notable weaknesses, which in turn motivate analyses of these shortcomings and guide the development of more practically effective LCR systems. PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI Srija Mukhopadhyay, Sathwik Reddy, and Shruthi Muthukumar (International Institute of Information Technology, Hyderabad); Jisun An (Indiana University); and Ponnurangam Kumaraguru (International Institute of Information Technology, Hyderabad) Abstract Abstract Personalized AI agents rely on access to a user's digital footprint, which often includes sensitive data like private emails and chats. This access creates a fundamental societal and privacy risk: systems lacking social-context awareness can unintentionally expose user secrets, threatening digital well-being. We introduce PrivacyBench, a benchmark with socially grounded datasets containing embedded secrets and a multi-turn conversational evaluation to measure secret preservation. Testing Retrieval-Augmented Generation (RAG) assistants reveals that they leak secrets in up to 26.56% of interactions, with the majority of failures emerging only after sustained multi-turn pressure. A privacy-aware prompt lowers leakage to 5.12%, but offers only partial mitigation. The retrieval mechanism continues to access sensitive data indiscriminately, which shifts the entire burden of privacy preservation onto the generator. This creates a single point of failure, rendering current architectures unsafe for wide-scale deployment. Our findings underscore the urgent need for structural, privacy-by-design safeguards to ensure an ethical and inclusive web for everyone. To ensure transparency, we release the complete codebase and the associated synthetic datasets. PilotBench: A Benchmark for General Aviation Agents with Safety Constraints Yalun Wu (School of Computing, National University of Singapore); Haotian Liu (School of Informatics, Xiamen University); and Boyang Wang and Zhoujun Li (CESS, Beihang University) Abstract Abstract As Large Language Models (LLMs) advance toward embodied AI agents operating in physical environments, a fundamental question emerges: can models trained on text corpora reliably reason about complex physics while adhering to safety constraints? We address this through PilotBench, a benchmark evaluating LLMs on safety-critical flight trajectory and attitude prediction. Built from 708 real-world general aviation trajectories spanning nine operationally distinct flight phases with synchronized 34-channel telemetry, PilotBench systematically probes the intersection of semantic understanding and physics-governed prediction through comparative analysis of LLMs and traditional forecasters. We introduce Pilot-Score, a composite metric balancing 60% regression accuracy with 40% instruction adherence and safety compliance. Comparative evaluation across 41 models uncovers a Precision-Controllability Dichotomy: traditional forecasters achieve superior MAE of 7.01 but lack semantic reasoning capabilities, while LLMs gain controllability with 86--89% instruction-following at the cost of 11--14 MAE precision. Phase-stratified analysis further exposes a Dynamic Complexity Gap—LLM performance degrades sharply in high-workload phases such as Climb and Approach, suggesting brittle implicit physics models. These empirical discoveries motivate hybrid architectures combining LLMs' symbolic reasoning with specialized forecasters' numerical precision. PilotBench provides a rigorous foundation for advancing embodied AI in safety-constrained domains. Flying Pigs, FaR and Beyond: Evaluating LLM Reasoning in Counterfactual Worlds Anish R Joishy, Ishwar B Balappanawar, and Vamshi Krishna Bonagiri (IIIT Hyderabad); Manas Gaur (University of Maryland, Baltimore County); Krishnaprasad Thirunarayan (Wright State University); and Ponnurangam Kumaraguru (IIIT Hyderabad) Abstract Abstract A fundamental challenge in reasoning is navigating hypothetical, counterfactual worlds where logic may conflict with ingrained knowledge. We investigate this frontier for Large Language Models (LLMs) by asking: Can LLMs reason logically when the context contradicts their parametric knowledge? To facilitate a systematic analysis, we first introduce CounterLogic, a benchmark specifically designed to disentangle logical validity from knowledge alignment. Evaluation of 11 LLMs across six diverse reasoning datasets reveals a consistent failure: model accuracy plummets by an average of 14% in counterfactual scenarios compared to knowledge-aligned ones. We hypothesize that this gap stems not from a flaw in logical processing, but from an inability to manage the cognitive conflict between context and knowledge. Inspired by human metacognition, we propose a simple yet powerful intervention: Flag & Reason (FaR), where models are first prompted to flag potential knowledge conflicts before they reason. This metacognitive step is highly effective, narrowing the performance gap to just 7% and increasing overall accuracy by 4%. Our findings diagnose and study a critical limitation in modern LLMs’ reasoning and demonstrate how metacognitive awareness can make them more robust and reliable thinkers. Tuesday Virtual Room 7 IJCNN Paper LLM Evaluation and Benchmarking IV Session Chair: Xinhai Chen (National University of Defense Technology), Nan Wang (Lenovo) DoChartDes: Fine-Grained, Automated Benchmark for Chart Description in Document Contexts Nan Wang, Na Wang, Zeyang Yang, Yuxuan Xu, Fangzhen Peng, and Gefang Ma (Lenovo) Abstract Abstract Chart understanding is crucial for evaluating multimodal large language models' (MLLMs) capability. Mainstream benchmarks rely on closed-ended QA formats for rapid evaluation. Meanwhile, how to evaluate long-form natural language chart description is growing vital in open-ended tasks. Current open-ended evaluation suffers from vague and coarse-grained metrics, while most assess charts in isolation, ignoring interleaved document layouts. To address these, we introduce DoChartDes, a specialized benchmark designed for open-ended chart description task. (i) Diverse dataset: Featuring 18 diverse chart types across 6 languages and 7 document layouts, the dataset spans scenarios from simple charts to slide documents, supporting description evaluation for chart in document contexts. It includes 1.6K curated chart images, each annotated with over 25 meta-elements (42.3K total) in structured format. (ii) Fine-grained evaluation suite: This suite assesses chart description in terms of faithfulness and coverage, across 3 critical dimensions of perception and reasoning: layout understanding, factual correctness and analytical reasoning. It provides interpretable scores enabling detailed weakness analysis and hallucination detection. (iii) Intelligent judgement: The evaluation process is optimized into an intelligent way, reaching high consistency and low systemic bias with human experts while reducing time cost. With high-quality annotations and intelligent evaluation, DoChartDes provides more reliable benchmark for chart description. Event Knowledge Graph Narrative Planning for Faithful and Coherent Multi-Document Summarization Yukang Huang (School of Computer Engineering and Science, Shanghai University); Tongshu Liu (College of Arts and Sciences, The Ohio State University); and Wei Liu and Weimin Li (School of Computer Engineering and Science, Shanghai University) Abstract Abstract Multi-Document Summarization (MDS) remains challenging for Large Language Models (LLMs) due to difficulties in maintaining factual consistency, temporal coherence, and event-level faithfulness across sources. Existing approaches often rely on graph linearization, requiring LLMs to implicitly reconstruct structure and leading to suboptimal narrative quality. We propose Event Knowledge Graph Narrative Planning (EKGNP), a neuro-symbolic framework that separates eventlevel planning from surface realization. EKGNP constructs an event-centric graph and formulates summarization as a planning problem, where a Neural MCTS planner generates coherent event sequences. We further introduce a Sinkhorn-based metric to measure structural faithfulness. Experiments on Multi-News, WCEP, and EventSum show that EKGNP improves factual consistency, temporal accuracy, and structural alignment over strong baselines. Results demonstrate that explicit event-level planning provides an effective solution for faithful and coherent multi-document summarization. Granularity is the Hidden Bias: Revisiting the Factual Consistency Evaluation for Summarization Weina Ma (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences); Jiayang Gao (Beijing Institute of Computer Technology and Application); Zimo Nie, Zhiwei Yang, and Chaodong Tong (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences); Chengqing Guo (National Computer Network Emergency Response Technical Team/Coordination Center of China); and Lei Jiang (Chinese Academy of Sciences, Institute of Information Engineering) Abstract Abstract Factual consistency evaluation for summarization has attracted increasing attention as large language models (LLMs) increasingly generate hallucinations. Existing evaluation methods implicitly obey the fixed-fact-granularity assumation (verifying summaries under a uniform claim segmentation granularity) implicitly when decomposing and verifying summary content. This assumption inevitably introduces a hidden but systematic evaluation bias in real-world scenarios: logically independent facts may be merged to mask errors, while semantically interdependent complex facts may be over-splitting and break relational semantics. We revisit this problem from the perspective of granularity control, revealing that granularity mismatch is a fundamental failure mode that affects the reliability and interpretability of the evaluation. To address this problem, we propose a factual consistency evaluation framework FAIR, which achieves explicit modeling of factual granularity by performing logically aware fact segmentation on the summary and semantically aware evidence segmentation on the source documents. Furthermore, FAIR employs a retrieval-enhanced verification mechanism to perform precise alignment and verification of atomic content units and outputs concrete error types. Experiments on benchmark datasets demonstrate that FAIR achieves stronger correlation with human fidelity judgments and reduces evaluation bias, delivering both interpretable and reliable assessment of LLM-based summarization. A Fine-Grained Intelligent Mesh Quality Evaluation Algorithm Based on Neural Networks Ke Yang, Xinhai Chen, Zhichao Wang, Qinglin Wang, and Jie Liu (National University of Defense Technology) Abstract Abstract Mesh quality is a fundamental factor in large-scale scientific computing, serving as a critical bridge between physical modeling and numerical simulation which dictates the accuracy and stability of numerical solutions. Traditional mesh quality evaluation primarily relies on handcrafted geometric metrics and expert experience, which often requires extensive manual intervention and high preprocessing costs. While recent intelligent methods offer efficient and automated evaluation, they are restricted to mesh-level evaluation—judging only the overall quality of a mesh—and lack the capability to precisely locate individual low-quality elements. To bridge this gap, we first introduce NACA-FineGrade, the pioneering mesh dataset featuring element-level annotations. Building upon this benchmark, we further propose a fine-grained intelligent evaluation algorithm, termed FG-MeshNet, which enables the precise localization and identification of localized mesh defects. Specifically, FG-MeshNet employs a dedicated three-channel mesh representation, integrated with multi-branch segmentation heads and a collaborative attention mechanism to effectively capture multi-scale structural features. Experimental results demonstrate that FG-MeshNet attains 89.3% accuracy, surpassing existing methods by enabling reliable element-level diagnostics that precisely identify and localize critical geometric defects, such as poor orthogonality, insufficient smoothness, and uneven density. Tuesday Virtual Room 8 IJCNN Paper LLM Evaluation and Benchmarking V Session Chair: Jie Zhou (Changsha University of Science and Technology, Hunan Provincial Key Laboratory of Mathematical Modeling and Analysis in Engineering), Ruiting Dai (University of Electronic Science and Technology of China) Robust High-Dimensional Classification via Granular Balls-based $L_{0/1}$-SVM Guangming Lang, Xinyu Du, Jie Zhou, and Qimei Xiao (Changsha University of Science and Technology, Hunan Provincial Key Laboratory of Mathematical Modeling and Analysis in Engineering) Abstract Abstract High-dimensional data frequently exhibit feature redundancy, noise sensitivity, and complex underlying distributions, all of which severely impair classification performance. To address these challenges, we propose a novel classification framework termed granular ball $L_{0/1}$ support vector machine (GB-$L_{0/1}$SVM), which synergistically integrates principal component analysis (PCA), granular ball (GB) computing, and $L_{0/1}$ soft-margin SVM ($L_{0/1}$-SVM). In this paper, we apply PCA to reduce dimensionality and eliminate redundant features, and generate GBs in the transformed feature space to form compact and homogeneous information granules, thereby enhancing structural interpretability and robustness to noise. After that, GBs are fed into an $L_{0/1}$-SVM classifier to enable sparse and discriminative decision boundaries. Finally, we present an algorithm for the GB-$L_{0/1}$SVM classifier and conduct extensive experiments on six benchmark datasets. Experimental results demonstrate that GB-$L_{0/1}$SVM effectively bridges the gap between robust non-convex optimization and large-scale data mining, offering a superior trade-off between sparsity and predictive accuracy. Granular Ball Computing-Based Large Margin Distribution Machine Guangming Lang, Tingyao Di, Jie Zhou, and Qimei Xiao (Changsha University of Science and Technology, Hunan Provincial Key Laboratory of Mathematical Modeling and Analysis in Engineering) Abstract Abstract The Large Margin Distribution Machine (LDM) and its variants have received considerable attention due to their superior generalization performance. However, they encounter significant hurdles regarding computational efficiency and noise robustness when handling complex, large-scale data. To overcome these limitations and further enhance model performance, this paper proposes the Granular Ball Large Margin Distribution Machine (GBLDM). First, we develop a dual-strategy granular ball generation mechanism designed to reconcile label consistency with geometric compactness. Specifically, a coarse-grained strategy extracts the topological backbone using high-purity ''Important Granular Balls,'' while a fine-grained strategy refines complex boundaries via ''Edge Granular Balls.'' Second, we construct an adaptive weighted optimization model integrating granular ball geometric information, effectively extending margin distribution theory into the granular ball space. By dynamically allocating weights based on the scale and type of granular balls, GBLDM explicitly optimizes global margin statistics at the granular ball level. Finally, extensive experiments on eight UCI benchmark datasets demonstrate that GBLDM significantly outperforms four baselines in both classification accuracy and training efficiency, exhibiting exceptional robustness particularly in noisy scenarios. Adaptive State Space Experts: Dynamic Expert Selection with State-Aware Routing for Efficient Sequence Modeling Yue Guo, Huiping Ma, Qingkang Tang, and Aoyu Li (Nanjing University of Aeronautics and Astronautics) Abstract Abstract Abstract-State space models (SSMs) provide a principled route to linear-time sequence modeling and have recently reached competitive accuracy via selective, input-dependent dynamics. In parallel, mixture-of-experts (MoE) scaling decouples model capacity from per-token compute through sparse routing. Yet, existing SSM+MoE designs largely reuse token-only routers, ignoring the fact that an SSM already maintains a learned summary of the prefix in its hidden state. We introduce Adaptive State Space Experts (ASSE), a framework that couples SSM dynamics and expert selection through (i) state-conditioned routing that conditions expert choice on a compact summary of the running state, (ii) adaptive state retention that lets experts learn distinct temporal horizons, and (iii) a convergence-based earlyexit signal derived from state dynamics for input-adaptive depth. On language modeling, ASSE achieves 13.04 perplexity on the Pile benchmark, outperforming both SSM+MoE baselines (13.38) and hybrid Transformer-SSM architectures (13.21), while trading merely 0.06 PPL for 30% compute reduction via early exit. On the Long Range Arena, ASSE achieves 83.01% average accuracy, outperforming Mamba (82.14%) and other SSM baselines. Analysis reveals that experts naturally specialize into distinct temporal scales, with routing patterns showing 73% contextual coherence compared to 58% for token-only baselines. Our results demonstrate that tight integration between SSM state dynamics and MoE routing yields substantial gains in both quality and efficiency. Preset No More: Pattern-Aware Graph Routing and Meta-Adaptive Prompt Distillation for Incomplete Multimodal Learning Lisi Mo, Zihan Yang, Ruofan Zhou, Letian Xu, and Ruiting Dai (University of Electronic Science and Technology of China) Abstract Abstract The efficacy of multimodal learning emanates from the delicate orchestration of cross-modal interactions. However, pervasive modality absence in real-world scenarios shatters established alignment pathways, triggering a structural collapse of the joint representation space. Existing research remains primarily constrained by a preset interaction logic: treating modality absence as a passive feature reduction or a static parameter index, while lacking the capacity to actively reshape reasoning topologies to accommodate specific informational gaps. To this end, we propose Pattern-Aware Graph routing and Meta-Adaptive Prompt Distillation (PAG-MPD), a framework that shifts the missing-modality learning paradigm from static aggregation to a pattern-aware adaptive process. Specifically, we introduce a pattern-aware graph encoder to capture the shifting structural dependen-cies among available signals, complemented by a differentiable sparse routing mechanism that generates meta-adaptive prompts to dynamically re-calibrate the information flow. Furthermore, we develop an uncertainty-aware distillation strategy to mitigate epistemic uncertainty via confidence-weighted supervision from a complete-modality teacher. Extensive evaluations on MM-IMDb and CMU-MOSEI benchmarks demonstrate that PAG-MPD significantly outperforms state-of-the-art methods, achieving a notable improvement of up to 4.5% on MM-IMDb. Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V8 Session Chair: Li Cao (China University of Geosciences) Model Compression of Vision Transformer for Electric Motor Cogging Torque Prediction Jian Chen, Toshiaki Koike-Akino, and Ye Wang (Mitsubishi Electric Research Laboratories); Tatsuya Yamamoto and Yusuke Sakamoto (Mitsubishi Electric Corporation); and Bingnan Wang (Mitsubishi Electric Research Laboratories) Abstract Abstract Vision Transformer (ViT) models have recently demonstrated strong performance in predicting electric motor cogging torque from visual representations of motor geometry. However, the high computational cost and memory footprint of ViT architectures hinder their deployment in resource-constrained environments such as embedded motor drives and real-time digital twins. This paper investigates post-training model compression techniques for a ViT-based cogging torque prediction model. Specifically, we evaluate some activation-aware pruning methods, which leverage activation statistics to identify redundant weights, as well as low-rank factorization applied to attention and feed-forward layers. Experimental results demonstrate that substantial reductions in parameter count and inference cost can be achieved with negligible degradation in prediction performance. These findings highlight the feasibility of deploying compressed ViT models for efficient and scalable electric motor analysis. Cost-Efficient Power Orchestration in Heterogeneous EV-Integrated VPPs: A Peak Demand Repair Scheme Ahmed Selim, Mohamed Abido, and Ali Al-Awami (King Fahd University of Petroleum and Minerals) Abstract Abstract Vehicle-to-Grid (V2G) technology enables Electric Vehicles (EVs) to function as Virtual Power Plants (VPPs) which provide essential grid support services. The operational process becomes highly difficult because the system needs to handle electric vehicles which have different battery capabilities while it tries to achieve three different goals. This investigation resolves the major ”horizon effect” problem which occurs in realtime rolling-horizon optimization because short-term economic planning cannot prevent peak charges from occurring within the monthly period. The research demonstrates that standard heuristic methods which use V2G arbitrage for their operations cannot handle extended periods of load density because they result in critical constraint breaches which create additional demand expenses. The research presents a Peak-Repair Particle Swarm Optimization (PR-PSO) algorithm which uses LaxityBased Pruning for its operation. The PR-PSO system successfully controlled total peak demand during extreme usage patterns which were tested with 10 commercial EV models throughout a 30-day billing cycle using 11-run Monte Carlo statistical analysis. The proposed mechanism established physical search boundaries through its conversion of soft economic penalties which resulted in 100% session completion and a Net Total Cost decrease of approximately 40% compared to standard rolling-horizon and uncoordinated benchmarks. TE Model: A Transformer Encoder Only Based Model for Prefetch and Replacement in Solid-state Drivers Jianming Ge (BeiHang university) Abstract Abstract Memory technology offers access latency that is three orders of magnitude lower than that of solid-state drives (SSDs); prefetching and replacement, which rely on predicting future access sequences, are commonly used methods to bridge this speed gap, but previous research typically treats prefetching and replacement as independent tasks: the Stride prefetcher is only suitable for memory with strong locality, the LSTM prefetcher supports only short input sequences, and replacement strategies mostly rely on heuristic rules; in this work, we formulate prefetching as a time-series prediction classification problem, utilizing long historical sequences as features to predict future sequences for guiding prefetching and replacement decisions, propose the TE model, a Transformer encoder-only framework dedicated to learning and predicting complex access patterns in SSDs, which can forecast multi-step future request sequences and derive optimal prefetching and replacement strategies by integrating current cache blocks with predicted future access blocks; to the best of our knowledge, this is the first caching management framework for SSDs that unifies prefetching and replacement under a holistic design, and experimental results show that in most scenarios, the TE-base framework—which only predicts future sequences to guide prefetching—outperforms three baseline prefetching models, while the TE-advance model, which uses future sequence predictions to guide both prefetching and replacement, achieves even better performance, demonstrating that the proposed model effectively learns complex access patterns and that an integrated perspective of prefetching and replacement holds great promise for future research. Multi-Agent Collaborative Search Algorithm with Adaptive Scalarization Function and its Application in Aerospace Multi-Objective Optimization Problems Li Cao (China University of Geosciences), Ziyu Song (Sichuan University), Yirui Wang (University of Strathclyde), Bo Zhang (SLB Houston Digital Technology Center), and Wei Li (Faculty of Engineering University of Nottingham) Abstract Abstract Multi-Agent Collaborative Search Algorithm (MACS) is an effective method in handling multi-objective optimization problems. It is also regularly used in solving aerospace multi-objective optimization problems and demonstrates competitive problem-solving capabilities. The algorithm is a decomposition-based method. Tchebycheff scalarization function and dominant relationship are mixed to select individuals. The mixture with fixed p value represents constant characteristics. It may fail to meet the requirements of different stages of evolutionary algorithm. In order to prioritize convergence in early stages and diversity in later stages which may lead to final results with better quality, a new individual quality measurement parameter which adaptively emphasizes different features is introduced. To appropriately adjust p, reinforcement learning method (Upper Confidence Bound, UCB) is utilized to select p on the base of the new parameter. Building upon these improvements, a new algorithm Multi-Agent Collaborative Search Algorithm with Adaptive Scalarization Function (MACS-AS) is proposed. Two standard benchmark sets and two aerospace multi-objective problems are selected to evaluate algorithm performance. The new algorithm is compared with MACS and shows competitive results. Tuesday Virtual Room 1 IJCNN Paper LLM Evaluation and Benchmarking VI Session Chair: Wenpeng Lu (Qilu University of Technology, Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing; Shandong Fundamental Research Center for Computer Science), Zijiao Zhang (Zhengzhou University) SafeLoc: Token-Level Hallucination Localization in LLM-Generated Code via CodeBERT Buwei Wang, Liping Wen, Hao Jiang, Fanglin Xu, and Zijiao Zhang (Zhengzhou University) Abstract Abstract Large Language Models (LLMs) have revolutionized the code generation paradigm, yet their inherent "hallucination" issue—generating plausible but logically incorrect code or invoking non-existent libraries—poses severe security risks. Existing detection methods are mainly limited to coarse-grained classification or rely on expensive external execution environments and LLM self-reflection, lacking low-cost and high-precision token-level localization capabilities.To address this, we propose SafeLoc, a fine-grained hallucination localization framework. Utilizing CodeBERT as the semantic backbone, the framework integrates Bidirectional Long Short-Term Memory (BiLSTM) to enhance long-range context modeling and employs a Focal Loss-based training strategy to address the class imbalance caused by sparse hallucination tokens. Experimental results demonstrate that SafeLoc achieves an F1-score of 97.14% on an independent test set. Notably, in terms of micro-localization precision, it achieves a semantic tolerance accuracy of 86.04%, significantly outperforming GPT-5 Mini and Claude 4.5 Haiku. Furthermore, entropy-based uncertainty analysis confirms the interpretability and trustworthiness of the proposed method. LCIC: A Unified Hallucination Detection-and-Mitigation Method Using LLM’s Tail-Pooled Internal Signals Jiahang Li, Jiafei Niu, Yidong Ding, and Ping Yi (Shanghai Jiao Tong University) Abstract Abstract Hallucination remains a critical reliability bottleneck for large language model (LLM) applications, where models may generate fluent yet unverifiable or incorrect statements when internal evidence is insufficient. Most governance pipelines treat hallucination handling as two loosely coupled stages: detection, which estimates the reliability of a generated response, and mitigation, which suppresses hallucinations through inference-time control or training-time alignment. Such designs often incur substantial overhead: external evidence pipelines increase system complexity and latency, while sampling-based consistency approaches raise inference cost proportionally to the number of generations. Hybrid Late Fusion for UAV RGB Soybean Disease Detection and Spray Decision Support Kishan Kumar Yadav and Aruna Tiwari (Indian Institute of Technology Indore); Rajesh Dwivedi (Indian Institute of Information Technology,Ranchi); Pavan Kumar Kankar (Indian Institute of Technology Indore); Satya prakash Kumar and Shashi Rawat (ICAR - Central Institute of Agricultural Engineering, Bhopal); and Milind Ratnaparkhe (ICAR-National Soybean Research Institute,Indore) Abstract Abstract Low-cost UAV RGB imagery provides a practical solution for large-area soybean health monitoring, but disease detection remains difficult due to class imbalance, background clutter, and illumination variation. To address these challenges, we propose a hybrid late-fusion framework that combines an EfficientNet B3 image encoder with lightweight, interpretable numeric cues derived from each canopy tile, including vegetation color statistics, brightness and blur measures, and a leaf-area prior. The model is trained using imbalance aware learning to reduce bias towards majority classes. Experiments conducted on 2,842 UAV canopy tiles across four classes, i.e., Healthy, Mosaic, PestAttack, and Rust show that fusion improves robustness and achieves a test macro-F1 of 0.98, with strong per class F1-scores (0.99/0.97/0.98/0.97). Beyond disease recognition, we introduce an RGB only severity proxy based on leaf masking and class- specific color and texture cues. This enables an action-oriented decision policy that outputs SPRAY, MONITOR, or NO SPRAY, along with pixel-level overlays to enhance interpretability. Using conservative thresholds, the system achieves a SPRAY precision of 1.00 and a recall of 0.83 on the test set. Precision recall analysis reveals that recall can be increased with minimal loss in precision. The proposed RGB only pipeline, in general, provides a comprehensive path from UAV-based disease detection to practical spray decision support for precision crop protection. EHR-TASR: Leveraging Time-Aware Scoring and LLM Reasoning for Clinical Prediction Xiaoyang Meng, Xiaoding Zhou, Yang Liu, Weiyu Zhang, and Hongjiao Guan (Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences), Jinan, China; Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, China); Xueping Peng (University of Technology Sydney); and Wenpeng Lu (Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences), Jinan, China; Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, China) Abstract Abstract Clinical prediction tasks are critical for supporting medical decision-making in healthcare. Recent advances in large language models (LLMs) have opened new opportunities to improve predictive accuracy and clinical reasoning. However, existing methods are limited in capturing both short-term fluctuations and long-term trends in physiological signals, and insufficient interpretability remains a key barrier to clinical adoption. To address these challenges, we propose EHR-TASR, a novel framework that integrates temporal dynamics with interpretable reasoning for clinical prediction. Specifically, we segment patient visits into 12-hour windows and compute multi-scale time-aware scores to capture deviations in vital signs, thereby supporting the modeling of physiological changes across different time scales. These temporal features are combined with structured electronic health record (EHRs) data and processed by a Bi-LSTM with self-attention to enhance sequential representation. In parallel, a locally fine-tuned LLM generates structured reasoning chains from medical concepts, ensuring interpretability and privacy preservation. Finally, the outputs of the sequence model and the LLM are integrated through a linear combination to generate robust predictions. Experimental results on the public MIMIC-III and MIMIC-IV datasets demonstrate that EHR-TASR consistently outperforms state-of-the-art baselines. Tuesday Virtual Room 2 IJCNN Paper LLM Reasoning and Planning III Session Chair: Linna Zhou (Beijing University of Posts and Telecommunications), Jaeeun Jang (Hanwha Systems) CoSu-UQ: Uncertainty Quantification for Reasoning via Confidence and Semantic Support Yinglun Feng, Qinhong Lin, Yuhao Zhang, Zhongliang Yang, and Linna Zhou (Beijing University of Posts and Telecommunications) Abstract Abstract Estimating uncertainty for long-form reasoning in large language models remains challenging, as existing methods based on token probabilities or semantic diversity often break down in multi-step reasoning scenarios. We propose CoSu-UQ, a multi-sample uncertainty quantification framework that jointly models generation confidence and cross-sample semantic support. CoSu-UQ aggregates token-level probabilities into response-level confidence and measures cross-sample semantic support via sentence-level entailment. These complementary signals are integrated through bidirectional entailment-based clustering over final answers to produce a robust uncertainty estimate. We conduct systematic evaluations on GSM8K, MATH, HotpotQA, 2WikiMultiHopQA, and MedQA across large language models of different scales. Results show that CoSu-UQ consistently outperforms representative baselines (e.g., Predictive Entropy, Semantic Entropy, LUQ) in terms of AUROC for distinguishing correct from incorrect generations. Ablation studies further confirm the complementarity between generation confidence and cross-sample semantic support, with particularly pronounced gains in long-chain reasoning and multi-hop question answering scenarios. Label-Confidence-Aware Uncertainty Quantification in Natural Language Generation Qinhong Lin, Yinglun Feng, Yuhao Zhang, Zhongliang Yang, and Linna Zhou (Beijing University of Posts and Telecommunications) Abstract Abstract Large Language Models (LLMs) demonstrate remarkable capabilities in generative tasks but pose potential risks due to their tendency to generate hallucinatory responses. Therefore, Uncertainty Quantification (UQ), which aims to distinguish the validity of answers, is crucial for ensuring the safety and robustness of AI systems. However, existing methods primarily rely on measuring the entropy of multiple stochastic samples to represent uncertainty, often overlooking the specific uncertainty information associated with the candidate answer under evaluation. This oversight can lead to biased classification outcomes. In this paper, we investigate the discrepancy between global entropy from multiple samples and local confidence of candidate answer, and propose a Label-Confidence-Aware Uncertainty Quantification (LCA-UQ) method based on Pointwise Kullback-Leibler (PKL) divergence. Our method effectively bridges the gap between the consistency of sampled outputs and the calibration of the candidate answer, thereby enhancing the reliability and stability of uncertainty assessments. Empirical evaluations across a range of popular LLMs and NLP datasets reveal that label sources significantly impact classification. Furthermore, our approach effectively captures the nuances between sampling results and label sources, demonstrating superior performance in uncertainty estimation. Information Geometric Uncertainty Modeling for Agentic Reinforcement Learning Zhao Weikang (Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences); Hua Jing, Wang Xiujuan, and Wang Haoyu (Institute of Automation, Chinese Academy of Sciences); and Kang Mengzhen (Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Abstract Abstract Recent agentic reinforcement learning (RL) methods have made strong progress in training tool-using large language models (LLMs). A common approach is to trigger extra rollouts at uncertain decision points, often detected by high entropy. However, entropy is only a scalar measure. It often fails to capture the geometric structure of the predictive distribution, particularly after tool interactions. As a result, entropy struggle to tell informative uncertainty from random noise. To address this, we propose GUA-RL (information Geometric Uncertainty modeling for Agentic Reinforcement Learning), an information geometric framework for adaptive sampling in agentic RL. Instead of relying on entropy, GUA-RL tracks fisher-based geometric signals after each tool interaction. It allocates sampling budget to the decision points that are most informative for exploration. We evaluate GUA-RL on 13 long-horizon reasoning tasks spanning competition-level math, knowledge-intensive multi-hop QA, and deep search. Across these benchmarks, GUA-RL delivers the best overall results and remains consistently effective. Overall, our findings show that information geometric uncertainty provides a practical replacement for entropy-only rollout in tool-augmented agentic reinforcement learning. Preserving the SPINE: How a Few Critical Tokens Prevent Global Entropy Collapse in LLM Reasoning Jaeeun Jang, Hansle Lee, and Sangmin Kim (Hanwha Systems) Abstract Abstract Reinforcement Learning with Verifiable Rewards (RLVR) and Reinforcement Learning from Internal Feedback (RLIF) often fail to benefit from test-time compute due to entropy collapse and the resulting loss of reasoning diversity. Through a token-level analysis of GRPO-style policy updates, we show that this collapse is driven not by uniform entropy decay, but by premature overconfidence at a small set of structurally critical, high-entropy decision points where reasoning trajectories branch (i.e., forking tokens). Based on this insight, we propose SPINE (Structural Preservation of Information at Nodal Elements), a structurally targeted entropy-preserving optimization method that selectively preserves entropy at these high-impact tokens via partial KL regularization. SPINE identifies forking tokens via offline entropy profiling and then applies online collapse-pressure scoring to regularize the policy only for the most collapse-prone subset, resulting in entropy control over just 2-3% of tokens. Across multiple scales of Qwen2.5 Base models and math reasoning benchmarks, SPINE consistently improves single-sample accuracy and test-time scaling performance (e.g., multi-sample inference) under RLVR and RLIF, demonstrating that targeted entropy preservation at key branching points sustains reasoning diversity without sacrificing optimization efficiency. Tuesday Virtual Room 3 IJCNN Paper LLM Reasoning and Planning IV Session Chair: Zhiwei Yu (Tsinghua University), Rui Yan (Zhejiang University of Technology) ASCII-Reasoning: A Structural Spatial Reasoning Benchmark for Large Language Models Zhiwei Yu (Tsinghua University), Xiangyu Wang and Chengze Du (Beijing University of Posts and Telecommunications), and Lin He and Jiahai Yang (Tsinghua University) Abstract Abstract Spatial reasoning is a critical component of general intelligence, yet evaluating this capability in Large Language Models (LLMs) without visual encoders remains a significant challenge. Existing benchmarks either rely on multimodal inputs or lack structural complexity in pure text. To bridge this gap, we introduce the ASCII Reasoning Benchmark (ARB), a comprehensive framework comprising 2,400 samples across eight tasks, such as mechanical transmission and 3D folding. ARB utilizes discrete ASCII grids to represent topology, compelling models to construct internal mental models solely based on symbolic alignments. A systematic zero-shot evaluation of 16 LLMs reveals a significant capability gap, with the top-performing model achieving only 37.4\% accuracy. Crucially, our analysis uncovers counter-intuitive insights: a Parameter Efficiency Paradox, where lightweight, code-optimized models frequently outperform larger counterparts, suggesting symbolic reasoning derives more from code synthesis than parameter scale; and distinct Capability Specialization, where models achieve state-of-the-art performance in specific niches like pathfinding despite lower overall rankings. Code and Data are available at https://anonymous.4open.science/r/ASCII-Based-Spatial-Reasoning-Benchmark-4C2B. Learning to Reason across Viewpoints: Multi-Viewpoint Spatial Reasoning for Multimodal Large Language Models Xiaojian Huang (University of Science and Technology of China); Jie Zhao and Xin Liu (Noah's Ark Lab, Huawei Technologies Ltd.); and Jin Xu, zhang zhihong, Luo Zhuodong, and Xuejin Chen (University of Science and Technology of China) Abstract Abstract Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding and reasoning, yet they still struggle with spatial reasoning---especially when it requires reasoning across viewpoints. To address this limitation, we propose a scalable data generation pipeline that combines a text-to-image generative model with a scene generator. The pipeline enables annotation-free construction of high-quality multi-viewpoint spatial reasoning data, spanning diverse task formulations and viewpoint types. Using this pipeline, we build MVSR-Data, a dataset of 205K question--answer (QA) pairs covering both single-image and multi-image settings. We further introduce MVSR-Bench, a manually verified benchmark that provides a more comprehensive evaluation of multi-viewpoint spatial reasoning than prior benchmarks. Fine-tuning on MVSR-Data leads to strong performance across multiple multi-image spatial reasoning benchmarks. Unsupervised Mutual Learning among Comparable Large Language Models with Small Parameters Jianjun Zeng, Weiyan Zhang, and Yan Zhou (East China University of Science and Technology); Chao Wang (Shanghai University); Tong Ruan (East China University of Science and Technology); and Jingping Liu (Sun Yat-sen University) Abstract Abstract Recent progress in Large Language Models (LLMs) has increased interest in collaborative learning, typically grouped into two types: inference-based and training-based ones. However, both face challenges in a low-resource and small models (e.g., 7B-parameter LLMs) setting. The former contains instructions too complex for small models, while the latter require costly data. To this end, we propose Unsupervised Mutual Learning (UML), a novel framework that allows multiple small and comparable models to learn mutually using minimal unlabeled data. By leveraging inherent model and response diversity, UML establishes an iterative cycle where the models generate diverse outputs and learn from mutual input-output pairs. Experiments show that UML surpasses 8 baselines on 4 diverse tasks. Efficient ANN-to-SNN Conversion for Billion-Scale Language Models via Optimized Nonlinear Operator Approximation Shuyang Chen, Jiayu Zhang, and Chenhao Ye (Zhejiang University); Rui Yan (Zhejiang University of Technology); and Huajin Tang (Zhejiang University) Abstract Abstract Large Language Models (LLMs) have achieved remarkable performance across a wide range of tasks, but their high computational and energy costs remain a critical challenge. At the same time, Spiking Neural Networks (SNNs) offer exceptional energy efficiency through event-driven asynchronous computation, yet scaling them towards LLMs remains an open problem. We present an optimized ANN-to-SNN conversion framework that successfully scales to 7B-parameter LLMs by addressing the bottleneck of nonlinear operator conversion and quantization error in Transformer architectures. Our method introduces a piecewise optimization method for unary nonlinear operators, replacing gradient-based optimization with interpretable analytical method that achieve superior stability and 65% lower computational overhead compared to baseline methods. Integrated with modern LLM quantization techniques, our converted SNN maintains performance comparable to quantized ANNs while reducing theoretical matrix multiplication energy consumption to 22% of conventional ANNs. This work demonstrates a successful application of ANN-SNN conversion at billion-scale LLM parameters, providing a pathway toward energy-efficient neuromorphic language models. Tuesday Virtual Room 4 IJCNN Paper LLM Reasoning and Planning V Session Chair: Yujia Huo (School of Data Science and Information Engineering, Guizhou Minzu University, China; Guizhou Provincial Key Laboratory of Applied Mathematics and Computing Power & Algorithms, Guizhou, China), wen Zhang (Zhejiang University) Multi-stage Rationale Generation and Contrastive Learning for Chain-of-Thought Distillation Juan Li, Wen Zhang, Mingchen Tu, and Shijian Li (Zhejiang University) Abstract Abstract Recently, knowledge distillation (KD) has gained significant attention for its ability to transfer stronger reasoning capabilities from large language models (LLMs) to small language models (SLMs). Among the KD methods, chain of thought distillation (CoTD) has been explored widely, distilling the intermediate rationales of LLMs to bolster the reasoning performance of SLMs. Existing CoTD works use LLMs as teachers to generate rationales and fine-tune student SLMs based on correct rationales, neglecting the crucial value of incorrect rationales. To address the issue, we propose a novel CoTD framework MGCL that leverages Multi-stage rationale Generation and dual-level Contrastive Learning to enhance the reasoning ability of SLMs by improving the quality of rationales. Specifically, we first employ a self-correction pipeline to achieve multi-stage generation, which diagnoses errors and synthesizes refined and high-quality rationales. Then, we apply a dual-level contrastive learning mechanism to align sub-questions and rationales across different stages. Finally, the SLM is optimized via multi-task joint learning. Extensive experiments demonstrate the superiority of our method on both commonsense and biomedical reasoning tasks. Divide, Verify, and Conquer: Verification-Guided DAG Reasoning for Large Language Models Yigeng Jiang and Tingjun Su (Xiamen University); Tong Wu (Du Xiaoman Finance); Shumeng Zhang, Yanxu Zhao, Zhaohong Huang, Tianyu Xie, Yuhang Wu, and Yisheng Lin (Xiamen University); Yuze Liu (Shanghai AI Lab); and Fei Chao and Xiawu Zheng (Xiamen University) Abstract Abstract Chain-of-Thought (CoT) prompting is widely adopted for mathematical and scientific reasoning by decomposing problems into stepwise chains. However, the lack of intermediate verification leads to error propagation, where early mistakes can invalidate the entire reasoning process. To address this limitation, we introduce Divide, Verify, and Conquer (DVC), a training-free framework that models problem-solving as a DAG rather than a linear chain. DVC employs meta-rules to guide two key transformations: Problem Equivalence (PE) for logically equivalent reformulations and Problem Decomposition (PD) for breaking problems into manage-able subproblems with explicit dependency modeling. We design dedicated verification templates combined with a multi-round reflection mechanism to ensure each transformation step is reliable, effectively reducing error accumulation. Extensive experiments conducted on challenging mathematical and scientific reasoning benchmarks clearly demonstrate the effectiveness of the proposed DVC framework, outperforming existing methods with improved accuracy and reliability. ToT-ST: Tree-of-Thought Reasoning with Large Language Models for Zero-Shot Speech Translation Dandan Xiao, Zhenbei Guo, Kaidi Zhu, Fei Ouyang, Yan Xiang, and Zhengtao Yu (Kunming University of Science and Technology) Abstract Abstract In recent years, combining pre-trained speech foundation models(SLM) with large language models (LLMs) has emerged as an effective paradigm for improving speech-to-text translation performance. However, most existing LLM-based speech translation methods rely on linear Chain-of-Thought (CoT) reasoning and single-path decoding strategies, which limits their ability to fully exploit diverse translation hypotheses and makes them prone to error propagation, especially under zero-shot settings. In this paper, we propose a novel Tree-of-Thought reasoning framework for speech translation task, termed ToT-ST, which reformulates speech translation as a structured reasoning problem. Instead of being restricted to a single reasoning trajectory, ToT-ST enables LLMs to generate, evaluate, and select multiple candidate translation paths in parallel, thereby explicitly modeling diverse translation strategies. This structured reasoning mechanism allows the model to better explore the translation hypothesis space effectively improving translation performance in zero-shot speech translation scenarios. We evaluate ToT-ST on multiple benchmark datasets, including FLEURS, CoVoST-2, and MuST-C, covering zero-shot and cross-lingual scenarios. Experimental results demonstrate that ToT-ST achieves competitive improvements over end-to-end models and LLM-based baselines in terms of BLEU and COMET scores. CAD: Curriculum-Enhanced Agent Distillation via Trajectory Refinement and Difficulty-Aware Scheduling Xiang Yu (School of Data Science and Information Engineering, Guizhou Minzu University, China); Yujia Huo (School of Data Science and Information Engineering, Guizhou Minzu University, China; Guizhou Provincial Key Laboratory of Applied Mathematics and Computing Power & Algorithms, Guizhou, China); Yongbin Qin, Fujian Feng, and Gan Liu (School of Data Science and Information Engineering, Guizhou Minzu University, China); and Dawen Xia (College of Microelectronics and Artificial Intelligence and College of Big Data Engineering, Kaili University, China) Abstract Abstract Large Language Model (LLM) agents exhibit robust problem-solving capabilities through tool usage; however,their deployment is constrained by prohibitive computational costs. Knowledge distillation offers a compression pathway; however, transferring dynamic agent behaviors remains challenging due to the complexity of interaction trajectories and the risk of training instability. We present Curriculum-enhanced Agent Distillation (CAD), a framework designed to progressively transfer agent capabilities to lightweight student models. Unlike static chain-of-thought distillation, CAD orchestrates a dynamic, easy-to-hard learning curriculum governed by a novel multi-dimensional metric quantifying tool frequency, code complexity, and interaction depth. Furthermore, we introduce trajectory optimization techniques, including a "First-thought Prefix" mechanism, to refine teacher supervision. Extensive evaluations across eight benchmarks (e.g., mathematical reasoning, factual QA) demonstrate that CAD-trained students (1.5B) significantly outperform standard distillation baselines and achieve parity with 3B-scale models, exhibiting superior generalization in open-ended tasks. Tuesday Virtual Room 5 IJCNN Paper LLM Reasoning and Planning VI Session Chair: Zhe Cui (Beijing University Of Posts and Telecommunications, Beijing Key Laboratory of Network System and Network Culture), a a (none) Medium-GRPO: Online Difficulty-Adaptive Calibration for Enhancing LLM Reasoning Yihang Feng (Beijing University Of Posts and Telecommunications); Zhe Cui (Beijing University Of Posts and Telecommunications, Beiing Advanced Innovation Center for Future Blockchain and Privacy Computing); and Fei Su (Beijing University Of Posts and Telecommunications) Abstract Abstract Group Relative Policy Optimization (GRPO) is widely adopted in the post-training of Large Language Models (LLMs) to enhance mathematical reasoning performance. However, its effectiveness depends critically on the variance of rewards within a group of sampled responses. When the difficulty of the dataset is misaligned with the model’s capability, the model often yields groups of responses that are uniformly correct or incorrect, leading to vanishing gradient signals that hinder optimization. Addressing the limitations of existing augmentation methods, we propose Medium-GRPO, an online difficulty-adaptive calibration framework. Leveraging the model's intrinsic capabilities, Medium-GRPO utilizes an expression-based strategy to shift problems that are overly hard or easy for the model into the medium-difficulty range. We further introduce a fine-grained difficulty metric based on response length statistics to improve modification success rates, and integrate the modification process into the training objective for joint optimization. Experiments on five mathematical benchmarks show that Medium-GRPO consistently outperforms GRPO. Notably, on the challenging AIME-24 and AIME-25 benchmarks, our method achieves relative performance gains of 16.4\% and 11.2\% with Qwen3-4B. ACPR: Proactive Recommendation under User Acceptance Constraints Yexuan Che, Lei Zhang, Siyue Sheng, Zhendong Mao, and Yongdong Zhang (University of Science and Technology of China) Abstract Abstract Recommender systems typically prioritize immediate satisfaction, inadvertently reinforcing the Filter Bubble and limiting exploration. Proactive Recommendation addresses this by enhancing interest in a target item through a sequence of intermediate items. Prior works typically assume that users will passively accept all intermediate items, and primarily focus on the interest uplift toward the target item. However, in practice, if an intermediate item deviates significantly from a user's interest, the user is likely to actively abandon the current session, resulting in the complete failure of the target item's exposure. To address this, we propose the Acceptance-Constrained Proactive Recommendation (ACPR) framework, formulating proactive recommendation as a Constrained Markov Decision Process (CMDP) with user acceptance as a safety constraint. Instead of optimizing cumulative rewards related to the target, ACPR employs a step-wise greedy strategy to maintain every intermediate item within the user's current acceptance region, securing a feasible sequence for effective target exposure. Specifically, we instantiate the policy using a Large Language Model (LLM) to leverage its rich semantic knowledge to infer natural and logical connections between diverse items. To solve the constrained optimization problem, we integrate Reward Constrained Policy Optimization (RCPO) with Group Relative Policy Optimization (GRPO) to balance the proactive objective and acceptance constraint. Extensive experiments demonstrate that ACPR effectively improves users' interest in target items while maintaining high intermediate acceptance compared to state-of-the-art methods. The advantage of ACPR is particularly evident in scenarios with larger interest gaps, especially in sparse data scenarios where limited user history makes preference understanding more difficult. Code is available at: https://github.com/YexuanChe00/ACPR Motion Generation from Fine-grained Textual Descriptions With AI Feedback Ruilin Wang and Keze Wang (Sun Yat-sen University) Abstract Abstract Text-to-motion generation aims to generate 3D human motion sequences that align with given textual descriptions. Most existing methods achieve remarkable performance under coarse-grained text conditions, but their applicability is limited in scenarios requiring fine-grained descriptions with rich temporal and spatial details. Prior studies on fine-grained text conditions leverage large language models(i.e., ChatGPT) to transform coarse-grained descriptions into fine-grained ones. But the contents usually include redundant, missing, or erroneous details due to the absence of visual input, reducing the text-motion alignment performance of the generative models. To address these challenges, we first investigate the performance of three mainstream generative models conditioned solely on the fine-grained descriptions. Secondly, to mitigate the above limitation of large language models, based on the group relative policy optimization(GRPO) algorithm, we propose a reinforcement learning framework to optimize the generated fine-grained descriptions with feedback from the specific evaluation models(AI Feedback). Specifically, we employ modality alignment scores as reward functions to steer the policy model toward generating fine-grained descriptions that enable the generative model to synthesize motions with higher fidelity and better text-motion alignment. Furthermore, we incorporate a contrastive mechanism: generated samples are compared against existing fine-grained annotations based on their rewards. Samples achieving rewards greater than or equal to the rewards of existing annotations are regarded as higher-advantage(positive reward) candidates within the group, whereas those receiving lower rewards are considered lower-advantage(negative reward) candidates. This mechanism explicitly guides the model’s exploration toward more precise directions: surpassing existing fine-grained annotations, further enhancing the effectiveness of reinforcement learning. Extensive experiments on public datasets demonstrate the effectiveness of our approach. Notably, we observe that fine-grained descriptions generated by the large language model optimized in the proposed reinforcement learning framework provide consistent gains across the three generative models, highlighting the strong generalizability of our method. History-Aware Two-Stage Mathematical Verifier: A Coarse-to-Fine Method for Step-Level Mistake Finding Marcus Lee (none) Abstract Abstract Large language models (LLMs) are strong reasoners yet prone to hallucinations, which limits their reliability in error-sensitive applications such as education. The existing mistake finding methods typically provide two output labels such as correct and incorrect, and struggle to integrate historical data without substantial computational overhead, limiting both interpretability and scalability. We propose a novel two-stage, coarse-to-fine framework that couples confidence-based binary detection with rich, explanatory outputs and incorporates simulated interaction history through token-thrifty user embedding techniques. Our method achieves competitive performance in weighted average F1 and binary classification accuracy scores on challenging benchmarks by iteratively refining errorful solution steps and conditioning re-evaluation on targeted corrections. Our approach offers a practical and effective solution for reducing LLM hallucinations and enhancing the reliability of LLM reasoning. Tuesday Virtual Room 6 IJCNN Paper Learning Paradigms and Model Efficiency IV Session Chair: Ashish Anand (Indian Institute of Technology Guwahati), Yan Wan (DongHua University) Gated State-Space Model for Event Trigger Detection: A Linear-Time Alternative to Transformers Chaitanya Kirti (Indian Institute of Technology Guwahati), Harsh Kumar (Indian Institute of Technology Madras), and Ashish Anand and Prithwijit Guha (Indian Institute of Technology Guwahati) Abstract Abstract Event trigger detection aims to identify sparse yet decisive tokens that signal event occurrences in text. Most recent approaches rely on Transformer-based architectures with self-attention, which incur quadratic complexity and recompute global interactions at each layer. In this work, we revisit event trigger detection from a state evolution perspective and propose a gated state-space modeling framework. Our approach builds on a selective state-space model backbone and introduces a gated multi-head state refinement mechanism. The gating selectively amplifies event-relevant state updates while suppressing background tokens, enabling persistent contextual modeling without explicit attention. The model is trained end-to-end using a standard token-level classification objective. We evaluate the proposed models on four benchmark datasets: MAVEN, RAMS, OntoEvent, and Vrittanta-EN. Experimental results show that both vanilla and gated state-space models achieve performance comparable to strong Transformer-based baselines, including Gemma-2B, SPEECH, and EDM3. Across datasets, the proposed approach matches or slightly exceeds prior baselines in F1 score without relying on attention mechanisms. This work demonstrates that gated state-space modeling provides a principled and scalable alternative to attention-based architectures for event trigger detection, highlighting state evolution as an effective inductive bias for event trigger detection and related event-centric text understanding. EFSA-TD: Event-Level Financial Sentiment Analysis with Trigger Detection Shanhong Liu, Hongde Liu, Xingren Wang, Feiyang Meng, Senbin Zhu, Chenyuan He, Jiayang Luo, and Yuxiang Jia (Zhengzhou University) Abstract Abstract Event-level financial sentiment analysis aims to identify sentiment polarities associated with financial events mentioned in text. Existing approaches typically represent events as abstract labels without explicitly modeling the textual evidence that signals event occurrences, which limits interpretability and robustness, especially in texts containing multiple events. To address this limitation, we propose EFSA-TD, a joint modeling framework for event-level financial sentiment analysis with trigger detection. EFSA-TD simultaneously detects financial events and their corresponding trigger spans while performing event-level sentiment analysis. By grounding event recognition and sentiment prediction in explicit textual evidence, EFSA-TD enables more accurate, robust, and explainable sentiment modeling. We further construct FinEvent-Trigger, a newly annotated dataset consisting of 9,120 financial news articles, in which 18,158 events are annotated with explicit event trigger spans. Experimental results demonstrate that jointly modeling events and triggers through EFSA-TD consistently improves both event classification and sentiment classification performance compared with methods without considering event triggers. Mitigating Negative Gradient Bias in Test-Time Learning for Few-Shot Object Detection Taijin Zhao, Heqian Qiu, Lanxiao Wang, Yu Dai, Qingbo Wu, Fanman Meng, and Hongliang Li (University of Electronic Science and Technology of China) Abstract Abstract Test-Time Learning for Few-Shot Object Detection (TTL-FSOD) aims to adapt few-shot detectors to unlabeled test data, typically via teacher-student pseudo-labeling. However, this paradigm is undermined by the intrinsic sensitivity of detectors to pseudo-labeling noise. In this work, we identify three critical challenges: (1) semantic corruption, where false negative pseudo-labels cause the model to distort learned correct class features; (2) long-tailed class imbalance, where dominant base classes suppress novel class learning; and (3) foreground-background proposal imbalance, where excessive background proposals overwhelm sparse foreground proposals. Accordingly, we introduce feature orthogonality to decouple objectness from classification, shielding class semantics from corrupting gradients. Additionally, we design a Prototype-based Exponential-Moving-Average (P-EMA) classifier update to provide stable supervision for novel classes, and employ background resampling to rebalance foreground-background proposals. Experiments demonstrate our method effectively mitigates these core biases, improving TTL-FSOD performance. Few-Shot Object Detection via Hierarchical Self-Calibrating and Synergistic Feature Aggregation Yan Wan and Zhiheng Pan (Donghua University) Abstract Abstract Few-shot object detection (FSOD) addresses data scarcity by learning to detect novel classes from limited examples. Meta-learning has advanced FSOD by enriching class-level representations from scarce support data. However, existing representative methods that improve robustness via distributional feature modeling still face critical bottlenecks: large inter-class bias, weak support–query interaction, and limited generalization capability of the detection head. To address these issues, we propose a Hierarchical Self-Calibrating and Synergistic Feature Aggregation framework. It integrates four calibration stages: 1) centroid-driven feature calibration to suppress intra-class variance; 2) self-attentive Variational Auto-Encoder with Hellinger prior for more reliable class distributions; 3) lightweight bidirectional support–query attention for mutual adaptation; and 4) adaptive prediction logic calibration. To our knowledge, this work is among the first meta-learning-based FSOD frameworks that enable end-to-end collaborative optimization across feature, distribution, interaction, and decision levels. Comprehensive experiments on the PASCAL VOC and MS-COCO benchmarks demonstrate that our model achieves state-of-the-art performance among meta-learning methods, achieving superior performance with only a minor increase in parameters. Tuesday Virtual Room 7 IJCNN Paper Learning Paradigms and Model Efficiency V Session Chair: Xufei Zhang (Beijing XingYun Digital Technology Co., Ltd), Cong Hu (Jiangnan University) Interpretable Logical Anomaly Classification via Constraint Decomposition and Instruction Fine-Tuning Xufei Zhang, Xinjiao Zhou, Ziling Deng, Dongdong Geng, and Jianxiong Wang (Beijing XingYun Digital Technology Co., Ltd) Abstract Abstract Logical anomalies are violations of predefined constraints on object quantity, spatial layout, and compositional relationships in industrial images. While prior work largely treats anomaly detection as a binary decision, such formulations cannot indicate which logical rule is broken and therefore offer limited value for quality assurance. We introduce Logical Anomaly Classification (LAC), a task that unifies anomaly detection and fine-grained violation classification in a single inference step. To tackle LAC, we propose \emph{LogiCls}, a vision–language framework that decomposes complex logical constraints into a sequence of verifiable subqueries. We further present a data-centric instruction synthesis pipeline that generates chain-of-thought (CoT) supervision for these subqueries, coupling precise grounding annotations with diverse image-text augmentations to adapt vision language models (VLMs) to logic-sensitive reasoning. Training is stabilized by a difficulty-aware resampling strategy that emphasizes challenging subqueries and long tail constraint types. Extensive experiments demonstrate that \emph{LogiCls} delivers robust, interpretable, and accurate industrial logical anomaly classification, providing both the predicted violation categories and their evidence trails. Simplifying user identity linkage model via user modeling and two-phase instruction fine-tuning Jiaqi Gao, Kangfeng Zheng, Minjiao Yang, and Yulin Yao (Beijing University of Posts and Telecommunications) Abstract Abstract User identity linkage (UIL) is a fundamental task in social network analysis, yet most existing approaches rely on computationally expensive ensembles of deep neural networks, limiting their scalability and generalization. We propose UILLLM, a simple and unified framework that performs UIL using a single compact large language model. UILLLM introduces a Physical–Social–Cognitive (PSC) user modeling paradigm and adopts a two-phase instruction fine-tuning strategy. In the first phase, structured PSC representations are distilled from user-generated content to guide model learning. In the second phase, UIL is formulated as a binary classification task based on these representations, enabling effective identity linkage with low computational overhead. Experiments on two real-world datasets, TWIN and TWFQ, show that UILLLM consistently outperforms state-of-the-art methods, achieving accuracies of 96.56% and 99.14%, respectively. Anchor Then Align: Anomaly-Aware CLIP Adaptation with Disentangled Text Anchors for Weakly-Supervised VAD chuanhao Liu (LANZHOU JIAOTONG UNIVERSITY) Abstract Abstract Weakly-supervised video anomaly detection (VAD) aims to localize anomalies using only video-level supervision, but it is often challenged by temporal ambiguity and semantic noise. Recent CLIP-based approaches benefit from vision–language priors, yet most of them keep CLIP frozen and rely on generic prompt matching. We observe that CLIP is not sufficiently anomaly-aware in its pretrained semantic space: it struggles to clearly distinguish normal from anomalous prompts, resulting in an insufficient similarity margin and unstable segment-level scoring. To address this limitation, we propose an anomaly-aware CLIP adaptation framework that equips CLIP with lightweight residual adapters and a text-space disentanglement objective, explicitly constructing more separable normal–anomaly semantic anchors. On top of the adapted representations, we further introduce a parameter-efficient temporal adapter and a dual-branch weak supervision design that unifies binary MIL classification with clip-to-text MIL-Align learning, optimized via a three-stage curriculum for stable training. Extensive experiments on UCF-Crime and XD-Violence demonstrate that our method achieves state-of-the-art performance under the weakly-supervised setting with lightweight trainable modules. SAT:Semantic-Guided Adversarial Training on Long-Tailed Distributions Huaxing Feng, Yuanbo Li, Cong Hu, and Xiaojun Wu (Jiangnan University) Abstract Abstract Adversarial robustness remains a significant challenge in deploying deep neural networks (DNNs) for real-world applications. While adversarial training is widely acknowledged as a promising defense strategy, most existing studies primarily focus on balanced datasets, neglecting the fact that real-world data often exhibit a long-tailed distribution, which introduces substantial challenges to robustness. Existing adversarial training methods based on examples or loss balancing strategies, although able to alleviate class bias in long-tailed distributions to some extent, rely on data-driven balancing constraints that struggle to learn tail class features with sufficient discriminative information. To address this problem, we propose a novel framework based on adversarial training, termed Semantic-guided Adversarial Training (SAT). SAT utilizes class-level representations extracted by a text encoder as fixed semantic embeddings to guide model training. Specifically, we employ a vision–language model, CLIP, to obtain textual semantic representations for each class, exploit the semantic consistency of the text encoder, and introduce an alignment loss to mitigate the model’s bias toward head classes, thereby improving adversarial robustness for tail classes. Furthermore, to more effectively evaluate the class-balanced performance of our method under long-tailed distributions, we introduce a more comprehensive evaluation metric, termed Balanced Robustness. Extensive experiments on multiple datasets demonstrate that SAT significantly enhances adversarial robustness under long-tailed distributions compared to state-of-the-art methods. In particular, on the CIFAR-10 dataset, our method achieves more than a 5% absolute improvement over existing state-of-the-art approaches. Tuesday Virtual Room 8 IJCNN Paper Learning Paradigms and Model Efficiency VI Session Chair: Guofeng Zhang (Zhejiang University), Yulin Hu (Southwest University) An Efficient and Scalable Graph Condensation with Structure-Preserving Yulin Hu, Fuyan Ou, and Ye Yuan (Southwest University) Abstract Abstract Graph condensation (GC) is pivotal for enabling Graph Neural Networks (GNNs) deployment in resource-constrained scenarios by compressing large-scale graphs into compact synthetic counterparts. Existing GC methods commonly suffer from computational inefficiency due to coupled optimization as well as encountering poor generalization across GNN architectures. To address these challenges, this study proposes an Efficient and Scalable Graph Condensation with Structure-Preserving (SP-ESGC), which possesses a decoupled design that separates node condensation from graph structure generation. Specifically, it first employs heat kernel feature propagation to generate node representation via spectral graph theory-inspired diffusion. Further, a novel hybrid clustering strategy is designed to extracts discriminative intra-class centroids from the node representation. Finally, a pre-trained edge predictor infers transferable structural patterns from the original graph, ensuring accurate synthetic graph generation. Extensive experiments on real-world graph datasets demonstrate that the proposed SP-ESGC implementes a precise GC with significantly high computational efficiency. Moreover, SP-ESGC also generalizes well across diverse GNN architectures. Causal Discovery for Cross-Sectional Data Based on Super-Structure and Divide-and-Conquer Wenyu Wang and Yaping Wan (University of South China) Abstract Abstract This paper tackles a critical bottleneck in Super-Structure-based divide-and-conquer causal discovery: the high computational cost of constructing accurate Super-Structures—particularly when conditional independence (CI) tests are expensive and domain knowledge is unavailable. We propose a novel, lightweight framework that relaxes the strict requirements on Super-Structure construction while preserving the algorithmic benefits of divide-and-conquer. By integrating weakly constrained Super-Structures with efficient graph partitioning and merging strategies, our approach substantially lowers CI test overhead without sacrificing accuracy. We instantiate the framework in a concrete causal discovery algorithm and rigorously evaluate its components on synthetic data. Comprehensive experiments on Gaussian Bayesian networks, including magic-NIAB, ECOLI70, and magic-IRRI, demonstrate that our method matches or closely approximates the structural accuracy of PC and FCI while drastically reducing the number of CI tests. Further validation on the real-world China Health and Retirement Longitudinal Study (CHARLS) dataset confirms its practical applicability. Our results establish that accurate, scalable causal discovery is achievable even under minimal assumptions about the initial Super-Structure, opening new avenues for applying divide-and-conquer methods to large-scale, knowledge-scarce domains such as biomedical and social science research. SMC: Hessian-Aware One-Shot Dataset Distillation Heng Shu, Yuquan Wu, and Qiming Yang (Institute of Software Chinese Academy of Sciences) Abstract Abstract Dataset distillation addresses the resource-intensity of deep learning by compressing a large training set into a small, high-fidelity synthetic dataset, significantly reducing memory footprint and training time while preserving competitive performance. One-shot dataset distillation extends this concept by generating a single, versatile synthetic set adaptable to varying Images-Per-Class (IPC) budgets, eliminating the need for repeated runs in scenarios with diverse requirements. However, existing one-shot methods are hampered by optimization instability, manifested as persistent and spatially unbalanced pixel updates, and tend to yield synthetic samples with blurred contours. To overcome these geometric and visual limitations, we propose Stabilized Multi-dimensional Condensation (SMC), consisting of three synergistic components: a Hessian-aware curvature regularization for landscape smoothing, a dynamic hybrid loss enhanced by a learned sampler for a balanced integration of feature stability and distributional fidelity, and a staged one-shot condensation schedule to coordinate the global optimization process. Extensive experiments across multiple datasets demonstrate that SMC delivers significant and consistent improvements. Notably, on CIFAR-10, SMC achieves an average accuracy of 60.6%, outperforming the state-of-the-art baseline by 2.2%. Code will be made available upon publication. PureDiff: Purifying Sparse-View Reconstructions with Distractors via One-Step Diffusion Model Qirui Hu, Jingjing Wang, Chong Bao, Xiyu Zhang, Zhaopeng Cui, and Guofeng Zhang (Zhejiang University) Abstract Abstract Neural rendering from casual sparse-view captures with dynamic distractors is challenging. Existing sparse-view methods assume static scene and ignore distractors, thereby yielding artifacts, while dense-view distractor removal based on fitting residual assumption fails under sparse-view settings. We propose PureDiff, a novel one-step diffusion model enhanced with multi-view epipolar attention for high-quality, distractor-free reconstruction from casual sparse-view captures. PureDiff serves dual roles. As a Cleaner, it adopts a divide-and-conquer strategy with intra-set training and inter-set contrastive inference to effectively detect distractors. As a Refiner, it repairs severe rendering artifacts (e.g., holes and floaters) and feeds refined results back as additional supervision to improve reconstruction quality. To fit sparse-view settings, we further introduce an efficient few-shot fine-tuning scheme that significantly reduces training cost and data requirements. Experiments show that PureDiff consistently outperforms state-of-the-art methods in both rendering quality and distractor removal on casual sparse-view scenes with dynamic distractors. Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V9 Session Chair: Ru Lei (Northwestern Polytechnical University) MODRAL: Two-Phase PPO-Guided ALNS for Multi-Center Home Healthcare Routing with Time-Window Satisfaction YongQi Liu (Qilu University of Technology) Abstract Abstract Multi-center home healthcare routing under tight and heterogeneous time windows poses significant challenges: the feasible region is sparse and fragmented, and the effectiveness of destroy–repair operators is highly state-dependent, limiting static or weakly adaptive policies. We study a bi-objective variant where patient satisfaction is defined by arrival times relative to both hard and preferred time windows. To address these challenges, we propose MODRAL, a two-phase framework that decouples structure-aware initialization from learning-guided improvement. In Phase-1, a diverse pool of feasible solutions is constructed using time-window-prioritized heuristics that consider constraint tightness and spatial distribution. In Phase-2, Adaptive Large Neighborhood Search (ALNS) is enhanced with a Proximal Policy Optimization (PPO) agent that learns a state-aware operator scheduling policy, selecting destroy–repair pairs based on solution quality, search progress, and constraint satisfaction. A global external archive with dominance filtering preserves non-dominated solutions. Experiments on adapted Cordeau MDVRPTW instances show that the proposed initialization improves Pareto-front coverage and robustness, while the learned policy yields additional gains on tightly constrained instances by balancing exploration and exploitation. These results demonstrate the effectiveness of combining structured initialization with state-aware operator control for strongly constrained multi-objective routing. A Parameter-Free Joint-Space Metric for Multimodal Multiobjective Optimization Wenhua Li, Lida Zhang, Rui Wang, Tao Zhang, and Xin Lu (National University of Defense Technology) Abstract Abstract The evaluation of multimodal multi-objective evolutionary algorithms (MMEAs) is crucial for effective algorithm development and selection, yet current performance metrics often assess convergence and diversity in isolation within either the objective space or the decision space. Metrics such as Inverted Generational Distance (IGD) and Inverted Generational Distance in the Decision Space (IGDX) provide insights into convergence and solution diversity, respectively, but do not simultaneously capture performance across both spaces and frequently rely on subjective parameterization. This presents limitations in the comprehensive assessment of algorithm performance and adaptability to varying problem dimensions. To address these challenges, this paper proposes a new metric: the Dimension-Weighted Joint-Space Inverted Generational Distance (DWD). This metric integrates evaluations from both the decision and objective spaces. It exhibits robustness to variations in dimensionality and scale, operates without relying on adjustable parameters, and ultimately provides a more accurate reflection of an algorithm’s multifaceted capabilities. Cross-Wavelength Correlation Dynamic Multi-Objective Optimization for Optical Sensor Design Ru Lei, Lin Li, Kun Zhang, Yiqi Feng, Jiao Shi, and Chen Xia (Northwestern Polytechnical University) Abstract Abstract This paper proposes a cross-wavelength correlation based dynamic multi-objective optimization (DMO) framework for multilayer optical device design. Such structures require balancing optical performance and structural compactness. Optimal layer configurations at different wavelengths are not independent but exhibit intrinsic cross-wavelength correlations. To exploit this property, the proposed algorithm extracts information from previously optimized wavelengths to predict promising regions for the next wavelength and provides correlation-aware initialization, thereby reducing redundant exploration and improving convergence efficiency. The framework produces a diverse set of Pareto set(PS) that trade off near-perfect absorption and compact multilayer sensor structures. We further develop a prediction-driven DMO tailored to this design task and validate it through numerical simulations on representative multilayer design cases. The results show that the proposed approach can reliably obtain robust and structurally compact designs under multi-wavelength requirements, while maintaining strong absorption performance and efficient convergence. This study demonstrates an effective real-world deployment of DMO and facilitates efficient design of multilayer optical sensor devices and related sensing technologies. HydroFusion: A Multimodal Hydrological Data Model by Fusing Knowledge Graphs and Spatiotemporal Data Cubes yanghui he (Changsha University of Science and Technology, Yongzhou Industrial and Trade Secondary Vocational School) Abstract Abstract Hydrological analysis increasingly relies on het- erogeneous, multi-source spatiotemporal data, yet existing pipelines struggle to jointly support semantics, provenance, and high-throughput numerical analytics. This paper presents HydroFusion, a multimodal hydrological data model that couples a knowledge graph (KG) with a spatiotemporal data cube (STDC) in a two-layer architecture. HydroFusion con- tributes: (1) a dynamic, process-centric hydrological ontology with container–content separation; (2) a semantic-pointer- driven hybrid storage and two-stage execution that separates graph reasoning from parallel numerical analytics; and (3) an LLM+RAG natural language interface that translates requests into Cypher and reproducible Python analytics. Experiments on CAMELS-GB, ERA5-Land, USGS NWIS, and Landsat 8 ARD show that HydroFusion achieves 53.9× and 63.4× speedups over PostGIS on cross-modal association and complex attribution queries, reaches 523.7 QPS at 32-way concurrency, and delivers 80% end-to-end accuracy for natural language tasks. A drought attribution case study further demonstrates provenance-aware, reproducible analysis. Tuesday Virtual Room 1 IJCNN Paper Learning Paradigms and Model Efficiency VII Session Chair: Wenge Rong (School of Computer Science and Engineering, Beihang University, China; Engineering Research Center of Integration and Application of Digital Learning Technology Ministry of Education,China), Jia LIU (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences) Controlling Mutual Information for Generative Few-Shot Data Augmentation via Parametric Distribution Construction Zhiwei Sun, Zhuofan Chen, Jianfei Zhang, Wenge Rong, and Zhang Xiong (Beihang University) Abstract Abstract Data Augmentation (DA) remains a challenge for Natural Language Understanding in extreme low-resource scenarios. Although generative models can exploit auxiliary high-resource data to synthesize few-shot samples, current approaches typically oscillate between semantic drift and mode collapse. From an information-theoretic standpoint, we formally establish that the efficacy of generative DA is governed by the Mutual Information within the training distribution, and identify existing methods as operating at rigid entropy extremes of this spectrum. To address this, we formulate Parametric Distribution Construction, a principled framework for regulating conditional entropy via tunable hyper-parameters, and instantiate it through two complementary strategies: Hard (cluster-based) and Soft (attention-based). Extensive evaluations across three real-world NLU benchmarks show that these instantiations effectively balance the correctness-diversity trade-off, filling the Pareto gap left by static baselines. They outperform the state-of-the-art PromDA, particularly on coarse-grained tasks characterized by high intra-class variance. Furthermore, entropy regulation serves as a robust safeguard against sampling bias, ensuring superior generalization even when seed data is non-representative. SIR-LMM: Learning Domain-Invariant Rationales for Out-of-Distribution Hateful Meme Detection Yeqi Sun, Shizhong Peng, and Yuqing Li (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences); Huan Liu and Zheng Lin (Institute of Information Engineering, Chinese Academy of Sciences); Hao Cui (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences); and Huichuan Liu (State Grid Jiangsu Electric Power Co., Ltd) Abstract Abstract Explainable Hateful Meme Detection (EHMD) aims to classify multimodal content while providing natural language rationales, which is a critical capability for transparent content moderation. However, existing approaches typically operate under the independent and identically distributed (i.i.d.) assumption. Under real-world distribution shifts, Large Multimodal Models (LMMs) are prone to overfitting source-specific biases, such as recurring templates and visual styles, resulting in poor generalization on diverse domains. To address this, we propose SIR-LMM, a framework that learns a Sufficient Independent Representation (SIR) to decouple semantic hatefulness from domain variations. By imposing an information bottleneck, we conceptually partition the LMM into a feature extractor and a rationale generator. This structure is optimized via a two-stage paradigm: 1) conditional domain adversarial training on the extractor to ensure the representation is predictive of hatefulness (sufficiency) yet domain-invariant (conditional independence), and 2) fine-tuning the generator conditioned on the fixed SIR. This design ensures that high detection accuracy and rationale quality are maintained even under significant distribution shifts. Extensive experiments on three benchmarks (FHM, HarmC, and HarmP) demonstrate the robustness of our approach. Without requiring target domain rationale supervision, SIR-LMM outperforms state-of-the-art explainable baselines, improving average target-domain F1 score by 5.4% and BERTScore by 9.4%. DESD: A Dual-Encoder Model with Adaptive Training for TCM Syndrome Differentiation Han Guan and Chongjun Wang (State Key Laboratory for Novel Software Technology, Nanjing University, Nanjing, China) Abstract Abstract Accurate syndrome differentiation is essential for diagnosis and treatment in Traditional Chinese Medicine (TCM), yet real-world clinical data often contain complex and heterogeneous cases that challenge automated systems. To address this, we propose DESD, a novel dual-encoder model that separately processes clinical texts and detection results, improving feature extraction while maintaining architectural simplicity. Furthermore, we introduce an adaptive training strategy that identifies hard samples using an evaluation model, verifies their label confidence via a large language model (LLM), and quantifies sample difficulty using information entropy. A nonlinear mapping function then assigns dynamic weights to emphasize high-confidence hard samples during training. Experiments on the TCM-SD dataset show that DESD achieves state-of-the-art performance, with an accuracy of 84.85\% and significant gains in Macro-F1 and Macro-Precision, demonstrating its effectiveness in enhancing both classification accuracy and robustness in real-world TCM syndrome differentiation. Towards Principled Prompt-Based Continual Learning via Consistent Preservation Yunjie HAN (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences); Huiyun LI (Shenzhen University of Advanced Technology); Zhengmin JIANG (City University of Hong Kong); Shunran ZHANG (University of Macau); Ming SANG (Harbin Institute of Technology (Shenzhen)); and Tiantian XU and Jia LIU (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences) Abstract Abstract Continual Learning (CL) aims to learn from a stream of tasks without forgetting previously acquired knowledge, yet catastrophic forgetting remains a major challenge. Prompt-based CL has recently achieved strong empirical performance, but its theoretical foundations are still limited. In this work, we provide a unified and rigorous characterization of consistent performance preservation for prompt-tuned Transformers in CL. We show that preserving past-task performance can be decomposed into two jointly sufficient requirements, and we derive explicit sufficient consistency conditions under a full prompt-tuned Transformer parameterization. Importantly, these conditions admit a tractable orthogonal null-space projection rule, yielding our Consistent Null-Space Projection (CNSP) method. Extensive experiments across diverse benchmarks and Transformer backbones demonstrate consistent gains in accuracy and reduced forgetting, with particularly pronounced improvements in high-dimensional and domain-shift settings. Tuesday Virtual Room 2 IJCNN Paper Learning with Noisy and Limited Labels I Session Chair: Baixin Li (Northeastern University), Shuohao Li (National University of Defence Technology) From Texture to Semantics: Uncertainty-Gated Data Synthesis for Long-Tailed Recognition Xuantong Liu, Jiaxin Yang, Shuohao Li, Jun Zhang, and Jun Lei (National University of Defence Technology) Abstract Abstract Real-world visual recognition is frequently challenged by long-tailed distributions, where the scarcity of tail-class samples severely degrades model performance. While classical approaches like re-balancing and Logit Adjustment (LA) mitigate prediction bias, they fail to fundamentally address the information deficiency inherent in tail classes. Data synthesis offers a solution but often suffers from blind cropping (e.g., standard CutMix) or reliance on external pretrained segmentation models. To bridge this gap, we propose Dual-Prior Mix (DP-Mix), a novel data augmentation strategy that synthesizes diverse tail samples by transplanting precise tail foregrounds onto rich head-class backgrounds. Specifically, we leverage two intrinsic image priors: a Texture Prior for structural integrity and a Semantic Prior for class-discriminative localization. Central to our method is the Uncertainty-Gated Selection mechanism, which dynamically switch between these priors based on mask entropy. This ensures high-quality object extraction throughout the training process—from the “cold start” phase to convergence—without requiring any pre-trained models. Extensive experiments on CIFAR-10/100-LT and ImageNet-LT demonstrate that DP-Mix not only achieves competitive performance but also consistently enhances the efficacy of existing mixing-based strategies. Similarity-based Soft Supervision and Classifier Adjustment for Federated Long-Tailed Learning Peng Wang, Liangcheng Wang, Hanlei Li, and Jun Zhou (Southwest University) Abstract Abstract Federated learning usually assumes that global class distributions are balanced. However, real-world data exhibit a long-tail class distribution, which can significantly degrade the model performance on tail classes. Existing federated long-tailed learning methods do not fully utilize the semantic relationship between classes to enhance tail class learning and often suffer from representation bias caused by imbalanced distributions. To solve this issue, we propose a novel federated long-tailed learning framework called FedSSCA, which improves model performance on tail classes through an inter-class knowledge transfer mechanism and adaptive classifier adjustment. Specifically, we introduce a semantic association-based inter-class similarity soft supervision mechanism, which uses classifier weights as class prototypes to compute inter-class similarity and promote knowledge transfer among classes. Furthermore, we incorporate a dynamic classifier adjustment strategy based on estimated class distributions to optimize the feature representation space and mitigate representation bias . Extensive experiments on the CIFAR-10/100-LT and Tiny-ImageNet-LT long-tailed datasets demonstrate that FedSSCA achieves significant performance improvements compared to existing federated long-tailed learning methods. Feature Manifold-Aware Margin for Logits Calibration in Long-Tailed Recognition Jiaheng Chen and Chaopeng Guo (Northeastern University) Abstract Abstract Decoupled training has become a standard paradigm for long-tailed recognition by separating representation learning from classifier retraining. However, most existing second-stage retraining methods characterize class-wise learning difficulty primarily based on sample quantity. As a result, differences in how classes are structured in the learned feature space are often overlooked. In this work, we identify representation-level characteristics as an additional factor influencing classifier difficulty, and propose Feature Manifold-Aware Margin, a lightweight plug-and-play module for classifier retraining. FMAM estimates class-wise geometric complexity using intra-class feature dispersion. Based on this estimate, it applies adaptive margins to calibrate class logits during training. Experiments on CIFAR-100-LT and ImageNet-LT demonstrate that FMAM consistently improves strong second-stage baselines, including CRT, LOS, and ViT-based models, with notable gains on few-shot categories of up to 3.87%. In particular, FMAM achieves a 3.79% improvement in overall accuracy on ImageNet-LT with a ViT backbone. Alleviating Overconfidence in Long-Tailed Time Series Classification with Hierarchical Reverse Distillation Baixin Li, Peibo Duan, Zhipeng Liu, and Jialu Xu (Northeastern University); Fan Zhang (The University of Tokyo); and Tianlun Wang and Bin Zhang (Northeastern University) Abstract Abstract Time series classification is a fundamental task with wide-ranging applications. However, existing studies typically assume balanced datasets, which are idealized and overlook the long-tailed nature of real-world data. A key challenge in long- tailed time series data is model overconfidence, with models performing well on head classes but poorly on minority tail classes. Existing approaches attempt to mitigate this issue by enhancing tail class performance, which often comes at the expense of head class accuracy. To address this limitation, we propose a novel Hierarchical Reverse Distillation framework for long-tailed time series classification, named HRD-LT. Specifi- cally, HRD-LT integrates a frequency-based hierarchical data augmentation strategy, generating both weakly and strongly augmented variant series, which enhances the representation of tail samples and mitigates overfitting by generating increasingly challenging samples. Furthermore, we propose a dual-expert collaborative learning strategy with reverse distillation, where the two experts take turns as teacher and student, enabling the overconfident student model to be effectively guided. Extensive experiments on three public datasets, each constructed with versions of varying imbalance ratios, demonstrate that our HRD- LT outperforms state-of-the-art baselines. Compared to the best- performing baselines, HRD-LT achieves improvements of up to 10.01% in terms of accuracy. Tuesday Virtual Room 3 IJCNN Paper Learning with Noisy and Limited Labels II Session Chair: Jinao Li (Qilu University of Technology), Fei Ding (Peking University) Semi-supervised Label Distribution Learning from Incomplete Annotations via Multi-view Fusion and Reliability-aware Filtering Jinao Li, Jinyong Cheng, and Kening Cui (Qilu University of Technology) Abstract Abstract Label Distribution Learning (LDL) models label ambiguity by assigning each instance a soft distribution over labels. In many real applications, however, obtaining complete label distributions is costly: annotators typically provide explicit description degrees for only a few dominant labels, leaving the remaining entries unobserved. This leads to incomplete label distributions and substantially weakens supervision. In this paper, we study semi-supervised label distribution learning under incomplete annotations and propose CoFuse-LDL, an EMA teacher–student framework that leverages unlabeled data via multi-view probability fusion and reliability-aware pseudo-distribution filtering. Specifically, the teacher and student predict distributions from weak and strong augmentations, and we fuse multi-branch predictions to obtain a robust pseudo-distribution. We then apply a confidence-based gate to filter unreliable pseudo targets and enforce distribution consistency on unlabeled samples with a KL divergence loss, while labeled samples are optimized using a masked cross-entropy on observed entries. Experiments on six LDL benchmarks with an observable ratio of 0.5 show that CoFuse-LDL consistently outperforms competitive baselines, demonstrating the benefit of fusing and filtering distributionvalued pseudo supervision when label distributions are only partially observed. UD-Net: Uncertainty-Guided Label Distribution Learning for Facial Depression Severity Estimation Su Yu, Zhenyu Yang, Yu Zheng, and Jing Wang (Southeast University) Abstract Abstract Facial depression severity estimation is an important problem in affective computing, but learning reliable predictors is difficult when annotations are inherently subjective and available training data are limited, as in AVEC 2014. Standard regression pipelines optimize point-wise losses and implicitly assume that the provided scores are noise-free, which overlooks the heteroscedastic label ambiguity in the target labels. Although Label Distribution Learning (LDL) alleviates this issue by learning distributions instead of single values, most LDL-based methods adopt a constant distribution width for all samples, failing to reflect sample-dependent difficulty. In this work, we present UD-Net, a lightweight ambiguity-aware LDL framework that incorporates a novel Error-Driven Variance Rectification strategy. During training, the rectified variance serves as a self-calibrated, curriculum-style supervision signal, encouraging wider distributions for hard cases and sharper ones for easy cases. On the AVEC 2014 test set, UD-Net achieves an MAE of 6.88 and RMSE of 8.86. Despite using only a ResNet-18 backbone, it delivers competitive accuracy compared with heavier 3D-CNN approaches while significantly reducing computational cost. Learning from Noisy Label Proportions via Dynamic Noise-Aware Instance Selection Peimin Ning, Jing Chai, Gang Hu, and Shunfang Wang (Yunnan University) Abstract Abstract Learning from Label Proportions (LLP) is a challenging weakly supervised learning problem that training instances are grouped as bags with only the instance proportions belonging to each class in each bag being accessible, leaving the individual instance labels unknown. Most existing LLP methods are designed under the assumption that the label proportions are accurately annotated, despite that this assumption is typically unsatisfied in real applications due to measurement or human errors. To address the gap between assumption and practice, we proposed a new LLP method by resorting to Dynamic Noise-Aware Instance Selection (LLP-DNAIS). Specifically, we design a module that fuses probabilistic consistency, feature consistency, predictive self-certainty, and proportion matching to calculate a confidence score for each instance, and then develop an adaptive thresholding strategy to dynamically evaluate the cleanliness of individual instances based on confidence scores, enabling differentiated training of clean and noisy instances. Finally, the bag-level and instance-level losses are integrated to jointly train the networks. Experimental results on noisy benchmark datasets demonstrate that LLP-DNAIS significantly outperforms existing baselines in classification accuracy. Semantic-Aware Hierarchical Negative Label Selection for Zero-Shot Out-of-Distribution Detection Fei Ding, Feng Zhang, Wei Chen, Tengjiao Wang, and Yan Zhang (Peking University) Abstract Abstract Zero-shot out-of-distribution (OOD) detection aims to determine whether an input image belongs to any in-distribution (ID) category relying solely on textual labels. Existing negative label-based methods typically construct negative label sets by selecting labels that are semantically distant from ID labels within a large semantic pool to represent potential OOD categories. While effective, this strategy inherently favors easy negatives and overlooks hard negatives that lie close to the decision boundary. In addition, directly aggregating all negative labels may lead to erroneous predictions when multiple semantically similar negative labels collectively dominate the scoring function. To address these issues, we propose HiNeg, a novel framework for zero-shot OOD detection that jointly models negative label diversity and cluster-level relative dominance. HiNeg introduces a hierarchical negative label selection strategy HiSelect that captures both semantically distant easy negatives and boundary-adjacent hard negatives, providing complementary coverage of potential OOD semantics. Furthermore, a cluster-relative aggregated score is designed to evaluate negative evidence relative to the strongest ID response, suppressing redundant contributions from correlated negative labels within the same cluster. Extensive experiments on standard zero-shot OOD detection benchmarks demonstrate the effectiveness of the proposed approach. Tuesday Virtual Room 4 IJCNN Paper Machine Learning Methods and Applications III Session Chair: Jianbin Jiao (University of Chinese Academy of Sciences), Keji Mao (Zhejiang University of Technology) Dual-Branch Noise-Aware Representation Learning for Fault Diagnosis Duerguli Maimaititusun (University of Chinese Academy of Sciences), Gen Ge and HaiChuan Hu (China Chemical Safety Association), and Jianbin Jiao (University of Chinese Academy of Sciences) Abstract Abstract Robust fault diagnosis under mixed noise and low signal-to-noise ratio (SNR) remains challenging because weak fault signatures can be submerged by non-stationary disturbances. Existing approaches typically operate on single feature domains and employ relatively simple fusion strategies, failing to fully exploit complementary information from multi-domain representations and dual-sensor deployments under low SNR conditions. To address these limitations, we propose a dual-branch noise-aware representation learning (\model) framework that jointly learns from two synchronized acquisition channels and extracts complementary representations in both time and time-frequency (TF) domains. First, differentiable denoising modules, including learnable wavelet soft-thresholding and adaptive Wiener filtering with frequency attention, are embedded for domain-specific noise suppression. Then, domain-specific feature extraction and hierarchical fusion mechanisms are proposed to capture and integrate complementary cues from dual-channel measurements. The time-domain branch employs a lightweight bidirectional depthwise convolution block that enables efficient 1D sequence modeling with long-range temporal context aggregation, while the TF-domain branch utilizes 2D convolutions to capture joint time-frequency patterns from spectrogram representations. Leveraging dual-channel spatial redundancy, a gated cross-channel fusion mechanism adaptively fuses dual-position features, with an explicit difference pathway for channel mismatch and propagation cue capture. At last, the time-domain and TF-domain representations are jointly fused for robust fault classification. Experiments on the two benchmark datasets for bearing fault diagnosis demonstrate the effectiveness of the proposed method under challenging noise conditions. TFFNet: A Time–Frequency Fusion Network for Automatic Modulation Recognition Qifan Zhang (University of Chinese Academy of Sciences; The Institute of Software, Chinese Academy of Sciences) and Lixiang Liu and Xin Zhou (The Institute of Software, Chinese Academy of Sciences) Abstract Abstract The rapid proliferation of wireless communication devices has intensified spectrum congestion, highlighting the need for intelligent spectrum monitoring and management. Automatic modulation recognition (AMR) plays a critical role in this context by enabling modulation identification for demodulation, interference detection, and dynamic spectrum access. Traditional likelihood- and feature-based AMR methods suffer from high computational cost and limited adaptability, while recent deep learning approaches, though effective, often process different signal representations in isolation and struggle under low-SNR conditions. To address these limitations, we propose TFFNet, a time–frequency fusion network that jointly exploits temporal and spectral characteristics of modulation signals. TFFNet adopts a dual-branch architecture: the Time-Mamba branch extends the Mamba state space model with a Frequency-aware State Space Model (FSSM) for efficient temporal–spectral sequence modeling, while the frequency-domain CNN (Freq-CNN) branch leverages Stationary Wavelet Transform (SWT) to extract discriminative time–frequency features. A cross-attention fusion module adaptively aligns and integrates domain-specific features, enhancing consistency and robustness. Extensive experiments show that TFFNet achieves 63.6% accuracy on RML2016.10a and 56.1% on RML2018.01a, demonstrating competitive performance compared to representative methods. Deep Covariance Denoising and Attention-Guided On-Grid Classification: A Dual-Stage Network for Robust Underwater DOA Estimation Bin Zhang, Jiawen He, Peishun Liu, Wenxu Wang, Ruichun Tang, and Liang Wang (Ocean University of China) Abstract Abstract The direction of arrival estimation (DOA) is a pivotal task in the domain of underwater acoustic signal processing. Some traditional existing methods and recently emerged deep learning algorithms still face issues such as sensitivity to low signal-to-noise ratio (SNR) and high model redundancy. Moreover, the models are often regarded as black boxes, lacking interpretability. To address the aforementioned challenges, this paper utilizes a uniform sonar linear array, employing the array covariance matrix as dual-channel input features. Under different SNR conditions, a denoising autoencoder (DAE) is designed as a pre-filter to reconstruct spatial spectrum features, enhancing the decoupling capability of SNR features. Heatmaps are applied to visualize and elucidate the reconstruction alignment process, enhancing model feature interpretability. Furthermore, a convolutional neural network with a channel attention mechanism is constructed as a post processing classification model to achieve underwater acoustic DOA estimation. The proposed dual-stage algorithmic framework demonstrates superior accuracy with reduced computational overhead across diverse signal scenarios. Evaluated using simulated signals (model size: 1.59 MB), it achieves peak accuracies of 94.95% with an RMSE of 1.603 (frequency: 200 Hz, array elements: 10). Under real-world operating conditions with field measurement data (model size: 2.05 MB), the model attains 99.99% accuracy with an RMSE of 0.121, confirming its practical efficacy. We still conducted extensive sensitivity analyses, establishing that this performance—whether in terms of speed, size, or accuracy—is superior to that of traditional DOA estimation algorithms and other deep learning models, with an emphasis on interpretability. MTCA: A Multi-Scale Temporal and Cross-Channel Attention Model for Depression Recognition Kancheng Shen, Yan Ling, Tianyu Zhou, Ruiji Xu, Zhihu Zhou, and Keji Mao (Zhejiang University of Technology) Abstract Abstract Spatiotemporal modeling using functional near-infrared spectroscopy (fNIRS) has emerged as a promising approach for objective depression diagnosis, leveraging distinct haemodynamic responses in task-evoked cortical regions between patients with major depressive disorder (MDD) and healthy controls (HC). Prior studies have shown that deep learning models based on fNIRS can effectively assist in the diagnosis of depression. However, mainstream approaches often fail to capture multi-scale temporal dependencies or overlook the importance of interactions across multiple channels. To address this issue, this paper proposes a multi-scale temporal and cross-channel attention (MTCA) model for depression recognition. Our framework integrates: (i) a multi-resolution temporal-aware module that captures short- and long-range dependencies via parallel dilated convolutions; (ii) a channel-wise attention mechanism inspired by iTransformer, which models cross-channel functional connectivity; and (iii) a pseudo-sequence augmentation strategy that synthesizes temporally coherent samples to mitigate data scarcity. Evaluated on a clinically collected fNIRS dataset (N=50, 30 MDD + 20 HC), MTCA achieves 90.8% weighted accuracy and 89.4\% unweighted accuracy.These results demonstrate MTCA’s effectiveness for depression diagnosis. The code will be released. Tuesday Virtual Room 5 IJCNN Paper Machine Learning Methods and Applications IV Session Chair: Zhi Wang (Southwest University), SS L (South China University of Technology) Prior-Driven Tensor RPCA: A New Regularization Approach for High-Dimensional Data Representation Li Zhou (Southwest University), Dong Hu (Chongqing University of Technology), and Yao Fu and Zhi Wang (Southwest University) Abstract Abstract Tensor Robust Principal Component Analysis (TRPCA) is an important technique in computer vision and machine learning, with numerous applications such as color image denoising and video background separation. Tensor nuclear norm minimization (TNNM) is a classical approach for TRPCA that has been shown to recover low-rank and sparse tensors under certain conditions. However, TNNM penalizes all rank components equally, ignoring their relative importance, which often leads to suboptimal recovery results. To overcome this limitation, we propose a novel truncated nonconvex fractional function to approximate the tensor multi-rank function. Benefiting from the truncation strategy and the nonconvex fractional structure, the proposed function preserves dominant singular values while adaptively penalizing less significant ones. Based on this function, we formulate a TRPCA model for separating low-rank structures and sparse noise from tensor data. An efficient algorithm based on the alternating direction method of multipliers (ADMM) is developed, and its computational complexity is analyzed. Extensive experiments on both synthetic and real-world datasets demonstrate that the proposed method outperforms several state-of-the-art approaches. Adaptive Dynamic Convolution Module Based Low-Rank Tensor Factorization Framework for Multi-Dimensional Image Recovery Dan Xu and Xiaoli Sun (Shenzhen University) and Zimeng Li and Xiujun Zhang (Shenzhen Polytechnic University) Abstract Abstract Due to sensor limitations, noise interference, and other factors, tensor data exhibit random or structured missing scenarios during acquisition and transmission. Thus, recovering the original high-dimensional tensor from incomplete observations has become a challenging problem. Although tensor recovery methods based on tube transform and low-rank decomposition have achieved certain progress, they remain insufficient in exploring local neighborhood structures and characterizing cross-channel interactions, making it difficult to effectively preserve spatial-spectral consistency. To address these issues, we propose an adaptive dynamic convolution module based low-rank tensor factorization framework for multi-dimensional image recovery, named LRADC. First, we propose a non-linear multilayer neural network guided by the adaptive dynamic convolution (ADC) module. The ADC module dynamically generates convolution kernels based on input content, boosting the network's capacity for local detail reconstruction and cross-channel feature interaction. Then, using this learned non-linear network, we define a low-rank tensor factorization framework. The newly proposed framework leverages low-rank factorization to model global structural priors and uses the ADC module to dynamically enhance local details and cross-channel feature interaction. Therefore, when addressing challenging tube missing scenarios, the model has obvious advantages in global structure mining and detailed information preservation. Extensive experiments on public datasets show that LRADC outperforms state-of-the-art methods in both visual and quantitative metrics. A Brain-Inspired Deep Separation Network for Single Channel Raman Spectra Unmixing Gaoruishu Long and Jinchao Liu (College of Artificial Intelligence, Nankai University); Bo Liu (Guangdong Laboratory of Chemistry and Fine Chemical Industry, Smekal Tech (Shantou) Ltd.); Jie Liu (College of Artificial Intelligence, Nankai University); and Xiaolin Hu (Department of Computer Science and Technology, Tsinghua University) Abstract Abstract Raman spectra obtained in real world applications are often a noisy combination of several spectra of various substances in a tested sample. Unmixing such spectra into individual components corresponding to each of the substances is of great value and has been a longstanding challenge in Raman spectroscopy. Existing unmixing methods are predominantly designed to invert an overdetermined mixed model and therefore require multiple mixed spectra as input. However, open domain and/or non-cooperative detection applications in Raman spectroscopy such as controlled substance detection, call for single-channel solutions which can identify individual components from thousands of candidates by analyzing only a single noisy mixed spectrum. To our knowledge, sparse regression is the only existing solution which can cope with this scenario, yet it has very low tolerance to noises and can hardly be applicable in practice. To address these limitations, we introduce a novel neural approach for single-channel Raman spectrum unmixing inspired by speech separation. It aims at solving underdetermined systems and can decompose a noisy mixed spectrum from a library of thousands of components (substances). The core of our method is a deep separation neural network (RSSNet) which takes a mixed spectrum as input and outputs spectra of pure components. We created two synthetic datasets of single-channel Raman spectra unmixing and demonstrated feasibility and superiority of RSSNet on these datasets (outperform competing methods by >4dB). Furthermore, we verified that RSSNet, trained solely on synthetic data, can successfully unmix real-world mixed spectra of mixtures of mineral powders, exhibiting strong generalization. Our approach represents a new paradigm for Raman unmixing and enable new possibilities for fast detection of Raman mixtures. Dual-Path Attention-Enhanced Mamba Network for Efficient Long-Sequence Speech Separation Shasha Long, Qianyu Yan, Junmei Yang, and Delu Zeng (South China University of Technology) Abstract Abstract This paper proposes a Dual-path Attention-enhanced Mamba Network (DAMamba) for efficient long-sequence speech separation. To address Mamba's limitations in local detail modeling and channel redundancy, we introduce three key components: a Multi-Scale Dilated Convolution module (MSDC) for local feature capture, a Segment-wise Channel Attention-enhanced Mamba Block (SAEMB) for channel calibration, and a dual-path enhancement architecture for global-local collaboration. Experiments on LRS2-2Mix and Libri2Mix show that DAMamba achieves SI-SNRi of 17.5 dB with only 26.9 percent parameters of SepFormer, confirming its efficiency and superiority. Tuesday Virtual Room 6 IJCNN Paper Machine Learning Methods and Applications V Session Chair: Zhiming Wang (Imperial College London), Vivek B S (Tata Consultancy Services) When Constraints Go Both Ways: Dynamic Hierarchical Multi-label Classification Networks Gelsomina Di Palma (Università di Genova), Zhiming Wang (Imperial College London), and Eleonora Giunchiglia (Imperial College) Abstract Abstract Hierarchical Multi-label Classification (HMC) problem assigns multiple labels to each instance under the requirement that predictions satisfy a set of hierarchical subclass constraints among labels. The current state-of-the-art models tend to enforce coherence by exploiting the sufficient direction of each constraint by propagating positive evidence from subclasses to superclasses. In this paper, we extend that idea by additionally exploiting the necessary direction by suppressing a class when its superclass is not supported. The resulting model, Dynamic Hierarchical Network (DHN) , is thus able to leverage both sufficient and necessary conditions while preserving coherence by construction. An extensive evaluation across 30 real-world benchmark datasets shows that DHN consistently improves over competitive baselines. Image-set classification using Discriminant Neighborhood Preserving Embedding on Symmetric Positive Definite Manifold Ying Hu (Chongqing Normal University, College of Computer and Information Science); Benchao Li (Southwest Jiaotong University, School of Computing and Artificial Intelligence); and Ruisheng Ran and Ji Feng (Chongqing Normal University, College of Computer and Information Science) Abstract Abstract Symmetric positive definite (SPD) manifolds provide a fundamental representation for high-dimensional image sets and covariance descriptors by capturing second-order statistics and intrinsic geometric structures. However, in high-dimensional scenarios, SPD manifold data commonly suffer from the curse of dimensionality and geometric distortion, while conventional Euclidean dimensionality reduction methods fail to preserve Riemannian consistency. To address these challenges, this paper proposes a discriminative neighborhood preserving embedding on SPD manifolds (SPD-DNPE) with dual-mode neighborhood weighting strategies. By extending neighborhood preserving embedding(NPE) to the SPD Riemannian manifold and incorporating class label information, the proposed method enhances discriminability in the low-dimensional tangent space while preserving local manifold geometry. Extensive experiments on multiple benchmark datasets demonstrate that SPD-DNPE achieves superior classification accuracy and robustness compared with existing Grassmann and SPD manifold methods, kernel-based approaches, and several deep manifold models. RGANN: Advancing CPC Patent Classification with Novel Multi-Label Datasets Yuan Meng, Shen Nie, Qiao Jia, Ye Yuan, and Xuhao Pan (Shanghai University of International Business and Economics) Abstract Abstract As the number of patent applications is increasing sharply, the CPC system is gradually replacing the IPC system as the new standard. However, due to the lack of public datasets, there are still some research gaps in the study of automatic patent classification based on the CPC system. In this work, we first propose two new patent datasets on multi-label classification using CPC, named Patent14K and Patent2M. Then, we introduce a series of baseline models based on BERT-for-Patents with the latest optimization methods on our dataset. Finally, we propose an innovative model, named Relation Graph Attention Neural Network (RGANN) based on the syntactic graph structure to improve the baseline models. Experimental results demonstrate that RGANN achieves a large F1-score improvement (3.01% on Patent14K dataset, 4.22% on Patent2M dataset) compared to the baseline models. Our all codes, datasets and the way to reproduce our results are available at https://github.com/qingtian5/RGANN4Patent2M. MIDAS: Manifold-Inspired Diverse Feature Acquisition for Simplicity Bias Mitigation Bhartendu Kumar (Microsoft) and Vivek B S, Jayavardhana Gubbi, and Arpan Pal (Tata Consultancy Services) Abstract Abstract A comprehensive understanding of what Deep Neural Networks (DNNs) learn remains elusive. While the Manifold Hypothesis offers a compelling geometric lens, existing theories often assume smooth manifolds, frequently overlooking non-smooth components such as ReLU and max-pooling. In this work, we present a theoretical framework modeling DNN layers as piecewise-smooth mappings between stratified spaces. Our derivation suggests that activations from piecewise-smooth nonlinearities naturally admit a stratified-space structure, extending standard smooth manifold assumptions. Theoretically, we establish a non-increasing layerwise upper bound on the Intrinsic Dimensionality (ID) of representations. We identify excessive geometric compression as the mechanism behind Simplicity Bias (SB): networks shed complexity, latching onto simple, low-dimensional features (low ID) at the expense of task-relevant ones. Leveraging these insights, we propose MIDAS (Manifold-Inspired Diverse Feature Acquisition). First, we introduce an annotation-free bias metric based on the ID gap (analogous to the Generalization Gap), quantifying the deficit between learned representation dimensionality and the data’s intrinsic complexity. Second, we propose a regularization objective that explicitly counteracts ID collapse to promote high-complexity feature learning. Empirically, MIDAS achieves significant gains (up to 10%) on CelebA, Waterbirds, Colored MNIST, and BAR over label-free baselines, approaching fully-supervised performance. Our framework establishes intrinsic dimension as both a diagnostic and a corrective tool for representation quality in deep networks. Tuesday Virtual Room 7 IJCNN Paper Machine Learning Methods and Applications VI Session Chair: Jungang Xu (University of Chinese Academy of Sciences), Min Li (Institute of Information Engineering, Chinese Academy of Sciences,State Key Laboratory of Cyberspace Security Defense; School of Cyber Security, University of Chinese Academy of Sciences) SSEC: Semantic Self-Enhanced Classification for Few-Shot Tabular Data via Adversarial Reasoning Min Li (Institute of Information Engineering, Chinese Academy of Sciences,State Key Laboratory of Cyberspace Security Defense; School of Cyber Security, University of Chinese Academy of Sciences); Yifei Zhang (Alibaba Group); Ming Yin (Institute of Information Engineering, Chinese Academy of Sciences,State Key Laboratory of Cyberspace Security Defense; School of Cyber Security, University of Chinese Academy of Sciences); and Neng Gao and Jia Peng (Institute of Information Engineering, Chinese Academy of Sciences,State Key Laboratory of Cyberspace Security Defense) Abstract Abstract Few-shot classification of tabular data is challenging due to the limited labeled samples and unstable decision boundaries. Most existing approaches rely on statistical pattern matching or shallow metric learning, which often fail to capture the underlying semantic structure of tabular features, such as domain constraints and column meanings, under extreme data sparsity. As a result, models are prone to overfitting or boundary collapse in low-shot regimes. To address this problem, we propose SSEC (Semantic Self-Enhanced Few-Shot Classification), a learning method that explicitly incorporates semantic reasoning into few-shot tabular classification. SSEC constructs structured classlevel semantic representations from limited labeled samples and iteratively refines them through an adversarial reasoning process, which sharpens decision boundaries and suppresses spurious correlations. Unlike purely statistical methods, SSEC leverages semantic constraints to provide a stronger inductive bias for fewshot learning. Experiments on multiple public tabular benchmarks demonstrate that SSEC performs competitively across varying few-shot settings, with particularly strong results in the 4-shot and 8-shot regimes. These results indicate that semantic enhancement is an effective strategy for improving robustness and generalization in few-shot tabular classification. Causal-Enhanced and Regime-Adaptive Hybrid Framework for Asset Return Prediction Xinze Fan, Zhongliang Yang, and Linna Zhou (Beijing University of Posts and Telecommunications) Abstract Abstract Forecasting asset returns presents a formidable challenge for neural architectures, characterized by low signal-to-noise ratios, non-stationary concept drift, and the structural misalignment between semantic and numeric modalities. Existing paradigms often necessitate a compromise between interpretability and capacity, struggling to reconcile statistical rigor with the expressiveness of deep learning. To resolve these limitations, we propose a Causal-Enhanced, Regime-Adaptive Neuro-Symbolic Framework. Departing from monolithic black-box approaches, our architecture imposes a disciplined causal information bottleneck: a consensus-based structural discovery mechanism first filters spurious correlations to control the False Discovery Rate (FDR). Subsequently, a continuous regime-embedding network dynamically modulates factor attention, enabling smooth adaptation to evolving market distributions. Finally, a calibrated LLM-numeric fusion head aligns the semantic reasoning of Large Language Models with rigorous regression targets, rectifying modality-specific miscalibration. Extensive out-of-sample evaluations on U.S. equities (1926-2024) and Chinese credit bonds demonstrate that this architecture achieves superior generalization, attaining an annualized Sharpe Ratio of 1.63 and boosting out-of-sample R^2 by over 25% relative to state-of-the-art tabular foundation models. The framework exhibits robust out-of-distribution (OOD) performance, effectively mitigating downside risk during structural market turbulence. MODE: Subtractive Causal Representation via Manifold Orthogonality on Graphs Xiao Han (Beihang University); Mengyao Zhou (Academy of Mathematics and Systems Science, Chinese Academy of Sciences; University of Chinese Academy of Sciences); and Wei Wei (Beihang University) Abstract Abstract Graph out-of-distribution (OOD) generalization remains a central challenge in graph learning. Existing invariant learning methods often fail under strong feature interdependencies, where causal and trivial features become topologically entangled on collapsed latent manifolds, hindering effective disentanglement. To address this, we present MODE (Manifold Orthogonal Decoupling for Environments), which utilizes a subtractive representation learning approach. Rather than directly extracting causal features, the framework identifies and subsequently removes trivial features. Theoretically, the method employs Radon-Nikodym reweighting to achieve Conditional Barycenter Alignment of trivial features. This geometric alignment facilitates orthogonality in the tangent space between causal and trivial features, allowing for their decoupling. The causal subspace is then recovered by subtracting the orthogonal trivial subspace. Experimental results indicate that MODE performs competitively compared to established baselines across several datasets. AnyTemplate: A Zero-shot Template Matching Approach for Arbitrary-Resolution and Scale-Agnostic Images Dongjie Chen and Zhiwei Lv (University of Chinese Academy of Sciences), Bowen Liu (City university of Hong Kong), and Chengjie Fang and Jungang Xu (University of Chinese Academy of Sciences) Abstract Abstract Zero-shot Template Matching is challenging due to scale variations, aspect ratio distortions and background clutter. While foundation models like SAM3 and DINOv3 provide powerful features, cascading them often results in suboptimal performance due to geometric distortion and shallow feature interactions. To address this, we propose AnyTemplate, a novel framework unifying foundation models with deep feature interaction and geometry-aware refinement. Firstly, we introduce Aspect-Ratio-Preserving Normalization to maintain structural integrity and a learnable Cross-Attention Interaction Module to dynamically model semantic dependencies and suppress noise. Secondly, we propose a Multi-task Decoding Head capable of simultaneously predicting confidence and refining localization. Thirdly, we present DEAMPB dataset, a new multi-scale template matching benchmark. Extensive experiments show that AnyTemplate achieves SOTA performance on DEAMPB and standard benchmarks, significantly outperforming existing models in zero-shot scenarios. Tuesday Virtual Room 8 IJCNN Paper Machine Learning Methods and Applications VII Session Chair: Dan Qu (Information Engineering University), Zhikui Chen (Dalian University of Technology) CVCAM-K: A Specific Emitter Identification Method with Complex-Valued Convolutional Attention Mamba and KNN Jiajun Yu, Hao Zhang, and Dan Qu (Information Engineering University) Abstract Abstract Specific emitter identification (SEI) aims to achieve unique identification of emitters through subtle signal characteristics. Recently, Mamba-based methods have gained attention for selective state-space modeling, which achieves linear complexity—improving efficiency over Transformer's square complexity in long signal processing. However, Mamba primarily models temporal correlations and lacks cross-dimensional selectivity for multi-channel features from Complex-Valued Neural Network (CVNN), limiting feature representation. Moreover, existing methods mostly use trained parametric models for classification, ignoring intrinsic data correlations. To address these issues, this paper proposes the CVCAM-K method. The method constructs the Complex-Valued Convolutional Attention Mamba (CVCAM) model, introducing the Convolutional Block Attention Module (CBAM) into CVNN to enhance the signals fed to Mamba. Further, the method innovatively integrates the K-nearest neighbor (KNN) algorithm to make decisions based on both parametric models and data correlations. Experiments on an open-source automatic dependent surveillance-broadcast (ADS-B) dataset for 10-class recognition show that our method achieves 85.3% average accuracy using only 10% of training set, outperforming mainstream approaches. With full training set, accuracy reaches 99.3%, matching the best benchmarks. Physio-xLSTM: Unraveling Hemodynamic Dynamics via Dual-Domain KAN-xLSTM Networks for Robust PPG Biometrics Yanchao Xiao and Chunxiao Wang (Qilu University of Technology (Shandong Academy of Sciences)); Yuwen Huang (Heze University); Yue Zheng (Linyi University); and Wenhao Li, Hao Wang, Wenzhe Zhang, Ziqiang Liu, and Zhiwei Zhou (Qilu University of Technology (Shandong Academy of Sciences)) Abstract Abstract Photoplethysmography (PPG) offers a secure and non-invasive modality for wearable biometrics. However, robust identification in the wild is hindered by a fundamental theoretical conflict: the non-linear generative dynamics of hemodynamic signals versus the linear modeling constraints of existing efficient architectures. While State Space Models (SSMs) and Transformers have advanced long-range dependency modeling, they exhibit critical limitations: standard SSMs (like Mamba) compress history into fixed-size vector states, creating an information bottleneck that "averages out" fine-grained morphological details. Conversely, standard Multi-Layer Perceptrons (MLPs) rely on piecewise linear activations, which fail to efficiently approximate the continuous, smooth manifolds of vascular elasticity. To resolve these bottlenecks, we propose Physio-xLSTM, a physics-informed framework that synergizes the Matrix Memory of Extended LSTM (xLSTM) with the functional approximation power of Kolmogorov-Arnold Networks (KANs). Our approach is founded on three innovations: (1) Biomechanical Phase Space Reconstruction: Instead of raw waveform matching, we embed the signal's velocity and acceleration derivatives, extracting topological invariants that remain stable despite temporal warping. (2) High-Fidelity Matrix Memory: We leverage xLSTM's matrix-based state to store high-order pairwise correlations, effectively solving the resolution loss problem of vector-based models. (3) Vascular Compliance Modeling: We replace fixed linear layers with learnable B-splines (KANs), which mathematically mirror the non-linear stress-strain relationship of arterial walls. Extensive experiments on four public benchmarks (CAPNOBASE, BIDMC, PPG-DaLiA, and PTTP-PPG) demonstrate that Physio-xLSTM achieves state-of-the-art performance, surpassing recent Mamba and Transformer baselines in motion-corrupted scenarios. A Dynamic-Length Single-Step Sampling Network for Adaptive Inference in Efficient Radio Frequency Fingerprint Identification Kaicong Yu (Academy of Military Science), Qiaoyun Sun (Beijing City University), Yi Fang (Laboratory of Electromagnetic Space Cognition and Intelligent Control), Kai Xie (Academy of Military Science), and Jian Yang (Laboratory of Electromagnetic Space Cognition and Intelligent Control) Abstract Abstract Radio Frequency Fingerprint Identification (RFFI) has emerged as a promising physical layer authentication tech- nique for Internet of Things (IoT) security. While deep learning- based methods have achieved impressive accuracy, their high computational costs pose significant challenges for deployment on resource-constrained platforms. To address this limitation, we propose a Dynamic-Length Single-Step Sampling Network (DSSN) that adaptively allocates computational resources based on sample difficulty. Our method employs a lightweight policy network to process and extract signal features in a single for- ward pass and compute importance scores for signal fragments. Through an adaptive threshold mechanism, we achieve adaptive variation in fragment selection, dynamically determining the optimal number of fragments for each sample. To better utilize information from discarded fragments, we introduce a fragment condensing module that employs a one-way nearest neighbor algorithm to merge pruned fragments into retained ones through weighted aggregation. The policy network is trained via reinforce- ment learning with a composite reward function incorporating contrastive reward, cost penalty, and accuracy reward. It achieves 81.36% accuracy with 1.42× speedup on ADS-B and 93.71% accuracy with 1.60× speedup on LTE-RFF, surpassing state-of- the-art lightweight methods. RCDIB: Robust CLIP Distillation via Information Bottleneck Zhikui Chen, Yilong Lin, Yuzhe Li, Xingheng Wan, Zhengyang Tang, and Meng Liu (Dalian University of Technology) Abstract Abstract Recently, distillation methods are proposed to balance performance and parameters of large-scale contrastive vision-language models, which gains encouraging progress. However, these distillation methods primarily align the features between the student network and the teacher network with a simple projection layer. This approach is prone to overfitting when fine-tuning on small datasets, leading to degraded distillation performance. To address the limitation discussed above, this paper proposes a robust CLIP distillation with information bottleneck (RCDIB) via cascading a contrastive distillation module and an information bottleneck distillation module. RCDIB explicitly constrains the student network to learn a compact and smooth representation of the teacher’s output by imposing an approximate distribution from the true posterior over the latent space, effectively filtering out task-irrelevant noise and enhancing distillation robustness. Extensive experiments demonstrate that RCDIB achieves cutting-edge performance on both cross modal retrieval and zero-shot classification scenarios. The code will be available upon acceptance. Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V10 Session Chair: Jing-Yu JI (Hong Kong) Constrained Optimal Pulse-Width Modulation Using Dual-Population Differential Evolution Shengqi Gui, Zusheng Tan, Jing-Yu Ji, and Billy Chiu (Lingnan University) Abstract Abstract Synchronous optimal pulse-width modulation (PWM) is an effective means of improving multilevel inverter operation, thereby supporting steady and uninterrupted power delivery in consumer-electronics applications. The associated design task is inherently difficult because PWM optimization is commonly formulated with coupled nonlinear equality and inequality constraints, under which many classical optimizers become unreliable. In this paper, we develop a cooperative constrained evolutionary algorithm with two interacting populations. The first population is primarily guided by a constraint-handling technique to rapidly approach feasibility, while the second population coevolves to preserve search diversity and assist exploration. After the feasible region is sufficiently identified, the first population shifts its emphasis toward objective minimization and adopts an alternative constraint-handling strategy. The two populations are linked through an elite-offspring selection mechanism so that information exchange promotes coordinated progress. The proposed method is assessed in two stages. We first validate both overall performance and the benefit of the cooperative mechanism on 18 standard constrained benchmarks. We then apply the algorithm to practical synchronous optimal PWM design for multilevel inverters with different numbers of voltage levels. Statistical comparisons show that the proposed approach consistently outperforms several state-of-the-art algorithms in producing high-quality PWM control solutions. Evolutionary computation based parameter estimation algorithm for state space models Shrey Verma and Ankush Sharma (Indian Institute of Technology Kanpur) and Binh Tran and Damminda Alahakoon (Latrobe University Melbourne) Abstract Abstract State-space models (SSMs) provide a compact and interpretable framework for representing dynamic systems governed by differential equations. However, as system complexity increases, particularly in single-input, multiple-output (SIMO) configurations, traditional parameter estimation methods face challenges due to the denser state and parameter matrices. These issues often lead to reduced accuracy and limited scalability. This paper presents an evolutionary state space parameter estimator (ESSPE) based on differential evolution (DE) to address these limitations. The proposed method estimates the parameters of canonical SSMs directly from input–output data without requiring prior assumptions about the model. A new fitness function is introduced to simultaneously solve multiple coupled differential equations within the SSM framework. Simulation studies on systems with 6 and 12 parameters demonstrate that ESSPE achieves estimation errors below 0.1\% while maintaining robustness across varying data lengths. The results confirm that ESSPE provides improved accuracy, scalability, and interpretability compared to conventional recursive and iterative approaches, making it a reliable tool for parameter identification in complex dynamic systems. Optimization Methods for Generating Saliency Masks in Image Classifiers Rheidner Achiles Couto Silva Fernandes, Leonardo Nogueira Matos, André Britto, Marcelo Henrique Lima Barreto, and Flávio Arthur Oliveira Santos (Fundação Universidade Federal de Sergipe) Abstract Abstract The interpretability of artificial intelligence models is fundamental for confidence in their decisions, which is essential for the acceptance and adoption of the technology. In this context, this study investigates the performance of optimization-based methods to generate saliency masks in image classifiers, compared to approaches that use gradients. Several agnostics algorithms are proposed incorporating three main hypotheses: use of continuous pixel regions, application of inpainting techniques, and multi-objective optimization. Empirical results demonstrate that optimization-based methods generally outperform gradient approaches and that the three hypotheses, with emphasis on continuous regions, bring new potential to this line of research. A Surrogate-Assisted Differential Evolution Using Fuzzy Logic for Expensive Constrained Optimization Shengqi Gui, Zusheng Tan, and Jing-Yu Ji (Lingnan University) and Sanyou Zeng (China University of Geosciences, Wuhan) Abstract Abstract In recent years, surrogate-assisted evolutionary algorithms have demonstrated strong capability for solving expensive constrained optimization problems. Although much of the existing literature concentrates on cases dominated by inequality constraints, expensive optimization with equality constraints remains comparatively underexplored, despite equality constraints being common in classical constrained formulations. To address this issue, this paper develops a new framework that couples a multilayer perceptron regression surrogate with a rule-based fuzzy inference mechanism and differential evolution (DE) for efficient equality-constrained search under limited evaluation budgets. Specifically, we design a type-1 Mamdani fuzzy inference system to realize a self-adaptive penalty strategy without prescribing an explicit functional form for the penalty mapping. The MLP surrogate and the fuzzy rule-based penalty function are embedded into the DE process to jointly steer the population toward feasibility. With surrogate guidance, the algorithm can allocate more search effort to exploring and refining promising feasible regions. Experimental studies further verify the effectiveness and practical potential of the proposed surrogate-assisted evolutionary approach for challenging expensive equality-constrained optimization tasks. Tuesday 0.01 London FUZZ-IEEE Paper FUZZ 4: Hybrid systems of computational intelligence techniques Session Chair: Chang-Shing Lee (National University of Tainan) FuzzyKG: A Neuro-Symbolic Fuzzy Neural Network for Symbolic Knowledge Extraction Paulo Vitor de Campos Souza and Liah Rosenfeld (NOVA IMS) Abstract Abstract This paper introduces FuzzyKG, a neuro-symbolic framework for extracting interpretable symbolic knowledge from fuzzy neural networks. FuzzyKG transforms fuzzy rules learned by Gaussian membership functions and logical neurons into human-readable fuzzy axioms and organizes them into a structured Knowledge Graph (KG), explicitly representing relationships between features, linguistic terms, rule antecedents, and consequents. The framework preserves the numerical semantics of the underlying fuzzy model—such as feature relevance, rule strength, and class polarity—while exposing them in a symbolic, graph-based representation. Experimental results on benchmark datasets show that FuzzyKG maintains competitive predictive performance while producing coherent axioms and meaningful graph structures, enabling transparent analysis and supporting future neuro-symbolic extensions. Fuzzy Co-Clustering and Kernel Fuzzy Co-Clustering for Interval-Valued Data José Nataniel Andrade de Sá (Centro de Informática - Universidade Federal de Pernambuco), Marcelo Rodrigo Portela Ferreira (Departamento de Estatística - Universidade Federal da Paraíba), and Francisco de Assis Tenório de Carvalho (Centro de Informática - Universidade Federal de Pernambuco) Abstract Abstract Interval-valued data are a type of Symbolic Data, a field of statistics and artificial intelligence that deals with multivalued data representations. Such representations arise from the need to preserve information about data variability during data aggregation processes, such as aggregating a continuous feature according to a categorical attribute that often uses summary statistics like the mean or median. This paper proposes co-clustering algorithms for interval valued data. These algorithms simultaneously group samples and features and have demonstrated competitive performance in several applications. We introduced one variant without kernel functions and two variants employing the Gaussian kernel, making them suitable for grouping nonlinear clusters. Experiments conducted on synthetic and benchmark interval valued datasets demonstrated the effectiveness of the proposed algorithms, particularly the kernel-based variants. Progressive Training of Generative Diffusion Models with Fuzzy Growth for Higher Fidelity Michał Wieczorek, Alicja Polowczyk, Agnieszka Polowczyk, and Jakub Silka (Silesian University of Technology) Abstract Abstract Training diffusion models for high-resolution image generation remains computationally intensive and data-demanding due to the quadratic scaling of complexity with image dimensions. We present a progressive training framework that systematically increases spatial resolution during training with fuzzy based growth module, enabling efficient synthesis of high-quality images while significantly reducing computational overhead. Our approach employs a conditional U-Net architecture operating in a latent diffusion space, augmented with resolution-specific modules that are gradually activated through smooth fuzzy transition mechanisms. The model initiates training at low resolutions to establish structural priors, then progressively grows to higher resolutions while transferring learned representations through weight inheritance and controlled blending based on fuzzy logic gating. This progressive strategy reduces total training time by up to 60% compared to direct high-resolution training while maintaining comparable loss scores on the same dataset. Furthermore, the method demonstrates improved data efficiency by leveraging low-resolution pretraining as a regularizer for high-frequency detail learning. We provide comprehensive experimental validation on resolutions from 8x8 to 256x256 pixels, along with theoretical analysis of the progressive learning dynamics. Our work offers a practical pathway for scalable high-resolution image generation with diffusion models, balancing computational efficiency with sample quality. Fuzzy Reliability Redundancy Allocation Problem Using Immigrants-based Multifactorial Evolutionary Algorithm Md. Abdul Malek Chowdury (South Asian University); Rahul Nath (University of Bergen); Amit K. Shukla (University of Vaasa, South Asian University); and Amit Rauniyar and Pranab K. Muhuri (South Asian University) Abstract Abstract High system reliability is essential for failure-free operation of engineering systems. The Reliability Redundancy Allocation Problem (RRAP) aims to maximize system reliability through optimal component selection and redundancy allocation under resource constraints. As an NP-hard problem, RRAP is typically solved using metaheuristics, which often suffer from premature convergence, especially in constrained environments. Additionally, uncertainty in system parameters motivates the use of fuzzy RRAP (FRRAP) formulations. To address these challenges, this paper proposes an immigrant-based multifactorial evolutionary algorithm (I-MFEA) with elitist, random, and hybrid immigrant strategies for solving FRRAPs. The approach incorporates a penalty-based constraint-handling mechanism and is evaluated on multitasking fuzzy RRAP benchmarks, which includes series and complex bridge systems. Taguchi method is used for the parameter tuning. The performance is assessed using average, best, and standard deviation of reliability, and compared with baseline MFEA. Further, TOPSIS based ranking method is utilized to rank the approaches based on their achieved reliability. The experimental results demonstrate that the elitist I-MFEA consistently achieves superior performance, highlighting the effectiveness of the proposed framework. Fuzzy Admission Control Using Non-Intrusive QoS Inference for Resource-Constrained Environments Abdullah Muslim, Ali Beiti Aydenlou, and Stephan Recker (University of Applied Sciences and Arts Dortmund) Abstract Abstract Edge and fog nodes operate under limited resources and often host multiple workloads simultaneously, making it difficult to maintain application quality of service (QoS) under resource stress. Admission control is therefore essential, especially for black-box workloads where application-level metrics are unavailable. However, many existing approaches rely on such metrics or fixed resource thresholds. Fuzzy Inference–Driven Knowledge Modeling for LLM-Assisted Human–Machine Co-Learning Chang-Shing Lee, Mei-Hui Wang, Chao-Cyuan Yue, and Sheng-Chi Yang (National University of Tainan) and Yusuke Nojima, Naoki Masuyama, Takeru Konishi, and Ryosuke Saga (Osaka Metropolitan University) Abstract Abstract This paper introduces a fuzzy inference–driven knowledge modeling framework for evaluating LLM-assisted human–machine co-learning performance in multilingual speech-based learning environments. The proposed framework integrates Taiwanese-Japanese Open AI Whisper–based Automatic Speech Recognition (ASR), fuzzy Large Language Model (LLM)–based translation using the Trustworthy AI Dialogue Engine (TAIDE), and a personalized fine-tuned Meta AI Massively Multilingual Speech (MMS)-based Taiwanese Text-To-Speech (TTS) model. The framework has been successfully deployed in real-world human-machine co-learning scenarios in Taiwan and Japan. Based on the outputs of these components, three fuzzy evaluation variables are derived, including fuzzy ASR similarity (ASRsim), fuzzy LLM translation similarity (LLMsim), and Mean Opinion Score (MOS) of the fuzzy Taiwanese TTS model (TTSmos). These fuzzy variables are further utilized to construct a knowledge model and a quantum fuzzy inference model for assessing human–machine co-learning performance and generating a simulated quantum circuit. To evaluate the performance of the Taiwanese Whisper–based ASR model, the complete text of the Heart Sutra, together with selected passages from the Diamond Sutra and the Tao Te Ching, is used as benchmark data across different training checkpoints. In particular, the complete text of the Heart Sutra serves as the primary benchmark for evaluating the performance of the fine-tuned OpenAI Whisper Taiwanese-Japanese ASR model. Experimental results demonstrate that the proposed framework effectively supports students’ learning processes and is capable of accurately recognizing Taiwanese-Japanese speech and mapping it to corresponding English and Chinese texts. These findings indicate that the proposed fuzzy inference–driven framework provides a practical and effective approach for multilingual human–machine co-learning evaluation. Tuesday 0.04 Brussels IJCNN Paper IJCNN SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models Session Chair: Thorben Markmann (Bielefeld University), Valerie Vaquet (Bielefeld University) A Fuzzy Pause-Aware Framework for Alzheimer’s Disease Detection from Spontaneous Speech Rishabh . and Kuldeep Singh (University of Delhi), Dhirendra Kumar (Delhi Technological University), and Yogendra Meena (Jawaharlal Nehru Univeristy) Abstract Abstract Early detection of Alzheimer’s disease via speechbased biomarkers offers a scalable, non-invasive alternative to traditional diagnostics. Prior studies have examined semantic or prosodic features separately, but the unified integration of pause cues to the semantic representation remain under-explored. The paper introduces a fuzzy pause-aware framework for the detection of Alzheimer’s disease (AD) using spontaneous speech. The suggested technique improves language representations by directly modeling pause dynamics taken from word-level transcriptions based on Whisper. Using fuzzy C-means clustering, pause durations are sorted into groups that make sense semantically. The resultant pause information is added using pause tokens and fuzzy membership-based pause embeddings. These are then combined with frozen BERT-based lexical embeddings to capture both semantic and prosodic signals associated to cognitive disfluency. A Transformer-based model with attention pooling and pause statistics fusion is then used to arrange the merged representations. Experiments conducted on the ADReSSo dataset show that the suggested method yields an accuracy of 77.46%, which is better than several state-of-the-art baselines. The findings validate that adaptive, data-driven modeling of pause patterns substantially improves the discriminative efficacy of speech-based Alzheimer’s disease detection systems. This work underscores the importance of integrating prosodic information with linguistic representations to maintain a robust cognitive. Hmsanet: a Hierarchical Multi-scale Attention Network for Precise Retinal Layer Segmentation in Covid-19 Oct Images RADHWAN ALI ABDULGHANI SALEH and Lorena Álvarez-Rodríguez (University of A Coruña), Beatriz Cordon and Elena Garcia-Martin (Miguel Servet University Hospital), Erchan Aptoula (Sabanci University), and Joaquim de Moura and Marcos Ortega (University of A Coruña) Abstract Abstract The COVID-19 pandemic has underscored the systemic impact of the virus, including retinal abnormalities that may affect vision. Precise segmentation of retinal layers in optical coherence tomography (OCT) images is essential for reliable diagnosis and monitoring. We present the Hierarchical Multi-Scale Axial Attention Network (HMSANet), a novel deep learning architecture specifically designed for retinal layer segmentation in COVID-19 OCT data. HMSANet integrates two main modules: the Multi-Path Inception Attention Module (MPIAM) for multi-scale feature extraction, and the Multi-Scale Attention Module (MSAM) for adaptive attention weighting across spatial scales. This design enables the model to capture both fine-grained and large-scale retinal changes associated with COVID-19. Experimental results show that HMSANet consistently outperforms U-Net and SCAttNet by at least 1.2% in several retinal layers, demonstrating improved sensitivity to subtle structural alterations. By automating retinal layer segmentation, HMSANet enhances diagnostic speed and consistency, with potential applicability to other retinal diseases. Stride-Net: Fairness-Aware Disentangled Representation Learning for Chest X-Ray Diagnosis Darakshan Rashid (Indian Institute of Technology Delhi), Raza Imam (Mohamed bin Zayed University of Artificial Intelligence), Dwarikanath Mahapatra (Khalifa University), and Brejesh Lall (Indian Institute of Technology Delhi) Abstract Abstract Deep neural networks for chest X-ray classification achieve strong average performance, yet often underperform for specific demographic subgroups, raising critical concerns about clinical safety and equity. Existing debiasing methods frequently yield inconsistent improvements across datasets or attain fairness by degrading overall diagnostic utility, treating fairness as a post hoc constraint rather than a property of the learned representation. In this work, we propose Stride-Net (Sensitive aTtribute Resilient learning vIa Disentanglement and learnable masking with Embedding alignment), a fairness-aware framework that learns disease-discriminative yet demographically invariant representations for chest X-ray analysis. Stride-Net operates at the patch level, using a learnable stride-based mask to select label-aligned image regions while suppressing sensitive attribute information through adversarial confusion loss. To anchor representations in clinical semantics and discourage shortcut learning, we further enforce semantic alignment between image features and BioBERT-based disease label embeddings via Group-Optimal Transport. We evaluate Stride-Net on the MIMIC-CXR and CheXpert benchmarks across race and intersectional race–gender subgroups. Across architectures including ResNet and Vision Transformers, Stride-Net consistently improves fairness metrics while matching or exceeding baseline accuracy, achieving a more favorable accuracy–fairness trade-off than prior debiasing approaches. Controlling Attention Dynamics for Factual Grounding in Neural Text Generation Roseline Mary Rozario, Philip O. Ogunbona, Khin Than Win, and Jie Yang (University of Wollongong) Abstract Abstract Attention mechanisms govern internal information routing in neural encoder-decoder architectures, yet their dynamics are typically optimized only implicitly under likelihood-based objectives. In this work, we argue that unconstrained attention dynamics contribute to factual errors in neural text generation and that attention should be treated as a controllable architectural mechanism rather than as an emergent alignment artifact. We propose Attention Bias Augmentation (ABA) , a lightweight intervention that injects a bounded, head-specific bias into the decoder cross-attention logits to explicitly steer information flow toward factually salient source content. The bias is derived from a rule-based token-level salience signal, enabling domain-informed guidance without introducing additional supervision. ABA operates entirely within the attention computation and preserves architectural complexity by avoiding auxiliary objectives and prediction heads while introducing only a minimal set of per-head scaling parameters. We evaluate ABA on echocardiography report summarization using the MIMIC-III dataset in a controlled setting and observe consistent improvements in entailment-based factuality metrics, including SummaC, while maintaining fluency and content coverage. These results demonstrate that the controlled modulation of attention dynamics provides an effective and scalable mechanism for improving factual grounding in neural text generation. RG-DermNet: A Multimodal Attention-Based Model with Residual Block Usage for Skin Lesion Classification Wyctor Rocha and Pedro Bouzon (Federal University of Espı́rito Santo), Lucas Ramos (NHL Stenden University of Applied Sciences), and André Pacheco and Luis Souza Jr. (Federal University of Espı́rito Santo) Abstract Abstract Skin cancer accounts for nearly one-third of all diagnosed tumors worldwide, making early and accurate recognition critical for improving patient outcomes. In this work, we propose RG-DermNet, a multimodal deep learning framework that integrates skin lesion images with structured clinical metadata through a residual gated-attention (RG-ATT) fusion mechanism. The architecture combines CNN- and Transformer-based visual backbones with a lightweight one-hot encoding pipeline for metadata, enabling effective cross-modal interaction. The proposed model is evaluated using a patient-wise cross-validation protocol across four dermatological datasets with heterogeneous metadata. On PAD-UFES-20, using Caformer-B36 as the visual backbone, RG-DermNet achieves an accuracy of 0.75 ± 0.05, balanced accuracy of 0.78 ± 0.03, F1-score of 0.77 ± 0.04, and AUC of 0.95 ± 0.01, outperforming existing multimodal baselines under the same evaluation setting. In addition, a SHAP-based analysis provides insights into the contribution of clinical metadata to the model’s predictions, supporting both performance gains and interpretability. Explainable Vision Transformers for Candida Species Identification from Brightfield Microscopy Rodrigo Sá, Bernardete Ribeiro, and Luís Torres (Department of Informatics Engineering of the University of Coimbra) and Catarina Pimentel (Yeast Molecular Biology Lab NOVA) Abstract Abstract The high mortality of invasive fungal infections necessitates rapid, automated differentiation between Candida albicans and Candida glabrata to overcome the bottlenecks of manual microscopy. We propose a classification framework using the SimpleViT architecture, trained on a novel dataset of 3,731 brightfield microscopy patches derived from 60 slide preparations. To prevent data leakage, we enforced a strict biological partition between training and testing replicates. Results show that aggregating predictions from the patch level (84.39%) to the slide level significantly improves decision stability, achieving 92.86% balanced accuracy on a held-out test set. A rigorous explainable (XAI) analysis using Uniform Manifold Approximation and Projection (UMAP) and Attention Rollout exposes two diverging error regimes. In contrast to the entropic signal loss characterizing false negatives, the predominant false positives exhibit the intense attentional focality of correct predictions. This indicates the model suffers not from stochastic noise, but from confident misidentification of ambiguous yeast-like textures. We conclude that while global attention effectively localizes cellular targets, resolving fine-grained phenotypic mimicry requires architectural enhancements for stronger local feature modeling. Tuesday 0.05 Paris IJCNN Paper IJCNN SS19 Advances in Trustworthy XAI: Novel Methodologies, Benchmarking, and Diverse Data Modality Contexts Session Chair: Imen Jdey (REGIM Lab, Sfax University), M. Tanveer (Indian Institute of Technology Indore, India) Explanations Leak: Membership Inference with Differential Privacy and Active Learning Defense Fatima Ezzeddine (Universita della Svizzera italiana, Scuola universitaria professionale della Svizzera italiana); Osama Zammar (Lebanese University); and Silvia Giordano and Omran Ayoub (Scuola universitaria professionale della Svizzera italiana) Abstract Abstract Machine Learning as a Service (MLaaS) has become increasingly central to modern AI pipelines by simplifying the training and deployment of machine learning models. This convenience, however, amplifies privacy attacks such as membership inference attacks (MIAs), which aim to determine whether a specific data record was part of a model's training set, compromising data privacy. In parallel, MLaaS providers increasingly expose Explainable AI outputs, most notably counterfactual explanations (CFs), to improve transparency and accountability. Because CFs can implicitly encode properties of the training distribution and the learned decision function, their release may inadvertently increase vulnerability to MIAs. In this paper, we investigate the shadow–based MIA that leverages CFs as the attacker's primary source of knowledge of an attacker. Unlike prior approaches that rely on thresholding rules, our method trains shadow models directly on CF instances paired with the target model's observed outputs, enabling the adversary to faithfully approximate the target model's local decision boundary and, in turn, improve MIA. To counter this emerging attack surface, we introduce a defense framework that combines Active Learning with Differential Privacy to prevent memorization and minimize the number of training points exposed. Such a setup creates a three-way tension between privacy, performance, and explainability. Across two datasets, we systematically evaluate this trade-off, quantifying how AL and DP interact and how their joint deployment affects MIA, model utility, and CF quality. BiasIG: Benchmarking Multi-dimensional Social Biases in Text-to-Image Models Hanjun Luo and Zhimu Huang (New York University Abu Dhabi); Haoyu Huang, Ziye Deng, and Ruizhe Chen (Zhejiang University); Xinfeng Li (Nanyang Technology University); Zuozhu Liu (Zhejiang University); and Hanan Salam (New York University Abu Dhabi) Abstract Abstract Text-to-Image (T2I) generative models have revolutionized content creation, yet they inherently risk amplifying societal biases. While sociological research provides systematic classifications of bias, existing T2I benchmarks largely conflate these nuances or focus narrowly on occupational stereotypes, leaving the multi-dimensional nature of generative bias inadequately measured. In this paper, we introduce BiasIG, a unified benchmark that quantifies social biases across a curated dataset of 47,040 prompts. Grounded in sociological and machine ethics frameworks, BiasIG disentangles biases across 4 dimensions to enable fine-grained diagnosis. To facilitate scalable and reliable evaluation, we propose a fully automated pipeline powered by a fine-tuned multi-modal large language model, achieving high alignment accuracy comparable to human experts. Extensive experiments on 8 T2I models and 3 debiasing methods not only validate BiasIG as a robust diagnostic tool, but also reveal critical insights: interventions on protected attributes often trigger unintended confounding effects on unrelated demographics, and debiasing methods exhibit a persistent tendency toward discrimination rather than mere ignorance. Our work advocates for a precise, taxonomy-driven approach to fairness in AIGC, providing a theoretical framework for using BiasIG's metrics as feedback signals in future closed-loop mitigation. The benchmark is openly available at https://github.com/Astarojth/BiasIG. Automated Screening of Cognitive Skills Impairment via Handwriting Kinematics Yasir Hussain (Brno University) Abstract Abstract The detection of cognitive impairment has traditionally relied on manual neuropsychological tests that are time-consuming, subjective, and resource-intensive. To address this limitation, we propose an automated framework that leverages handwriting kinematics to identify domain-specific cognitive impairments. The method extracts stroke-level features such as speed, acceleration, pressure, jerk, and temporal dynamics from digitized handwriting samples and maps them to cognitive domains using independent XGBoost classifiers. For evaluation, we used the publicly available DARWIN data set of 175 participants, applying 5-fold stratified cross-validation to distinguish healthy from impaired subjects in seven domains: Fine Motor Coordination, Sensory Integration, Processing Speed, Visual Planning, Executive Function & Sequencing, Working Memory load, and Attention & Inhibition. Experimental results demonstrate AUC scores ranging from 0.75 to 0.91, with the best performance in Sensorimotor Integration (0.91) and the lowest in Visuospatial Planning (0.75). This study contributes to a scalable and interpretable handwriting-based assessment framework that explicitly links digital biomarkers to cognitive domains, offering a promising tool for early and domain-specific cognitive screening. Evidential Concept Bottleneck Models Haifei Zhang (UMR-CNRS 5516 Laboratoire Hubert Curien, Université Jean Monnet Saint-Étienne); Weixuan Xiao (College of Computer Science, Shandong Xiehe University); Benjamin Quost (UMR-CNRS 7253 Heudiasyc, Université de Technologie de Compiègne); and Marie-Hélène Masson (UMR-CNRS 7253 Heudiasyc, Université de Technologie de Compiègne; IUT de l’Oise, Université de Picardie Jules Verne) Abstract Abstract Concept Bottleneck Models (CBMs) are attractive for interpretable image classification, but they often (i) ignore concept uncertainty at the label prediction stage and (ii) suffer when jointly training for concepts and labels, since label accuracy can improve at the expense of concept fidelity. We propose an Evidential Concept Bottleneck Model (EVCBM) that formalizes the concept-to-label phase within the Dempster-Shafer framework. Each concept is mapped to a learnable concept-for-class mass function, discounted according to its calibrated reliability, so that uncertain concepts contribute to ignorance rather than sharp class support. Discounted masses are pooled with a tunable linear combination of Dempster and Yager rules, and converted to class probabilities via the pignistic transform, along with an explicit uncertainty signal that supports abstention and risk control. Experiments show that under joint training, EVCBM achieves label accuracy comparable to jointly trained CBMs while consistently matching or surpassing sequentially trained CBMs in concept accuracy, improving the balance between predictive performance and concept fidelity. Tuesday 0.10 Sydney IJCNN Paper IJCNN SS04 Tiny Machine Learning Session Chair: Massimo Pavan (Politecnico di Milano) BioTrain: Sub-MB, Sub-50mW On-Device Fine-Tuning for Edge-AI on Biosignals Run Wang, Victor Jung, Philip Wiese, Sebastian Frey, and Giusy Spacone (ETH Zurich); Francesco Conti (University of Bologna); Alessio Burrello (Politecnico di Torino); and Luca Benini (ETH Zurich) Abstract Abstract Biosignals exhibit substantial cross-subject and cross-session variability, inducing severe domain shifts that degrade post-deployment performance for small, edge-oriented AI models. On-device adaptation is therefore essential to both preserve user privacy and ensure system reliability. However, existing sub-100 mW MCU-based wearable platforms can only support shallow or sparse adaptation schemes due to the prohibitive memory footprint and computational cost of full backpropagation. In this paper, we propose BioTrain, a framework enabling full-network fine-tuning of state-of-the-art biosignal models under milliwatt-scale power and sub-megabyte memory constraints. We validate BioTrain using both offline and on-device benchmarks on EEG and EOG datasets, covering Day-1 new-subject calibration and longitudinal adaptation to signal drift. Experimental results show that full-network fine-tuning achieves accuracy improvements of up to 35% over non-adapted baselines and outperforms last-layer update by approximately 7% during new-subject calibration. Furthermore, on the GAP9 MCU platform, BioTrain enables efficient on-device training throughput of 17 samples/s for EEG and 85 samples/s for EOG models with a power envelope below 50 mW. In addition, BioTrain's efficient memory allocator and network topology optimization enable the use of a large batch size, thereby reducing peak memory usage. For a completely on-chip backpropagation on GAP9, BioTrain reduces the memory footprint by 8.1×, from 5.4 MB to 0.67 MB, compared to conventional full-network fine-tuning using batch normalization with batch size 8. BFAMod: FPGA Acceleration of Binary Fuzzy ART for Rapid Inference on Embedded MPSoCs Jacob Schroll, Jian Liu, Niklas Melton, Sasha Petrenko, Micah Renfrow, Iwan Sandjaja, and Donald Wunsch II (Missouri University of Science and Technology) Abstract Abstract Adaptive Resonance Theory (ART) enables stable learning in non-stationary environments, but deploying ART models on embedded platforms is often limited by floating-point operations and the cost of repeatedly scanning many learned categories. We present BFAMod, a lightweight CPU--FPGA co-design that accelerates inference for Binary Fuzzy ART (BFA) on embedded MPSoCs. BFAMod streams magnitude-coded inputs and category prototypes via DMA, performs on-the-fly magnitude-to-thermometer expansion, and implements BFA matching and category selection using word-level logic and popcount accumulation with division-free cross-multiplication. An on-chip input cache reuses the current input across the full category scan to reduce host--device traffic. Implemented on a Kria KV260 platform at 100\,MHz, BFAMod achieves more than an order-of-magnitude higher end-to-end inference rate than a quad-core ARM Cortex-A53 baseline on Fashion-MNIST, HAR, and Mini Speech Commands, and exhibits predictable scaling with category count and encoded dimensionality in controlled sweeps. The small resource footprint indicates feasibility as an embedded accelerator alongside other real-time pipelines, while optional online updates can remain host-managed when enabled. Tiny CNNs or Tiny Transformers? A Systematic Analysis of Latency, Accuracy and Domain Shift Robustness in Extreme-Edge Image Classification Călin Diaconu, Luca Bompani, Alberto Dequino, Davide Nadalini, and Francesco Conti (Università di Bologna) Abstract Abstract Vision Transformers have recently surpassed convolutional networks on several benchmark tasks, but their benefits on extreme-edge IoT hardware remain unclear due to higher computational and memory demands of attention-based blocks. This paper provides a systematic comparison of tiny CNN, Transformer, and hybrid architectures for image classification under resource constraints. We evaluate models using hardware-agnostic metrics (FLOPs, parameter count, and accuracy) and hardware-specific measurements (working-memory footprint and on-device inference latency) on an 8-core RISC-V microcontroller with a 2 MiB on-chip memory budget. Accuracy is assessed in-domain and under two out-of-distribution transfer settings: a dataset shift at fixed resolution and a combined dataset-and-resolution shift. Under hardware-agnostic profiling, attention-based models achieve higher accuracy than CNNs at the same input resolution (up to +3.81% in-domain and +6.67% under dataset shift) while requiring up to 10× more operations. When the input resolution is reduced, this advantage is not preserved and CNNs outperform by up to 16.4%. On-device measurements show that, among models that fit within the memory budget, the best attention-based model (CaiT) provides 21% higher accuracy than the fastest CNN, at a 7x latency cost, while also exhibiting a more uniform layer-wise working-memory profile. Overall, the results indicate that selecting tiny models for edge IoT deployment requires hardware-aware profiling that jointly accounts for accuracy under distribution and resolution changes, working-memory behavior, and latency, with CNNs remaining competitive when latency is the primary constraint. Binary Image Classification on 2D Partitioned Linear Hybrid Cellular Automata on FPGA Naoki Sawahashi and Jose Principe (University of Florida) Abstract Abstract The cellular automaton (CA) is an ideal processor for ultra-low-power machine learning. It is based on logic operators, and its computational locality allows massively parallel computing. However, its application to machine learning has been limited due to a lack of training methodologies. In this paper, we extend a partitioned linear hybrid CA (P-LHCA) to a two-dimensional lattice to preserve the input image space geometry. We show that a 2D P-LHCA improves image classification accuracy by up to 17.8% on test sets, compared to a 1D P-LHCA with a similar parameter size. We also propose the first hardware design for a hybrid CA and benchmark it on the AMD Alveo V80 FPGA. Our prototype achieves over 300× speed-up compared to the GPU-based implementation while using less than 1 watt of power, demonstrating significant energy efficiency of non-arithmetic computation. The genetic algorithm remains a primary limitation of P-LHCA architecture, especially in 2D, because of its exponential growth in parameter size and slower training due to a higher degree of freedom in information propagation. Supernet NAS for Hybrid TinyML CNN-Vision Transformers with Mixed Precision Quantization Mikhael Djajapermana (Technical University of Munich), Daniel Mueller-Gritschneder (TU Wien), and Ulf Schlichtmann (Technical University of Munich) Abstract Abstract TinyML aims to bring intelligence to resource-constrained edge devices, with computer vision being a key application. Conventional strategies such as manual design, pruning, or uniform quantization often yield suboptimal trade-offs or large accuracy losses under strict resource budgets. While Vision Transformers (ViTs) have shown promise in large-scale scenarios, their potential remains largely unexplored in the TinyML regime, where CNN-based networks continue to dominate. In this work, we present a Neural Architecture Search (NAS) framework that jointly optimizes architectures and Mixed-Precision Quantization (MPQ) policies for hybrid CNN–ViT models in a single supernet training run. We further incorporate feature-based knowledge distillation to reduce quantization noise during the MPQ-aware supernet training. After the supernet training, we can search for a number of smaller, quantized models without retraining. Unlike prior approaches that either treat NAS and MPQ separately or focus exclusively on either CNNs or ViTs, our framework co-designs quantized hybrid CNN–ViT models specifically for TinyML. Experiments on CIFAR10 and Tiny-ImageNet demonstrate that our method discovers CNN–ViTs with superior accuracy to state-of-the-art tiny CNNs and ViTs, achieving 63.89% on Tiny-ImageNet at 1 MB and 87.48% on CIFAR10 at 75 kB. Our results demonstrate the first sub-100kB CNN–ViT models with competitive performance and efficiency, highlighting their potential for extremely constrained edge devices. TRAPTI: Time-Resolved Analysis for SRAM Banking and Power Gating Optimization in Embedded Transformer Inference Jan Klhufek (Brno University of Technology); Alberto Marchisio (New York University (NYU) Abu Dhabi, UAE); Vojtech Mrazek and Lukas Sekanina (Brno University of Technology); and Muhammad Shafique (New York University (NYU) Abu Dhabi, UAE) Abstract Abstract Transformer neural networks achieve state-of-the-art accuracy across language and vision tasks, but their deployment on embedded hardware is hindered by stringent area, latency, and energy constraints. During inference, performance and efficiency are increasingly dominated by the Key--Value (KV) cache, whose memory footprint grows with sequence length, straining on-chip memory utilization. Although existing mechanisms such as Grouped-Query Attention (GQA) reduce KV cache requirements compared to Multi-Head Attention (MHA), effectively exploiting this reduction requires understanding how on-chip memory demand evolves over time. This work presents TRAPTI, a two-stage methodology that combines cycle-level inference simulation with time-resolved analysis of on-chip memory occupancy to guide design decisions. In the first stage, the framework obtains memory occupancy traces and memory access statistics from simulation. In the second stage, the framework leverages the traces to explore banked memory organizations and power-gating configurations in an offline optimization flow. We apply this methodology to GPT-2 XL and DeepSeek-R1-Distill-Qwen-1.5B under the same accelerator configuration, enabling a direct comparison of MHA and GQA memory profiles. The analysis shows that DeepSeek-R1-Distill-Qwen-1.5B exhibits a 2.72× reduction in peak on-chip memory utilization in this setting compared to GPT-2 XL, unlocking further opportunities for power-gating optimization. Tuesday 0.11 Cape Town IJCNN Paper Causal, Probabilistic, and Uncertainty-Aware Learning Session Chair: Varun Sampath Kumar (University of Southern Denmark), Pranab K. Muhuri (South Asian University) Uncertainty Quantification in Radioactivity Level Estimation via Heteroscedastic Count Regression Arthur Roblin (ASNR, Mines Paris PSL); Santiago Velasco-Forero (Mines Paris PSL); and Jean Baccou and Grégoire Dougniaux (ASNR) Abstract Abstract In nuclear facilities, dedicated instruments are used to monitor airborne radioactivity in real-time. Current nuclear measurement analyses rely on simple statistical methods that exhibit significant limitations in atypical or low-count situations. To address these shortcomings, we propose a deep neural network approach for count regression, enabling more accurate estimation of transuranic alpha emissions. Uncertainty quantification is integrated through parametric distributional prediction combined with the theoretical guarantees of conformal prediction. In addition, we introduce a novel optimization strategy for heteroscedastic predictive distributions, which mitigates the inflated variance behavior commonly observed in existing approaches. Causal-INSIGHT: Probing Temporal Models to Extract Causal Structure Benjamin Redden, Hui Wang, and Shuyan Li (Queen's University Belfast) Abstract Abstract Understanding directed temporal interactions in multivariate time series is essential for interpreting complex dynamical systems and the predictive models trained on them. We present Causal-INSIGHT, a model-agnostic, post-hoc interpretation framework for extracting model-implied (predictor-dependent), directed, time-lagged influence structure from trained temporal predictors. Rather than inferring causal structure at the level of the data-generating process, Causal-INSIGHT analyzes how a fixed, pre-trained predictor responds to systematic, intervention-inspired input clamping applied at inference time. Granular Ball Computing-Based Large Margin Distribution Machine Guangming Lang, Tingyao Di, Jie Zhou, and Qimei Xiao (Changsha University of Science and Technology, Hunan Provincial Key Laboratory of Mathematical Modeling and Analysis in Engineering) Abstract Abstract The Large Margin Distribution Machine (LDM) and its variants have received considerable attention due to their superior generalization performance. However, they encounter significant hurdles regarding computational efficiency and noise robustness when handling complex, large-scale data. To overcome these limitations and further enhance model performance, this paper proposes the Granular Ball Large Margin Distribution Machine (GBLDM). First, we develop a dual-strategy granular ball generation mechanism designed to reconcile label consistency with geometric compactness. Specifically, a coarse-grained strategy extracts the topological backbone using high-purity ''Important Granular Balls,'' while a fine-grained strategy refines complex boundaries via ''Edge Granular Balls.'' Second, we construct an adaptive weighted optimization model integrating granular ball geometric information, effectively extending margin distribution theory into the granular ball space. By dynamically allocating weights based on the scale and type of granular balls, GBLDM explicitly optimizes global margin statistics at the granular ball level. Finally, extensive experiments on eight UCI benchmark datasets demonstrate that GBLDM significantly outperforms four baselines in both classification accuracy and training efficiency, exhibiting exceptional robustness particularly in noisy scenarios. Scalable Perturbation-Based Explanations via Tradeoff-Conditioned Smooth Masking Mehdi Naouar, Jens Rahnfeld, Yannick Vogt, and Joschka Boedecker (University of Freiburg) and Gabriel Kalweit and Maria Kalweit (University of Freiburg, Collaborative Research Institute Intelligent Oncology) Abstract Abstract Understanding and interpreting the predictions of deep image classifiers remains a fundamental challenge in explainable artificial intelligence. Among existing approaches, model-based removal methods train an auxiliary explainer to approximate a classifier’s response to masked inputs and have recently achieved remarkable performance in attribution. Despite their effectiveness, these methods rely heavily on large-scale Monte Carlo sampling, which incurs significant computational cost and severely limits scalability as the perturbation space increases. In this work, we first analyze these limitations and show that, as the perturbation resolution grows, sampling-based removal methods exhibit degraded attribution quality. Motivated by this observation, we introduce a sampling-free, perturbation-based training framework based on continuous and differentiable masking. By replacing discrete perturbations with smooth masks, the proposed framework eliminates the need for sampling while preserving a causal interpretation grounded in feature suppression. Finally, we adopt a tradeoff-conditioned formulation that enables a single model to produce explanations at multiple levels of sparsity and granularity. Through experiments on several datasets and classifier architectures, we show that the proposed training strategy achieves competitive or superior attribution faithfulness compared to strong sampling-based baselines, while dramatically reducing computational cost and enabling substantially improved scalability. Uncertainty-Aware Data Imputation Using Bayesian Network Guided Bayesian GANs Thomás de los Santos Verrijp (Maastricht University, Eindhoven University of Technology) and Chang Sun (Maastricht University) Abstract Abstract Deep generative models such as Generative Adversarial Networks show strong performance in data imputation, but typically ignore uncertainty and struggle to preserve conditional dependencies among variables. Probabilistic graphical models encode structure and uncertainty, yet scale poorly to complex, high-dimensional data. To address this gap, we propose BN-BGAN, a Bayesian Network–guided Bayesian Generative Adversarial Network for uncertainty-aware data imputation. The method learns a Bayesian network from data and integrates it as a structural prior into a Bayesian GAN, enabling stochastic imputations that respect learned inter-feature dependencies while modelling both aleatoric and epistemic uncertainty. We evaluated the method on benchmark clinical datasets under missing completely at random, missing at random, and missing not at random mechanisms. Results show that the proposed method consistently achieves lower reconstruction error, improved robustness to increasing missingness, and substantially reduced uncertainty dispersion compared with unstructured Bayesian Generative Adversarial Networks and classical imputers. Ablation studies further demonstrate that both the Bayesian Network prior and uncertainty-aware loss components are essential to stable and reliable performance. Robust Regression Trees Based on the Maximum Correntropy Criterion Goktug T. Cinar and Ryan Burt (University of Florida), Aslihan Demirkaya (University of Hartford), and Jose C. Principe (University of Florida) Abstract Abstract Standard regression trees optimize squared loss and can be sensitive to outliers and heavy-tailed noise. We propose a Maximum Correntropy Regression Tree (MCRTree), in which the Maximum Correntropy Criterion (MCC) governs both leaf estimation and split selection. Each leaf prediction is obtained by a correntropy-weighted fixed-point update, yielding a robust estimator with redescending influence. We define the corresponding node impurity and split gain, and train the tree using a global bandwidth selected from a small MAD-based grid. On synthetic regression tasks, MCRTree matches standard MSE trees under Gaussian noise and improves absolute-error metrics under heavy-tailed contamination, in some cases slightly outperforming LAD trees. On the Beijing PM2.5 dataset, MCRTree improves median error relative to CART while remaining competitive with LAD. A bandwidth sensitivity study shows stable behavior across a broad range of kernel scales. Tuesday 0.15 Washington IJCNN Paper IJCNN SS11 Engineering Trust: Ethical, Legal, and Societal Impacts of Computational Intelligence on Human Agency Session Chair: Keeley Crockett (Manchester Metropolitan University, Dalton Building), Robert G. Reynolds (Wayne State University), Tayo Obafemi-Ajayi (Missouri State University, Missouri, USA) Bridging the Accountability and Responsibility Gap: The Voices of SMEs and Local Authorities Shi Yun Ng, Keeley Crockett, Annabel Latham, and Mohammed Kaleem (Manchester Metropolitan University) Abstract Abstract Business adoption of artificial intelligence (AI) has become more significant due to the capacity to potentially enhance business productivity and operational efficiency. However, rapid advancement of AI technology also brings the question of “who is accountable when things go wrong?” While current AI legislation such as the EU AI Act and GDPR provide high-level principles, the practical interpretation of accountability in the context of AI decision making varies significantly between stakeholders. This paper presents the findings of a mixed-method study designed to capture and compare the distinct “voices” of two different groups in this ecosystem: the small-medium enterprise (SME) and the local government authority. The study facilitated two workshops, which aimed to (1) trial a prototype of the AI Accountability framework and (2) explore accountability and responsibility gaps within organizations by analyzing a recruitment workflow scenario. The results revealed that SMEs are prioritizing the integration of AI systems into their business workflows to enhance productivity; however, this focus may result in some ethical considerations being neglected. In contrast, local government authorities emphasized the need for systemic oversight, accessibility, and fairness, viewing accountability through the lens of institutional equity to safeguard public trust. The findings highlight the divergence in AI accountability approaches and reinforce the expressed need for practical tools and clear governance structures. From Participants to Co‑Researchers: How AI Literacy Builds Community Trust and Power in Participatory AI Design Keeley Crockett, Annabel Latham, Rochelle Taylor, and Kaleem Mohammed (Manchester Metropolitan University) Abstract Abstract This paper explores the use of participatory AI as a mechanism for increasing public trust in AI by involving community members in the co-creation and co-production of AI projects. It challenges the traditional top-down approach to AI development, arguing that public confidence and trust is directly influenced by the opportunity for citizens to actively shape and understand the technologies that impact their lives. We present a synthesis of research and case studies illustrating the efficacy of this approach, specifically in community-based initiatives. The results demonstrate that public participation in both the conceptualization and the design/development lifecycle of AI can lead to a measurable increase in confidence and trust, fostering a more equitable and socially responsible future for artificial intelligence. Trustworthy Large Language Models in Organisations: A Competence‑Based Alignment Framework Daniel Pascal Hefti and Tianxiang Lu (IU International University of Applied Science) Abstract Abstract Large language models (LLMs) increasingly influence strategic decision-making in organisations, yet their alignment with organisational culture, values, and goals (OCVG) remains underexplored. Existing bias mitigation techniques primarily address societal categories—such as gender, race, and religion—offering limited relevance for small and medium-sized enterprises (SMEs) seeking to embed their unique ethos into AI systems. Moreover, SMEs often lack the resources to implement complex alignment strategies. This study proposes a novel, human-centred approach: adapting methods from competence management — traditionally used to align human behaviour with organisational expectations — to guide and evaluate LLM outputs. Results show statistically significant improvements in alignment scores for treatment models across different datasets and languages (Mann–Whitney U, p<.001; Rank Biserial Correlation r = .36). Human validation of 8.73% of outputs confirmed strong agreement with automated evaluations (Pearson r = .71, p<.001), underscoring the reliability of our rubric-driven assessment. These findings demonstrate that competence-based instructions offer a scalable, interpretable, and ethically grounded method for aligning LLMs with organisational values. Artificial Agency and the Ethics-Law Divide in AI Governance: A Philosophical Reflection ALEXANDER RICHARD KRIEBITZ (Ludwig-Maximilian University, Technical University of Munich); Ali Hessami (City University London); Nell Watson (European Responsible Artificial Intelligence Office); Amanda Horzyk (University of Edinburgh); and Patricia Shaw (Beyond Reach Consulting Limited) Abstract Abstract AI governance commonly treats ethics and law as complementary normative instruments, often framed through the distinction between “soft” and “hard” law. This paper challenges that view by arguing that ethics and law constitute a fragile normative equilibrium destabilized when AI encodes macro-level societal preferences into micro-level individual moral reasoning spaces, with adverse consequences for personality rights. Tuesday 2.18 Mekong IJCNN Paper Human Activity Recognition and Motion Understanding Session Chair: Siyuan Yang (KTH Royal Institute of Technology), Gaurvi Goyal (Maastricht University) Pseudo-Rehearsal for Audio-Visual Continual Learning on Embedded Systems Jonathan Grienay (Univ. Grenoble Alpes, CEA, LIST, Grenoble, France; LEAT, Université Côte d'Azur / CNRS UMR 7248, Sophia Antipolis, France); Martial Mermillod (Univ. Grenoble Alpes, Univ. Savoie Mont Blanc, CNRS, LPNC, Grenoble, France); Laurent Rodriguez and Benoit Miramond (LEAT, Université Côte d'Azur / CNRS UMR 7248, Sophia Antipolis, France); and Marina Reyboz (Univ. Grenoble Alpes, CEA, LIST, Grenoble, France) Abstract Abstract Continual learning for multimodal audio-visual systems on embedded devices faces dual challenges: catastrophic forgetting and privacy constraints prohibiting raw data storage. We present AV-DreamNet, the first noise-based pseudo-rehearsal method for audio-visual continual learning. By generating pseudo-samples from noise rather than storing real data, our approach enables privacy-preserving learning without auxiliary generative models or large buffers. On CL-VGGSound, AV-DreamNet achieves 58.1% accuracy versus 26.3% for state-of-the-art buffer-based methods. Considering embedded deployment, working in latent space enables offloading frozen encoders to dedicated inference accelerators and avoids heavy decoder architectures, while latent-space buffers provide ~360x memory reduction compared to raw data storage. Motion-Adaptive Multi-Scale Temporal Modelling with Skeleton-Constrained Spatial Graphs for Efficient 3D Human Pose Estimation Ruochen Li, Shuang Chen, Wenke E, Farshad Arvin, and Amir Atapour-Abarghouei (Durham University) Abstract Abstract Accurate 3D human pose estimation from monocular videos requires effective modelling of complex spatial and temporal dependencies. However, existing methods often face challenges in efficiency and adaptability when modelling spatial and temporal dependencies, particularly under dense attention or fixed modelling schemes. In this work, we propose \textbf{MASC-Pose}, a Motion-Adaptive multi-scale temporal modelling framework with Skeleton-Constrained spatial graphs for efficient 3D human pose estimation. Specifically, it introduces an Adaptive Multi-scale Temporal Modelling (AMTM) module to adaptively capture heterogeneous motion dynamics at different temporal scales, together with a Skeleton-constrained Adaptive GCN (SAGCN) for joint-specific spatial interaction modelling. By jointly enabling adaptive temporal reasoning and efficient spatial aggregation, our method achieves strong accuracy with high computational efficiency. Extensive experiments on Human3.6M and MPI-INF-3DHP datasets demonstrate the effectiveness of our approach. HiRA-CAM: Preserving Fine-Grained Spatial Relevance in Gradient-Based Visual Explanations Manasi Nerurkar and Ali A. Minai (University of Cincinnati) Abstract Abstract Deep Learning models can include billions of parameters or more, making it difficult to explain their internal transformations and outputs. However, explainability is increasing in importance due to the use of AI in crucial applications. This paper focuses on the interpretability of convolutional neural networks (CNNs). Building on the popular gradient based method LayerCAM for extracting internal features in CNNs, we propose an improved method named HiRA-CAM, and show that it outperforms both LayerCAM and Grad-CAM on creating useful saliency maps for object classification. The main feature of HiRA-CAM is its adaptive use of activation maps from all the layers of the CNN to arrive at a more focused saliency map. Beyond Addition: HDC Binding for Transformers’ Position Encoding applied to Time Series Classification Kenny Schlegel (Chemnitz University of Technology); Jose J. Peña Gomez (Luleå University of Technology); Dmitri A. Rachkovskij (Luleå University of Technology, Institute of Information Technologies and Systems Ukraine); Stefan Streif (Chemnitz University of Technology); and Evgeny Osipov (Luleå University of Technology) Abstract Abstract Transformers integrate positional information by adding Positional Encodings (PE) to token embeddings, which is particularly relevant for time series classification (TSC). We revisit PE from a Hyperdimensional Computing (HDC) perspective and cast positional encoding as a representation design problem with two axes: (i) the operation used to integrate token and position information and (ii) the similarity structure of the positional vectors. We compare additive superposition to HDC binding, instantiated by Circular Convolution, and evaluate sinusoidal, random, and Fractional Power Encoding (FPE) positional vectors. Across the UEA multivariate time series classification benchmark, binding consistently improves performance over addition, and circular-convolution binding combined with similarity-aware FPE achieves the best average performance (64.43% mean accuracy vs. 57.74% for the additive baseline). These findings suggest that positional integration is a first-class design choice for time-series Transformers. VAug-CLIP: Pose-Guided Augmentation for CLIP-based Video Person Re-identification Guquan Jing, Peng Gao, and Yujian Lee (Beijing Normal-Hong Kong Baptist University, Hong Kong Baptist University) and Hui Zhang (Beijing Normal-Hong Kong Baptist University) Abstract Abstract Video-based person re-identification (Re-ID) aims to retrieve particular pedestrians from video sequences over non-overlapping cameras. While existing methods attempt to improve performance with designed networks or multi-modal fusion, they often struggle with misalignment, occlusion, and cross-modal domain gaps. To address these, we propose VAug-CLIP, a novel CLIP-based network that leverages 2D pose priors to alleviate data challenges. Our method features three key components. Before training stage, Pose-guided Augmentation that contains a Pose-based Alignment Module (PAM) rectifies misaligned frames by clustering pose features and recalculating bounding boxes, and a occlusion masking strategy that mines and applies realistic occlusion patterns from the dataset. Furthermore, we design an Augmentation Memory (AM) with a coarse-to-fine update to enhance CLIP-based sequence embeddings, complemented by a Temporal Interaction (TI) module for fine-grained temporal modeling. Extensive experiments on four benchmarks demonstrate that VAug-CLIP consistently achieves state-of-the-art performance, validating the effectiveness of our method. Tuesday 2.1 Volga IJCNN Paper IJCNN SS22 Self-organizing Clustering for Continual Learning and its Applications Session Chair: Naoki Masuyama (Osaka Metropolitan University), Yuichiro Toda (Okayama University) Adaptive Synapse Generation in Cortical Learning Algorithm with Dynamic Range Takeru Aoki (Tokyo University of Science, The University of Tokyo) and Tomoaki Tatsukawa (Tokyo University of Science) Abstract Abstract The Cortical Learning Algorithm (CLA) is a brain-inspired framework modeled after the human neocortex. By representing and predicting time-series data using sparse distributed representations, CLA achieves lower computational costs and superior adaptability to shifting trends compared to conventional artificial neural networks, making it well-suited for online time-series forecasting. However, CLA-DR (Dynamic Range), which eliminates predefined input range constraints by dynamically adding synapses, often suffers from decreased prediction accuracy due to underutilization of the allocated predictor. To address this issue, this study proposes an adaptive synapse generation method for the CLA-DR based on bias in the input data distribution. Experimental results using both synthetic and real-world datasets demonstrate that the proposed method achieves superior prediction accuracy compared to existing CLA models and a related LSTM approach by utilizing the predictor more effectively. Cluster Merging in Adaptive Resonance Theory-based Clustering Guided by Mass and Distance Criteria Shoki Inoue, Naoki Masuyama, and Yusuke Nojima (Osaka Metropolitan University); Yuichiro Toda (Okayama University); Zongying Liu (Dalian Maritime University); and Chu KiongLoo and Wei ShiungLiew (University of Malaya) Abstract Abstract There is a demand for clustering methods that can efficiently handle large-scale data. Adaptive Resonance Theory (ART)-based clustering algorithms adaptively generate prototype nodes that represent local data regions during learning. This mechanism allows ART-based methods to scale to large datasets. However, ART-based methods tend to generate an excessive number of nodes, which reduces scalability and often degrades clustering performance. This paper addresses these limitations by extracting an explicit cluster structure from the generated nodes through node grouping with a criterion that combines node importance and similarity. Experiments on 19 real-world datasets show that the proposed algorithm achieves superior clustering performance compared with state-of-the-art clustering methods. GROWL: Self-Organizing Memory with Active Semantic Consolidation for Lifelong Language Agents Weihong Chin and YUCHEN GUO (Tokyo Metropolitan University) Abstract Abstract Large Language Models (LLMs) cannot effectively update their beliefs when facts change. If a user’s preference evolves from coffee to tea, standard Retrieval-Augmented Generation (RAG) retrieves both statements, leaving the model to reconcile contradictory information. We introduce GROWL (Graph-based Retrieval Organized for Working Long-term memory), built on Anchored Prototype Networks (APN). GROWL maintains an immutable Episodic Log of all observations while organizing them into a navigable Prototype Graph. When concepts merge, their index pointers combine—not the underlying data—preserving distinct episodes for temporal re-ranking at retrieval time. A key innovation is the Hippocampal Buffer, a biologically-inspired short-term memory cache that ensures candidate inclusion of the most recent episodes regardless of graph topology. On our FluxQA benchmark, GROWL achieves 75.2% accuracy on high-velocity state tracking compared to Temporal RAG’s 69.6%, outperforming or matching Temporal RAG in 3 of 4 domains while maintaining 27.9× retrieval speedup at scale. Self-Organizing Maps with Optimized Latent Positions Seiki Ubukata, Akira Notsu, and Katsuhiro Honda (Osaka Metropolitan University) Abstract Abstract Self-Organizing Maps (SOM) are a classical method for unsupervised learning, vector quantization, and topographic mapping of high-dimensional data. However, existing SOM formulations often involve a trade-off between computational efficiency and a clearly defined optimization objective. Objective-based variants such as Soft Topographic Vector Quantization (STVQ) provide a principled formulation, but their neighborhood-coupled computations become expensive as the number of latent nodes increases. In this paper, we propose Self-Organizing Maps with Optimized Latent Positions (SOM-OLP), an objective-based topographic mapping method that introduces a continuous latent position for each data point. Starting from the neighborhood distortion of STVQ, we construct a separable surrogate local cost based on its local quadratic structure and formulate an entropy-regularized objective based on it. This yields a simple block coordinate descent scheme with closed-form updates for assignment probabilities, latent positions, and reference vectors, while guaranteeing monotonic non-increase of the objective and retaining linear per-iteration complexity in the numbers of data points and latent nodes. Experiments on a synthetic saddle manifold, scalability studies on the Digits and MNIST datasets, and 16 benchmark datasets show that SOM-OLP achieves competitive neighborhood preservation and quantization performance, favorable scalability for large numbers of latent nodes and large datasets, and the best average rank among the compared methods on the benchmark datasets. Tuesday 0.01 London IEEE CEC (Evolutionary Computation) CEC 9 - Algorithms II Session Chair: Kalyanmoy Deb (MIchigan State University) Modified GUESS Approach Using Machine Decision-makers for an Efficient Interactive Multi-criterion Decision-Making Procedure Deepanshu Yadav and Kalyanmoy Deb (Michigan State University) Abstract Abstract Interactive multi-criterion decision-making (iMCDM) procedures generally incorporate the human decision maker’s (DM’s) preferences to iteratively compute preferred Pareto-optimal (PO) solutions. DM preferences often involve objective classification indicating their improvement, relaxation, satisfaction, and indifference at the current PO solution. In addition, the DM provides one or more preference information in the form of a reference point, objective weights, and reference direction. These preferences are used to reformulate the original multi-objective optimization (MOO) problem into a single-objective optimization problem using the achievement scalarization function (ASF). The iterative preferences are then used to solve the ASF problem to compute a preferred PO solution. In this paper, we propose a modified version of a well-known iMCDM procedure: GUESS. In order to compare the modified GUESS with the original GUESS, we first propose a Machine Learning (ML)–based decision-maker (Machine-DM), which uses two Artificial Neural Networks (ANNs) to replace human DM steps in GUESS. The two trained ANNs predict the objective class and bounding parameters of objectives from the decision variable vector, without explicit information on the location and value of the preferred point, thereby enabling a fair performance evaluation. Following the Machine-DM approach, three variants of GUESS: MachDM-GUESS (original), MachDM-GUESS-I and MachDM-GUESS-II (modified), are developed and compared using newly proposed performance metrics. The implementation of Machine-DM with the original and modified GUESS procedures reveals superiority of one of the modified GUESS procedures. Moreover, the study conducted in this paper opens up new avenues for benchmarking existing iMCDM procedures and proposing new and competitive ones, paving the way and encouraging EMO and MCDM researchers to conduct more such studies. Explainable Random Forests: a Set-Based Particle Swarm Optimization approach Anje Erasmus and Andries Engelbrecht (Stellenbosch University) Abstract Abstract By extracting human-understandable rules from black-box ensemble models, it is possible to reconcile their strong predictive performance with the interpretability required for user trust. Random forests are highly accurate models, yet classify as black-box models. This study proposes a novel rule extraction technique from random forests using a set-based particle swarm optimization framework. The method constructs a universal rule set from induced random forests, then employs set-based particle swarm optimization in a separate-and-conquer approach to optimize a rule subset. This optimization is guided by a fitness function that effectively balances accuracy with interpretability criteria. Empirically validated across ten diverse datasets, the findings highlight the potential of SBPSO for rule extraction. An Evolutionary Method for Joint Optimization of UAV Path and Camera Direction Planning with Optional Viewpoints Yifeng Sheng, Ning Xue, and Jiayi Dong (University of Nottingham Ningbo China); Libin Hong (School of Information Science and Technology, Hangzhou Normal University, Hangzhou, PR China; University of Nottingham Ningbo China); and Ruibin Bai and Heshan Du (University of Nottingham Ningbo China) Abstract Abstract This paper addresses the Unmanned Aerial Vehicle (UAV) path planning problem with optional viewpoints, formulated as a joint optimization that considers surface coverage quality and flight energy consumption. The problem involves strong interdependence among viewpoint selection, path sequencing, and camera orientation, resulting in a large-scale combinatorial challenge that cannot be efficiently solved by off-the-shelf linear programming solvers. To tackle this, we develop a dual-layer chromosome genetic algorithm (GA) that encodes both UAV paths and camera directions. The GA employs adjacency-preserving edge-recombination crossover, feasibility-aware node replacement mutation, and a coverage repair operator, enabling constraint-preserving search while satisfying the sample point coverage requirement. Experiments on urban 3D scenes confirm that the GA consistently produces feasible solutions, achieving near-optimal quality on the majority of instances and solving large-scale problems intractable for Gurobi with a 94% time reduction, thereby demonstrating superior scalability and efficiency for UAV photogrammetric planning. Projected Hypercube Sampling with Parallel Reference Vectors for Uniform Coverage of Irregular Pareto-Optimal Fronts Kannan Sekar, Ankush Kapoor, Hemant Kumar Singh, and Tapabrata Ray (University of New South Wales) Abstract Abstract Evolutionary algorithms commonly utilize search along reference vectors to solve multi/many-objective optimization problems. However, it is now well known that their effectiveness strongly depends on the geometric compatibility between reference vectors and the geometry of the Pareto front (PF). Conventional simplex-based systematic sampling (SS) implicitly assumes simplex-aligned PFs, which may lead to missing intersections and poor coverage when the front is highly curved, inverted or disconnected. Adaptation of reference vectors during the search is one potential approach explored in the literature to counter this. This paper investigates an alternate approach, referred to as Projected Hypercube Sampling (PHS). It is a reference vector generation strategy intended to work across a wider range of PF shapes. It uniformly samples in a unit hypercube and projects the samples onto a reference hyperplane, producing a set of well-spaced parallel reference vectors that span the full PF-relevant region. PHS is evaluated on analytically defined convex and non-convex PFs with controllable curvature across up to ten objectives. Numerical experiments show that PHS consistently achieves better coverage and inverted generational distance compared to SS over a range of front geometries. The performance of PHS also improves monotonically with increasing reference-point density, unlike SS where the performance saturates if the shape is not aligned to the simplex. These observations demonstrate that PHS provides a reliable and scalable reference vector generation mechanism to deal with many-objective optimization problems. A Strongly Typed Genetic Programming Approach to Feature Learning in Time Series Classification Xuanhao Yang, Bing Xue, and Mengjie Zhang (Victoria University of Wellington) Abstract Abstract Learning discriminative yet interpretable features remains a challenge in time series classification (TSC). Unlike traditional methods that rely on fixed hand-crafted features or deep neural network methods that lack interpretability, this paper proposes a strongly typed genetic programming approach to learning discriminative and interpretable features in TSC. Specifically, a new program structure is introduced, comprising three layers: domain transformation, feature extraction, and feature concatenation. The layered design enforces an explicit order of operations that mimics human expert logic, while still permitting flexible compositions. This enables the evolutionary search to learn discriminative representations tailored to the specific data. Experiments on 31 UCR datasets demonstrate the effectiveness of the learned features. In addition, visualisation of the evolved programs highlights the interpretability of the resulting solutions. Multi-Strategy Enhanced Hippopotamus Optimization for Robot Path Planning Lizhen Wu and Xudong Li (National University of Defense Technology), Xiang Fang and Tingting Zhang (Hohai University), Shaofei Chen (National University of Defense Technology), Chaoyang Chen (Hunan University of Science and Technology), and Chang Wang (National University of Defense Technology) Abstract Abstract As a core technology for autonomous driving and robot navigation, robot path planning often exhibits slow convergence and a tendency to become trapped in local optima in complex obstacle environments. The Hippopotamus Optimization (HO) algorithm can be used for path planning by simulating hippopotamus herd behavior. However, it suffers from insufficient population diversity and a fixed search step size. In this paper, we propose a Multi-Strategy Enhanced Hippopotamus Optimization (MSEHO) algorithm. Firstly, a Logistic chaotic mapping initializes the population, and a similarity-based classification mechanism divides the population into elite groups that search in different directions to optimize resource allocation. Secondly, a Levy-Logistic chaotic disturbance operator is introduced in early iterations to enhance global exploration, and a dynamic step-size Brownian motion strategy is employed in later stages to improve local exploitation precision. The simulation results demonstrate that MSEHO outperforms the comparative algorithms in both convergence speed and solution diversity across 29 CEC2017 benchmark functions and three path-planning scenarios. Tuesday 0.02 Berlin IEEE CEC (Evolutionary Computation) CEC 10 - Evolutionary Machine Learning II Session Chair: Efrén Mezura-Montes (University of Veracruz) A Bi-Surrogate Scheme for Black-Box Optimization with Unrevealed Constraints Chi-En Tang, Chen Chien, and Tian-Li Yu (National Taiwan University) Abstract Abstract Surrogate-assisted evolutionary algorithms (SAEAs) are widely used on black-box optimization problems involving computationally intensive function evaluations, such as physics-based simulations that embed complex constraints. In such scenarios, the simulator usually integrates a repair mechanism that modifies the candidate solution handed by the optimizer prior to evaluation. Under these conditions, traditional SAEAs cannot accurately capture the complex constraints in the simulators since some of them are unrevealed to the EA. To address this opacity, this paper proposes a bi-surrogate architecture that explicitly models both the repair operator and the fitness evaluation. This framework characterizes the repair behavior to capture unrevealed constraint information overlooked by traditional approaches. Experiment results demonstrate that the proposed architecture significantly reduces the number of function evaluations while maintaining competitive success rates. Notably, this advantage becomes more pronounced as simulator costs increase, making the approach well-suited for computationally intensive applications. LoRA-NAS: Optimizing Spatial Adapter Placement in Large Language Models via Evolutionary Search Alejandro Rosales-Perez (CIMAT) Abstract Abstract Parameter-Efficient Fine-Tuning is essential for adapting Large Language Models to downstream tasks under limited computational resources. While Low-Rank Adaptation (LoRA) is widely adopted, conventional implementations typically use uniform placement heuristics, inserting adapters into all layers without considering task relevance. Dynamic methods such as AdaLoRA optimize adapter rank but often retain dense topologies and overlook the spatial distribution of adapters. This study introduces LoRA-NAS, a framework that employs evolutionary search to optimize the spatial allocation of adapters. Moreover, LoRA-NAS also optimizes the hyperparameters for LoRA adapters, customized by each layer. By formulating adapter placement as an optimization problem, LoRA-NAS autonomously identifies the most effective subset of layers for adaptation and hyperparameters. Extensive experiments on the GLUE benchmark with LLaMA-3-8B demonstrate that LoRA-NAS consistently surpasses both standard uniform baselines and leading dynamic-rank methods. The results indicate that selecting the placement og adaptation layers has a greater impact on generalization than adaptation capacity. PNKL-DE: A Positive-Negative Knowledge Learning Framework for Differential Evolution Kanchan Rajwar, Nelishia Pillay, and Thambo Nyathi (University of Pretoria) Abstract Abstract Knowledge-Learning Evolutionary Algorithms represent an emerging paradigm designed to transfer acquired information for more effective search the solution space. However, existing literature typically focuses on extracting and sharing only "positive" knowledge—information derived exclusively from successful search regions. In this study, we propose a novel approach that utilizes negative knowledge as well to guide the evolutionary process. We introduce Positive-Negative Knowledge-based Learning Differential Evolution (PNKL-DE), a novel algorithm that explicitly exploits search history from both successful (positive) and unsuccessful (negative) outcomes to enhance optimization performance. The proposed algorithm is rigorously evaluated on the IEEE CEC 2017 benchmark suite and compared against state-of-the-art and CEC competition winners, including JADE, SHADE, L-SHADE, jSO, NL-SHADE-RSP, and MadDE. Furthermore, PNKL-DE is compared with the recently propped evolutionary transfer learning method Knowledge Learning DE (KLDE) and Knowledge Learning PSO (KLPSO). Experimental results demonstrate that PNKL-DE achieves statistically significant improvements, validating the hypothesis that negative knowledge is a critical, underutilized resource in evolutionary computation. Evolutionary Design of Generative Adversarial Networks using Deep Evolution Hana Derouiche and Maha Elarbi (SMART Lab, Computer Science Department, ISG, University of Tunis, Tunis, Tunisia); Slim Bechikh (COL Lab, School of Computer Science, University of Nottingham); and Carlos Artemio Coello Coello (Computer Science Department, CINVESTAV-IPN, Av Instituto Politécnico Nacional 2508, San Pedro Zacatenco, Gustavo A. Madero, 07360, Mexico City, Mexico; Faculty of Excellence of the School of Engineering and Sciences, Tecnologico de Monterrey, Monterrey, Mexico) Abstract Abstract Designing effective architectures for Generative Adversarial Networks (GANs) remains a challenging task due to the strong interdependence between generator and discriminator structures, the instability of adversarial training, and the sensitivity of performance to architectural choices. Existing evolutionary Neural Architecture Search (NAS) approaches for GANs often optimize a single component or evaluate candidate architectures under a fixed generator–discriminator pairing, in spite of their distinct roles in adversarial training, which can lead to biased evaluations and poor generalizable designs. To address these limitations, we propose here an evolutionary NAS framework for GANs based on a Genetic Algorithm (GA). Each individual represents a full adversarial model and is evaluated at the GAN level using standard image quality metrics. Generator architectures are recombined using a standard crossover operator, while discriminator architectures are optimized via a Decuple Crossover Scheme (DCS), which is a recently introduced structured crossover scheme that generates multiple discriminator candidates through hierarchical recombination. Each offspring generator is paired with all discriminator candidates, and the resulting GANs are jointly evaluated to select the best-performing architectures for the next generation. This strategy enables efficient exploration of generator designs alongside deep, memetic refinement of discriminators, improving stability, robustness, and generative performance. Experimental results on standard image generation benchmarks show that our proposed framework consistently outperforms baseline GANs and existing evolutionary approaches. MODE-based Oblique Decision Tree Induction: Addressing the Accuracy-Complexity Trade-off through a Normalized Complexity Measure Ulises-Ramsés Prado-Valderrábano and Efrén Mezura-Montes (University of Veracruz); Rafael Rivera-López (TecNM, Veracruz Institute of Technology); and Nancy Pérez-Castro (University of Veracruz) Abstract Abstract Oblique Decision Trees improve predictive performance over univariate trees at the cost of increased computational complexity and reduced interpretability. While most approaches formulate their induction as a single-objective problem favoring accuracy, this work proposes a Multi-Objective Differential Evolution framework to jointly minimize classification error and tree size. A normalized complexity measure is introduced to capture the structural properties of the induced models and mitigate discontinuities in the Pareto Fronts. The approach is evaluated using three Multi-Objective Differential Evolution variants across twelve datasets, considering hypervolume, classification accuracy, and tree size as performance indicators. Results show that the proposed method achieves comparable predictive performance to a single-objective approach while significantly reducing model complexity, demonstrating the benefits of explicitly addressing the accuracy–complexity trade-off. Leader-Enhanced Optimization for Robust Neural Ensemble Learning in Medical Audio Diagnosis Ziang Zhao (Royal Holloway, University of London); Arjun Panesar (DDM Health Ltd.); Li Zhang (Royal Holloway, University of London); and Yonghong Yu (Nanjing University of Posts and Telecommunications) Abstract Abstract Medical audio analysis offers a non-invasive and scalable solution for disease screening, but its predictions are often affected by substantial uncertainty arising from inter-subject variability, recording conditions, and limited data availability. While ensemble models are commonly used to improve performance, their effectiveness critically depends on how individual models are combined, particularly under noisy evaluation settings. This paper presents a leader-enhanced evolutionary search framework for robust ensemble weight optimization in medical audio analysis. We formulate ensemble construction as a constrained swarm-based optimization problem with noisy objectives and propose four new Particle Swarm Optimization (PSO) variants for base classifier weight identification. where Genetic Algorithm (GA), Differential Evolution (DE), Simulated Annealing (SA), and Covariance Matrix Adaptation Evolution Strategy (CMA-ES) are used for swarm leader enhancement. Five audio deep networks and transformer models are adopted as base classifiers, including ResNet, RegNet, MobileViT, Whisper, and wav2vec2, with either spectrograms or waveforms as inputs to increase base model diversity. Experiments conducted on diabetes and Parkinson’s disease voice/speech datasets demonstrate that the proposed framework consistently outperforms individual base models and conventional ensemble baselines. Ablation studies further reveal that leader-enhancement strategies using GA, DE, SA and CMA-ES effectively balance exploration and stability in adaptive ensemble weight search. These empirical results indicate that evolutionary ensemble optimization provides a practical and reliable approach for robust medical audio classification. Tuesday 0.04 Brussels IJCNN Paper Neural Learning and Optimization II Session Chair: Benjamin Redden (Queen's University Belfast), Fabian Hinder (Bielefeld University) A Novel Approach to Training Deep Neural Networks Via Multiple Output Heads and Boosting Mohamed Hamdy (QatarUniversity) and Ponnuthurai N Suganthan and Abdulaziz Al-Ali (Qatar University) Abstract Abstract Deep neural networks have demonstrated strong performance across diverse learning tasks by learning hierarchical feature representations through end-to-end optimization. Nevertheless, conventional training relies on a single output head attached to the final network layer, leaving intermediate representations underutilized and limiting the potential benefits of internal supervision. Although some existing approaches attempt to exploit intermediate outputs through greedy layer-wise training or multi-head ensembles, these methods often constrain the representational capacity of deeper layers or rely on auxiliary heads that are introduced post hoc and do not actively shape the training of the backbone. In this paper, we propose a boosting-guided progressive training and freezing strategy that explicitly leverages multiple intermediate output heads within a single neural network to guide representation learning across depth. The proposed approach progressively incorporates intermediate predictors and employs boosting-inspired sample reweighting to focus the training of deeper layers/blocks on harder examples. We evaluate the proposed method on both vision and tabular classification tasks. Our results show that intermediate output heads and boosting improve the performance of convolutional and transformer-based backbones on four image classification datasets. Multilayer perceptrons, operating as an ensemble using intermediate output layers, outperform state-of-the-art methods on ten tabular classification datasets. A Dynamic Framework for Grid Adaptation in Kolmogorov–Arnold Networks Spyros Rigas, Thanasis Papaioannou, Panagiotis Trakadas, and Georgios Alexandridis (National and Kapodistrian University of Athens) Abstract Abstract Kolmogorov–Arnold Networks (KANs) have recently demonstrated promising potential in scientific machine learning, partly due to their capacity for grid adaptation during training. However, existing adaptation strategies rely solely on input data density, failing to account for the geometric complexity of the target function or metrics calculated during network training. In this work, we propose a generalized framework that treats knot allocation as a density estimation task governed by Importance Density Functions (IDFs), allowing training dynamics to determine grid resolution. We introduce a curvature-based adaptation strategy and evaluate it across synthetic function fitting, regression on a subset of the Feynman dataset and different instances of the Helmholtz PDE, demonstrating that it significantly outperforms the standard input-based baseline. Specifically, our method yields average relative error reductions of 25.3% on synthetic functions, 9.4% on the Feynman dataset, and 23.3% on the PDE benchmark. Statistical significance is confirmed via Wilcoxon signed-rank tests, establishing curvature-based adaptation as a robust and computationally efficient alternative for KAN training. On the Difficulty of Training Non-negative Neural Networks Manos Kirtas (Centre for Nanosciences and Nanotechnologies, Aristotle University Of Thessaloniki) and Nikoalos Passalis, Nikos Pleros, and Anastasios Tefas (Aristotle University of Thessaloniki) Abstract Abstract Although non-negative neural networks are equipped with intrinsic properties that can be utilized in emerging analog accelerators and interpretable models, their use is limited due to the inherent constraints of the parameters. In fact, existing work fails to generalize across a wide range of architectures, resulting in significant performance degradation compared to their regular counterparts. In this work, we study the difficulties of non-negative optimization, claiming that without the proper initialization of non-negative parameters, the model is a priori limited before even the optimization process takes place. To this end, we propose a post-initialization method that can be applied on top of any existing initialization scheme, preserving the variance of activation units and, as a result, the expressivity of the non-negative models. The proposed method can be applied directly to convolutional and recurrent architectures, with the work providing experimental results in demanding scenarios involving easily saturated neural layers. FlowAdam: Implicit Regularization via Geometry-Aware Soft Momentum Injection Devender Singh and Tarun Sheel (Memorial University of Newfoundland) Abstract Abstract Adaptive moment methods such as Adam use a diagonal, coordinate-wise preconditioner based on exponential moving averages of squared gradients. This diagonal scaling is coordinate-system dependent and can struggle with dense or rotated parameter couplings, including those in matrix factorization, tensor decomposition, and graph neural networks, because it treats each parameter independently. We introduce FlowAdam, a hybrid optimizer that augments Adam with continuous gradient-flow integration via an ordinary differential equation (ODE). When EMA-based statistics detect landscape difficulty, FlowAdam switches to clipped ODE integration. Our central contribution is Soft Momentum Injection, which blends ODE velocity with Adam's momentum during mode transitions. This prevents the training collapse observed with naive hybrid approaches. Across coupled optimization benchmarks, the ODE integration provides implicit regularization, reducing held-out error by 10-22% on low-rank matrix/tensor recovery and 6% on Jester (real-world collaborative filtering), also surpassing tuned Lion and AdaBelief, while matching Adam on well-conditioned workloads (CIFAR-10). MovieLens-100K confirms benefits arise specifically from coupled parameter interactions rather than bias estimation. Ablation studies show that soft injection is essential, as hard replacement reduces accuracy from 100% to 82.5%. NysReg-Gradient:~Regularized Nystr\"om-Gradient for Large-Scale Unconstrained Optimization and Its Applications Hardik Tankaria (IIT Mandi), Makoto Yamada (OIST), and Dinesh Singh (IIT Mandi) Abstract Abstract We develop a regularized Nystr\"om method for solving large-scale unconstrained optimization problems that arise in high-dimensional machine learning. In such settings, classical second-order methods are often impractical due to prohibitive computational and memory costs. While quasi-Newton methods rely exclusively on first-order information and Newton-sketch approaches require large embedding matrices, the proposed method achieves scalability by efficiently exploiting partial curvature information. Specifically, we construct a regularized Nystr\"om approximation of the Hessian using a small subset of columns, yielding an accurate and computationally efficient second-order approximation. The method is compatible with both deterministic and stochastic optimization frameworks, including gradient descent and stochastic gradient descent. To further reduce per-iteration complexity, the regularized-Nystr\"om inverse–gradient product is computed directly without explicit matrix inversion. We provide theoretical convergence guarantees and validate the proposed approach through extensive experiments on benchmark datasets. The results demonstrate clear computational advantages over randomized subspace Newton and Newton-sketch methods in high-dimensional regimes. We further illustrate the practical effectiveness of the method on a brain tumor detection task, highlighting its potential for real-world learning applications. Tuesday 0.05 Paris IJCNN Paper Explainable, Fair, and Responsible AI Session Chair: Monowar Bhuyan (Umeå University), M. Tanveer (Indian Institute of Technology Indore, India) FairBound: Boundary-Preserving Fair Downsampling for Accurate and Equitable Imbalanced Classification Chrysostomos Kalousios, Kenji Kobayashi, and Virginia Ghiara (Fujitsu Research of Europe) and Hiroya Inakoshi (Fujitsu Japan) Abstract Abstract We present a fair downsampling classification method applicable to a wide range of imbalanced datasets. The imbalance appears both at the class (majority vs. minority) as well as at the group level (e.g. gender, age, race). In traditional downsampling methods, such as Near Miss, there is usually a trade-off between accuracy and fairness, where in most cases downsampling often deteriorates fairness. We propose FairBound, a novel downsampling method, designed to prevent underfitting while balancing fairness and accuracy, by retaining samples near the decision boundaries. We demonstrate the utility of our method in several examples. Post-Hoc Feature Selection Layer for Neural Networks Interpretability João Kenji Suwa, João Pedro Silveira e Silva, and Bruno Iochins Grisci (Universidade Federal do Rio Grande do Sul) Abstract Abstract The interpretability of complex neural networks remains a critical challenge, especially for models already deployed in high-stakes domains. To address this, we adapted the Feature Selection Layer (FSL) to optimize feature weights post-training, a variant we term post-hoc FSL. Our approach reframes the FSL as a lightweight, trainable module that integrates with already frozen pre-trained models on tabular datasets to highlight the features the original model considers most important. This post-hoc FSL learns feature relevance by fine-tuning its weights based on the original model's learned outputs. Crucially, this process is non-invasive, operating without altering the original model's architecture or its learned parameters. We conducted our experiments using both statistical and visual metrics, including accuracy, F1 score, recall, precision, weighted t-SNE and silhouette score, and also analyzed the stability of the post-hoc FSL on high-dimensional synthetic and real-world tabular datasets. We compared the post-hoc FSL feature weighting method using these metrics against the original embedded FSL and other post-hoc interpretability methods, such as Integrated Gradients, Noise Tunnel, DeepLIFT, Gradient SHAP, and Feature Ablation. Experimental results demonstrate that the post-hoc FSL feature weighting method successfully identified relevant features across the different datasets, maintaining the predictive power of the original neural network while enhancing its interpretability. Multi-feature Procedural Fairness in Machine Learning Tianze Zhu (Lingnan University), Ziming Wang (Southern University of Science and Technology), Hao Tong and Jialin Liu (Lingnan University), Changwu Huang (Beijing Normal-Hong Kong Baptist University), and Xin Yao (Lingnan University) Abstract Abstract Machine learning models may reinforce existing social biases if fairness is not properly considered. Prior research on fair machine learning has largely focused on distributive fairness, which requires models to produce fair prediction outcomes across different demographic groups. By contrast, procedural fairness, which focuses on model decision logic, requiring consistent treatment of similar individuals across demographic groups, has been relatively understudied. Existing limited studies examine procedural fairness under a single sensitive feature with two groups (e.g., female and male), leaving more complex settings with multiple sensitive features and multiple groups unexplored. To address this limitation, we propose a novel metric for evaluating overall procedural fairness across multiple sensitive features and multiple groups. We further propose a learning method that incorporates procedural fairness objectives while preserving predictive performance. Experimental results show that our method achieves high overall procedural fairness across all sensitive features and groups while maintaining a competitive accuracy. A Comparative Study of Improved Disparate Impact Remover and Fair Adversarial Learning for Bias Mitigation JUNYU ZHOU, Shuojingrui He, JINGYU HU, and Weiru Liu (University of Bristol) Abstract Abstract Machine learning models in high-stake domains can inherit and amplify historical biases from training data, deepening social inequality. We compare the Improved Disparate Impact Remover (IDIR) with adversarial learning (AL) for comprehensive bias mitigations. IDIR enhances bias detection using a dual-metric combination of KL divergence and Total Variation distance. KL divergence captures local distributional anomalies while TV distance measures macroscopic probability deviations, enabling robust identification of indirect bias across structural differences and population shifts. We apply controlled data repair with a tunable strength coefficient that selectively adjusts non-sensitive group distributions toward sensitive groups through targeted augmentation, preserving model discriminability while achieving effective mitigation. Adversarial learning addresses biases that emerge during training through an encoder-predictor-adversary architecture, which uses gradient reversal layers to learn representations that minimize the correlation between predictions and sensitive attributes. IDIR is a preprocessing technique that reduces historical biases before training, whilst in-processing AL maintains fairness constraints during optimization. Experiments on three datasets demonstrate consistent reductions in group-level disparities by both methods while maintaining predictive accuracy. FINU: Fisher-Informed Noise Injection for Efficient Zero-Shot Unlearning Varun Sampath Kumar, Esmaeil S Nadimi, and Vinay Chakravarthi Gogineni (University of Southern Denmark) Abstract Abstract Machine unlearning aims to remove the influence of specific training data from a deployed model in order to satisfy privacy and regulatory requirements. While exact unlearning methods such as retraining provide strong guarantees, they are computationally expensive and impractical at scale. Recent approximate unlearning approaches improve efficiency but typically rely on access to retained training data, limiting their applicability in privacy-constrained settings. This paper studies the more challenging zero-shot unlearning setting, where only the trained model and the data to be forgotten are available. We propose FINU, a Fisher-guided noise injection-based zero-shot unlearning framework that induces selective forgetting without using the retain dataset. FINU is motivated by the hierarchical structure of deep neural networks (DNNs) and employs an adaptive, layer-wise masking strategy based on Fisher Information to identify parameters most influential for the forget set. Controlled, learnable noise is then injected into the selected parameters to maximize the loss on forget samples, effectively removing their influence while preserving generalizable knowledge. We evaluate FINU across class-level, subclass-level, and sample-level unlearning scenarios on CIFAR-100, CIFARSuper20, and ImageNet-1k using prominent DNNs. Tuesday 0.10 Sydney IJCNN Paper Edge, Federated, and Privacy-Preserving Learning Session Chair: Run Wang (ETH Zurich), Niklas Melton (Missouri University of Science and Technology) FTTE: Enabling Federated and Resource-Constrained Deep Edge Intelligence Irene Tenison, Anna Murphy, Charles Beauville, and Lalana Kagal (MIT) Abstract Abstract Federated learning (FL) enables collaborative model training without sharing raw data, but its deployment in edge-dominated networks is fundamentally limited by feasibility under device constraints and straggler-induced delays. Existing synchronous and semi-asynchronous FL frameworks implicitly assume that all participating clients can store, train, and communicate the full model, an assumption that breaks in real-world federated systems dominated by resource-constrained edge devices. We present FTTE (Federated Tiny Training Engine), an FL framework that enables semi-asynchronous training under hard per-device memory limits. FTTE enforces a global feasibility constraint induced by the most resource-limited client through server-side memory-aware parameter selection, sparse client updates, and sparse model communication in FL. To ensure stable optimization under sparse updates and heterogeneous client behavior, FTTE integrates buffered semi-asynchronous aggregation with an age–variance staleness weighting mechanism. Extensive experiments across diverse models, datasets, data distributions, and straggler regimes—including up to 500 clients and 90\% stragglers—show that FTTE consistently reaches target accuracy in substantially fewer communication rounds than synchronous and semi-asynchronous baselines, while reducing on-device training memory by up to 80\% and communication payload by up to 69\%. These results establish FTTE as a practical and scalable solution for federated learning on heterogeneous, resource-constrained edge devices. InteFL: Framework for AI-Assisted Design of Trustworthy and Efficient Federated Learning Applications Dmitrii Korobeinikov, Arnaldo Barea, Leon Reznik, and Raman Zatsarenko (Rochester Institute of Technology) and Sergei Chuprov and Angel Peredo (University of Texas Rio Grande Valley) Abstract Abstract Although Federated Learning (FL) enhances clients’ security and privacy by retaining data locally, the decentralized nature of this paradigm exposes it to diverse malicious actions as well as technological and reliability issues that can degrade data quality, hinder convergence, and reduce global model accuracy. Accounting for these destabilizing factors is essential in FL design for its integration in robust and reliable practical applications. To address this need, we introduce InteFL, a software tool engineered for the systematic design and evaluation of efficient and secure FL. InteFL provides functionality for benchmarking of FL performance under diverse conditions, including static and temporally dynamic adversarial attacks, and supports configuring and finetuning of robust aggregation algorithms and their parameters. Built as an extension to existing FL tools, InteFL possesses a higher practicality by incorporating intelligent agents that enhance tool usability. We present several use cases demonstrating how our tool supports the search and investigation of more resilient aggregation techniques for diverse attack models, and how parameter fine-tuning accelerates convergence and improves the accuracy of resulting ML models. KoalaMamba: A Federated Method with Vision Mamba UNet for Privacy-Preserving Medical Segmentation Haocheng Kan (Peking University; FOXCONN (Hon Hai) Precision Industry Co., Ltd.) and Mingpei Cao and Yuesheng Zhu (Peking University) Abstract Abstract Distributed machine learning holds strong promise for medical image segmentation in clinical diagnosis and treatment planning. However, collaborative training across multiple medical centers is fundamentally constrained by strict privacy requirements, which limits data sharing and hinders large-scale deployment. This paper presents KoalaMamba, a practical federated learning (FL) method designed for privacy-preserving medical segmentation while remaining compatible with multiple FL optimizers. KoalaMamba integrates client-side differential privacy (DP) into federated training, making it statistically difficult to infer whether any individual record participated in learning, thereby mitigating privacy leakage risks and protecting institutional confidentiality. To improve segmentation accuracy and generalizability, KoalaMamba incorporates Vision Mamba UNet-based architectures, including VMUNet, HVMUNet, and VMUNetV2, as configurable backbones. On the ISIC 2018 dermoscopic segmentation benchmark, KoalaMamba achieves competitive performance across diverse FL settings. In particular, under the SCAFFOLD Non-IID configuration, KoalaMamba is evaluated with DP enabled using the Moment Accountant (MA) at $\epsilon=10$ and $\delta=10^{-5}$, reaching 0.7822 IoU and 0.8778 DSC. This DP-constrained result substantially outperforms the UNet baseline and remains competitive with strong Mamba backbones evaluated without DP. Moreover, in separate privacy evaluation experiments, applying DP with a tighter budget ($\epsilon=1$) reduces the attack success rate (ASR) from 0.92 to 0.19 and decreases the membership inference AUC from 0.89 to 0.51, demonstrating that KoalaMamba effectively mitigates privacy leakage while preserving practical segmentation utility. Efficient Random Forests under TFHE via Topology & Bit-width Co-design Nunzio Licalzi, Erich Malan, Valentino Peluso, Andrea Calimera, and Enrico Macii (Politecnico di Torino) Abstract Abstract Fully Homomorphic Encryption enables privacy-preserving Machine Learning (ML) in third-party cloud environments by processing inference directly on encrypted data. Among the available approaches, Fully Homomorphic Encryption over the Torus (TFHE) has attracted significant interest due to its efficient support for low-noise non-linear operations, which are at the core of many ML models. However, deploying ML models under TFHE is extremely challenging, as encrypted arithmetic operations incur substantial computational overhead and integer approximations introduce non-negligible precision loss. This work focuses on the concurrent optimization of these aspects for a widely adopted class of ML models, namely Random Forests (RFs). Existing TFHE-based RF deployment workflows typically rely on a two-stage pipeline in which RF hyper-parameters are first optimized in plaintext to maximize accuracy, while cryptographic precision is adjusted in the second step to meet latency constraints. This approach implicitly assumes weak coupling between model topology and ciphertext bit-width, an assumption that does not hold in practice and leads to suboptimal accuracy-latency trade-offs and inefficient implementations. To overcome this limitation, we propose a co-design methodology that jointly optimizes RF topology and TFHE arithmetic precision. The proposed approach systematically explores a unified design space encompassing the number of trees, the maximum tree depth, and the ciphertext bit-width, thereby capturing cross-layer interactions between model architecture and cryptographic parameters. Experimental results on standard classification benchmarks show that the proposed co-design achieves up to a 100.38x faster inference compared to state-of-the-art sequential strategies, without loss of accuracy. Furthermore, the methodology finds a broader and previously unexplored set of operating points that substantially extends the accuracy-latency Pareto frontier, improving the scalability of trustworthy RF-based inference services. Edge-Ready 3D Scene-Language Reasoning with Compact Object Graphs Yatharth Agarwal and Arghadip Das (Purdue University), Arnab Raha (Intel Corporation), and Vijay Raghunathan (Purdue University) Abstract Abstract The proliferation of assistive robotics and Artificial Intelligence of Things demands indoor spatial reasoning capabilities that are semantically rich yet bounded by the latency and energy envelopes of edge devices. Addressing the fundamental conflict between the complexity of spatial reasoning and these resource constraints, we present an edge-ready 3D scene–language pipeline that decouples perception from reasoning through a compact, object-centric relational representation. Rather than relying on computationally expensive dense geometry or 2D feature extraction, we distill instance-segmented scenes into a sparse graph of object identities and local spatial relations, explicitly designed for efficient LLM prompting. This representation is serialized into a deterministic prefix, enabling a small, 4B parameter LLM to resolve complex spatial queries in a zero-shot manner without task-specific fine-tuning. By constraining both input structure and output schema, we bound prompt length and decoding cost while maintaining sufficient relational context for compositional indoor language. To enable practical deployment on edge hardware, we introduce a holistic end-to-end system design: we eliminate host-side bottlenecks via a fully GPU-resident superpoint partitioning routine and improve inference efficiency through prefix KV-cache reuse and low-precision FP4 quantization. Evaluated on an NVIDIA Jetson platform, our system achieves a 3.2$\times$ end-to-end speedup and a 1.6$\times$ improvement in energy efficiency. At the same time, it delivers state-of-the-art zero-shot performance, surpassing strong prior methods by +15.95\% Acc@0.50 on complex 3D reasoning segmentation and by +2.6\% Acc@0.50 on ScanRefer. Our results demonstrate that compact object-level relations provide an effective and efficient substrate for spatial reasoning, enabling high-fidelity 3D scene understanding within edge device resource budgets. A Lightweight Binary ART Model for Resource-Constrained Online Learning Niklas Melton, Sasha Petrenko, Jian Liu, Jacob Schroll, Micah Renfrow, Iwan Sandjaja, and Donald Wunsch (Missouri University of Science and Technology) Abstract Abstract Online learning models offer clear advantages over static alternatives when operating in dynamic settings. Unfortunately, this class of models is often too computationally demanding for the embedded and edge platforms where such adaptability is most valuable. This work introduces a new model architecture derived from the Adaptive Resonance Theory family that exploits binary data representations and exclusively hardware-efficient integer operations, eliminating the need for floating-point arithmetic or division. In contrast to existing ART architectures that rely on floating-point arithmetic, the proposed design enables fast, memory-efficient, and hardware-optimized online learning, making it suitable for deployment on resource-constrained devices. Through comparisons with analog counterparts, we demonstrate that the proposed model is capable of achieving faster training and inference while reducing memory usage by orders of magnitude, without any significant trade-offs in accuracy. Tuesday 0.11 Cape Town IJCNN Paper Graph and Relational Learning Session Chair: DANIELE ZAMBON (Università della Svizzera italiana), Luca Virgili (Marche Polytechnic University) GTCD: Gravity-Tension based Overlapping Community Detection Swetha Balasubramanian and Pranab K. Muhuri (South Asian University) Abstract Abstract Uncovering communities in large-scale networks remains a fundamental challenge due to the tradeoff between computationally expensive global optimization methods and efficient but unstable local heuristic-based methods. In this paper, we propose GTCD (Gravity-Tension based overlapping Community Detection), a deterministic framework that models the community detection problem as a discrete gradient flow over a density-based manifold. By treating community centers as local density peaks or “gravitational anchors”, the proposed GTCD moves beyond modularity-based methods, using Topological Tension to determine the membership of nodes. We quantify the divergence of local flow pointers in a node’s neighborhood, which provides a geometry-based strategy for multi-community assignment which evades the resolution limit problem faced by global optimization methods. GTCD achieves O(m + n log n) time complexity, thereby scaling to million-node networks on standard hardware. Extensive experiments performed on LFR benchmarks and diverse large-scale real-world networks demonstrate that GTCD outperforms state-of-the-art community detection methods in both accuracy and quality. Unsupervised Learning of Local Updates for Maximum Independent Set in Dynamic Graphs Devendra Parkar, Anya Chaturvedi, and Joshua Daymude (Arizona State University) Abstract Abstract We present the first unsupervised learning model for Maximum-Independent-Set (MaxIS) in dynamic graphs where edges change over time. Our method combines structural learning from graph neural networks (GNNs) with a learned distributed update mechanism that, given an edge addition or deletion event, modifies nodes' internal memories and infers their MaxIS membership in a single, parallel step. We evaluate our model against a mixed integer programming solver and a breadth of unsupervised and supervised learning models for combinatorial optimization on static graphs. Across dynamic graphs of 200-1,000 nodes, our model achieves approximation ratios that are competitive with the state-of-the-art models while running 1.91-6.70x faster. When generalizing to graphs with 100x more nodes than those used for training, our model produces MaxIS solutions 1.00-1.18x larger than all other unsupervised models, but is outperformed by the state-of-the-art supervised model. These results demonstrate that this novel, unsupervised, update-based learning approach to dynamic combinatorial optimization is a viable alternative to the naive reapplication of analogous models for static graphs, leveraging temporal information to improve neural methods for combinatorial optimization. A self-supervised joint embedding predictive approach for bipartite graphs to solve linear programs Joachim Verschelde, Daniel Stanley Tan, and Stefano Bromuri (Open Universiteit) Abstract Abstract Linear programs (LPs) are traditionally solved using iterative algorithms such as the simplex method or interior-point method, which enforce feasibility and optimality but treat each problem instance independently. In contrast, end-to-end neural solvers aim to learn a direct mapping from an LP instance to its solution via a single forward pass. However, most existing ap- proaches depend on large datasets of solver-generated optimal so- lutions or require embedding a classical solver within the training loop. This paper presents a fully self-supervised framework that learns optimization-aware representations and predicts primal- dual LP solutions without access to ground-truth solutions. First, we introduce BiJEPA, a joint-embedding predictive pretraining objective for bipartite variable-constraint graphs of LPs. BiJEPA learns instance representations by predicting global embeddings from heavily masked local views, enabling pretraining on large collections of unsolved LP instances. Second, we fine-tune a bipartite graph neural network using a differentiable residual- based loss that requires no access to optimal solutions. The loss is defined entirely in terms of the optimization structure of the LP using the Karush-Kuhn-Tucker(KKT) conditions. Across four families of linear programs, we show that the proposed approach learns an end-to-end mapping with low optimality gaps. Overall, our results highlight how optimization structure can replace external supervision, enabling neural networks to learn to solve LPs from first principles. GravSpec: Spectral-Gravitational flow on Density Manifolds for Deterministic Overlapping Community Detection Swetha Balasubramanian and Pranab K. Muhuri (South Asian University) Abstract Abstract The discovery of overlapping community structures in large-scale networks remains a crucial aspect of graph mining. However, the existing methods face a fundamental trade-off between computational efficiency and structural representation. While traditional Density Peak Clustering (DPC) constitutes a robust center selection strategy, it is constrained by O(N^2) complexity due to its pairwise distance calculations and limited topological awareness. In this paper we propose GravSpec, a sub-quadratic framework that reformulates DPC for non-Euclidean graph manifolds. GravSpec comprises two key ideas: 1) A spectral density estimation technique derived from Personalized PageRank and 2) A Jaccard-weighted gravitational flow that bounds the search space to the local adjacency manifold. We demonstrate that GravSpec runs in near-linear time complexity O(M+NlogN) and in O(M+N) space complexity. Results from extensive experiments on LFR benchmarks and massive real-world networks show that GravSpec outperforms its competitors. Notably, in high-noise environments where traditional methods degrade, our approach achieves up to a 34.8% improvement in Overlapping Normalized Mutual Information (ONMI) and superior structural resolution on large-scale graphs while ensuring scalability for graphs containing millions of edges. KANNER Zhe Li and K. L. Eddie Law (Macao Polytechnic University) Abstract Abstract Grid-based methods have become the dominant paradigm for Named Entity Recognition (NER) by explicitly enumerating all possible spans. However, these methods face a fundamental trade-off between global context aggregation and local boundary precision. Conventional convolutional encoders effectively capture long-range dependencies but inadvertently act as low-pass filters, blurring the sharp boundaries required for accurate span detection. Furthermore, traditional Multi-Layer Perceptrons (MLPs) often struggle to disentangle highly coupled semantic features in nested structures without excessive parameter growth. To resolve these issues, we propose KANNER, a parameter-efficient architecture that integrates Kolmogorov-Arnold Networks (KANs) into the NER framework. Specifically, we design a parallel enhancement mechanism: (1) Multi-scale Adaptive Gradient Sharpening (MAGS), which employs learnable difference operators to recover high-frequency boundary signals from the smoothed context; and (2) Masked Spectral KAN Attention (MaskedSpeKAN), which leverages the learnable activation functions of KANs to capture non-linear channel dependencies with fewer parameters than standard MLPs. Experiments on CADEC, CoNLL2003, and GENIA benchmarks demonstrate that KANNER achieves state-of-the-art performance. Notably, it surpasses strong baselines by +1.01 F1-score on CADEC while maintaining competitive inference speeds, offering a superior balance between accuracy and efficiency. Graph-Based Approaches to Learning Epileptogenic Zone Localization Using Stereo-EEG Recordings Daniel Wendelken (University of Cincinnati), Brian Ervin and Ravindra Arya (Cincinnati Children's Hospital), and Ali Minai (University of Cincinnati) Abstract Abstract The epileptogenic zone (EZ) is the brain region that generates seizures in an individual, and is the target of epilepsy surgery. Localizing the EZ from stereo-EEG (sEEG) recordings supports surgical planning, but manual interpretation is time-consuming and focuses on seizure recordings. Graphical learning models of resting-state functional connectivity among the recorded brain regions are an attractive alternative, but depend crucially on the network topology chosen for the model. Tuesday 0.15 Washington IJCNN Paper Trustworthy and Safe Language Models Session Chair: Robert G. Reynolds (Wayne State University), Ihsen Alouani (Queen's University Belfast, Upper Bound) Do Language Models Trust Their Own Justifications? A Study on Functional Consistency Alisson Rosa Pereira and Bruno Iochins Grisci (Universidade Federal do Rio Grande do Sul) Abstract Abstract Large Language Models (LLMs) have been widely adopted in text classification tasks, where they not only output class predictions but also generate explanations that highlight the tokens deemed most relevant for reaching the predicted label. Yet it remains unclear whether these highlighted elements are behaviorally consistent with the model’s own predictions under targeted interventions, which we refer to as functional auto-consistency. While much of the literature evaluates the textual plausibility of such explanations, few studies assess their functional consistency with the model’s actual behavior. In this work, we propose an experimental framework based on the principle of auto-consistency: if a model identifies certain tokens as decisive, then isolating, removing, or semantically inverting them should produce systematic and interpretable changes in its predictions. We operationalize this evaluation through sufficiency, comprehensiveness, and counterfactuality metrics, and conduct experiments on IMDB and Steam reviews across both closed-source (GPT-4o-mini) and open-source LLMs (Gemma3, Granite8B, DeepSeek). Results show that GPT-4o-mini follows the expected progression across all metrics, Gemma3 and Granite8B maintain coherence under sufficiency but lose consistency under more demanding interventions, while DeepSeek variants display structural deviations, either failing to preserve sufficiency or overreacting under comprehensiveness and counterfactuality. These findings show that explanation reliability varies across LLM families and scales, with smaller models displaying contradictions and larger ones exhibiting over-sensitivity. By combining sufficiency, comprehensiveness, and counterfactuality, our approach provides a systematic methodology for assessing the functional consistency of LLM self-explanations. ReaSafe: Reasoning-Centric Safety Alignment for Multimodal LLMs via Preference Optimization Xiaoning Dong (Tsinghua University), Kangshuai Zhao and Wenbo Hu (Hefei University of Technology), and Hang Su (Tsinghua University) Abstract Abstract Multimodal Large Language Models (MLLMs) power a wide range of applications, but remain vulnerable to jailbreak attacks. While prior work has improved defense effectiveness, MLLMs still lack robust safety reasoning capabilities, making them struggle to identify unsafe queries concealed within sophisticated jailbreak attacks and discern seemingly risky yet benign ones. In this work, we first show that existing defenses either fail to effectively block harmful queries in subtle jailbreak attacks or unnecessarily refuse benign queries that appear risky. Motivated by these observations, we propose ReaSafe, a reasoning-centric safety alignment framework that explicitly activates MLLMs’ safety reasoning capability via preference optimization. Specifically, we first curate a high-quality preference dataset and then propose a variant of DPO to train MLLMs on this dataset, teaching them to engage in inferring user intent. We evaluate ReaSafe against recent defense baselines under various jailbreak attacks, and further assess its sensitivity and utility. Experimental results show that ReaSafe achieves a favorable balance between safety and helpfulness, outperforming strong baselines by a clear margin. For example, ReaSafe reduces the overall attack success rate (ASR) to 3.1% under advanced attacks across three datasets, while maintaining a low refuse-to-answer rate (RAR) of 9% on MOSSBench. ECRT: Evidence-Centric Cross-Modal Reasoning with Trust-Aware Gating for Multimodal Fake News Detection Meghna Gade and Sriman Narayana (IIITDM Jabalpur); Rakesh Sanodiya (IIT Ropar); Lekshmi R (Muthoot Institute of Technology and Science, Kochi, Kerala, India); and Jimson Mathew (IIT Patna) Abstract Abstract Fake news detection is becoming difficult due to unreliable visual evidence and inconsistent interactions between textual claims, images, and social signals. Existing methods primarily rely on feature-level fusion or static attention, lacking explicit modeling of evidence reliability. We propose \textbf{ECRT (Evidence-Centric Cross-Modal Reasoning with Trust-Aware Gating)}, a novel framework that treats images as primary evidence signals and performs structured reasoning over claims, visual content, caption-derived semantics, and social context. ECRT introduces evidence-centric tokenization, trust-aware dynamic gating to regulate modality contributions, and a lightweight cross-modal consistency loss that improves generalization without external supervision. Experiments on the Fakeddit dataset show strong performance across multiple label granularities, achieving \textbf{93.05\%}, \textbf{92.95\%}, and \textbf{89.17\%} macro-F1 for 2-way, 3-way, and 6-way classification, respectively. Also on the Weibo dataset, attains \textbf{96.03\%} macro-F1 and \textbf{98.60\%} AUC. Extensive ablation studies confirm improved robustness and training stability, highlighting ECRT as an interpretable and computationally efficient alternative to many heavy multimodal approaches. A Two-Stage LLM Framework for Accessible and Verified XAI Explanations Georgios Mermigkis and Dimitris Metaxakis (Department of Computer Engineering and Informatics, University of Patras, Patras, Greece; Archimedes Unit, Athena Research Center, Athens, Greece); Marios Tyrovolas (Department of Informatics and Telecommunications, University of Ioannina, Arta, Greece; Industrial Systems Institute, Athena Research Center, Patras, Greece); Argiris Sofotasios (Department of Computer Engineering and Informatics, University of Patras, Patras, Greece; Archimedes Unit, Athena Research Center, Athens, Greece); Nikolaos Avgeris (Department of Informatics and Telecommunications, University of Ioannina, Arta, Greece; Archimedes Unit, Athena Research Center, Athens, Greece); Panagiotis Hadjidoukas (Department of Computer Engineering and Informatics, University of Patras, Patras, Greece; Industrial Systems Institute, Athena Research Center, Patras, Greece); and Chrysostomos Stylios (Department of Informatics and Telecommunications, University of Ioannina, Arta, Greece; Industrial Systems Institute, Athena Research Center, Patras, Greece) Abstract Abstract Large Language Models (LLMs) are increasingly used to translate the technical outputs of eXplainable Artificial Intelligence (XAI) methods into accessible natural-language explanations. However, existing approaches often lack guarantees of accuracy, faithfulness, and completeness. At the same time, current efforts to evaluate such narratives remain largely subjective or confined to post-hoc scoring, offering no safeguards to prevent flawed explanations from reaching end-users. To address these limitations, this paper proposes a Two-Stage LLM Meta-Verification Framework that consists of (i) an Explainer LLM that converts raw XAI outputs into natural-language narratives, (ii) a Verifier LLM that assesses them in terms of faithfulness, coherence, completeness, and hallucination risk, and (iii) an iterative refeed mechanism that uses the Verifier’s feedback to refine and improve them. Experiments across five XAI techniques and datasets, using three families of open-weight LLMs, show that verification is crucial for filtering unreliable explanations while improving linguistic accessibility compared with raw XAI outputs. In addition, the analysis of the Entropy Production Rate (EPR) during the refinement process indicates that the Verifier’s feedback progressively guides the Explainer toward more stable and coherent reasoning. Overall, the proposed framework provides an efficient pathway toward more trustworthy and democratized XAI systems. Tuesday 2.1 Volga IJCNN Paper Generative Models and Visual Synthesis Session Chair: Van Huyen Dang (Paderborn University), Dawid Połap (Silesian University of Technology, Poland) FaceComposer: Learning Coherent and Realistic Face Composition from References LING LI (Nanyang Technological University); Lanqing Guo (The University of Texas at Austin); Siyuan Yang (KTH Royal Institute of Technology); Yufei Wang (SparcAI, Inc.); Qian Zheng (Zhejiang University); Alex C. Kot (Shenzhen MSU-BIT University); and Weisi Lin and Lihui Chen (Nanyang Technological University) Abstract Abstract Reference-based face composition enables precise editing of individual facial components in a source portrait using a reference image, with applications in forensic sketching, facial surgery simulation, and digital media. Despite recent advances in diffusion models, achieving realistic and structurally consistent component-level editing remains challenging. To fill this gap, we propose Context-Aware Face Fusion, a generative framework that seamlessly integrates selected facial components into a source portrait while preserving both structural coherence and texture harmony. To handle mismatches in pose, scale, and cross-ethnic appearance variations, we introduce a Component-Aware Augmentation strategy that adaptively aligns reference features to the source. In addition, we design a Multi-Stage Training pipeline that mitigates the lack of paired training data by progressively transitioning from reconstruction to composition tasks. Extensive experiments demonstrate that our method produces high-fidelity composites that faithfully incorporate reference components while maintaining realism and global coherence. Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation Ryugo Morita (DFKI, RPTU); Stanislav Frolov, Brian Bernhard Moser, and Ko Watanabe (DFKI); Riku Takahashi (Hosei University, DFKI); and Andreas Dengel (DFKI) Abstract Abstract Diffusion models have demonstrated high-quality performance in conditional text-to-image generation, particularly with structural cues such as edges, layouts, and depth. However, lighting conditions have received limited attention and remain difficult to control within the generative process. Existing methods handle lighting through a two-stage pipeline that relights images after generation, which is inefficient. Moreover, they rely on fine-tuning with large datasets and heavy computation, limiting their adaptability to new models and tasks. To address this, we propose a novel Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation (LGTM), which manipulates the initial latent noise of the diffusion process to guide image generation with text prompts and user-specified light directions. Through a channel-wise analysis of the latent space, we find that selectively manipulating latent channels enables fine-grained lighting control without fine-tuning or modifying the pre-trained model. Extensive experiments show that our method surpasses prompt-based baselines in lighting consistency, while preserving image quality and text alignment. This approach introduces new possibilities for dynamic, user-guided light control. Furthermore, it integrates seamlessly with models like ControlNet, demonstrating adaptability across diverse scenarios. TPSO: Training-Free Diverse Image Generation via Semantic Prompt Embedding Optimization Debin MENG (Queen Mary University of London), Chen Jin (AstraZeneca), Zheng Gao (Queen Mary University of London), Yanran Li (University of Bedfordshire), and Ioannis Patras and Georgios Tzimiropoulos (Queen Mary University of London) Abstract Abstract Image diversity remains a fundamental challenge for text-to-image diffusion models. Low-diversity generation often leads to repetitive outputs, increasing sampling redundancy and hindering both creative exploration and downstream applications. A key factor is the tendency of diffusion models to collapse toward strong modes in the learned distribution. Existing attempts to improve diversity, such as steering-based guidance, often introduce distortions that degrade image quality. To address this issue, we propose Token-Prompt Embedding Space Optimization (TPSO), a training-free and model-agnostic module. TPSO introduces learnable parameters to explore underrepresented regions of the token embedding space, reducing the tendency to repeatedly sample from strong modes of the distribution. Meanwhile, a prompt-level semantic constraint regulates distribution shifts, preventing quality degradation while preserving semantic fidelity. Extensive experiments on MS-COCO across three representative diffusion backbones demonstrate that TPSO substantially improves diversity, boosting performance from 1.10 to 4.18, while maintaining image quality with only a modest inference-time overhead of 3.6 to 8.9. Code is available at: https://github.com/Open-Debin/TPSO. Disentangled Motion Diffusion for Long-Form Voice-Driven Portrait Animation Yi-Chun Chang, Po-Hsuan Fan Chiang, and Jen-Tzung Chien (National Yang Ming Chiao Tung University) Abstract Abstract Long-form portrait animation with high-fidelity and low-latency has been highly-demanding when building the human-computer interactive systems. The existing methods often suffer from inherent trade-off between effectiveness and efficiency while improving one objective but compromising another. This paper presents a novel audio-driven portrait animation, which mitigates the compromise through enhancing the representation by flow-divergence matching in sequence animation under an orthogonal latent space for motions. To strengthen the facial motion controllability, the latent space is decomposed into a transferable motion subspace for identity-agnostic dynamics and a preservable subspace for portrait-specific traits. Furthermore, a sparse motion dictionary is constructed to fulfill the fine-grained disentanglement, allowing an independent control over semantic motion factors and improving both realism and identity consistencies in portrait animation. To speed up the generation, we employ a transformer-based vector field predictor that leverages flow-based temporal fusion to integrate intra-clip and inter-clip audio perception, resulting in temporally-coherent and contextually-aware facial dynamics. Extensive experiments demonstrate that our method achieves stable long-term video synthesis and outperforms state-of-the-art baselines in terms of video quality, generation speed, and temporal coherence. An Analysis of Regularization and Fokker–Planck Residuals in Diffusion Models for Image Generation Onno Niemann, Gonzalo Martínez-Muñoz, and Alberto Suárez Gonzalez (Universidad Autónoma de Madrid) Abstract Abstract Recent work has shown that diffusion models trained with the denoising score matching (DSM) objective often violate the Fokker--Planck (FP) equation that governs the evolution of the true data density. Directly penalizing these deviations in the objective function reduces their magnitude but introduces a significant computational overhead. It is also observed that enforcing strict adherence to the FP equation does not necessarily lead to improvements in the quality of the generated samples, as often the best results are obtained with weaker FP regularization. In this paper, we investigate whether simpler penalty terms can provide similar benefits. We empirically analyze several lightweight regularizers, study their effect on FP residuals and generation quality, and show that the benefits of FP regularization are available at substantially lower computational cost. Our code is available at https://github.com/OnnoNiemann/fp_diffusion_analysis. Self-Regulating Annealing in Heavy-Tailed Diffusion Models Keito Wakatsuki and Hideaki Shimazaki (Kyoto University) Abstract Abstract Diffusion models have emerged as a leading framework for deep generative modeling. While the standard Gaussian formulation is theoretically convenient, its suitability for heavy-tailed datasets remains unclear. To address this, heavy-tailed diffusion models (HTDMs) extend the standard formulation by replacing the Gaussian distribution with a Student's $t$-distribution, thereby improving tail fidelity on heavy-tailed datasets. Although stochastic differential equation (SDE)-based sampling is possible in HTDMs, it has not been fully explored. In this paper, we propose an SDE-based sampler for HTDMs that explicitly incorporates a state-dependent diffusion coefficient. This state dependence naturally induces a self-regulating annealing mechanism by adaptively modulating the effective noise scale. We theoretically explore this mechanism and experimentally verify its necessity for reproducing samples from a heavy-tailed distribution. Tuesday Brightlands Foyer IJCNN Paper, FUZZ-IEEE Position Paper, CEC Late Breaking Paper, CEC Paper, FUZZ J2C Presentation, CEC J2C Presentation, FUZZ-IEEE Paper, CEC Position Paper, IJCNN J2C Presentation, IJCNN Position Paper, IJCNN Late Breaking Paper, FUZZ-IEEE Late Breaking Paper LBR Posters IJCNN / CEC / FUZZ Session Chair: Maximilian Krentzien (University of Rostock, Institute of Applied Microelectronics and Computer Engineering), Mihail Popescu (University of Missouri) Generalized complex numbers properties applied to A-Linearly Correlated fuzzy numbers Felipe Longo, Estevão Esmi Laureano, and Laécio Carvalho de Barros (University of Campinas) Interval Type-3 Fuzzy Evidence Fusion for Large Language Model Disagreement Mihail Popescu (University of Missouri) Abstract Abstract Large language models (LLMs) frequently exhibit output variability across repeated queries, disagreement across models, and instability under prompt perturbation. Standard ensemble methods such as majority voting, mean-score aggregation, or probability averaging often improve accuracy, but they typically collapse multiple layers of uncertainty into a single number. This paper proposes a theoretical and empirical framework that combines Dempster–Shafer (DS) evidence theory with Interval Type-3 fuzzy sets to model hierarchical uncertainty in LLM ensembles. We also introduce an interval-width-based reliability metric and evaluate the framework on two biomedical tasks: UMLS concept verification and gene–gene interaction verification. Type-n Fuzzy Control for Robotics Venkata Subba Reddy Poli (Rohith-Medhaj Computers Private Limited) Abstract Abstract Robotics has to handles the problem with inexact to slowly approach the Object without impact. This inexactness is fuzy. The slow is defined by fuzziness. Fuzzy logic made inexact to exact. For this fuzzy type-n of approach. Robotics has to handle with inexact or fuzzy information and object should be placed exact position. The inexact information is fuzzy. Fuzzy algorithms are designed for Robotics to deal with inexact information. In this presentation, Type-n fuzzy Algorithms are designed for controlling Robotics. And reach Robotics exact position without impact Fuzzy Blackboard GPT: Generative Pre-trained Transformer Venkata Subba Reddy Poli (Rohith-Medhaj Computers Private Limited) Abstract Abstract Blackboard System is shared the knowledge, controlling knowledge sources and data sources of problem solving concept. Blackboard System con access knowledge sources simultaneously by transforming the rule based into nested rule-based. In this presentation fuzzy nested conditional inference is studied. Blackboard system is studied using fuzzy nested conditional inferences and controlling knowledge sources. The knowledge sourcess are pre-trained for Blackboard system. Blackboard Generative Pre-Trained Transformer is studied usin question answering system. Blackboard System will reduce the search, integrate and reuse the knowledge sources (KSs). The Fuzzy Blackboard GPT Shell is given. Fuzzy Fractal Image Processing Venkata Subba Reddy Poli (ROHITH-MEDHAJ COMPUTERS PRIVATE LIMITED) Abstract Abstract Fractals geometry is introduced by Mandelbrot [ as” the Geometry of Nature”. Clouds are not spheres, mountains are not cones, coastlines are not circles and nor does lightning travel in a straight line. All Clouds, Coastline, walking and Mountains are having scale invariance and self-similarity. These structures are imprecise or fuzzy In this presentation the fuzzy fractal structures are studied. Fuzzy Fractals are structures, have the property of scale invariance or self-similarity.. Fuzzy fractal structures allow dimension of fractals. These structures need computer assistance for a generation. Some methods and techniques are studied for Computer generation/. Fuzzy Security for Blockchain Technology Venkata Subba Reddy Poli (Rohith-Medhaj Computers Private Limited) Abstract Abstract Blockchain Technology is decentralized and secured. The Blockchain is transaction is Blocks of transactions. Blockchain is transaction processes which avoid third party intervention. Blockchain data sources are encrypted and code is send to second party. The code is intermediate node and is used to for transaction. In This presentation, fuzzy security code is studied for Blockchain, Steiner tree is studied as security node for Blockchain. Fuzzy security code is highly secured and irreversible. Breaking the Cycle: Stopping Overthinking in Language Models via Cyclic Reasoning Patterns Baban Gain, Anchal Dubey, and Asif Ekbal (Indian Institute of Technology, Patna) Whova Tag: Poster Presentation Classical Shadows for Privacy-Preserving Data Sharing Alexandru Ionita (AI Multimedia Lab, CAMPUS Research Institute, National University of Science and Technology Politehnica Bucharest, Romania) and Bogdan Ionescu (AI Multimedia Lab, CAMPUS Research Institute National University of Science and Technology Politehnica Bucharest) Whova Tag: Poster Presentation Electrical AI Copilot – A Framework for Empowering Power System Engineers Ahmed Saber (ETAP) Whova Tag: Poster Presentation On the Consistency of Graph Structure Learning in Spatiotemporal Modeling Daniele Zambon (Università della Svizzera italiana) and Cesare Alippi (Università della Svizzera italiana, Politecnico di Milano) Whova Tag: Poster Presentation Grid-Discretized Multi-Task Transformer for Joint Tropical Cyclone Path and Intensity Forecasting Thanh Nguyen (Vietnam National University) Whova Tag: Poster Presentation Automated Scoliosis Assessment via Dual-Stage YOLOv8 Detection and Attention U-Net Segmentation Miri Weiss Cohen (Braude College of Engineering) Whova Tag: Poster Presentation An Evidence-based Arbitration Framework for Zero-shot Medical Image Segmentation Trishita Trishita (MSRIT) Whova Tag: Poster Presentation Successful Adaptation of a Group EEG Generative Model to Individuals Depends on Participant Typicality Xinyu Li (University Medical Center Groningen), Marieke van Vugt (University of Groningen), and Natasha Maurits (University Medical Center Groningen) Whova Tag: Poster Presentation Abstract Abstract To investigate whether a group-pretrained EEG generative model can be stably adapted to individual participants, we first trained a group-level EEG generative model on on-task trials using an event-related potential (ERP) Wasserstein generative adversarial network with gradient penalty (ERP-WGAN-GP), and then fine-tuned it to individual participants to obtain more participant-specific trials. We examined how stable this participant-specific fine-tuning is across subjects. The model was pretrained using on-task trials from 25 participants (4753 trials in total) and subsequently fine-tuned separately for each participant using full fine-tuning. Adaptation quality was assessed using principal component analysis (PCA)_overlap between real and fake trials, the Frechet Distance (FD16) calculated from 16 EEG features capturing temporal, spectral, and complexity characteristics, and waveform-level comparisons (ERP and power spectral density (PSD)). We observed variability in the fine-tuning outcomes: ten participants achieved good adaptation (PCA_overlap ≥ 0.60), six were borderline (0.50—0.60) and nine showed poor adaptation (<0.50). These results indicate that participant-specific fine-tuning is possible but it exhibits variability across participants, motivating participant-specific fine-tuning and evaluation strategies for EEG generation under limited per-participant data. A Formalization of the On-device Learning Problem Massimo Pavan, Xenofon Fafoutis, and Luca Pezzarossa (DTU) Whova Tag: Poster Presentation Abstract Abstract There is little agreement on what doing On-Device Learning (ODL) means, and many ODL works don’t consider the constraints of real-world conditions: limited memory, computation, and label availability, data as a stream, no model validation. We propose a formalization of the ODL problem that works with different environments and constraints, with the goal of designing reliable solutions in real-world conditions. Physiological Time Series Prediction Model Based on Multivariate State Graphs Dongxun Jiang (Tongji University) and Jiyang Wu (University of Science and Technology of China) Whova Tag: Poster Presentation Abstract Abstract Physiological variables exhibit state-dependent and dynamically evolving correlations that are not explicitly modeled by conventional time-series architectures. We propose a physiological time series prediction model that explicitly captures state-dependent variable interactions for multivariate physiological time-series forecasting. The temporal CNN module is introduced for extracting multi-scale temporal features from each physiological variable independently. The multivariate state graph construction module is designed for explicitly modeling state-dependent and dynamically evolving correlations among variables. The temporal prediction module based on LSTM is exploited for integrating temporal dynamics of individual variables with instantaneous inter-variable interactions. Experiments on PhysioNet 2012 show consistent improvements over a baseline, highlighting the benefit of explicit variable interaction modeling. Resource Constrained Software Development Paradigm Shift to Neural-Network-Based Energy and Runtime Estimation of Software Execution Maximilian Krentzien and Marc Reichenbach (University of Rostock, Institute of Applied Microelectronics and Computer Engineering) Whova Tag: Poster Presentation Abstract Abstract Resource-constrained software development traditionally relies on executing code on target hardware to measure energy consumption and runtime, which hampers rapid prototyping and platform independence. This poster proposes a data-driven paradigm shift that replaces hardware-bound profiling with artificial-neural-network models, specifically multi-layer perceptrons, to predict execution energy and latency from program-level features. A reproducible dataset covering diverse microarchitectures is used to train and evaluate several network topologies and feature-selection strategies. Preliminary experiments on ESP32-C6 (HP) and Banana Pi BPI-F3 platforms achieve R² scores above 0.95, demonstrating the feasibility of accurate, hardware-agnostic estimation. Human-Robot Collaborative Manipulation via Switchable Teleoperation and Autonomous Grasping Jia Guo (Nagoya University); Jiacheng Li (Kanagawa University); and Kenta Urano, Takuro Yonezawa, and Nobuo Kawaguchi (Nagoya University) Whova Tag: Poster Presentation Abstract Abstract Teleportation of robotic arms over long distances places a significant cognitive and physical burden on human operators, particularly during the precision-demanding final phase of grasping. This work proposes a human-robot collaborative control framework that integrates large-range teleportation with proximity-triggered autonomous grasping. In the proposed system, the operator maintains full control during coarse positioning, while an automatic mode transition is activated when the end-effector reaches a predefined distance threshold from the target object. This work employs a distance-aware switching mechanism to seamlessly hand over control from the human operator to an autonomous grasping module, enabling efficient and reliable object acquisition without requiring continuous high-precision input from the operator. Experimental results demonstrate that the proposed approach significantly reduces operator workload while improving grasping success rate and overall task efficiency. This work provides a practical and scalable solution for human-robot collaboration in remote manipulation scenarios. Tuesday 0.01 London FUZZ-IEEE Paper FUZZ 5: FUZZ-IEEE SS03 Information fusion techniques based on aggregation functions, preaggregation functions and their generalizations & Main:Mathematical and theoretical foundations of fuzzy sets, measures and integrals Session Chair: Giancarlo Lucca (Universidade Federal de Pelotas), Graçaliz Dimuro (Universidade Federal do Rio Grande) A Unified Framework For Regularizing The Choquet Integral Godswill Ikwan and Muhammad Aminul Islam (University of New Haven) Abstract Abstract The Choquet Integral (ChI) through its fuzzy measure (FM) provides a powerful means for non-linear aggregation. Existing work on ChI regularization primarily focuses on either optimizing towards a target FM or minimizing the FM parameters, similar to classical machine learning. The latter approach is essentially goal-oriented optimization towards the minimum operator. These works are limited in the sense of the part of the parameter space they explore. In this article, we enhance the exploration by combining regularization towards two goals, the minimum and the maximum, thus facilitating searching over the entire FM spectrum, which we call the operator regularization. Additionally, we introduce model complexity regularization via additivity conditions that guides the parameters towards linear model (termed as the complexity regularization). We provide a unified framework synthesizing these two regularizers, allowing the ChI to effectively optimize over the FM spectrum. We conducted experiments on real-world datasets demonstrating the effectiveness of the proposed method for limited data. Power Choquet-Inspired Integral: A Family of Tunable Aggregation Functions with Application to Fuzzy Rule-Based Classification Systems Giancarlo Lucca (Universidade Federal de Pelotas); Tiago Asmus, Bruno Dalmazo, Rafael Berri, and Miqueias Amorim (Universidade Federal do Rio Grande); Cedric Marco-detchart and Humberto Bustince (Universidad Pública de Navarra); and Graçaliz Dimuro (Universidade Federal do Rio Grande) Abstract Abstract The Choquet integral is a well-established aggregation operator capable of modeling interactions among criteria, and it has been successfully employed in fuzzy rule-based classification systems through the use of adaptive fuzzy measures such as the Power Measure. More recently, Choquet-inspired aggregation functions have been proposed as an alternative formulation that reduces the amount of required information while preserving the structural principles of the Choquet integral. In this paper, we introduce the Power Choquet-Inspired Integral (PCII), a novel family of aggregation operators that combines the adaptability of the Power Measure with the Choquet-inspired framework. The proposed approach led to a new aggregation mechanism with reduced informational requirements and improved computational efficiency. An extensive experimental study on 33 benchmark datasets demonstrates that PCII consistently outperforms the classical Power Measure in terms of classification accuracy, while also achieving lower execution times. Statistical analyses confirm the significance of the observed improvements, supporting the effectiveness of the proposed operator in classification tasks. Characterisation of the Quantum Fuzzy Measure Based on a Quantum Circuit Yanhao Huang and Christian Wagner (University of Nottingham, Lab for Uncertainty in Data and Decision Making) Abstract Abstract The fuzzy measure (FM) is a powerful means to capture the worths of combinations of sub-sources, offering potential for applications from sensor fusion to ensemble classification. In most approaches, worths are treated as numeric values, which may not reflect real-world settings, as they may contain underlying uncertainty. Intervals and fuzzy sets have been explored to capture the uncertainty in the worths. Building on these ideas, this paper leverages quantum computing, specifically qubits, to interpret the uncertainty in the worths. The approach puts forward a new way of understanding the FM in the sense of quantum mechanics. We define the boundary conditions and the monotonicity constraint, and articulate that a measurement of a given quantum FM (QFM) can be interpreted as a binary FM (BFM). Compared to a BFM, which encodes a single 0/1 configuration of worths, a QFM captures information of all configurations: measuring it can yield any configuration, each occurring with a certain probability. A quantum circuit is proposed to ensure that the measurement of a QFM, i.e. a BFM, always follows the monotonicity of the numeric FM. Going one step further, we apply the Choquet fuzzy integral (CFI) based on a QFM and show that it can be treated as the probability-weighted sum, i.e. the expectation, of the CFI over all possible measurement outcomes of this QFM. We conclude by discussing limitations and charting the steps for future work on real-world applications of the QFM, in particular for ensemble classifiers. Generalizing the Antecedent Layer of ANFIS: A Performance Analysis Under Classes of Fuzzy Aggregators Bruna Camily Domingues Novack, Gabriel Rosa de Oliveira Silva, and Juliano Buss (Universidade Federal de Pelotas); Giancarlo Lucca and Helida Santos (Universidade Federal do Rio Grande); and Renata Reiser (Universidade Federal de Pelotas) Abstract Abstract Machine Learning, a branch of Artificial Intelligence, allows systems to infer patterns and make data-driven decisions without explicit programming. Among its approaches, supervised learning and Neural Networks are fundamental to solving complex problems, from medical diagnoses to banking security. Fuzzy Logic complements this area in decision-making involving uncertain and imprecise information formalized by linguistic variables, from modeling to computational simulation of human calculation. Neuro-fuzzy systems foster the integration of these technologies by exploring hybrid models that combine fuzzy inference with neural network learning. A prominent example is ANFIS (Adaptive Neuro-Fuzzy Inference System), which uses the Takagi-Sugeno method to model complex, non-linear systems using processing layers and automatic rule adjustment. In this context, this work proposes a generalization of the ANFIS architecture. The generalized architecture component analyzes the network's performance by applying different classes of fuzzy aggregators during the input data fuzzification step. For validation, we use specific metrics to evaluate the results of the proposed case studies. Activation-Based Grouping and Overlap Operators: Construction and Properties SWATI RANI HAIT and ARKA PRABHA DAS (BIRLA INSTITUTE OF TECHNOLOGY AND SCIENCE, PILANI, HYDERABAD CAMPUS) Abstract Abstract Grouping and overlap operators are widely exploited in aggregation theory and information fusion to model disjunctive and conjunctive characteristic among the data inputs. In this study, we propose a new class of fusion operators, termed as the activation-based grouping and overlap operators. These operators are framed by introducing a threshold-based activation mechanism into the classical ideology of general grouping and overlap operators. The threshold parameter $\tau$ controls the point at which the activation mechanism becomes functional in the aggregation process and is chosen according to the necessities of the decision system. The main idea of the proposed framework is to retain the original behavior of the grouping or overlap operator in the non-activated region, while modifying the response of output near the boundary values by using a suitable activation function. Several construction methods for such operators are discussed, and their mathematical properties are examined. The class of activation-based overlap operators have been studied as a negation of the class of activation-based grouping operators. The proposed framework offers a flexible way to design a class of fusion operators that are sensitive towards the threshold parameter chosen close to the boundary points, and can be efficiently adapted to different decision scenarios. Due to their adaptive mechanism, such operators can be utilized in a wide range of applications, including decision-making, information fusion, classification problems, and reliability analysis. Measure derivatives induced by the Maximal chain-based integral Yasuo Narukawa (Tamagawa University), Zuzana Ontkovicova (Slovak University of Technology), and Vicenc Torra (Umea University) Abstract Abstract Fuzzy measures and integrals are used in a wide range of applications related to decision making and information fusion. A key concept of measure theory is the notion of derivatives. For example, the Radon-Nikodym derivative between two additive measures is the function that, when integrated with respect to one measure, produces the second one. In this paper, we consider the generalized concept of derivatives for set-valued integrals. More particularly, we study the derivatives induced by a Maximal chain-based integral. Tuesday 0.02 Berlin FUZZ-IEEE Paper FUZZ 6 : FUZZ-IEEE SS04 Fuzzy Foundation Models & Main:Fuzzy system Session Chair: Javier Andreu-Perez (University of Essex), Hani Hagras (university of essex) An Explainable Fuzzy Based Approach for Embeddings in Generative AI Transformer Models Jera Makar, Hani Hagras, and Javier Andreu-Perez (University of Essex) and Thomas Golden and Scott Payton (Bowen Craggs & Co Ltd) Abstract Abstract Transformer models achieve remarkable performance across language, vision, and multimodal tasks, yet their internal representations remain largely opaque. In particular, the semantic structure of embedding spaces, central to how transformers operate, remains difficult to interpret. This paper introduces a framework that converts embedding vectors into interpretable, linguistically grounded fuzzy semantic representations. The approach integrates external lexical knowledge with data-driven analysis, using WordNet-derived semantic categories as reference signals. Dimensionality reduction first compresses the embedding space while preserving major variance patterns, followed by a fuzzy modelling layer that combines Fuzzy C-Means clustering with multiscale Gaussian membership functions. This produces a transparent semantic signature for each word, capturing graded associations with semantic concepts across multiple levels of abstraction. We evaluate the approach on six human similarity benchmarks (WordSim353, SimLex999, MEN, RareWords, MTurk, and SimVerb3500) across five transformer models (GPT-2, BERT, Gemini, Qwen3, and MiniLM). The proposed representations consistently achieve higher agreement with human judgments, with improvements of up to 448% in Spearman correlation, supported by paired bootstrap significance testing. These findings show that improving interpretability can also enhance alignment with human semantic intuition, supporting more trustworthy AI systems. A Fuzzy Attribution Based LoRA Fine-Tuning for Vision-Language Models within Telecom Inspections Mazen Ahmed and Hani Hagras (University of Essex), Hugo Leon-Garza (Brtitish Telecom), and Anasol Pena Rios (British Telecom) Abstract Abstract Automating the inspection of telecommunication Customer Service Point (CSP) boxes remains a challenge due to small-scale components, clutter and the subtle nature of installation errors. This paper introduces an explainable hybrid framework that leverages parameter-efficient fine-tuning of a Vision-Language Model (VLM) and a fuzzy attribution layer for interpretable audit decision-making. The proposed system adapts LLaMA 3.2 Vision-Instruct (11 B) model using Low-Rank Adaptation (LoRA) with 4-bit quantization to achieve domain-specific auditing of completeness and correctness across 37 controlled installation setups. Each setup is annotated with explicit Chain-of-Thought (CoT) reasoning, enabling the model to learn structured inspection logic. The fine-tuned model achieved a completeness F1-score of 0.86 and a correctness F1-score of 0.80, representing more than 60 % improvement over the zero-shot baseline (0.21 F1 for both). Moreover, BLEU-4 increased from 0.12 to 0.44, and BERTScore from 0.41 to 0.85, demonstrating substantial gains in reasoning fidelity and language–vision alignment. The fuzzy layer further enhanced interpretability by reducing attribution entropy by 34 %, fusing VLM confidence with rule-based linguistic reasoning to produce transparent audit verdicts (PASS, WARN, FAIL) and quantitative attribution scores reflecting each input’s influence on the final decision. The results confirm that the integration of parameter-efficient tuning and fuzzy interpretability enhances both scalability and transparency. A Neuro-Symbolic System for Interpretable Multimodal Physiological Signals Integration in Human Fatigue Detection Mohammadreza Jamalifard, Yaxiong Lei, Parasto Azizinezhad, Javier Fumanal Idocin, and Javier Andreu-Perez (University of Essex) Abstract Abstract We propose a neuro-symbolic architecture that learns four interpretable physiological concepts, oculomotor dynamics, gaze stability, prefrontal hemodynamics, and multimodal, from eye-tracking and neural hemodynamics, functional near-infrared spectroscopy, (fNIRS) windows using attention-based encoders, and combines them with differentiable approximate reasoning rules using learned weights and soft thresholds, to address both rigid hand-crafted rules and the lack of subject-level alignment diagnostics. We apply this system to fatigue classification from multimodal physiological signals, a domain that requires models that are accurate and interpretable, with internal reasoning that can be inspected for safety-critical use. In leave-one-subject-out evaluation on 18 participants (560 samples), the method achieves 72.1% +/- 12.3% accuracy, comparable to tuned baselines while exposing concept activations and rule firing strengths. Ablations indicate gains from participant-specific calibration (+5.2 pp), a modest drop without the fNIRS concept (-1.2 pp), and slightly better performance with Lukasiewicz operators than product (+0.9 pp). We also introduce concept fidelity, an offline per-subject audit metric from held-out labels, which correlates strongly with per-subject accuracy (r=0.843, p < 0.0001). Interpretable Fuzzy Rule-Based Regression Extension for Ex-Fuzzy Library Cayan Deniz Kucuktopana, Javier Andreu-Perez, Javier Fumanel Idocin, and Richard Pitts (University of Essex) Abstract Abstract Machine learning models achieve high predictive accuracy in regression tasks, but their deployment in safety-critical and regulated domains requires interpretability. While fuzzy rule-based systems offer transparent, linguistically explicit interpretable models, Mamdani-style fuzzy regression remains underrepresented in modern machine learning software libraries. This paper presents an interpretable regression extension for the Ex-Fuzzy library, enabling Mamdani fuzzy inference with scalar consequents learned directly from data. For this, a target-aware partition initialisation strategy based on Fuzzy C-Means clustering is introduced, in which linguistic variables are derived from an augmented input–output space to emphasize output-relevant regions of the feature space. The proposed extension is evaluated on ten regression datasets from the KEEL repository, comparing Gaussian and trapezoidal partition strategies against standard baselines including linear regression, multilayer perceptrons, and random forests. Experimental results show that Gaussian partitions consistently outperform uniform trapezoidal partitions, achieving a mean coefficient of determination of approximately 0.86 while producing compact rule bases of 10–15 human-readable rules. The proposed implementation provides a transparent and competitive alternative to black-box regression models, supporting practical interpretability with competitive predictive performance. Affine Transformation-Based Lightweight Random Vector Functional Link with Fuzzy Attention for Real-Time Concept Drift Handling Aanand Upadhayay (South Asian University); Amit Shukla (University of Vaasa, South Asian University); and Pranab K. Muhuri (South Asian University) Abstract Abstract Online stream data exhibit concept drift that make hard for models to continuously adapt to evolving data distributions. Although many approaches exist but most relay on costly updates, unbounded memory growth, and instability. Random Vector Functional-Link (RVFL) network fits well under this streaming scenario due to their computational efficiency through closed-form output-weight estimation. But their performance degrades in streaming classification because of activation saturation and ineffective drift-handling strategies. To confront these constraints, this paper proposes a lightweight RVFL framework with affine feature calibration to prevents activation saturation and a fuzzy-attention-based forgetting mechanism to adaptively handle drift. Affine calibration adopts a lightweight approach by taking the output of two randomly selected optimized hidden nodes rather than all of them. This incurs minimal overhead and sufficiently preserves feature diversity. Concurrently, the fuzzy attention module also uses a compact rule base with (low, medium, high) linguistic terms to map drift signals and adaptively modulate forgetting setup. Experimental results on both real-world sensor datasets confirms that the proposed framework attains outstanding classification accuracy with competitive runtime. Overall, it indicates that the proposed framework consistently outperforms state-of-the-art drift-aware learning models in real-time adaptation, robustness, and stability under concept drift. Tuesday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 11 - Optimization II Session Chair: Frank Neumann (Adelaide University) TC-MFEA: Task Clone-Based Multifactorial Evolutionary Algorithm for Constrained Reliability Redundancy Allocation Problem Md. Abdul Malek Chowdury (South Asian University); Rahul Nath (University of Bergen); Amit Rauniyar (South Asian University); Amit K. Shukla (University of Vaasa, South Asian University); and Pranab K. Muhuri (South Asian University) Abstract Abstract The Reliability Redundancy Allocation Problem (RRAP) is a challenging NP-hard optimization problem aimed at maximizing system reliability under resource constraints. While evolutionary algorithms perform well for single-task RRAP, evolutionary multitask optimization (EMTO) has recently shown promise in solving multiple RRAPs simultaneously. However, the constrained and nonlinear nature of RRAP often leads to premature convergence, necessitating effective constraint handling strategies. To address this, we propose a task-cloning based EMTO framework using the multi-factorial evolutionary algorithm (TC-MFEA). It formulates a primary constrained RRAP alongside an auxiliary unconstrained task generated via task cloning and treats them as joint optimization. It enables exploration and knowledge transfer, which improves solution quality for the primary task. The proposed method is validated on both series and bridge system RRAPs. Further, Taguchi method is used for the parameter tuning, and the experimental results demonstrate that TC-MFEA-RRAP outperforms existing EMTO approaches in achieving higher system reliability. Evolutionary Optimization of Hospital Bed Allocation to Reduce Mortality from Severe Acute Respiratory Syndrome Beatriz da Silva Vieira, Luiz Gustavo Almeida Martins, and Murillo Guimarães Carneiro (Federal University of Uberlândia) Abstract Abstract Hospital bed availability is a critical factor in reducing mortality from Severe Acute Respiratory Syndrome (SARI), particularly under high-demand conditions. This study presents a framework that integrates machine learning–based mortality prediction with a genetic algorithm (GA) to optimize hospital bed allocation using a multi-criteria scalar objective function. Tree-based models, especially the Random Forest Regressor, capture variability in SARI mortality under overdispersion and zero inflation. The GA explores trade-offs between mortality reduction, equity, and redistribution cost. Results show that performance depends on hospital occupancy. In high-occupancy scenarios (98%), small redistributions reduce mortality, whereas in lower occupancy scenarios (60% and 75%), gains are limited. These findings indicate greater effectiveness under high occupancy and limited impact in lower-demand scenarios. OMCDE: Opposition-based Multi-Stage Matrix-Based Competitive Differential Evolution Mohammad Sarhangzadeh and Doruk Oner (Bilkent University) and Shahryar Rahnamayan (Brock University) Abstract Abstract Differential Evolution has proven to be an effective algorithm for continuous global optimization, but it's modern high-performing variants often introduce additional complexity and computational cost. We propose OMCDE, a matrix-based competitive Differential Evolution algorithm that aims to improve optimization performance while maintaining computational efficiency. The proposed approach integrates competition-driven search dynamics, adaptive behavior between optimization stages, and opposition-based initialization within a parallelizable framework. Empirical results on widely used CEC 2017 functions show that OMCDE achieves strong performance compared to several state-of-the-art DE algorithms, while remaining efficient and scalable. These findings indicate that OMCDE is a promising alternative for solving complex numerical optimization problems. On the Use of Survival Selection Methods for Evolutionary Diversity Optimisation Adel Nikfarjam (Adelaide University), Jakob Bossek (Paderborn University), and Aneta Neumann and Frank Neumann (Adelaide University) Abstract Abstract Generating a diverse set of high quality solutions for an optimisation problem has been studied extensively in recent years by the evolutionary computation community. A paradigm that has received increasing attention is evolutionary diversity optimisation (EDO), where the goal is to maximise the diversity of a solution set subject to quality constraints. Since the contribution of each solution to the diversity of the population depends on other solutions and can change dramatically if several solutions in the population are modified simultaneously, most EDO approaches generate a single new solution per generation and discard the solution with the least contribution to diversity, ensuring a steady increase in population diversity over successive generations until convergence. In this study, we aim to answer two questions: (1) Is generating multiple solutions in each generation beneficial for EDO? (2) How can this be achieved efficiently, given that conventional survival selection methods do not work well in EDO due to the dependency of a solution's contribution to diversity on other solutions? We propose three survival selection methods. Empirical investigations show that the evolutionary algorithm (EA)-based selection method outperforms the other approaches, including the generation of a single solution per generation. Cross-Docking Operation Scheduling with Maximized Product Diversity Requirements: Model Formulation and First Results David Hutter, Thomas Steinberger, and Michael Hellwig (Vorarlberg University of Applied Sciences) Abstract Abstract Cross-docking is a logistics strategy in which in-bound goods are directly transferred to outbound shipments with little or no intermediate storage. The resulting scheduling problem is an NP-hard combinatorial optimization problem. Motivated by a real world usecase, an atypical variant of the cross-docking problem is considered in which, in addition to efficient scheduling the outbound shipments require highest possible product diversity. The paper presents a problem formulation as Mixed-Integer Linear Program (MILP) as well as a problem encoding suitable for evolutionary algorithms. It compares the performance of different solvers on example instances provided by the industrial partner. The methods under consideration comprise the MILP solver GUROBI and three evolutionary Algorithms, namely a Genetic Algorithm (GA) variant, the Covariance-Matrix Adaptation Evolution Strategy (CMA-ES), and the Matrix Adaptation Evolution-Strategy (MA-ES). The results on the problem instances show that all methods represent feasible approaches, i.e. none significantly dominates all other methods. MA-ES and GUROBI in particular show the most promising performance across all instances. Tuesday 0.10 Sydney IEEE CEC (Evolutionary Computation) CEC 12 - Related Topics II Session Chair: Andy Tyrrell (University of York) Exploring the Adjacent Possible in Engineering Design via Graph Operation Programs Babis Peteinarelis, Andy Tyrrell, and Simon Hickinbotham (University of York) Abstract Abstract Nature-inspired algorithms and mechanisms have been used for many years to assist the design of engineering structures. However, in most fitness landscapes in engineering, Evolutionary Algorithms converge and the result is accepted as-is. Design tasks on the other hand often prioritise novelty and diversity in solutions, hoping for innovation, not simply optimisation, needs that are satisfied by Generative Design methods usually at the cost of optimality, transparency or feasibility. Inspired by Kauffman’s Adjacent Possible, this paper attempts to reimagine the popular concept in contexts that evolve structural paradigms, aiming to combine the glass-box rationale of design with the generative power of Evolutionary Methods. The presented formulation uses the graph space to represent structures and the formalism of Graph Automata to express transition rules that can reach adjacent structures based on a set of elementary graph operations. Sequences of such graph operations (programs) are evolved to reveal a set of fit adjacent structures. This evolutionary logic is demonstrated on a Multi-Objective Topology Optimisation scenario where a seed truss structure experiences environmental forces. The evolved programs transform the seed structure into a variety of truss designs that respond to localised strain in different ways in order to address trade-offs between the objectives. The paper illustrates the method on truss (chassis) design and discovers a range of Pareto-optimal adjacent designs within a program length radius, maintaining the main design intent. Genetic Programming with Transformer-Based Mutation for Approximate Circuit Design Ondrej Galeta and Lukas Sekanina (Brno University of Technology) Abstract Abstract A recent trend is to leverage machine learning models to improve the evolutionary design and optimization process. We propose a novel transformer-based mutation operator for Cartesian genetic programming (CGP) for the automated design of approximate arithmetic circuits. We introduce a hybrid scheme for CGP in which the proposed mutation operator is switched with the standard mutation operator to prevent stagnation of the circuit approximation process. We also develop a new training scheme for the underlying transformer that utilizes training vectors composed of thousands of CGP chromosomes representing various approximate multipliers. For several target error constraints, the approximate multipliers evolved with CGP utilizing the transformer-based mutation achieve better trade-offs than the highly optimized designs available in the state-of-the-art EvoApproxLib library of approximate circuits. Although both training and evolutionary processes are computationally demanding, they appear to be necessary steps for improving existing approximate circuits and producing new, potentially patentable circuit designs. Adapted Genetic Algorithm for the Design Optimization of the Multi-Pulse Solid Rocket Motor Tiago Neves (Bayern-Chemie, Instituto Superior Técnico); Karim El Sioufy, Stephan Hase, Pedro C. Pinto, and Christoph Bauer (Bayern-Chemie); and João M. C. Sousa (IDMEC, Instituto Superior Técnico) Abstract Abstract Design optimization of solid rocket motors can be formulated as an optimization problem. The multi-pulse variant of this design offers many advantages compared to other variants, making it a valuable design choice. However, its design problem lacks a known formulation, to the best of our knowledge. Therefore, this paper proposes a mixed integer programming formulation for this problem. The formulation contains equality and inequality constraints, which bring forth many challenges, often causing results yielded by evolutionary algorithms to have feasibility issues, without specialized constraint handling. This paper proposes a constraint handling technique that is suitable for the design optimization problem of the multi-pulse solid rocket motor, and serves as the backbone for a novel gradient-based algorithm, the hybrid random restart hill climbing algorithm, to be compared with an adapted genetic algorithm. A brief overview of this problem is given in terms of modelling and simulation. On average across all seeds, the hybrid random restart hill climbing algorithm yields better solutions compared to the genetic algorithm. However, in terms of best across all seeds, the latter outperforms the former. This indicates that the solution space of the multi-pulse solid rocket motor design problem is nonconvex, thus justifying the usage of metaheuristic algorithms. A Co-Evolutionary Multi-Agent Path Planning Framework Using NURBS for Non-Holonomic Robots Elias J. R. Freitas (Institute Federal of Minas Gerais - IFMG, Universidade Federal de Minas Gerais - UFMG); Miri Weiss-Cohen (Braude College of Engineering); and Frederico G. Guimarães and Luciano C. A. Pimenta (Universidade Federal de Minas Gerais - UFMG) Abstract Abstract Multi-agent path planning in continuous spaces poses significant challenges due to inter-agent coupling, kinematic constraints, and temporal collision avoidance requirements. This paper proposes a cooperative co-evolutionary planning framework in which each agent evolves its own path and a constant tracking speed while implicitly coordinating with other agents through shared evolutionary information. Robot paths are represented using Non-Uniform Rational B-Splines, enabling smooth and kinematically feasible trajectories for non-holonomic robots. Simulation and real-world experiments demonstrate that the proposed approach reliably generates coordinated, collision-free solutions in complex and cluttered environments, achieving stable convergence within a small number of evolutionary epochs. The results further demonstrate safe and robust path execution under bounded disturbances, ensured by safety margins derived from the employed vector field. Spark-Enabled Binary Particle Swarm Feature Selection Algorithm Evaluated on Genomic Breast Cancer Data Simone Ludwig (North Dakota State University) Abstract Abstract In large-scale machine learning applications, feature selection is essential for reducing dimensionality and computational overhead while maintaining or improving classification performance. Feature selection aims to identify an optimal subset of features to improve classification performance, with evaluation criteria and search strategies forming its two fundamental components. This paper implements a Spark-Enabled Binary Particle Swarm Optimization (BPSO-P) approach applied to classification of Genomic Breast Cancer Data. BPSO-P aims to enhance classification accuracy while efficiently exploring the feature search space. The BPSO-P algorithm is implemented using Spark, leveraging distributed computing to handle large-scale datasets. The performance of the proposed approach is evaluated across multiple dataset sizes using speedup and scaleup as the primary performance metrics. All datasets consist of 5,000 features and the sizes vary between 6.2 GB to 31 GB. Minor Embedding for Max-Cut: Empirical Analysis of Quantum Annealing's Combinatorial Bottleneck Ron Lavi (Migal Institute); Alberto Moraglio (University of Exeter); and Ofer Shir (Tel-Hai College, Migal Institute) Abstract Abstract Quantum optimization by annealing constitutes a stochastic global search over QUBO-formulated combinatorial optimization problems. Its performance depends on pre-solving the Minor Embedding (ME) challenge - the NP-hard problem of mapping logical graphs representing QUBOs onto the quantum hardware. Despite determining solution feasibility and quality, ME remains understudied as a standalone optimization challenge. Current understanding is largely limited to synthetic graphs on isolated topologies, leaving the generalizability of structural predictors across hardware generations unclear. We conduct the first cross-topology study analyzing the MINORMINER heuristic - the de facto industry standard - over ~550 benchmark graph instances across D-Wave's Chimera, Pegasus, and Zephyr architectures. Our analyses reveal a functional dissociation between embedding feasibility and cost: topological independence (via Maximum Independent Set ratio) predicts feasibility by identifying structural `slack', while edge density - reflecting clique-like proximity - governs the resource `toll' (chain length). Since ME is a purely graph-theoretic problem, agnostic to the combinatorial origin of the source graph, these structural predictors apply to any QUBO whose logical graph falls within the studied structural regime. This dissociation remains invariant across all architectures, establishing ME as a benchmark for Evolutionary Computation, where the local distribution of edges, rather than global density alone, determines algorithmic difficulty. Tuesday 0.11 Cape Town IEEE CEC (Evolutionary Computation) CEC 13 - Algorithms III Session Chair: Diederick Vermetten (Sorbonne Université, CNRS, LIP6) Evaluating Multi-Objective Algorithms for 3D Basement Topography Inversion from Gravity Data Arthur Anthony da Cunha Romão e Silva, Bruno Motta de Carvalho, and Francisco Márcio Barboza (Federal University of Rio Grande do Norte) Abstract Abstract The inversion of basement topography from gravimetric data aims to determine the depths of subsurface density contrasts from gravitational anomalies measured by terrestrial or marine instruments. Such inverse problems are inherently ill-posed and highly sensitive to instability. To address these challenges, this study applies and compares three multi-objective optimization algorithms—MOWGN, NSGA-II, and MOPSO—to identify the most effective strategies for extracting geological information from real gravimetric observations. Each algorithm minimizes two competing objectives: the data misfit between observed and calculated anomalies, and a smoothness (SM) regularization term that stabilizes the inversion process. Implemented in MATLAB, the methods were applied to Bouguer anomaly data from the Recôncavo Basin, northeastern Brazil. The evaluation combined quantitative metrics—hypervolume, C-metric, and computational time—with qualitative analyses of Pareto front distributions, dominance regions, and the geological plausibility of the reconstructed models. Because the data distributions deviated from normality and exhibited heteroscedasticity, non-parametric statistical tests were used: the Mann–Whitney \textit{U} test for pairwise C-metric comparisons, and the Kruskal–Wallis test with Dunn–Bonferroni correction for multiple-group analyses. Overall, the study provides a comprehensive assessment of these algorithms, emphasizing their strengths, limitations, and practical relevance for real-world gravimetric inversion. E-CREM: A Specialized Evolutionary Algorithm for the Multiple Bay Container Relocation Problem Camila Diaz and Maria Cristina Riff (Universidad Técnica Federico Santa María) Abstract Abstract The Container Relocation Problem for multiple bays consists of finding the minimum number of moves to load a set of stacked containers on a ship according to a given loading sequence and minimizing the crane’s working time for a yard of multiple bays. In this paper, we propose an Evolutionary Metaheuristic Algorithm to solve this problem. We use a heuristic encoding representation that allows us to solve the problem of solution's reconstruction that arises when applying mutation and crossover operations on traditional movements representations. In order to validate our approach, we use a large set of well-known instances, as well as a statistical analysis of our results. Our experiments show that our approach obtains high quality solutions, minimizing both goals in a reduced computational time especially for large size instances. 2D Basement Relief Inversion: A Comparative Study Among Multi-Objective Algorithms Arthur Anthony da Cunha Romão e Silva, Bruno Motta de Carvalho, and Francisco Márcio Barboza (Federal University of Rio Grande do Norte) Abstract Abstract The gravimetric inversion of basement relief estimates the depth of density-contrast interfaces in the subsurface from measured gravitational anomalies. Such geophysical inverse problems are inherently ill-posed and unstable, requiring both regularization and efficient parameter estimation techniques. In this study, we apply and compare three multi-objective optimization algorithms — MOWGN, NSGA-II, and MOPSO — to minimize two conflicting objectives: (i) the misfit between observed and calculated gravity anomalies, and (ii) a smoothness regularization term (SM) to stabilize the inversion. The methods were implemented in MATLAB and tested on a real two-dimensional dataset from the Poem Bridge region. The comparison includes quantitative metrics — hypervolume, C-metric, and computational time — as well as qualitative assessments of Pareto front distribution, dominance, and smoothness of the estimated solutions. Statistical analyses were performed using the Mann–Whitney U test for pairwise comparisons and the Kruskal–Wallis test with Dunn–Bonferroni correction for multiple groups. The results reveal important differences in efficiency, robustness, and solution quality, providing practical insights into selecting the most suitable algorithm for real gravimetric inversion applications. Evolution Strategies' Genetic Operators Revisited Over Integer Search-Space Jacob de Nobel (Leiden University), Diederick Vermetten (Sorbonne Université), Ofer Shir (Tel-Hai College), Michael Emmerich (University of Jyväskylä), and Thomas Bäck (Leiden University) Abstract Abstract Evolution Strategies (ESs) are highly effective for continuous black-box optimization, yet their standard operators do not transfer cleanly to integer lattices: rounding introduces artifacts, step sizes have a natural lower bound, and recombination can break feasibility and distort selection. We empirically isolate how \emph{mutation distributions} and \emph{recombination operators} jointly shape the behavior of classical self-adaptive $(\mu,\lambda)$-ES variants on convex quadratic benchmarks over $\mathbb{Z}^n$, including $\ell_2$- and $\ell_1$-based landscapes up to a dimensionality of $n=100$, and we compare against an integer-handling CMA-ES baseline. Across all integer settings, Double Geometric (DG) mutation consistently dominates truncated-normal (rounded Gaussian) mutation and, when paired with uniform discrete recombination, often matches and sometimes surpasses the covariance-based baseline despite using only a single global step size. In contrast, intermediate recombination systematically degrades DG-based search, revealing a strong operator--domain mismatch. Finally, an ablation on \emph{zero mutations} shows that permitting null steps can improve DG-based ES performance, suggesting a step-size adaptation lag specific to lattice-native mutations. Importantly, several methods exhibit a pronounced “last-mile” slowdown on $\mathbb{Z}^n$, where late-stage progress reduces to rare single-coordinate corrections; We hypothetise that DG mutation with discrete recombination mitigates this effect and improves final-phase reliability. Replacing rounded-Gaussian mutation with lattice-native Double Geometric mutation and using uniform discrete recombination yields consistently strong results on quadratic $\mathbb{Z}^n$ benchmarks, often matching or surpassing integer-handling CMA-ES. Turning Parent Degradation into an Advantage: Archive-Directed Mating for Archive-Based Evolutionary Multi-Objective Optimization Kazuma Sato and Minami Miyakawa (The University of Electro-Communications), Keiki Takadama (The University of Tokyo), and Hiroyuki Sato (The University of Electro-Communications) Abstract Abstract In multi-objective evolutionary algorithms, archives are widely used to preserve non-dominated solutions during the search process. However, introducing an archive can make parent degradation more explicit, where the parent population is dominated by archive solutions, which has traditionally been regarded as an undesirable phenomenon. This degradation, however, implies a clear dominance relationship between parents and archive solutions and can be viewed as informative for search improvement. In this paper, we propose Archive-Directed Mating, which exploits parent degradation as a source of search guidance rather than suppressing it. When the first parent is dominated by archive solutions, the second parent is selected from the archive, directly incorporating information from existing high-quality solutions into mating. Experiments on multi-objective optimization problems with two to six objectives show that the proposed method consistently outperforms conventional mating and re-selection approaches in terms of Hypervolume, particularly in many-objective settings. Further analysis of dominance relations between offspring and archive solutions reveals that the archive-directed mating increases the number of promising offspring while suppressing poor ones. These results demonstrate that parent degradation can be effectively leveraged for search improvement, and that archive-directed mating provides a general and complementary strategy to existing mating and re-selection mechanisms. Proposal of Guidance Markers to Reduce the Time to Reach the Pareto-optimal Front Yuji Sato (Hosei University), Mikiko Sato (Tokai University), and Ruhul Sarker (University of New South Wales Canberra) Abstract Abstract Companies often develop products with similar trends, such as improving existing products or creating product series. This paper proposes using non-dominated solutions in an objective function space corresponding to the design variables of past similar products as a Guidance Marker (GUM) to rapidly guide a population of individuals toward their desired destination. More specifically, we propose a method that uses difference vectors derived from the GUM to correct the population's search positions as a supplement to the genetic operations of the underlying evolutionary multi-objective optimization algorithm. Using the standard benchmark problem, the ZDT series, and a Human-Powered Aircraft (HPA) problem, we compare the solution search accuracy of the original NSGA-II with that of Hyper Volume, IGD, and IGD+, as well as the convergence speed to the Pareto-optimal front (PF), and the process of the population's convergence to the PF. The results show that the proposed method improves convergence speed and solution search accuracy compared to the original NSGA-II, without reducing the diversity of the solution distribution. Tuesday 0.15 Washington IEEE CEC (Evolutionary Computation) CEC 14 - Evolutionary Machine Learning III Session Chair: Keiki Takadama (The University of Tokyo) CEGARE: A Coevolutionary Approach for Interpretable Rule Ensembles Simon De Lange and Matthias Bogaert (Ghent University, Flanders Make); Koen W. De Bock (Audencia Business School); and Dirk Van den Poel (Ghent University, Flanders Make) Abstract Abstract Interpretability is an essential aspect of machine learning for high-stakes decision-making. Rule ensembles, which combine simple decision rules in a linear model, offer a balance between the interpretability of an additive model and the flexibility to capture non-linear relationships and interactions of tree-based ensembles. However, current rule ensemble methods rely on a two-step approach that first generates a large set of rules from decision trees and then selects a subset of these rules using a regularized linear model. This decoupling can lead to suboptimal rule ensembles, as the rules are not optimized for their combined performance in the linear model, which can result in redundant rules in the final model. To address this limitation, we propose CEGARE, a CoEvolutionary Genetic Algorithm for Rule Ensembles, which simultaneously optimizes the rules and their coefficients in a bilevel optimization framework. By coevolving both components, CEGARE aims to discover more compact and relevant rule ensembles that achieve competitive predictive performance while providing a sparser model. We benchmark CEGARE against RuleFit and other models on multiple real-world datasets, demonstrating its ability to produce sparser rule ensembles without sacrificing predictive performance. Continuous Efficient Rank Evolution Strategy for Model Compression Cheng-Yu Sie, Ching-Chi Lee, Che-Rung Lee, and Chuan-Kang Ting (National Tsing Hua University) Abstract Abstract Model compression constitutes a complex optimization problem characterized by vast search spaces. While low-rank decomposition has been widely adopted for model compression, its performance is limited by the rank selection. Previous studies formulated rank selection as a one-shot neural architecture search (NAS) problem but suffered from high search costs due to full-dataset evaluations and limited optimization potential caused by the discrete rank representation. To address these issues, this study proposes the continuous efficient rank evolution strategy (CERES), which transforms rank selection into a continuous optimization problem to enable fine-grained search and utilizes a minibatch strategy for rapid evaluation. To mitigate the stochastic noise introduced by minibatch evaluation, CERES employs the covariance matrix adaptation evolution strategy (CMA-ES) as the core solver. Experimental results on ImageNet-1k demonstrate that CERES outperforms state-of-the-art techniques; specifically, CERES reduces search time by approximately 12 times while maintaining competitive accuracy. LLM-Guided Analogical Reasoning Particle Swarm Optimization for Neural Architecture Search Tianshui Li, Keiki Takadama, and Hitoshi Iba (The University of Tokyo) Abstract Abstract The integration of Large Language Models (LLMs) into Neural Architecture Search (NAS) has introduced a new paradigm that leverages LLMs' domain knowledge to generate valid candidate architectures efficiently. However, a key limitation of existing methods is that their mutation operations often lack explicit guidance from the evolutionary history. They are mainly driven by the LLM's internal pre-trained priors without explicitly analyzing the structural trends within the optimization landscape. This mismatch often leads to suggestions that are poorly aligned with the problem-specific search landscape and can reduce search efficiency. To address this issue, we propose LLM-Guided Analogical Reasoning Particle Swarm Optimization (LARPSO). The proposed framework integrates the learning dynamics of particle swarm optimization into LLM-driven mutation by providing the model with explicit improvement trajectories derived from the swarm search history. This mechanism encourages the LLM to infer structural improvement patterns and apply analogous transformations to the current architecture, which mitigates stagnation and promotes more directed exploration. We evaluate LARPSO on the NAS-Bench-201 benchmark across three datasets. Experimental results show that LARPSO consistently outperforms representative LLM-based NAS baselines and remains competitive with other state-of-the-art NAS methods. The source code of LARPSO is publicly available at https://github.com/LiTianshui/LARPSO. Evolutionary Malware Data Augmentation Callum Musselwhite (UCT); Bingle Kruger (Frankfurt School of Finance & Management, 60322 Frankfurt am Main); and Geoff Nitschke (UCT) Abstract Abstract Machine learning benefits malware detection and classification, yet scarce labelled data and rapidly evolving threats hinders progress. We investigate generative, evolutionary, and hybrid data augmentation in static malware detection given data scarcity and distributional shifts. We evaluate a generative adversarial network, evolutionary method, and hybrid method, with static features and Random Forest classifiers. Experiments across random scarcity, temporal holdout, and leave-one-familyout regimes (dataset treatments) indicate evolutionary augmentation provides consistent improvements if positive samples are extremely scarce and under temporal drift. Results also indicate the effects of leave-one-family-out evaluation are family dependent and as real data availability increases, augmentation benefits diminish, where training on real data alone becomes sufficient. Improving Mutation-Based Evolving Artificial Neural Network via Population Partitioning for Topology and Parameter Mutations Kenta Okumura and Motoaki Hiraga (Kyoto Institute of Technology), Daichi Morimoto (Kyushu Institute of Technology), Kazuhiro Ohkura (Hiroshima University), and Nanako Miura and Arata Masuda (Kyoto Institute of Technology) Abstract Abstract Topology and Weight Evolving Artificial Neural Networks (TWEANNs) can simultaneously optimize neural network structures and parameters. Mutation-Based Evolving Artificial Neural Network (MBEANN) is a TWEANN algorithm that uses only mutations for genetic variation. In conventional MBEANN, the rapid growth of network topology can outpace parameter optimization, leading to bloated networks and an inefficient search of network topologies. This study proposes a population partitioning strategy that divides a population into two subgroups. One subgroup focuses on parameter optimization, whereas the other focuses on topological exploration, thereby balancing the two search processes. Experimental results on MuJoCo locomotion benchmarks demonstrate that the proposed approach achieves higher fitness values than the conventional MBEANN algorithm. The Robustness-Generalization Dilemma: Trade-off via Multi-Objective Optimization Ran Wang, Haojie Zhai, and Wenhui Wu (Shenzhen University); Le Ou-Yang (Shenzhen MSU-BIT University); and Wing W. Y. Ng (South China University of Technology) Abstract Abstract Existing studies reveal an inherent conflict between the generalization ability and robustness of deep models, known as the robustness-generalization dilemma. Specifically, in the mid-to-late stages of model training, the improvement in robustness is often accompanied by a decrease in generalization ability, and vice versa. This dilemma has been extensively discussed for more than a decade but remains incompletely addressed, indicating the hardness of improving the two performance indicators simultaneously. Existing works usually purse one aspect at the expense of the other, e.g., by optimizing single objectives (such as adversarial loss or clean loss) or a scalar combination of them, causing difficulties to achieve an effective trade-off between them. In this paper, for the first time, we introduce the multi-objective optimization (MOO) technique to deal with the robustness-generalization dilemma, and propose the Robustness-Generalization Trade-off framework based on MOO (RGTro-MOO). The proposed RGTro-MOO is incorporated into adversarial training, which dynamically and stage-wisely optimizes the attack strategies to trade-off robustness and generalization. Non-dominated sorting genetic algorithm II is used to evolve and select the Pareto-optimal solutions, and Pareto front is visualized to intuitively show the dilemma. Experiments on benchmark image classification datasets validate the feasibility of applying MOO to trade-off robustness and generalization, and comparisons with some state-of-the-art methods demonstrate its superior performance. Tuesday 2.1 Volga IEEE CEC (Evolutionary Computation) CEC 15 - SS17:Computational Intelligence in Power Electrical Engineering Session Chair: Maurizio Repetto (Politecnico di Torino) Systematic Evaluation of Spatial Interpolation Methods for Reconstructing Missing Solar Irradiance Data for Energy Applications Lilla Barancsuk (HUN-REN Centre for Energy Research) and Paolo Lazzeroni (Politecnico di Torino Dipartimento Energia "Galileo Ferraris") Abstract Abstract This paper presents a systematic comparison of spatial interpolation methods for reconstructing missing solar irradiance data on quasi-regular latitude–longitude grids. Using PVGIS-derived datasets, we evaluate interpolation performance under realistic missing-data scenarios, including (M1) random removal, (M2) contiguous gaps, and (M3) regular-grid thinning. A reproducible evaluation pipeline is used to apply k-nearest neighbors, inverse distance weighting, and analytical grid-based methods (linear, nearest) across multiple regions, orientations, and hyperparameter settings. Reconstruction accuracy is assessed using standard pointwise error metrics. Across all scenarios, K-Nearest Neighbors consistently outperforms alternative methods, demonstrating lower reconstruction errors and higher explanatory power even at increased levels of data removal. The results highlight method-dependent performance trends and trade-offs under different missing-data patterns. The results provide practical guidance for interpolation method selection in solar resource assessment and renewable energy applications. Surrogate-Assisted Many-Objective Optimization for Traction Electric Motor Design Gianmarco Lorenti, Luigi Solimene, and Maurizio Repetto (Politecnico di Torino) Abstract Abstract This paper investigates surrogate-assisted many-objective optimization for traction electric motor design under computationally expensive multiphysics evaluations. The case study is a V-shaped interior permanent magnet motor described by eight design variables and assessed through electromagnetic, thermal, and structural metrics. We optimize four objectives (torque, power factor, copper mass, and magnet mass) under mechanical-stress and winding-temperature constraints. Starting from a public dataset, we train a multi-output multi-layer perceptron and embed it within NSGA-III to enable a large evaluation budget. Multiple optimization runs are used to create a best-so-far Pareto set, from which a compact, well-spread subset is selected via reference-direction proximity, which is then re-evaluated with the high-fidelity FEM workflow. Results show that, although FEM re-evaluation reveals some surrogate discrepancies, especially near constraint boundaries, the adopted workflow provides high-quality Pareto designs at a reduced computational cost. Nested-Adaptive Framework for Islanded DC Microgrid Secondary Control Over Wireless Links Mohammadamin Jarrahi and Kamyar Mehran (Queen Mary University of London) Abstract Abstract This paper proposes a nested-adaptive secondary control framework for islanded DC microgrids operating over wireless links with bounded delay and packet loss. The method combines five elements that are designed to work on different time scales: a distributed consensus secondary controller for voltage restoration and state of charge sharing, a hysteresis event trigger with minimum and maximum inter-event times to reduce transmissions and avoid event accumulation, a controller bank with eight prevalidated parameter modes, a lightweight regime classifier that proposes mode changes from recent electrical and communication features, and a supervisory gate that enforces dwell time, confidence checks, and rollback to a conservative safe mode. A fast Level-0 bounded trim is added to each converter reference at every control tick, with explicit magnitude and slew limits so its effect can be treated as a bounded input to the secondary loop. The approach is evaluated in experimental 380V DC microgrid hardware-in-the-loop (HIL) test-bench with wireless delays in the 10 to 200 ms range and packet loss up to 15\%. Across the tested scenarios, the proposed scheme reduces message rate by about 30\% to 50\% relative to periodic communication while maintaining improved voltage regulation and SoC sharing. GEvTO: A Hybrid Evolutionary Scheme for Large-Scale Multi-objective Binary Topology Optimization of Electromagnetic Devices Francesco Lucchini and Piergiorgio Alotto (University of Padova) Abstract Abstract This paper presents a hybrid multi-objective framework for large-scale topology optimization of electromagnetic devices characterized by binary decision variables. The proposed method, termed Gradient Evolutionary Topology Optimization, combines a differential evolution strategy with a local gradient-based search formulated as an integer linear programming problem. The evolutionary component is responsible for global exploration of the design space through an adaptive mutation operator, while a filtering strategy mitigates checkerboard patterns and improves topology connectivity. The local gradient search, activated at prescribed generations, refines the candidate solution. The effectiveness of the proposed approach is demonstrated through the constrained multi-objective topology optimization of the receiver ferrite plate in a wireless power transfer system. Numerical results demonstrate that GEvTO can generate a well-resolved Pareto front and produce manufacturable topologies. Multi-Objective Design and Optimization of Grid-Connected Photovoltaic Power Plants Considering Real-World Constraints Angelo Raeli and Marco Mussetta (Politecnico di Milano) Abstract Abstract The optimal design of utility-scale photovoltaic (PV) plants involves complex trade-offs between maximizing energy production and minimizing Capital Expenditures (CAPEX). This paper proposes the use and integration of a Multi-Objective Evolutionary Algorithm (MOEA) framework utilizing the well-known Non-Dominated Sorting Genetic Algorithm II (NSGA-II) to optimize the electrical and geometrical configuration of grid-connected PV systems with geographical and orographical contraints. An automated workflow is introduced, integrating MATLAB for the optimization logic, PVsyst for detailed carrier-grade energy simulations considering real-world environmental features and constraints, and AutoHotkey for interface automation. The methodology is applied to a real-world 26.6 MWp case study in Southern Italy. Two scenarios are analyzed: one reflecting historical market conditions and one reflecting modern pricing and bifacial technology trends. Results demonstrate that the proposed algorithm identifies configurations that dominate the actual project design in terms of Pareto optimality. Furthermore, the study highlights how recent market shifts significantly impact the optimal selection between monofacial and bifacial tracking systems. Tuesday Virtual Room 1 IJCNN Paper SS26 Brain Machine Intelligence: Models, Systems, and Translational Applications Session Chair: Dongrui Wu (Huazhong University of Science and Technology), Zhao Boshi (Beijing university of technology; The Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences) Unsupervised Alignment using Latent Dynamics in Spiking Neural Networks for Stable and Energy-Efficient iBCI Decoding Boshi Zhao (Beijing University of Technology; Center for Excellence in Brain Science and Intelligence Technology, State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Neuroscience, Chinese Academy of Sciences); Liyan Han (Center for Excellence in Brain Science and Intelligence Technology, State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Neuroscience, Chinese Academy of Sciences); Zhengtuo Zhao and Xue Li (Center for Excellence in Brain Science and Intelligence Technology, State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Neuroscience, Chinese Academy of Sciences; University of Chinese Academy of Sciences); Kebin Jia (Beijing University of Technology); and Tielin Zhang (Center for Excellence in Brain Science and Intelligence Technology, State Key Laboratory of Brain Cognition and Brain-inspired Intelligence Technology, Institute of Neuroscience, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences The Key Laboratory of Cognition and Decision Intelligence for Complex Systems, Institute of Automation,Chinese Academy of Sciences) Abstract Abstract Invasive brain–computer interface (iBCI) enables direct interaction between the brain and external devices by decoding neural signals. However, cross-day variations in neural signal distributions degrade decoding performance, requiring frequent recalibration with labeled data and limiting long-term usability. Moreover, traditional artificial neural networks (ANNs) used for decoding are energy-intensive, posing challenges for future edge computing and deployment in iBCI systems. To address these issues, we propose a stable and energy-efficient iBCI decoding framework based on a sequential variational autoencoder (seqVAE). Our approach combines multi-day pre-training with a lightweight bias-based unsupervised alignment module, enabling reliable cross-day generalization using only 60 s unlabeled neural data. We further develop an SNN-compatible implementation by incorporating LIF spiking neurons into all core modules to enable low-power inference. Evaluations on two non-human primate datasets show that our framework achieves more stable cross-day behavioral decoding than existing unsupervised baselines, and the spiking variant reduces the estimated computational energy by over 50 %, highlighting its potential for future neuromorphic iBCI applications. SMEEG: Spectral Mamba-Driven EEG Temporal and Frequency Representation Learning Bingru Lin, Hanran Lin, Bo Zheng, Shaojun Zhu, and Maonian Wu (Huzhou University, Zhejiang Province Key Laboratory of Smart Management and Application of Modern Agricultural Resources) Abstract Abstract Electroencephalogram (EEG) signals reflect the brain's functional state, providing a reliable physiological basis for applications such as disease classification, emotion detection and cognitive state assessment. However, EEG representation learning faces challenges such as low signal-to-noise ratio, inter-individual feature variability and long-sequence modeling. To address these critical challenges, this study proposes learning EEG temporal and frequency representations based on Spectral Mamba global driving (SMEEG). The SMEEG includes a temporal and frequency domain feature fusion architecture, which encodes the temporal-domain and frequency-domain signals independently, and achieves feature alignment through shared positional encoding and normalization, effectively alleviating the feature quality problem caused by low signal-to-noise ratio. Depthwise Separable Wavelet (DSW) analysis is introduced to process the multi-scale features of EEG. Through learnable wavelet filters and an inverse transform mechanism, the limitations of traditional convolution in processing EEG signals with large amplitude variations and high individual differences are effectively solved. The Spectral Mamba (SpeMamba) model is based on the linear complexity state-space model of Mamba, enabling efficient long-sequence modeling. Experiments on 5 EEG datasets show that SMEEG effectively learns temporal-frequency representations under a SpeMamba-driven self-supervised framework. The extracted features allow downstream classifiers to reach supervised-level accuracy, confirming their potential for generalizable brain-computer interface (BCI) and clinical EEG analysis. The code is available at https://github.com/huaidan666/SMEEG. An End-To-End Intelligent EEG Bad Channel Detection and Reconstruction System Using Deep Learning and Reservoir Computing Zhaojin Chen and Lijuan Duan (Beijing University of Technology), Changming Wang and Xixi Zhao (Capital Medical University), and Zijian Zhou (North China University of Science and Technology) Abstract Abstract Bad channels significantly compromise EEG analysis, yet existing solutions often struggle with generalization or rely on precise electrode coordinates that are frequently unavailable. This study proposes an automated framework that integrates deep learning and reservoir computing without requiring electrode coordinates. The system features a Bad Channel Long Short-Term Memory (BC-LSTM) network with multi-order differential features for precise anomaly detection, and a Correlation-Weighted Reservoir Computing (CWRC) module for signal reconstruction. We also introduce a detection-guided data selection strategy to ensure artifact-free training, alongside a Neural Electrophysiological Similarity-Driven Transfer Learning (NESTL) mechanism to address data scarcity. Experimental results across 19-, 32-, and 64-channel datasets demonstrate 92.4% detection accuracy and an 80% improvement in reconstruction quality over traditional interpolation. The proposed system provides an effective automated solution for EEG preprocessing, suitable for both clinical and research applications. Frequency-Aware Class-Incremental Learning for SSVEP-Based Brain–Computer Interfaces Jiayu An, Huanyu Wu, and Dongrui Wu (Huazhong University of Science and Technology) Abstract Abstract Steady-state visual evoked potential (SSVEP)–based brain–computer interfaces (BCIs) have achieved significant progress in recent years. However, most existing SSVEP classification methods are designed under offline settings with a fixed set of stimulus frequencies, limiting their flexibility in practical applications, where SSVEP systems are gradually extended from a small set of stimulus frequencies to denser multi-frequency paradigms. To address this limitation, this paper focuses on class-incremental learning (CIL) for SSVEP-based BCIs and proposes a unified framework to alleviate catastrophic forgetting. Specifically, we propose a frequency-aware rehearsal strategy that exploits the periodic structure of SSVEP signals to improve replay efficiency under limited memory, and feature compactness regularization that enforces intra-class feature consistency across incremental learning stages. Extensive experiments on two public SSVEP datasets demonstrated that the proposed method consistently outperforms representative baselines and effectively mitigates catastrophic forgetting. Tuesday Virtual Room 2 IJCNN Paper SS02 Advances in Machine Learning and Deep Learning for Hyperspectral Image Classification Session Chair: Hongmeng Lu (Xinjiang University), Yuchuan Chen (Changsha University of Science and Technology, School of Computer Science and Technology) CA-DEMAE: Content-Aware Diffusion-Based Masked Autoencoder for Hyperspectral Image Classification Hongguang Xiao, Yuchuan Chen, Ziping He, and Junjie Li (Changsha University of Science and Technology, School of Computer Science and Technology) Abstract Abstract Hyperspectral image (HSI) classification serves as a pivotal task in remote sensing, playing a critical role in precise land cover identification and environmental monitoring. However, due to the scarcity of annotated HSI samples and the inherent spectral redundancy and noise of HSI, extracting robust and discriminative features with limited label support remains a significant challenge in this field. Meanwhile, existing masked self-supervised learning methods typically treat all spectral bands equally, ignoring their varying discriminative power, and rely on fixed-scale reconstruction, which limits their ability to capture complex spatial structures. This paper proposes a content-aware diffusion-based masked autoencoder (CA-DEMAE) framework for HSI classification, which restores fine-grained spatial structures and extracts discriminative spectral features by adaptively reassembling multi-scale spatial details and dynamically calibrating spectral importance weights. This method integrates a multi-scale feature fusion strategy into the content-aware feature reassembly mechanism, the multi-scale content-aware reassembly (MSCAR) mechanism that effectively harmonizes local details and global contextual information. Furthermore, the spectral-weighted masking (SpWM) module is proposed to address the neglect of varying spectral contributions, achieving dynamic calibration of band importance and enhanced spectral discriminability. Experimental results on three public benchmark hyperspectral datasets demonstrate that the proposed method exhibits superior performance. DSFMNet: A Dual-Scale Frequency-Modulated Network With a Linear-Complexity Hybrid Architecture for Hyperspectral Image Classification Junjie Li, Ziping He, and Ningjuan Wang (Changsha University of Science and Technology, the School of Computer Science and Technology); Baokai Zu (Beijing University of Technology, the College of Computer Science); and Hongguang Xiao and Yuchuan Chen (Changsha University of Science and Technology, the School of Computer Science and Technology) Abstract Abstract Hyperspectral image (HSI) classification demands efficient modeling of spatial–spectral–frequency information, but existing methods often struggle to balance accuracy and computational cost. In this paper, we propose DSFMNet, a dual-scale frequency-modulated linear-complexity network that comprises a fine-scale branch and a coarse-scale branch. Specifically, for the fine-scale branch, we introduce the Cross-Flow Frequency-Modulated Vision Mamba (CF-FMViM) module, which integrates state space duality (SSD)-based hidden state mixing with local frequency modulation to preserve detailed spatial–spectral structures. Complementarily, for the coarse-scale branch, we design the Frequency-Modulated Token-Selective Statistical Attention (FM-TSSA) module, which combines token-selective statistical attention with frequency modulation to enable efficient global semantic modeling, reducing the complexity from quadratic to linear in the number of tokens. The fusion of these two branches enables synergistic optimization across spatial, spectral, and frequency domains. Extensive experiments on the Houston2013, Whu-Hi-HongHu, and Houston2018 datasets demonstrate that DSFMNet consistently outperforms state-of-the-art methods while significantly reducing computational costs. MEQNet: Measurement-Efficient Quantum-Inspired Network for Hyperspectral Image Classification Yuren Chen, Juan Xu, Sunqi Cai, Mingyuan Pang, and Jiawei Xu (College of Computer Science and Technology, Nanjing University of Aeronautics and Astronautics) Abstract Abstract Hyperspectral image (HSI) classification leverages the high-dimensional spectral information per pixel to achieve fine-grained identification of objects and land cover. However, its performance is significantly degraded by band redundancy, spectral variability, and class imbalance. Existing deep learning methods often struggle to simultaneously achieve robustness, computational efficiency, and effective suppression of spectral redundancy and variability. Although quantum-inspired learning introduces complex-valued state representations to enhance modeling capacity, it remains sensitive to imbalanced data distributions. To address these challenges, this paper proposes a Measurement-Efficient Quantum-inspired Network, denoted as MEQNet, for robust HSI classification. MEQNet consists of a spectral projection and spatial refinement (SPSR) module and an adaptive complementary measurement (ACM) module to improve numerical stability and reduce computational overhead in quantum-inspired feature fusion. Specifically, SPSR uses a lightweight feature stem to compress the spectral dimension and refine spatial features, reducing spectral redundancy and stabilizing complex-valued state construction. ACM integrates the global low-rank measurement (GLRM) and the local operator measurement (LOM) with a gating-based adaptive selection mechanism, to enhance global representation under class-imbalanced conditions while maintaining lightweight inference under benign scenarios. Experimental results on three HSI datasets demonstrate the superiority of MEQNet over state-of-the-art methods. Dynamic Dependency Convolution and Hybrid Attention Transformer Network for Hyperspectral Image Classification Hongmeng Lu, Liejun Wang, and Shaochen Jiang (Xinjiang University) Abstract Abstract Recently, Convolutional Neural Networks (CNNs) and Transformer have gradually become mainstream methods for hyperspectral image (HSI) classification tasks. However, due to the complex land cover scenes and high-dimensional spectral characteristics of HSI, existing methods are often disturbed by redundant information during the feature learning process, which increases the difficulty of effective feature extraction. To address this challenge, this paper proposes a network based on dynamic dependency convolution and hybrid attention Transformer (DDCHAT). Specifically, the model designs a multi-scale dynamic dependency convolution (MDDC)module to enhance the ability to model local spatial details, while incorporating a hybrid attention Transformer (HAT) module to capture global spectral-spatial dependencies. This enables the collaborative effect of dynamic convolution operators and key discriminative information enhancement strategies. Experimental results on three publicly available datasets demonstrate that the proposed DDCHAT achieves high classification accuracy with relatively low parameters and computational cost. Tuesday Virtual Room 3 IJCNN Paper SS20 Deep Edge Intelligence Session Chair: Xujiang Tang (guilin university of electronic technology), Hui Zhang (Tianjin University of Science and Technology) Lightweight Multi-Scale Cross-Fusion Network Model for Infrared Small Target Detection Zengpei Zhang, Xuekui Zhang, Min Li, Ruixuan Zhang, Jing Liu, Dexin Zhang, and Hui Zhang (Tianjin University of Science and Technology) Abstract Abstract Infrared small target detection (IRSTD) is challenging due to extremely small target sizes, low signal-to-noise ratios, and strong background clutter. Existing deep models often improve accuracy at the cost of prohibitive computation, whereas lightweight designs tend to lose weak target cues during downsampling. This paper presents LWMCNet, a lightweight multi-scale cross-fusion network for IRSTD. LWMCNet integrates a Lightweight Multi-Scale Feature Attention (LMFA) module with a Cascaded Feature Enhancement (CFE) strategy to preserve weak target information, and employs Residual Feature Fusion (RFF) together with Adaptive Fusion Upsampling (AFU) to strengthen cross-scale interaction and decoder reconstruction. On the IRSTD-1K dataset, LWMCNet achieves 70.25% IoU with only 0.21M parameters and 0.44G FLOPs at 256 × 256 input resolution. Moreover, it runs at 29.70 FPS on an NVIDIA Jetson NX edge platform, demonstrating its practicality for resource-constrained real-time deployment. Efficient Pest Detection via Low-Rank Additive Attention and Scale-Dynamic Loss for Deep Edge Intelligence in Smart Agriculture Zhiqiang Chen, Li Li, Shuai Zhou, Zhangjun Peng, Guoqiang Zheng, and Chuanhao Chang (Southwest University of Science and Technology) Abstract Abstract Addressing the challenges of feature blurring in sub-pixel objects and limited computing power of edge devices in field pest detection, this paper proposes APLS-YOLO, a lightweight detection framework oriented towards Deep Edge Intelligence. First, targeting the slender geometric features of pests and the need for effective receptive fields, we design the Asymmetric-Priority Pinwheel Convolution (AP-PConv). By approximating the large kernel receptive field through orthog- onal strip convolutions, it significantly reduces computational redundancy while suppressing background noise. Second, to resolve the high-entropy feature blurring problem of traditional Softmax attention, we propose Low-Rank Additive Attention (LR-AAtten). This module utilizes an Additive Token Mixing strategy to replace complex matrix multiplication, effectively preserving the texture polarity features of tiny pests with O(N) complexity. Furthermore, a Scale-based Dynamic Loss (SD-Loss) is introduced to solve the gradient collapse problem of small targets during training via an adaptive re-weighting mechanism. Finally, combined with Global Channel Pruning based on BN layer scaling factors, extreme model compression is achieved. On the IP102 dataset, APLS-YOLO achieves an mAP@0.5 of 61.80%, a 2.45% improvement over the YOLOv11n baseline. On the more challenging FieldPest-25 dense tiny pest dataset, the mAP stabilizes at 60.31%, demonstrating strong robustness. On this dataset, the model parameters are only 2.18 M (17% reduction from baseline), and the model size is compressed to 3.88 MB. It achieves a real-time inference speed of 79 FPS on the RK3588 edge platform, showing exceptional potential for edge- side deployment PrismKV: Harmonizing Heterogeneous Compression Strategies through Layer-adaptive Partition Scoring Dongcheng Shi, Yongxiang Cao, and Hongxu Jiang (Beihang University) Abstract Abstract Large Language Models (LLMs) encounter significant scalability bottlenecks in long-context scenarios due to the prohibitive memory overhead of the Key-Value (KV) cache. Current optimization paradigms often overlook the hierarchical heterogeneity of Transformer layers, leading to suboptimal resource allocation. To address this, we propose PrismKV, a novel KV cache compression framework that leverages adaptive partition scoring to orchestrate heterogeneous compression strategies. PrismKV employs an unsupervised, layer-adaptive scoring function that jointly evaluates compositional redundancy and attention sensitivity to dynamically bifurcate Transformer layers into distinct operational zones. For shallow regions characterized by high directional alignment, we implement a reconstructible codebook-based replacement strategy to achieve high-ratio structured compression by clustering redundant vectors. Conversely, for deep regions exhibiting sparse, task-specific attention, we transition to an importance-aware dynamic eviction policy that quantifies spatiotemporal saliency for precise information retention. Experimental results across three prevailing LLMs demonstrate that PrismKV reduces the KV cache to 10.8% of its original size with a marginal performance degradation of only 1.9%. Our approach consistently outperforms existing baselines in long-context benchmarks and achieves 65.8% decoding acceleration in 32K context scenarios. By precisely aligning heterogeneous compression primitives with the intrinsic evolution of layer-wise representations, PrismKV achieves a superior balance between hardware-level inference efficiency and high-fidelity generation. SOAR: Scene-Aware Adaptive Resolution for Efficient UAV Multi-Object Tracking Yutao Tang, Xujiang Tang, and Lingying Zhao (guilin university of electronic technology) Abstract Abstract Deploying Multi-Object Tracking (MOT) on resource-constrained edge devices is challenging due to the high computational cost of real-time video stream processing. While scene complexity and target dynamics vary significantly over time, traditional MOT systems rigidly operate at fixed frame rates, which inflexibility leads to substantial computational waste in simple scenes and potential latency in complex ones. In this paper, we propose \sys, a reinforcement learning-based framework that dynamically adjusts frame rates to optimize the trade-off between tracking performance and computational efficiency. Unlike prior adaptive approaches that focus on resolution tuning or require expensive online profiling, \sys formulates frame rate control as a Markov Decision Process (MDP). This formulation allows the agent to learn policies that account for the temporal dependencies of tracking, balancing immediate efficiency gains against long-term tracking stability. To enable effective training and address the challenge of sample scarcity, we introduce a multi-variant strategy generation method that systematically constructs diverse frame rate scheduling policies. Extensive experiments on MOT17 and MOT20 demonstrate that \sys achieves 28.53% and 54.44% computational savings with negligible MOTA (MOT Accuracy) degradation (1.7% and 0.1%, respectively), significantly outperforming both fixed frame rate baselines and rule-based adaptive strategies. Tuesday Virtual Room 4 IJCNN Paper SS30 Computational Intelligence and AI Applications for Sustainable Energy Management in Smart Grids and Energy Communities (2nd ed.) Session Chair: Xiaolu Xu (闽南师范大学), Yuan Qiu (SIT) Self-periodicity and environment based recurrent imputation for PV production data Xiaolu Xu (Minnan Normal University) and Muqing GE (the Hong Kong Polytechnic University) Abstract Abstract Missing measurement of PV production data is a common problem due to sensor breakdowns, server data corruption, and so on. Many imputation methods have been proposed, RNNs based methods are very popular among these methods due to the ability to capture temporal dependency, but they ignore periodicity features. Considering PV production data has a strong periodicity, we proposed an imputation method called SPERI, which combines Seasonal Trend decomposition using Loess (STL) with RNNs. Additionally, we adapt an encoder-decoder structure to introduce environment data. Specially, STL is used to compose periodicity features (such as daily and annual periods), and then the residual component is processed by RNNs. We evaluated our SPERI method on real-world datasets in point missing mode and consective missing mode, and SPERI achieves the highest scores in general. MS-ECFormer: A Multi-Scale Endogenous–Exogenous Coupling Transformer for Wind Power Forecasting Yuchi Liu (South-Central Minzu University) and Zheng Ye (Cyberspace Security University of China) Abstract Abstract Accurate wind power forecasting is critical for the reliable integration and operation of renewable energy systems, yet it remains challenging due to strong non-stationarity in wind dynamics and complex interactions between historical power outputs and meteorological variables. Wind power generation exhibits multi-scale temporal patterns ranging from short-term fluctuations to long-term trends. However, existing deep learning–based forecasting models often struggle to capture such multi-scale dependencies while effectively modeling the coupling between endogenous variables (historical power) and exogenous variables (meteorological factors).To address these challenges, we propose MS-ECFormer, a Multi-Scale Endogenous–Exogenous Coupling Transformer for wind power forecasting. MS-ECFormer learns multi-scale temporal dynamics by decoupling endogenous and exogenous signals and coupling them through structured cross-attention fusion. Specifically, a block-based encoder with a learnable global token captures both local fluctuations and global trends, while a multi-scale decomposition module extracts trend and seasonal components from key meteorological inputs such as wind speed and temperature. Furthermore, a dual cross-attention mechanism enables hierarchical interactions between endogenous and exogenous representations, enhancing representation capability. Extensive experiments on real-world datasets demonstrate that MS-ECFormer consistently outperforms state-of-the-art baselines in both predictive accuracy and stability. A Novel Bi-Level Framework for Electric Vehicle Charging Scheduling Using Reinforcement Learning and Transfer Learning Abdennour Azerine, Mahmoud Golabi, and Lhassane Idoumghar (IRIMAS, Université de Haute-Alsace) Abstract Abstract The growing adoption of electric vehicles (EVs) presents significant challenges for efficiently managing charging infrastructure, a critical component in the transition toward sustainable transportation. This study addresses the Electric Vehicle Charging Scheduling Problem (EVCSP), a computationally complex optimization problem aimed at maximizing the number of satisfied charging demands while adhering to grid and infrastructure constraints. To address this challenge, a novel hybrid bi-level optimization framework is proposed. At the upper level, advanced reinforcement learning (RL) methods, specifically Q-learning, Double Q-learning, SARSA, and Expected SARSA, are employed to generate adaptive EV-to-charger assignment policies by iteratively exploring and learning from decision environments. These assignments, which may not fully satisfy operational constraints, are subsequently refined at the lower level through a mathematical programming model that optimizes energy allocation under realistic grid capacity and charger availability constraints. Additionally, transfer learning is incorporated to enhance computational efficiency by transferring knowledge gained during the RL-based assignment process to the optimization phase. Extensive experimental evaluations demonstrate the framework’s scalability and effectiveness, with SARSA combined with transfer learning emerging as the most reliable method, particularly for larger and more complex problem instances. Interpretable Battery Aging without Extra Tests via Neural-Assisted Physics-based Modelling Yuan Qiu, Wei Li, Wei Zhang, and Yi Zhou (SIT); Fang Liu (SUSS); and Jianbiao Wang and Zhi Wei Seh (A*STAR) Abstract Abstract State of health (SoH) is widely used for battery management, but it is a single scalar and offers limited interpretability. Two batteries with similar SoH can exhibit very different degradation behaviors and the lack of interpretability hinders optimal battery operation. In this paper, we propose IBAM for interpretable battery aging modelling with a neural-assisted physics-based framework. IBAM outputs a 2-D aging fingerprint without extra diagnostic tests and uses only routine logs from the battery management system. The fingerprint offers great interpretability by capturing a battery's curve-wide polarization voltage loss and the tail loss near the end-of-discharge. IBAM first creates a physics-based battery model based on a fractional-order equivalent-circuit model, and then extracts per-cycle fingerprints from the model using a two-stage least-squares method. IBAM further anchors fingerprints on the SoH axis with physics-guided regression, where the per-cycle SoH is estimated via a bidirectional gated recurrent unit with customized multi-channel voltage features. Across batteries with short-, medium-, and long-lifespans, IBAM consistently yields the best physics-model fidelity at different aging stages, and provides clear interpretations of degradation mechanisms and fingerprint patterns about batteries of different lifespans. The resulting fingerprints support interpretable battery health assessment and can inform battery control choices. Tuesday Virtual Room 5 IJCNN Paper SS33 AICS: Artificial Intelligence for Complex Systems Session Chair: zeyu zhang (Inner Mongolia Normal University), Peiwen He (South China Normal University) DualCrossAD: Multivariate Time Series Anomaly Detection with Homogeneous-Heterogeneous Dual-Cross Attention zeyu zhang and gaizhi Guo (Inner Mongolia Normal University) Abstract Abstract Abstract—In recent years, research on multivariate time series anomaly detection has increasingly focused on leveraging Transformer models to address the challenges of long-range dependency modeling and complex variable interactions. However,conventional Transformer architectures struggle to simultaneously capture temporal patterns and variable-wise relationships within a unified representation space, which limits their expressive capacity. To overcome this, a dual-branch Transformer model was designed that generates temporally dominated and variable-dominated representations in parallel. Yet an effective fusion of these two homogeneous-but heterogeneous representations remains challenging: single-path fusion often suffers from insufficient interaction and static combination, while multi-path fusion can lead to coarse-grained interactions and isolated feature flows.Therefore, we propose a dual cross-attention mechanism to achieve deep fusion of homogeneous-heterogeneous data, and introduce a symmetric KL divergence term to mitigate potential distributional inconsistencies during fusion.Experiments conducted on five benchmark datasets validate the effectiveness of the proposed method. Two-Level Routed Sparse Mixture-of-Experts with Spatiotemporal Self-Attention for Next POI Recommendation Qiang Li, Lei Zhang, and Bailong Liu (China University of Mining and Technology) and Fuyang Zhang (Beijing University of Posts and Telecommunications) Abstract Abstract Next Point-Of-Interest (POI) recommendation predicts a user's next visited location from historical check-in sequences with spatio-temporal context, and serves as a core task in location-based services and urban computing. However, large-scale irregular check-ins still exhibit multi-scale spatio-temporal dependencies and heterogeneous long-tail mobility patterns, making dense shared Feed-Forward Networks (FFNs) prone to pattern entanglement under limited compute. To address these issues, we present \textbf{BiRtSTMoE}, which augments the two-stage spatio-temporal attention framework as follows: (i) We design multi-channel, multi-scale relation biases and inject stage-appropriate biases into the two attention stages: temporal, spatial, and semantic biases for history aggregation, and candidate-history spatial and semantic biases for candidate matching. (ii) We replace the dense FFN in each Transformer block with a two-level routed sparse Mixture-of-Experts (MoE) feed-forward network. A mechanism gate performs coarse routing based on transition dynamics and periodic rhythm features to select an expert group, and a scene gate performs fine-grained routing within the selected mechanism using POI scene semantics to sparsely activate sub-experts. We add hierarchical load-balancing regularization to stabilize training and avoid expert collapse. Experiments on common LBSN benchmarks (NYC and TKY) show consistent gains over strong baselines, and improved robustness on long-tail time-gap/distance slices; routing analyses also support interpretability. Multi-Type Incremental Local-to-Global Context Encoder for Video Boundary Detection Xiaoxuan Tian, Ruifan Zhao, Zhilong Ou, and Hongxing Wang (Chongqing University) Abstract Abstract Video boundary detection (VBD) aims to identify salient temporal transitions in videos, such as event and scene changes. Learning a shared model that can handle both boundary types remains challenging. To address this problem, we introduce a multi-type video boundary detection framework. The framework adopts a Local-to-Global Context Encoder as a shared backbone to jointly capture local dynamics and global contextual dependencies for both event and scene boundaries. To train the proposed framework, we design an incremental learning strategy that integrates a teacher–student paradigm, experience replay, and feature alignment to preserve prior knowledge while adapting to new data. Experimental evaluations validate that the proposed method generalizes well across event and scene boundary detection tasks, providing an effective solution for multi-type video boundary detection. Code is available at https://github.com/Xxxseventea/GLM. MC-SQL: A Multi-Agent Collaborative and Multi-Reasoning-Chain Framework for Text-to-SQL Peiwen He, Mingkai Liu, and Mengchi Liu (South China Normal University) Abstract Abstract In recent years, large language models (LLMs) have been increasingly adopted for Text-to-SQL tasks due to their strong emergent capabilities. However, most existing studies rely on single-strategy generation paradigms, which limits their ability to exploit the complementary strengths of different reasoning strategies, particularly in complex query scenarios. To address this limitation, this paper proposes MC-SQL, a Text-to-SQL framework that integrates multi-agent collaboration with multi–chain-of-thought generation. It consists of four cooperative agents: a Translator that alleviates column semantic ambiguity, a Pruner that compresses the database schema, a Generator that produces diverse SQL candidates through three distinct chains of thought, and a Post-processor that refines SQL based on execution feedback and selects the final output. By combining multi-strategy reasoning and agent collaboration, MC-SQL significantly improves the accuracy and robustness of SQL generation. Experimental results demonstrate that MC-SQL outperforms most single-strategy models across multiple benchmark datasets. Tuesday Virtual Room 6 IJCNN Paper SS34 Safety integration in neural networks for sensor systems in critical applications Session Chair: Miaoji Zheng (Foshan University), Yongtao Zhang (NPIC) DerNeXt: A Transformer-based Adversarial Robustness Enhanced Model for Palm Vein Recognition Miaoji Zheng, Yong Zhong, Fen Liu, Qiyuan Sun, Tufeng Xian, Weidong Wu, and Zhiliang Zhang (Foshan University) and Haitao Huang (Guangdong Chankong Big Data Technology Company) Abstract Abstract Palm vein biometrics has gained considerable attention due to its advantages of being difficult to forge, highly private, and well-accepted by users. With the success of deep neural networks in visual tasks, deep learning-based palm vein recognition models have also achieved remarkable performance. However, like many visual models, existing deep learning-based palm vein recognition systems are highly vulnerable to adversarial attacks. Attackers can cause a sharp decline in model performance by adding carefully designed subtle perturbations to input images. To mitigate this issue, we propose a novel adversarial robustness-enhanced palm vein recognition model named DerNeXt. In this study, we firstly introduce the TrsanNext architecture to the field of palm vein recognition. To further enhance the stability of the input distribution to the channel mixer in the above architecture, we propose NGLU, a normalized gated linear unit. We also design DEM, a defensive efficient attention module, placed at the end of the network to ensure the quality stability and defensive capability of palm vein features. Finally, we employ an orthogonal classifier with a dense orthogonal weight matrix to refine decision boundaries, combined with a newly designed orthogonal discriminant loss to enhance adversarial robustness while maintaining high identification accuracy. We evaluate DerNeXt on three public palm vein datasets in terms of identification accuracy, equal error rate, and white-box attack experiments. The results demonstrate that DerNeXt achieves state-of-the-art recognition performance while exhibiting superior defense capability against adversarial attacks compared to other baseline models. RT-CAM Enhanced: A Comparative Study of Multi-Task Brain Tumor Classification Using Explainable Attention Mechanisms Lissir Amani (Data Engineering and Semantics Unit Faculty of Sciences of Sfax Sfax), Rafik Khemakhem (Advanced Technologies for Medicine and Signals Laboratory (ATMS) National Engineering School of Sfax (ENIS) Sfax), and Tarek Frikha (Data Engineering and Semantics Unit Faculty of Sciences of Sfax Sfax) Abstract Abstract Brain tumor classification from magnetic resonance imaging (MRI) remains a challenging task due to the need for both high diagnostic accuracy and reliable interpretability to support clinical decision-making. This paper presents a sys tematic comparative study of three deep learning architectures for multi-task brain tumor classification and anatomical plane identification using the BRISC2025 dataset. We propose RT CAM Enhanced, a novel reliability-aware explainable artificial intelligence (XAI) framework that introduces an adaptive Trust Map fusion strategy. By dynamically weighting Grad-CAM, Grad-CAM++, and Layer-CAM, the framework generates robust and consistent visual explanations across different convolutional neural network architectures. The evaluated models include an EfficientNet-B0(DualAttention), EfficientNet-B0 and ConvNeXt Tiny(DualAttention). Experimental results demonstrate that attention-enhanced architectures outperform their non-attention counterparts, with EfficientNet-B0(DualAttention) achieving clas sification accuracies of 99.60% for tumor classification and 99.50% for anatomical plane classification. Furthermore, RT CAM Enhanced improves explainability quality in terms of faithfulness, localization accuracy, and consistency compared to individual CAM-based methods, effectively addressing the faithfulness-localization trade-off. These findings confirm that the proposed framework balances high predictive performance with interpretable and reliable visual explanations, supporting trustworthy automated brain tumor analysis. Index Terms—Brain Tumor Classification, Explainable AI, Class Activation Mapping, Multi-Task Learning, Attention Mech anisms, Comparative Architecture Study PhysDiff: Physics-Guided Diffusion for Manifold-Consistent Data Augmentation in Driving Style Classification BingHeng Wu and Bin Cai (Chongqing University) and Ruyi Tang and Dafei Huang (SERES Automobile Co., Ltd.) Abstract Abstract Driving Style Classification (DSC) is a pivotal component of Intelligent Transportation Systems; however, its real-world deployment is persistently hindered by the long-tail distribution of dangerous driving behaviors. While synthetic oversampling techniques like SMOTE effectively rebalance class distributions, they rely on deterministic linear interpolation that ignores the underlying physics of vehicle motion. We identify that this ”kinematic blindness” leads to Manifold Collapse, where synthetic trajectories drift into physically forbidden regions characterized by instantaneous acceleration jumps. To address this, we propose PhysDiff, a physics-guided diffusion framework designed for manifold-consistent data augmentation. Unlike traditional interpolation, PhysDiff models the minority class distribution as a probabilistic manifold, generating diverse samples via a conditional 1D-UNet. Crucially, we introduce a novel Coupled Kinematic Constraint into the reverse diffusion process. This mechanism explicitly enforces differential consistency between velocity and acceleration channels while penalizing unrealistic jerk energy to ensure rigorous physical plausibility. Empirical evaluations using repeated cross-validation on a realworld dataset demonstrate that PhysDiff achieves a superior Macro F1-score of 0.8156 (±0.0517), significantly outperforming state-of-the-art baselines (p < 0.05). Furthermore, our results show that strictly enforcing physical constraints acts as a robust regularizer rather than a hindrance. PhysDiff reduces the rate of physical limit violations by 64.5% compared to inherent sensor noise, ensuring that augmented data maintains both high discriminative utility and kinematic fidelity. KADF: Knowledge-Guided Aligned Multi-Source Data Fusion for ICS Anomaly Detection Yanqun Wu (College of Computer Science, Sichuan University; National Key Laboratory of Nuclear Reactor Technology, Nuclear Power Institute of China); Yongtao Zhang (National Key Laboratory of Nuclear Reactor Technology, Nuclear Power Institute of China); Mingxing Liu (College of Computer Science, Sichuan University; National Key Laboratory of Nuclear Reactor Technology, Nuclear Power Institute of China); and Quan Ma, Fan Wen, and Hengxu Shen (National Key Laboratory of Nuclear Reactor Technology, Nuclear Power Institute of China) Abstract Abstract Anomaly detection in industrial control systems (ICS)—a critical application safeguarding mission-critical infrastructure—requires the integration of multi-source signals such as network traffic, security logs, and physical process data. However, existing methods typically rely on single modality models or simplistic fusion approaches, making models vulnerable to noise or distributional biases from any single data source. This work re-examines the "alignment" step prior to fusion and proposes a knowledge-guided aligned multi-source data fusion method for ICS anomaly detection (KAFD). KAFD decouples the conventional end-to-end black-box paradigm into an “align first, then fuse” process: first, we convert formalizable asset information into learnable embedding vectors; subsequently, during the fusion stage, attention bias informed by security priors is injected to compel the model to focus on high-risk cross-domain correlations. We devise alignment loss and security prior knowledge loss to facilitate embed security priors. Experiments across public industrial control datasets show that KADF outperforms existing single‑modality models and recent benchmarks, achieving mean AUC and F1‑score values of 96.7% and 96.6%. This confirms the effectiveness of the knowledge‑guided alignment paradigm for network attack detection models. Tuesday Virtual Room 7 IJCNN Paper SS01 Privacy-Preserving Machine and Deep Learning / SS07 Quantum Machine Learning Algorithms and Applications Session Chair: Wenxia Wang (Information Engineering University), Peiying Xu (Institute of Information Engineering, CAS; School of Cyber Security, University of Chinese Academy of Sciences) Hybrid Quantum-Classical Graph Neural Networks for Solving Unit Commitment Problem Xuanzi Ma (Information Engineering University) and Junchao Cui, Wenxia Wang, Yizhen Huang, Bei Zhou, and Jinchen Xu (Information Engieering University) Abstract Abstract The Unit Commitment Problem (UCP) is a core optimization task for ensuring secure and economically efficient operation of power grids. However, existing methodologies face two major limitations: first, there is a lack of a general and efficient way to represent the complex relationships between UCP variables and constraints; second, exponential parameter growth with problem scale impedes characterization of complex dependencies in UCP, thereby limiting predictive accuracy. To address these issues, we propose a hybrid quantum-classical graph neural network framework (HQGNN) that establishes a novel paradigm for UCP by: (1) modeling UCP as a unified bipartite graph format, and (2) integrating quantum feature extraction modules to reduce the parameter size and enhance classical GNNs' capability in capturing high-dimensional structural characteristics. Extensive experiments on IEEE and RTS benchmark datasets demonstrate that HQGNN achieves improvements in UCP solution accuracy while validating the effectiveness of quantum modules in enhancing graph-structured feature learning. The quantum module in HQGNN is proven to be the key enabler of performance gains. Its integration cuts parameters by 16.8% and reduces MSE in optimal solution prediction by 24.4% across all benchmarks, while achieving full-task superiority on the IEEE-9 dataset.These results substantiate the proposed method's capacity to simultaneously enhance solution quality and computational efficiency through quantum-enhanced graph representation learning. HD-EQNN: High-dimensional Ensemble Quantum Neural Network for Multi-class UAV Signal Recognition Wenxia Wang (Information Engineering University); Hanyun Wang (Sun Yat-sen University); and Xin Zhou, Yaqi Chen, Yifan Hou, and Jinchen Xu (Information Engineering University) Abstract Abstract Under the constraints of current quantum hardware, there are challenges in using quantum neural networks (QNNs) to perform multi-classification for high-dimensional signal features. In this paper, a high-dimensional ensemble quantum neural network (HD-EQNN) is proposed for multi-classification of frequency-domain signals, achieving an end-to-end mapping from Doppler spectrum to recognition results. The core components of HD-EQNN are the feature extraction module, the feature fusion module, and the ensemble learning classification module. The feature extraction module is capable of simultaneously rotational encoding 128-dimensional features into the same model, capturing local dependencies in groups. The feature fusion module fuses independent intermediate features to maintain the integrity of frequency domain information. The ensemble learning classification module uses multiple QNNs to form a strong learner to effectively capture global complex relationships. The modules cooperate with each other to automatically extract features and realize multi-classification. Experiments on two real UAV radar echo signal datasets show that in multi-class UAV signal recognition tasks, the accuracy of HD-EQNN is superior to existing quantum models, and it uses fewer quantum resources than amplitude encoding quantum methods. This advancement provides an effective solution for applying quantum methods to handle complex signals in the real world during the NISQ era. Integrating Quantum Feature Maps into Graph Neural Networks for Robust Visual Localization Emma Bennett (University of Missouri), Ricky Massaro (US Army ERDC), and Filiz Bunyak (University of Missouri) Abstract Abstract Robust aerial localization in GNSS-denied environments is a critical challenge for autonomous navigation. This paper introduces a hybrid Quantum-Classical Graph Neural Network (QGNN) framework for visual localization, treating the task as an approximate subgraph matching problem robust to structural perturbations. We propose and evaluate two architectures for integrating quantum feature maps into a classical graph convolution backbone: a hardware-efficient global projection, where a single quantum circuit processes the pooled graph embedding, and a structurally expressive local projection, where quantum circuits are applied to individual node embeddings prior to aggregation. The models are trained via triplet loss on geospatial graphs constructed from the Massachusetts Buildings dataset. We demonstrate that while global methods are computationally tractable on current IBM Quantum hardware, the local quantum-classical hybrid significantly outperforms classical baselines in high-noise regimes. Specifically, the local quantum architecture improves Recall@1 by up to 20 percentage points under severe structural perturbations (e.g., node deletions and additions), suggesting that node-level quantum feature mapping provides improved regularization against topological noise. Accio: Secure and Lean Gradient Boosting Decision Tree Training via Communication Optimization Xu Peiying (State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, CAS; School of Cyber Security, University of Chinese Academy of Sciences) and Wang Li-Ping (State Key Laboratory of Cyberspace Security Defense, Institute of Information Engineering, CAS; School of Cyber Security, University of Chinese Academ) Abstract Abstract Privacy-preserving machine learning (PPML) offers a promising framework for collaborative model training without user privacy leakage. And secure multi-party computation (MPC) stands out as one of the key technologies, as it allows distrustful parties to compute a function without revealing their private inputs. However, existing GBDT training frameworks suffer from communication inefficiency, primarily due to two issues: (1) high round complexity from linear operations based on general MPC techniques, and (2) heavy communication overhead from argmax computation that involves numerous comparisons. In this paper, we present Accio, a secure two-party GBDT training framework on vertically partitioned dataset. The main contributions of Accio are two-fold: the first part includes a leaner argmax protocol that reduces the online communication by nearly half, while the second part includes a secure protocol to efficiently realize node split. The experimental results demonstrate that Accio outperforms the state-of-the-art framework (i.e., SiGBDT) in both LAN and WAN settings. Tuesday Virtual Room 8 IJCNN Paper SS03 Physics-Informed Neural Networks: Advancements and Applications / SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models Session Chair: Jiyan Qiu (Computer Network Information Center, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Yongwei Tang (Liaoning Technical University) UDCP: Uncertainty Driven Context-prompted nnUNet for Tertiary Lymphoid Structure Segmentation in H&E Images Meng Wang, Yuxin Liang, Yarong Feng, and Yongwei Tang (Liaoning Technical University); Tianyu Shi (Shenyang Ligong University); and Chao Lv (China Medical University) Abstract Abstract In real pathological scenarios, tertiary lymphoid structures (TLS) exhibit varying sizes, diverse morphologies, complex cellular compositions, and unclear boundaries. Traditional deep convolutional neural networks, such as nnU-Net, struggle to fully utilize contextual information in TLS segmentation tasks, often leading to missed and false detections. To address this issue, this paper proposes an uncertainty-driven contextual cueing nnU-Net based on a clinically collected TLS pathological image dataset. The method first trains an nnU-Net baseline model using pathological images. Then, it concatenates the foreground probability map output by the model with the original image and inputs it back into nnU-Net, explicitly incorporating high-level contextual information and constructing a two-stage auto-context framework. On this basis, to address potential uncertainties in the probability map, this paper further designs two parallel contextual information purification strategies: a region-level probability gating method based on an adaptive threshold of Otsu; and a pixel-level uncertainty modeling method constructed based on predicted probabilities. These two contextual cueing masks are fused to generate an uncertainty-driven cue. Experimental results show that on a TLS independent test set strictly divided by patients, the proposed method can effectively improve the segmentation performance of the baseline model without requiring additional annotations or complex structural modifications. SC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA Minrui Luo (Shanghai Qi Zhi Institute; Institute for Interdisciplinary Information Sciences, Tsinghua University); Fuhang Kuang (Institute for Interdisciplinary Information Sciences, Tsinghua University); Yu Wang (Institute of Information Engineering, Chinese Academy of Sciences); Zirui Liu (Institute for Interdisciplinary Information Sciences, Tsinghua University); and Tianxing He (Institute for Interdisciplinary Information Sciences, Tsinghua University; Shanghai Qi Zhi Institute) Abstract Abstract Parameter-Efficient Fine-Tuning (PEFT) methods, particularly Low-Rank Adaptation (LoRA), are indispensable for efficiently customizing Large Language Models (LLMs). However, vanilla LoRA suffers from slow convergence speed and knowledge forgetting problems. Recent studies have leveraged the power of designed LoRA initialization, to enhance the fine-tuning efficiency, or to preserve knowledge in the pre-trained LLM. However, none of these works can address the two cases at the same time. To this end, we introduce \textbf{S}ubspace-\textbf{C}onstrained LoRA (\textbf{SC-LoRA}), a novel LoRA initialization framework engineered to navigate the trade-off between efficient fine-tuning and knowledge preservation. We achieve this by constraining the output of trainable LoRA adapters in a low-rank subspace, where the context information of fine-tuning data is most preserved while the context information of preserved knowledge is least retained, in a balanced way. Such constraint enables the trainable weights to primarily focus on the main features of fine-tuning data while avoiding damaging the preserved knowledge features. We provide theoretical analysis on our method, and conduct extensive experiments including \textit{safety preservation} and \textit{world knowledge preservation}, on various downstream tasks. In our experiments, SC-LoRA succeeds in delivering superior fine-tuning performance while markedly diminishing knowledge forgetting, surpassing contemporary LoRA initialization methods. LUMF: A Deep Learning Based Accurate and Robust Liver Uptake Measurement Framework for SUV Evaluation in Whole-Body PET/CT Images Yarong Feng, Yongwei Tang, and Yuxin Liang (Liaoning Technical University); Chao Lv (China Medical University); Tianyu Shi (Shenyang Ligong University); and Meng Wang (Liaoning Technical University) Abstract Abstract Accurate localization and segmentation of the liver standard uptake value (SUV) reference region are prerequisites for calculating the liver SUV reference value, and also constitute a critical step in the quantitative analysis of PET images, directly affecting the accuracy and reproducibility of tumor grading and assessment of metabolic activity in tissues and organs. Traditional manual methods and existing semi-automatic methods suffer from low efficiency and high subjectivity. Current fully automatic methods, such as the maximum hyperellipsoid-based liver reference region segmentation (MHE-LRS) method, fit a maximum inscribed hyperellipsoid into the liver to segment the reference region. While efficient and applicable to most standard PET images, these methods lack robustness in complex scenarios. To address the above challenges, this paper proposes, for the first time, a deep learning framework named LUMF for achieving robust liver SUV reference region segmentation and quantitative assessment. This method innovatively combines a deep learning model with a machine learning model. First, a pre-trained TotalSegmentator is used to segment the liver mask from whole-body CT images, based on which paired liver PET/CT images are obtained. These paired images are then used as input to nnU-Net for supervised training. Finally, on the test set, the framework outputs probability predictions and localization information of the liver SUV reference region, which serve as a region-of-interest mask. The MHE-LRS method is subsequently applied to achieve robust segmentation of the liver SUV reference region. PARTNO: A Neural Operator based on Physically Adaptive Region Partitioning for Solving Large-Scale Partial Differential Equations Jiyan Qiu, Wu Yuan, and Jian Zhang (Computer Network Information Center, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Abstract Abstract Neural operators have shown strong potential for solving partial differential equations (PDEs). Among them, transformer-based models often achieve high accuracy, yet they face efficiency bottlenecks in large-scale settings because mesh discretization combined with global attention incurs quadratic computational cost. To address this limitation, we propose PARTNO, an efficient PDE solver based on physically adaptive region partitioning. Instead of operating on the full mesh globally, PART adopts a physics-informed region decomposition strategy that partitions the computational domain into physically similar and relatively independent subregions, and models region-specific properties. Within each subregion, we introduce a dedicated Halo-Attention mechanism to jointly model intra-region physical representations and interactions with neighboring subregions, while maintaining continuity of physical quantities across region boundaries. By restricting attention computation to within subregions, PART reduces the overall complexity to linear, improves parallelism, and delivers excellent scalability for large-scale inference. Extensive experiments show that PART achieves state-of-the-art performance on multiple benchmarks; in particular, it obtains the best results on DrivAerStar, the largest industrial-scale automotive dataset. We further demonstrate its practicality in real-world industrial simulations, including automotive and aerospace design. Tuesday Virtual Room 1 IJCNN Paper SS09 Multimodal Deep Learning in Applications / SS21 Novel Networks in Human-Machine Collaboration: Paradigms, Methods, and Applications Session Chair: Anming Dong (Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Industrial Network and Information System Security, \\ Shandong Fundamental Research Center for Computer Science, Jinan, China), Xiang Li (Wuhan University) A Self-Controllable Adaptive Block for Multispectral Object Detection Xiang Li and Jun Chen (Wuhan University) Abstract Abstract Multispectral object detection has garnered significant attention due to its enhanced accuracy and robustness in challenging scenarios, such as low-light environments and severe weather conditions. However, a critical challenge persists: achieving effective fusion of visible (RGB) and infrared (IR) features is difficult when there are substantial discrepancies between the two modalities. For instance, visible features may be obstructed by various factors (e.g., smoke, fog or occlusions), while infrared features may lack distinguishable temperature differences from the background, so both scenarios hinder reliable feature utilization. To address these limitations, we propose a self-controllable adaptive block (SCAB). This block is designed to preserve the efficient feature representation of the original modality during model fine-tuning, while simultaneously integrating as much valid feature information as possible from the other modality through gradual iterative optimization. Experimental results on multiple RGB-IR datasets verify that our method achieves state-of-the-art performance. A Sinkhorn Mixture-of-Experts Switch for Efficient RGB-Thermal Fusion Aditya Viswakumar and Rajalakshmi Pachamuthu (Indian Institute of Technology Hyderabad) Abstract Abstract In this work, we propose a computationally efficient RGB-Thermal (RGB-T) sensor fusion framework based on Sinkhorn Attention. The proposed method integrates a Sinkhorn-based alignment mechanism within a Mixture-of-Experts (MoE) architecture to enable effective and adaptive fusion of multi-modal features. The Sinkhorn module explicitly aligns features from RGB and thermal modalities, facilitating robust cross-modal interaction. The network adopts a multi-stage hierarchical fusion strategy, in which modality-specific and shared representations are progressively aligned and fused at multiple semantic levels. The proposed Sinkhorn fusion network performs competitively against existing methods while requiring reduced memory and computational resources. Owing to its lightweight design, the model is capable of real-time edge inference, making it suitable for resource-constrained environments. MAF-Net: Modality-Aligned Fusion with Modular Attention for Multimodal Sentiment Analysis Yanbo Wang and Qing-Dao-Er-Ji Ren (Inner Mongolia University of Technology) Abstract Abstract Multimodal Sentiment Analysis (MSA) aims to decipher human emotional states by synthesizing linguistic, acoustic, and visual signals. Despite recent advancements, effective cross-modal interaction remains hindered by two fundamental challenges: the inherent semantic gap between dense textual representations and redundant, noisy audio-visual signals, and the phenomenon of unimodal dominance, where models over-rely on the dominant modality (text) at the expense of subtle non-verbal cues. To address these issues, we propose MAF-Net (Modality-Aligned Fusion with Modular Attention). First, we introduce a Geometric Semantic Alignment (GSA) mechanism. Unlike traditional linear mappings, GSA leverages Gram matrix-based basis vectors to project non-verbal features into a unified textual manifold, effectively filtering low-level noise while preserving essential geometric structures. Second, to counter contribution imbalance, we design Modular Masked Interactive Fusion (MMIF). This strategy employs a decoupled attention masking scheme to independently regulate self-modal and cross-modal information flows, preventing the suppression of recessive modalities during fusion. Extensive experiments on standard benchmarks (CMU-MOSI, CMU-MOSEI) and a newly constructed Mongolian MSA dataset (MG-SIMS) demonstrate that MAF-Net achieves state-of-the-art performance. Our results highlight the framework’s robustness in complex interaction scenarios and its superior transferability to low-resource language settings. Mamba-based Multi-Scale Classification Framework for sEMG Gesture Recognition Yuhan Song, Anming Dong, and Yubing Han (Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputer Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences), Jinan, China; Shandong Provincial Key Laboratory of Industrial Network and Information System Security, Shandong Fundamental Research Center for Computer Science, Jinan, China); Jiguo Yu (School of Information and Software Engineering, University of Electronic Science and Technology of China, Chengdu, China); and You Zhou (Technology Department, Shandong HiCon New Media Institute Co., Ltd., Jinan, China) Abstract Abstract Surface electromyography signals (sEMG) are widely used in gesture recognition and human-computer interaction due to their ability to reflect muscle activity patterns. However, the non-stationarity, multi-scale variations, and noise interference of the signals severely limit modeling and recognition performance. Existing methods based on traditional machine learning rely on manually designed temporal-frequency features, making them difficult to adapt to complex and variable signal environments. Although deep learning methods can improve recognition accuracy through end-to-end modeling, their robustness is still insufficient in complex scenarios such as small sample sizes, signal noise, and multi-scale variations, and they are inefficient in long-sequence modeling, which limits their performance in sEMG tasks. To address the aforementioned issues, this paper proposes a multi-path Mamba multi-scale classification framework (MaMSC). This method decomposes sEMG signals into high-frequency, low-frequency, and temporal residual components using hybrid frequency-temporal decomposition (HFTD), and models these multi-scale components using multi-path selective Mamba units within the MaMSC block. Subsequently, an inverse hybrid frequency-temporal decomposition (IHFTD) module fuses the features into a unified temporal representation. This design fully leverages Mamba's efficient long-sequence modeling capabilities and enhances robustness through symmetric residual paths. Experiments on the Ninapro DB2 dataset demonstrate that this method achieves accuracies of 99.18%, 92.82%, and 99.35% on gesture recognition Exercise B, C, and D, respectively. Cross-subject experiments further demonstrate that a generalization accuracy of 93.56% can be achieved with minimal fine-tuning of target user data, exhibiting extremely low calibration costs and excellent potential for real-world interaction adaptation. The code is available at: https://github.com/yu-han-song/MaMSC. Tuesday Virtual Room 2 IJCNN Paper SS13 Neural Networks for Biodiversity / SS24 Advances in Hyperdimensional Computing and Vector Symbolic Architectures Session Chair: jinghao wen (Villanova University), Gabriel Dubus (Muséum National d'Histoire Naturelle; Sebitoli Chimpanzee Project, Great Ape Conservation Project, Sebitoli, Kibale National Park, Fort Portal, Uganda) DeepForestSound: a multi-species automatic detector for passive acoustic monitoring in African tropical forests, a case study in Kibale National Park Gabriel Dubus and Théau d'Audiffret (Eco-Anthropologie, Muséum National d'Histoire Naturelle, UMR7206, Paris, France; Sebitoli Chimpanzee Project, Great Ape Conservation Project, Sebitoli, Kibale National Park, Fort Portal, Uganda); Claire Auger (Eco-Anthropologie, Muséum National d'Histoire Naturelle, UMR7206, Paris, France; N'lab, Nitidae Association, Montpellier, France); Raphaël Cornette (Institut de Systématique, Evolution, Biodiversité (ISYEB), UMR7205 Centre National de la Recherche Scientifique, MNHN, SU, EPHE-PSL, UA, Paris, France); Sylvain Haupert (Centre d'Ecologie et des Sciences de la Conservation (CESCO), UMR CNRS Sorbonne Université 7204, Muséum national d'Histoire naturelle, Paris, France); Innocent Kasekendi and Raymond Katumba (Sebitoli Chimpanzee Project, Great Ape Conservation Project, Sebitoli, Kibale National Park, Fort Portal, Uganda); Hugo Magaldi (Eco-Anthropologie, Muséum National d'Histoire Naturelle, UMR7206, Paris, France; Sebitoli Chimpanzee Project, Great Ape Conservation Project, Sebitoli, Kibale National Park, Fort Portal, Uganda); Lise Pernel (Eco-Anthropologie, Muséum National d'Histoire Naturelle, UMR7206, Paris, France; Institut de Systématique, Evolution, Biodiversité (ISYEB), UMR7205 Centre National de la Recherche Scientifique, MNHN, SU, EPHE-PSL, UA, Paris, France); Harold Rugonge (Sebitoli Chimpanzee Project, Great Ape Conservation Project, Sebitoli, Kibale National Park, Fort Portal, Uganda); Jérôme Sueur (Centre d'Ecologie et des Sciences de la Conservation (CESCO), UMR CNRS Sorbonne Université 7204, Muséum national d'Histoire naturelle, Paris, France); John Justice Tibesigwa (Uganda Wildlife Authority, Kampala, Uganda); and Sabrina Krief (Eco-Anthropologie, Muséum National d'Histoire Naturelle, UMR7206, Paris, France; Sebitoli Chimpanzee Project, Great Ape Conservation Project, Sebitoli, Kibale National Park, Fort Portal, Uganda) Abstract Abstract Passive Acoustic Monitoring (PAM) is widely used for biodiversity assessment. Its application in African tropical forests is limited by scarce annotated data, reducing the performance of general-purpose ecoacoustic models on underrepresented taxa. In this study, we introduce DeepForestSound (DFS), a multi-species automatic detection model designed for PAM in African tropical forests. DFS relies on a semi-supervised pipeline combining clustering of unannotated recordings with manual validation, followed by supervised fine-tuning of an Audio Spectrogram Transformer (AST) using low-rank adaptation, which is compared to a frozen-backbone linear baseline (DFS-Linear). The framework supports the detection of multiple taxonomic groups, including birds, primates, and elephants, from long-term acoustic recordings. DFS was trained on acoustic data collected in the Sebitoli area, in Kibale National Park, Uganda, and evaluated on an independent dataset recorded two years later at different locations within the same forest. This evaluation therefore assesses generalization across time and recording sites within a single tropical forest ecosystem. Across 8 out of 12 taxons, DFS outperforms existing automatic detection tools, particularly for non-avian taxa, achieving average AP values of 0.964 for primates and 0.961 for elephants. Results further show that LoRA-based fine-tuning substantially outperforms linear probing across taxa. Overall, these results demonstrate that task-oriented, region-specific training substantially improves detection performance in acoustically complex tropical environments, and highlight the potential of DFS as a practical tool for biodiversity monitoring and conservation in African rainforests. Graph-Reasoning with Disentangled Attention: A Knowledge-Guided Approach for Fine-Grained Visual Classification Yaosen Huang (South China Normal University); Jiayao Xie (Yunnan University); and Juandong Huang, Tianrong Chen, and Hua Tang (South China Normal University) Abstract Abstract Fine-Grained Visual Classification confronts the dual challenge of distinguishing subtle inter-class differences and accommodating large intra-class variations. While existing methods address geometric and environmental variations effectively, they overlook Age-related Appearance Variation—drastic visual trait differences across life stages of the same species. To tackle this, we propose Graph-Reasoning with Disentangled Attention (GRDA), a unified framework integrating structural reasoning and explicit knowledge guidance. GRDA features a Disentangled Channel Attention module to decouple age-specific traits, and a Part-Aware Graph Attention Network that models invariant structural topology via semantic parts localized by Vision-Language Models. Additionally, we leverage Large Language Models to align visual features with stage-specific textual knowledge, bridging the semantic gap. We further curate the GuangDongBirds (GDB) dataset, which includes both adult and subadult instances. Extensive experiments show GRDA achieves competitive performance on GDB, CUB-200-2011, and NABirds, outperforming existing methods in handling complex biological variations. Consistent Soundscape Connectomes via Stability-Refined Graph Learning Maria J. Guerrero (Universidad de Antioquia, Rice University); Aref Einizade (Télécom SudParis, Institut Polytechnique de Paris); Jhony H. Giraldo (Télécom Paris, Institut Polytechnique de Paris); and César A. Uribe (Rice University) Abstract Abstract Passive Acoustic Monitoring (PAM) enables large-scale biodiversity assessment using networks of autonomous recorders. A natural way to summarize relationships across sites is to learn a weighted graph whose nodes are recorders and whose edges encode acoustic similarity. However, standard graph learning (GL) methods can be structurally unstable: removing a node (e.g., due to sensor failure) can trigger global rewiring, changing the graph’s connectivity and undermining ecological interpretation and longitudinal comparisons. HDFlow: A Multi-layer Design and Evaluation Framework for Hyperdimensional Computing Jinghao Wen (Villanova University); Dongning Ma (Mohamed bin Zayed University of Artificial Intelligence, Villanova University); and Sizhe Zhang and Xun Jiao (Villanova University) Abstract Abstract Brain-inspired hyperdimensional computing (HDC) is an emerging computing paradigm that imitates the deep behaviors of abstract brain circuit functions. HDC is a more memory-centric computing paradigm compared with existing machine learning algorithms. HDC has been increasingly applied to different domains as an energy-efficient and lightweight alternative to existing learning algorithms. Although promising, design and evaluation of HDC face significant obstacle due to a lack of a general user-friendly framework. Existing open source projects for HDC mostly focus on one or few applications and require notable manual effort for model configuration that become a bottleneck for the practical evaluation and deployment of HDC models. To address these challenges, this paper proposes a design and evaluation framework for HDC called HDFlow. It provides multi-layer access to the entire lifecycle of HDC processing, from the bottom hypervector to the top system layer, to enable exhaustive evaluations such as error injection and quantization. HDFlow is implemented entirely using python and only requires minimal libraries, allowing highly-flexible customization including encoding schemes, error models, quantization configurations, etc. HDFlow also provides flexible interface to read data from different applications to avoid building an entire HDC framework every time for a new application. HDFlow leverages Just-in-Time (JIT) compilation as acceleration to enhance the execution speed. Tuesday Virtual Room 3 IJCNN Paper SS17 Generative Foundation Models for Robotics: From Language and Vision to Embodied Action / SS26 Brain Machine Intelligence: Models, Systems, and Translational Applications Session Chair: Zhiguo Zhang (Harbin Institute of Technology, Shenzhen, China), Yinfeng Yu (Xinjiang University) Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction Jia Li and Yinfeng Yu (Xinjiang University) Abstract Abstract In Audio-Visual Navigation (AVN), agents must locate sound sources in unseen 3D environments using visual and auditory cues. However, existing methods often struggle with generalization in unseen scenarios, as they tend to overfit to semantic sound features and specific training environments. To address these challenges, we propose the Binaural Difference Attention with Action Transition Prediction (BDATP) framework, which jointly optimizes perception and policy. Specifically, the Binaural Difference Attention (BDA) module explicitly models interaural differences to enhance spatial orientation, reducing reliance on semantic categories. Simultaneously, the Action Transition Prediction (ATP) task introduces an auxiliary action prediction objective as a regularization term, mitigating environment-specific overfitting. Extensive experiments on the Replica and Matterport3D datasets demonstrate that BDATP can be seamlessly integrated into various mainstream baselines, yielding consistent and significant performance gains. Notably, our framework achieves state-of-the-art Success Rates across most settings, with a remarkable absolute improvement of up to 21.6 percentage points in the Replica dataset for unheard sounds. These results underscore BDATP's superior generalization capability and its robustness across diverse navigation architectures. Hierarchical Neuro-Semantic Control for UAV Swarm Formation: Bridging LLM Planning and Hard Safety Constraints Xianchang Wang and Jie Sun (Shenyang Aerospace University) Abstract Abstract The integration of high-level semantic planning with low-level physical constraints remains a critical challenge for UAV swarm formation. While Large Language Models (LLMs) demonstrate proficiency in task reasoning, their inherent non determinism and inference latency impede their utilisation in safety-critical systems. To address this issue, we propose Hier archical Neuro-Semantic Control (HNSC), an architecture that achieves non-authoritative semantic planning with conditional hard safety via HO-CBF-QP filtering when the QP remains fea sible. At the semantic layer, a contract-based logical framework constrains ambiguous intentions into verifiable command sets. At the Translation layer, a robust High-Order Control Barrier Function (HO-CBF), coordinated with a Reference Governor, performs real-time safety filtering under measurement uncertain ties. Moreover, the system incorporates a deadlock monitor that triggers strategy reconstruction when the swarm encounters local minima. Experiments on the gym-pybullet-drones platform show that, in a challenging narrow-gap traversal task, HNSC achieves a 95.4% success rate and reduces the deadlock rate from 41.0% to 1.8% compared with open-loop LLM baselines. Collision events are significantly reduced by HO-CBF-QP filtering during feasible QP execution, while sampled distance statistics remain close to the execution boundary under discrete-time reporting. Overall, the proposed architecture improves swarm adaptability and interpretability while providing conditional hard-safety guarantees through HO-CBF-QP filtering. RoboECC: Multi-Factor-Aware Edge-Cloud Collaborative Deployment for VLA Models Zihao Zheng (Peking University, School of Computer Science); Hangyu Cao (South China University of Technology); Jiayu Chen (Peking University, School of Computer Science); Sicheng Tian (Beijing Normal University); and Chenyue Li, Maoliang Li, Xinhao Sun, Guojie Luo, and Xiang Chen (Peking University, School of Computer Science) Abstract Abstract Vision-Language-Action (VLA) models are mainstream in embodied intelligence but face high inference costs. Edge-Cloud Collaborative (ECC) deployment offers an effective fix by easing edge-device computing pressure to meet real-time needs. However, existing ECC frameworks are suboptimal for VLA models due to two challenges: (1) Diverse model structures hinder optimal ECC segmentation point identification; (2) Even if the optimal split point is determined, changes in network bandwidth can cause performance drift. To address these issues, we propose a novel ECC deployment framework for various VLA models, termed RoboECC. Specifically, we propose a model-hardware co-aware segmentation strategy to help find the optimal segmentation point for various VLA models. Moreover, we propose a network-aware deployment adjustment approach to adapt to the network fluctuations for maintaining optimal performance. Experiments demonstrate that RoboECC achieves a speedup of up to 3.28x with only 2.55%~2.62% overhead. SSVEP-freeTUNE: Sample-wise Domain Gating with Dual-Template Priors for Calibration-Free SSVEP Decoding Kaisong Hu, Leilei Zhao, Fengcheng Wu, Yuxin Liu, Zhiguo Zhang, and Zhenxi Song (Harbin Institute of Technology, Shenzhen, China) Abstract Abstract Calibration-free decoding is essential for cross-subject use of steady-state visual evoked potential (SSVEP)-based brain–computer interfaces (BCIs). However, generalization is still constrained by inter-subject variability, likely reflecting insufficient exploitation of the EEG’s multi-domain characteristics. To address these challenges, we propose an adaptive domain-gated dual-template neural entrainment framework for calibration-free SSVEP decoding, termed SSVEP-freeTUNE. The gating mechanism adaptively balances spectral, temporal, and spatial representations. In parallel, neural entrainment priors implemented with dual templates combine explicit harmonic constraints and robust cross-subject response morphology, anchoring the model with physiologically grounded cues. At the decision level, frequency-locked correlation scores and fused multi-domain features are jointly integrated, bridging model-driven priors and data-driven representations. Experiments on the Benchmark and BETA datasets under leave-one-subject-out protocols demonstrate that SSVEP-freeTUNE achieves state-of-the-art accuracy and information transfer rate, advancing calibration-free SSVEP BCIs toward practical deployment. Tuesday Virtual Room 4 IJCNN Paper SS23 NeuroSNN: Neuromorphic Computing and Spiking Neural Networks / SS29 Graph-Based Solutions for Explainable and Efficient AI Session Chair: jia jian (河北工程大学), Jinsheng Xiao (Wuhan University) Improving Spiking Transformer Via Quantized Spiking Neurons with Shunting Inhibition Yunhua Chen, Chukuo Qu, and Zeqaun Xie (Guangdong University of Technology) and Jinsheng Xiao (Wuhan University) Abstract Abstract Spiking Transformers show promise in bridging the accuracy gap between SNNs and ANNs in complex visual tasks, yet their training efficiency is constrained by the high computational cost of time-step simulations. Quantized spiking neurons, which replace spiking activations with integer activations, improve training efficiency but overlook the critical temporal dynamics and excitatory-inhibitory mechanisms of biological neurons. This omission hinders the model's ability to capture deep spatio-temporal dependencies. To address this, we designed Quantized Inhibition Leaky Integrate and Fire (QI-LIF). By introducing learnable leakage factors and excitatory-inhibitory mechanisms, QI-LIF simulates neural temporal dynamics and refractory periods, preventing overexcitability under continuous stimulation to better capture deep spatio-temporal dependencies. Furthermore, to reduce the computational complexity of self-attention, we propose a Texture Saliency-based dynamic gated Spiking Attention (TSSA). This approach utilizes feature variance across channels as a texture saliency metric to construct dynamic gated masks. These masks block low-variance background noise while allowing high-frequency key features to pass through. Our approach achieves competitive accuracy across multiple datasets. In particular, on the CIFAR-100 dataset, our method achieves an accuracy of 80.92\% using only 4 time steps, representing a 3.4\% improvement over baseline. Bio-Inspired Robust Spiking Object Detection for Flapping-Wing UAVs En Lin, Huaning Li, and Guanghua Song (Zhejiang University); Rui Yan (Zhejiang University of Technology); and Huajin Tang (Zhejiang University) Abstract Abstract Resource-constrained visual perception is crucial for flapping-wing UAVs, where reliable object detection is required for autonomous flight. Spiking Neural Networks (SNNs) provide an attractive solution due to their event-driven computation and sparse activation. However, in practical flapping-wing flight, periodic wingbeat-induced vibration and rapid ego-motion inevitably introduce severe motion blur and vibration-induced jitter, which suppress high-frequency details and significantly degrade the feature representations of existing SNN-based object detectors. To address this challenge, we propose an Eagle-Eye Perception Interaction Module (E-PIM) for fully spiking object detection. E-PIM introduces a dual-branch interaction strategy to enhance blur-sensitive local cues while aggregating wider contextual information, thereby improving the robustness of spiking feature maps under motion degradation. We evaluate the proposed method on VisDrone, a motion-blurred VisDrone benchmark, and a real-world flapping-wing UAV dataset collected during flight with wingbeat-induced vibration and ego-motion blur. Results show consistent improvements under motion degradation while preserving spike-driven inference with minimal additional runtime overhead. Heterogeneous Graph In-Context Learning with Large Language Models jian jia (河北工程大学) and di jin (天津大学) Abstract Abstract Heterogeneous graph learning is crucial for modeling multi-typed entities and their interactions in real-world complex systems. However, existing methods face significant limitations: Heterogeneous Graph Neural Networks (HGNNs) underutilize textual attributes, while Large Language Models (LLMs) struggle to bridge the semantic gap between textual data and graph structures. To address these limitations, we propose HGICL, a novel framework that facilitates deep collaborative optimization between LLMs and HGNNs. The framework leverages a small set of actively selected exemplar nodes, fine-tunes HGNNs with LLM-generated scores, and applies the optimized HGNNs to node classification. This integration raises two key technical challenges. First, selecting appropriate nodes as in-context learning (ICL) examples is critical, since LLMs cannot inherently model heterogeneous graph structures. Second, bridging the semantic gap between textual and graph‑structural information is essential, given the inherent disparity between graph topology and textual modalities that demands explicit alignment. To overcome these challenges, we propose an optimization strategy that integrates HGNNs and LLMs. Specifically, we design a meta-path-based node selection mechanism to construct high-information-density ICL examples, use LLMs to assign confidence scores to different exemplar nodes, and leverage these scores as supervisory signals to fine-tune the HGNNs. Comparative experiments involving three HGNN architectures, two homogeneous graph architectures, and a variety of LLMs demonstrate HGICL's superior performance, establishing a new paradigm for collaborative optimization in heterogeneous graph representation learning. End-to-End 3D Reconstruction Network with Occlusion Detectors for Multi-View FPP Zuqiong Chen, Ruisheng Wang, and Yibin Tian (Shenzhen University) Abstract Abstract Most existing Deep learning (DL) –based 3D reconstruction methods for Fringe Projection Profilometry (FPP) rely on single-view inputs and suffer from occlusions and incomplete reconstruction when applied to complex geometries. To address these limitations, we propose an end-to-end DL framework for high-fidelity 3D reconstruction from fringe patterns acquired from four complementary views. The proposed network adopts a multi-task architecture built upon a shared 3D residual encoder–decoder backbone, augmented with an Atrous Spatial Pyramid Pooling (ASPP) module to capture multi-scale contextual information without sacrificing spatial resolution. To explicitly handle view-dependent occlusions arising from complex objects, a lightweight occlusion detector and an adaptive occlusion-aware fusion module are introduced to guide multi-view feature aggregation. In parallel, a self-supervised view reconstruction branch is jointly optimized with depth estimation, acting as a structural regularizer that promotes geometric consistency and preserves fine details in the latent feature space. Extensive experiments conducted on a large-scale synthetic dataset demonstrate that the proposed method achieves a Mean Absolute Error (MAE) of 0.0447 mm and a Root Mean Square Error (RMSE) of 0.1741 mm, representing a more than 50% reduction in reconstruction error compared with state-of-the-art single-view DL approaches. These results establish a new benchmark for robustness and completeness in FPP 3D measurement under challenging occlusion and complex geometry conditions. Tuesday Virtual Room 5 IJCNN Paper SS32 Deep Neural Networks and Generative AI for Multi-Agent Smart Vehicle Perceptron, Learning, Automation and Optimization / SS39 Computational Audio Intelligence for Perception & Representation Session Chair: Zejian Zhou (University of Wyoming), Qing Tian (University of Alabama at Birmingham) Full-Frequency Temporal Patching and Structured Masking for Enhanced Audio Classification Aditya Makineni, Baocheng Geng, and Qing Tian (University of Alabama at Birmingham) Abstract Abstract Transformers and State-Space Models (SSMs) have advanced audio classification by modeling spectrograms as sequences of patches. However, existing models such as the Audio Spectrogram Transformer (AST) and Audio Mamba (AuM) directly adopt square patching from computer vision, which disrupts continuous frequency patterns and produces an excessive number of patches, slowing training, and increasing computation. In this paper, we propose Full-Frequency Temporal Patching (FFTP), a patching strategy that better matches the time-frequency asymmetry of spectrograms by spanning full frequency bands with localized temporal context, preserving harmonic structure, and significantly reducing patch count and computation. We also introduce SpecMask, a patch-aligned spectrogram augmentation method that combines full-frequency and localized time-frequency masks under a fixed masking budget, improving temporal robustness while preserving spectral continuity. When applied on both AST and AuM, our method improves mAP by up to 6.76% on AudioSet-18k and accuracy by up to 8.46% on SpeechCommandsV2, while reducing computation by up to 83.26%, demonstrating both performance and efficiency gains. AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection Hai Son Nguyen Le and Hung Cuong Nguyen Thanh (University of Science, Ho Chi Minh City, Vietnam); Nhien An Le Khac (University College Dublin, School of Computer Science, Dublin, Ireland); Dinh Thuc Nguyen (University of Science, Ho Chi Minh City, Vietnam); and Hong Hanh Nguyen Le (University College Dublin, School of Computer Science, Dublin, Ireland) Abstract Abstract The rapid advancement of generative models has enabled highly realistic audio deepfakes, yet current detectors suffer from a critical bias problem, leading to poor generalization across unseen datasets. This paper proposes Artifact-Focused Self-Synthesis (AFSS), a method designed to mitigate this bias by generating pseudo-fake samples from real audio via two mechanisms: self-conversion and self-reconstruction. The core insight of AFSS lies in enforcing same-speaker constraints, ensuring that real and pseudo-fake samples share identical speaker identity and semantic content. This forces the detector to focus exclusively on generation artifacts rather than irrelevant confounding factors. Furthermore, we introduce a learnable reweighting loss to dynamically emphasize synthetic samples during training. Extensive experiments across 7 datasets demonstrate that AFSS achieves state-of-the-art performance with an average EER of 5.45%, including a significant reduction to 1.23% on WaveFake and 2.70% on In-the-Wild, all while eliminating the dependency on pre-collected fake datasets. Our code will be published after accepted. Uncertainty-driven Multi-View Feature Fusion for Spoofed Speech Detection Xinxin Luo (XinJiang University); Liangbo Zhu (Beijing University Of Technology); Zhongyu He (Shandong Jiaotong University); and Xiaolong Wu (Naval Aviation University, XinJiang University) Abstract Abstract Robust spoofed speech detection hinges on generalization to unseen spoofing attacks and adverse acoustic conditions (e.g., noise and channel mismatch), where the quality of different representations can vary substantially across utterances. We propose an uncertainty-driven multi-view framework that fuses raw-waveform and LFCC-based spectral representations by estimating per-view evidential uncertainty with a Beta distribution, converting it into utterance-level credibility weights, and using these weights to condition a lightweight multi-head self-attention fusion module; we further introduce a conflict-aware objective to explicitly discourage inconsistent view opinions. Experiments on ASVspoof 2019 LA and cross-dataset evaluation on ASVspoof 2021 LA show consistent improvements over strong single-view baselines and fixed fusion strategies, and controlled SNR corruption tests further confirm that uncertainty-guided credibility weighting yields more stable performance under severe noise. Fast Sub-optimal Mean-Field Control with Drift Imitation for Large-scale Heterogeneous Agents Zejian Zhou (University of Wyoming) Abstract Abstract We propose a fast and scalable framework for multiagent decision-making that eliminates the need for iterative fictitious play in large-scale systems. Leveraging the Fokker-Planck-Kolmogorov (FPK) equation, our method computes both the population-level control policy (drift) and the corresponding agent distribution in a single forward pass. This one-shot solution avoids the repeated equilibrium updates required by conventional continuous mean field control (MFC) approaches. Once the population dynamics are established, individual heterogeneous agents derive their actions via deep behavior cloning from the learned drift field. Our framework introduces two core innovations: (1) direct, non-iterative computation of the team policy and distribution, and (2) an imitation mechanism that seamlessly handles agent heterogeneity, capabilities not afforded by traditional MFC methods. Experimental results show that our approach achieves precise distribution control with large-scale heterogeneous agents. Tuesday Virtual Room 6 IJCNN Paper SS36 Human-Centered AI: Multi-modal Agent-based Systems (HCAI) / SS41 Randomization Based Deep and Shallow Learning Algorithms and/or Biomedical Applications Session Chair: zhao wang (Xidian University), Jingyi Lyu (Beijing Normal-Hong Kong Baptist University) Medical Image Classification via Integrated Attention Mechanisms and Dynamic Weighted Loss Function Jingyi Lyu, Benxiang Jiang, Songze Zhang, and Hongjian Shi (Beijing Normal-Hong Kong Baptist University) Abstract Abstract A concise and efficient network is the goal pursued in the field of medical image classification. In this study, we developed a lightweight medical image classification module for efficient classification of medical images. This module integrates spatial attention mechanism and channel attention mechanism to effectively capture channel and spatial information. In order to alleviate the problem of data imbalance in medical image classification without increasing the training data, we proposed a novel loss function DWDI. The loss function dynamically adjusts the influence of category weights during the training process, thereby enhancing the model's learning ability for minority categories. We validated our method on seven public datasets, and the results demonstrated that our method significantly improved classification accuracy while maintaining strong competitiveness in model complexity. Explainable Hybrid Transfer Learning for Infodemic Misinformation Detection Deepak kanneganti, Sajib Mistry, and Aneesh Krishna (Curtin Univeristy); Mufti Mahmud (King Fahd University of Petroleum and Minerals); and Monowar Bhuyan (Umeå University) Abstract Abstract Infodemic events such as pandemics, wars, and elections accelerate the spread of misinformation on social media, demanding accurate and transparent detection models. Recent advances in hybrid transfer learning models improve detection performance by combining multiple representation stages. Despite their effectiveness, their multi-stage architecture lacks a unified prediction interface, preventing the effective use of standard explainable artificial intelligence (XAI) techniques. This lack of compatible inference mechanisms restricts model interpretability in high-stakes decision-making settings. This paper proposes an explainable hybrid transfer learning framework that addresses this inference gap. We introduce a probability-based predictive inference mechanism that enables plug-and-play integration of widely used post-hoc XAI techniques, including SHAP and ELI5, without modifying the underlying model architecture. Experiments on multiple real-world infodemic datasets show that the proposed approach outperforms state-of-the-art baselines while achieving a balanced trade-off between detection accuracy and explainability. A GAN-Based MRI Denoising Model for Alzheimer’s Disease Using Free Attention and Wasserstein Loss Muhammad Afrizal Amrustian and Mufti Mahmud (King Fahd University of Petroleum & Minerals) Abstract Abstract Alzheimer's disease (AD) is a progressive neurodegenerative disorder that affects the physical brain. Magnetic Resonance Imaging (MRI) is the gold standard in medical image analysis for detecting and monitoring the brain structure changes. However, MRI data acquisitions are often degraded by noise and have low resolution due to short scan times, hardware constraints, or patient motion. Previous research has addressed the problem using approaches such as spatial filtering, wavelet thresholding, and interpolation-based denoising, which can over-smooth the image. The noisy MRI regions lack reliable local features, which makes restoration difficult. Therefore, this research proposes a GAN-based framework for MRI denoising in AD. The proposed method integrates a free-parameter attention mechanism into the generator, while a wavelet-informed UNet discriminator is used. The Wasserstein loss with the gradient penalty was adopted to ensure a stable training. Our proposed model achieves a PSNR of 34.755 and SSIM of 0.9180, while minimising reconstruction error with an NRMSE of 0.1492 and MAE of 0.0151. The experiment’s results demonstrate that the method surpasses existing GAN-based methods, such as RCA-GAN and OA-GAN. The quantitative and qualitative results confirm that the method effectively reduces noise while preserving anatomical details. A Hierarchical Multi-Agent framework for Automated Radar Signal Analysis zhao wang and jiahui yan (Xidian University), Maoguo gong (Inner Mongolia Normal University), and peng li and hao li (Xidian University) Abstract Abstract Radar signal analysis is pivotal for electromagnetic situational awareness, but faces a critical "perception-reasoning gap." While state-of-the-art deep learning models demonstrate superior performance in deterministic modulation classification, they lack the cognitive capability to perform multi-step reasoning and verify physical consistency. To address this limitation, this paper proposes RadarAgent, a novel hierarchical Multi-Agent framework designed for comprehensive and interpretable signal analysis. Governed by a rigorous Plan-and-Execute cognitive cycle, RadarAgent orchestrates a Supervisor-Worker architecture that integrates a deterministic toolchain with a Retrieval-Augmented Generation (RAG) system for domain knowledge injection. Notably, the toolchain is empowered by our proposed Dual-Scale Radar Transformer (DSRT) to ensure state-of-the-art modulation classification accuracy. Furthermore, a physics-based Critic Agent is introduced to enforce domain constraints, mitigating the hallucination risks inherent in large language models. To evaluate complex reasoning capabilities, we construct RadarQA, a specialized reasoning-intensive benchmark. Experimental results demonstrate that RadarAgent not only significantly outperforms baseline models on the DeepRadar dataset, particularly under low-SNR conditions ($-8$ dB), but also achieves a robust task success rate of 94.2\% on RadarQA. These findings validate the framework's ability to provide fully automated, physically consistent, and interpretable signal analysis reports. Tuesday Virtual Room 7 IJCNN Paper SS38 Deep Learning in Computational Biology and Biomedicine: from Biomedical Data to Drug Discovery / SS48 AI, Law and Regulation Session Chair: Jing Peng (Wuhan University of Technology), Raphaella Revis (UTS) Towards a Neurodata Protection Framework in Australia Raphaella Revis and Avinash Singh (University of Technology Sydney) Abstract Abstract Neurotechnology enables the recording, interpretation and modulation of neural activity, generating and extracting highly sensitive neural data with significant privacy implications. While jurisdictions worldwide have begun addressing neurodata governance and regulation through constitutional reform, sector- specific legislation and enhanced privacy protections, Australia currently lacks a framework to explicitly safeguard neurodata. This paper examines the current limitations of Australia’s existing privacy regime in addressing neurotechnology related risks, including re-identification, unauthorized secondary use, neuro-marketing and cybersecurity breaches. Drawing on comparative developments in the United States, Chile, and Mexico, the paper proposes the first Neurodata Privacy Protection Framework (NPPF) designed to Australia’s legal context. The NPFF adopts a more soft-law approach but aligned with the Australian Privacy Principles (APPs) embedded in statute, emphasizing informed consent, mental privacy, pseudonymisation and data security. To evaluate the framework’s effectiveness, a simulated case study is applied, demonstrating how the NPFF may mitigate harms in Australia identified. The proposed NPFF is practical and scalable model for neurodata governance without requiring strenuous Australian constitutional reform while also offering significant benefits for current and future neurodata protection. Assessing Defenses Against Prompt Injection Attacks in LLM Systems Fadime Zeliha Seyhan (BeyondGuard); Katira Soleymanzadeh (BeyondGuard, Istanbul Health and Technology University); Mustafa Oğuz Cengiz and Alp Ecevit (BeyondGuard); Roya Arkhmammadova (BeyondGuard, Koc Unıversıty); and Elif Tansu Kaya (BeyondGuard) Abstract Abstract Prompt injection poses a critical security and compliance risk for large language model (LLM)-integrated systems. This paper treats prompt injection as a system-level governance challenge and introduces a model-agnostic guardrail architecture that performs prompt-level risk assessment prior to LLM inference. As a core component, we develop a dedicated prompt injection detector trained using Adaptive Layer LoRA on an in-house annotated dataset. Experiments across Gemma and Qwen model families show that baseline prompted LLMs exhibit unstable precision-recall trade-offs and elevated false positive rates, while increasing prompt complexity alone does not improve robustness. In contrast, LoRA fine-tuned models achieve deployment-grade performance, reaching up to 99.05% accuracy with substantially reduced false positives, false negatives, and attack success rates. CaseSentinel: Retrieval-Augmented Multi-Agent Reasoning for Responsible Criminal Investigation Tianshuo Chen, Bosen Miao, Longfei Yin, Yushi Wang, Xinqi Wang, and Yueheng Sun (Tianjin University) Abstract Abstract Criminal investigations are increasingly overwhelmed by web-scale digital evidence, yet existing AI paradigms like retrieval-augmented generation (RAG) lack the stateful, auditable reasoning required for high-stakes sensemaking. This creates a gap between the potential of large language models and the accountability demanded by the legal system. To bridge this gap, we reframe investigation as a process of Bayesian inference and introduce CaseSentinel, a retrieval-augmented multi-agent system designed for responsible criminal investigation. CaseSentinel operationalizes a Hypothesis Space Graph (HSG) on a shared blackboard, where specialized agents collaboratively refine hypotheses to drive the investigation toward convergence. A human-in-the-loop web interface provides investigators with full control, enabling them to validate evidence, override agent suggestions, and inject domain knowledge, ensuring every step is auditable. Evaluations on twenty adjudicated cases --- with additional synthetic noise stress tests --- show CaseSentinel improves action-plan quality by +0.44 (relative +116%) and raises calibrated success probabilities by +0.48 absolute over single-agent RAG baselines. CaseSentinel offers a practical framework for building transparent, controllable, and effective AI-powered decision-support systems. Insertion Variant Calling in Long-Read Data Through a Transformer and Dual-Modal Fusion Model Junjie Zhang, Chunxiao Lai, Jing Peng, and Hao Wu (Wuhan University of Technology) Abstract Abstract Insertions represent a significant class of structural variations (SV), and their accurate detection is critical for elucidating the mechanisms of genetic diseases. Although current deep learning methods based on long-read sequencing data have demonstrated significant promise in detecting insertion structural variations, most rely on single-modal feature representations—such as CIGAR images or numerical matrices—which fail to comprehensively capture the multi-dimensional signals of SV, thereby limiting detection accuracy and robustness. To address this limitation, we propose TDSV, a Transformer-based dual-modal fusion model for analyzing insertion structural variations in sequencing data. TDSV constructs an image-sequencing dual-modal feature representation by simultaneously extracting CIGAR images from candidate regions and multi-dimensional numerical tensors of sequence statistics, providing complementary information to the model. We further design a bidirectional cross-modal attention mechanism to enable deep interaction, alignment, and enhancement between the two modalities, allowing the model to synergistically capture local fine-grained patterns in images and global statistical regularities in sequences. To improve the detection of cross-region long insertion structural variations, a Transformer encoder is introduced to perform global modeling of the fused features, capturing long-range dependencies across consecutive sub-regions. Experiments on multiple real datasets show that TDSV significantly outperforms existing mainstream methods in several evaluation metrics, demonstrating particularly robust performance in low-coverage genomic regions. Tuesday Virtual Room 8 IJCNN Paper SS11 Engineering Trust: Ethical, Legal, and Societal Impacts of Computational Intelligence on Human Agency / SS12 XSTASys: Explainability and Security in Trustworthy Artificial Intelligence Systems Session Chair: Ravi Kumar (Florida International University, Rowan University), Ali Karkehabadi (University of California, Davis) CatRAG: Functor-Guided Structural Debiasing with Retrieval Augmentation for Fair LLMs Ravi Kumar (Florida International University (FIU), Rowan University); Utkarsh Grover (University of South Florida, New York University); Mayur Akewar (Florida International University (FIU)); Xiaomin Lin (University of South Florida, John Hopkins University); and Agoritsa Polyzou (Florida International University (FIU), Georgetown University) Abstract Abstract Large Language Models (LLMs) are deployed in high-stakes settings but can show demographic, gender, and geographic biases that undermine fairness and trust. Prior debiasing methods, including embedding-space projections, prompt-based steering, and causal interventions, often act at a single stage of the pipeline, resulting in incomplete mitigation and brittle utility trade-offs under distribution shifts. We propose CatRAG Debiasing, a dual-pronged framework that integrates functor with Retrieval-Augmented Generation (RAG) guided structural debiasing. The functor component leverages category-theoretic structure to induce a principled, structure-preserving projection that suppresses bias-associated directions in the embedding space while retaining task-relevant semantics. On the Bias Benchmark for Question Answering (BBQ) across three open-source LLMs (Meta Llama-3, OpenAI GPT-OSS, and Google Gemma-3), CatRAG achieves state-of-the-art results, improving accuracy by up to 40% over the corresponding base models and by more than 10% over prior debiasing methods, while reducing bias scores to near zero (from 60% for the base models) across gender, nationality, race, and intersectional subgroups. Geometric Alignment Constraints for Mitigating Membership Inference in Federated Learning Ruijia Liu, Xiaochen Nie, Jianhua Wang, Yongle Chen, and Dan Yu (Taiyuan University of Technology) Abstract Abstract Federated learning avoids raw data sharing butremains vulnerable to membership inference attacks (MIAs). Recent generative federated learning improves utility under non-independent and non-identically distributed (non-IID) data bysynthesizing complementary data; however, the generation stagecan distort representation geometry, increasing member/non-member separability and strengthening both loss-based and representation-based MIAs. Existing defenses, such as differential privacy and gradient clipping, operate at the update level anddo not address this generation-induced leakage. We propose FedGAC, a lightweight geometric alignment strategy applied during complementary data generation. By imposing classifier-induced class-wise geometric constraints, FedGAC suppresses client-specific directions while preserving class-level semantics, without modifying the federated learning protocol or local training. Extensive experiments across diverse non-IID settings show that FedGAC substantially reduces MIA success with negligible accuracy loss, achieving a better privacy-utility trade-off. PointDetBA: Backdoor Attacks Against 3D Object Detection in Autonomous Driving Lei Zhang, cheng chen, and Linkun Fan (Henan University, Kaifeng, China); Weidong Zhang (Shanghai Jiao Tong University, Shanghai, China); and Bobo Wei (Henan University, Kaifeng, China) Abstract Abstract 3D point cloud object detection is fundamental to collaborative perception systems in autonomous driving and other cyber-physical systems (CPSs), enabling vehicles and infrastructure to perceive and share environmental information. However, these systems remain vulnerable to backdoor attacks that implant hidden malicious behaviors during the training stage. Once triggered at inference, such attacks can cause false, missing, or misclassified detections, threatening the safety and reliability of collaborative perception. Although several backdoor attacks against 3D point cloud detection models have been proposed, they suffer from limited attack objectives and poor stealthiness. To address these limitations, we propose PointDetBA, a multi-objective backdoor attack framework implemented under a unified poisoning pipeline for LiDAR-based 3D object detection. In detail, PointDetBA introduces two task-aware trigger mechanisms that exploit geometric and physical priors of point clouds. The point-injection mechanism achieves targeted misclassification and object generation by injecting compact reflective clusters, while the rotation–noise mechanism applies subtle rotations and Gaussian perturbations to induce misclassification and object disappearance with high stealth. Experiments on the KITTI dataset and five representative detectors demonstrate that PointDetBA achieves strong attack effectiveness with minimal impact on clean data, revealing critical security vulnerabilities in collaborative 3D perception and providing insights for the secure and trustworthy design of cyber-physical systems. SaliencyDecor: Enhancing Neural Network Interpretability through Feature Decorrelation Ali Karkehabadi (University of California, Davis); Jamshid Hassanpour (Georgia Institute of Technology); and Houman Homayoun and Avesta Sasan (University of California, Davis) Abstract Abstract Gradient-based saliency methods are widely used to interpret deep neural networks, yet they often produce noisy and unstable explanations that poorly align with semantically meaningful input features. We argue that a fundamental cause of this behavior lies in the geometry of learned representations: correlated feature dimensions diffuse attribution gradients across redundant directions, resulting in blurred and unreliable saliency maps. To address this issue, we introduce SaliencyDecor, a training framework that explicitly decorrelates intermediate feature representations through group-wise ZCA whitening integrated into the optimization process. By reshaping the feature space toward orthogonality, our approach promotes more concentrated gradient flow and improves the fidelity of saliency-based explanations. SaliencyDecor jointly optimizes classification, prediction consistency under feature masking, and a decorrelation regularizer, requiring no architectural changes or inference-time overhead. Extensive experiments across multiple benchmarks and architectures demonstrate that our method produces substantially sharper and more object-focused saliency maps while simultaneously improving predictive performance, achieving accuracy gains over the datasets. These results establish feature decorrelation as a principled mechanism for enhancing both interpretability and accuracy, challenging the conventional trade-off between explanation quality and model performance. Tuesday Virtual Room 1 IJCNN Paper SS14 Neuroevolutionary computation techniques for medical data understanding and pattern analysis / SS15 Agentic Edge Intelligence for the Internet of Things Session Chair: Xingpeng Zhang (Southwest Petroleum University), Ruiqin Bai (College of Artificial Intelligence, Taiyuan University of Technology, Taiyuan, 030024, P.R. China; Postdoctoral Workstation, China Railway 17th Bureau Group Co., Ltd., Taiyuan, 030006, P.R. China This work was supported by the Shanxi Provincial Young Scientists Research Project (202303021212081)) DSIB: Towards Robust Multimodal Survival Prediction via Discrete Self-Distillation Information Bottleneck Ye Tian (College of Computer Science and Technology (College of Data Science), Taiyuan University of Technology, Taiyuan, 030024, P.R. China) and Ruiqin Bai (College of Artificial Intelligence, Taiyuan University of Technology, Taiyuan, 030024, P.R. China; Postdoctoral Workstation, China Railway 17th Bureau Group Co., Ltd., Taiyuan, 030006, P.R. China) Abstract Abstract Multimodal survival prediction can improve prognosis by jointly modeling whole slide images (WSIs) and genomic profiles, yet practical models are brittle under multimodal heterogeneity. In particular, when training is constrained to extremely small batches (often required by WSI processing), modality-specific noise and distribution mismatch are amplified in the fused interaction space, making optimization unstable and prone to redundant correlations. We propose Discrete Self-Distillation Information Bottleneck (DSIB), a training-only algorithm that directly regularizes the fusion representation without modifying backbone architectures. DSIB inserts a stochastic binary bottleneck to explicitly control the information capacity of fused features, and injects hidden-space perturbations to expose heterogeneity-induced failure modes. Crucially, DSIB performs distributional self-distillation by minimizing the Jensen–Shannon divergence between the factorized Bernoulli posteriors of perturbed and unperturbed pathways, stabilizing learning under small-batch noise. We provide a theoretical connection showing that DSIB optimizes a tractable bound of the discrete information bottleneck objective. Experiments on three TCGA cohorts (BLCA, UCEC, LUAD) across multiple state-of-the-art multimodal survival predictors demonstrate consistent C-index gains and improved robustness under controlled noise and redundancy, with zero inference overhead. MMO-GFlowNet: Multi-objective Metahuristic Optimization of GFlowNet Parameters for Molecular Generation Olaide Nathaniel Oyelade (NCA&T University, Greensboro, USA) and Hui Wang (Queens University Belfast) Abstract Abstract Generative models have evolved in recent years, and are now capable of learn complex underlying structures of molecules. The structure-based drug design (SBDD) models are a class of generative models that learn the structure of target protein receptors to generate molecules capable of binding to the target. GFlowNets are stochastic-based machine learning models that sample complex objects like protein-protein interactions in drug design, and have recently merged with SBDD. Conditioned GFlowNets are characterized by combinatorial parameters needing carefully selected values, with impaired performance when poorly combined. To address this gap, we propose a novel hybrid of metaheuristics for multi-objective optimization of a subset of conditioned GFlowNet. We specifically adapted a biology-based optimizer to jointly maximize three objectives to learn the combined parameter values suitable to fully train GFlowNet conditioned to pockets of protein structure. Thereafter, drug-like molecules are derived from the model. Using the CrossDocked2020 and docking benchmark version 5 (DB5), we experimentally evaluated the proposed model. A Federated Learning Backdoor Attack Defense Method Based on Frequency Domain Perturbations and Model Repair Jinquan Zhang, Qingmian Wei, and Lina Ni (Shandong University of Science and Technology) Abstract Abstract Federated Learning (FL) enables privacy-preserving distributed modeling but remains vulnerable to backdoor attacks, where adversaries embed triggers to induce targeted misclassification while keeping normal accuracy intact. Existing defenses, such as robust aggregation or cosine-similarity–based gradient filtering, cannot reliably distinguish malicious clients from benign ones and often discard suspicious updates in a coarse-grained manner, which may inadvertently remove legitimate clients and degrade model performance. To address these limitations, we propose FedFdpm, a novel defense method based on frequency-domain perturbations and model repair. FedFdpm detects suspicious client models by analyzing their frequency spectrum responses and applies flexible repairs to align their distributions with normal patterns, reducing aggregation risks without discarding clients. It employs dynamic aggregation mechanism to assign weights to clients for robust aggregation. Experimental results on three public datasets, compared to five baseline algorithms, show that FedFdpm significantly reduces the success rate of backdoor attacks (up to 96%) and maintains high accuracy (98%) in the main task, outperforming existing methods. LABMamba: A Locally-Guided Alternating Bidirectional Mamba Method for Point Cloud Classification Yang Yu, Zili Yan, and Xingpeng Zhang (Southwest Petroleum University) Abstract Abstract The unordered nature and complex local structures of point clouds challenge accurate feature modeling. While state-space models (SSMs) like Mamba offer linear-time global modeling, they struggle with fine-grained local geometry and irregular 3D topologies. We propose LABMamba, a geometric-aware SSM framework for point cloud classification. Our key innovation is a geometric-aware group embedding (GAGE) module that integrates explicit geometric priors (normals, curvature, density) with learned features, addressing Mamba's local geometric limitation. We also introduce an alternating bidirectional Mamba strategy that alternately models sequence and channel dimensions, differing from parallel bidirectional designs. This improves performance, reduces sensitivity to point ordering, and minimizes reliance on heuristic serialization. Extensive experiments show state-of-the-art or competitive results, with ablations validating each component. Tuesday Virtual Room 2 IJCNN Paper SS16 Bayesian Neural Networks / SS18 Synergizing Multimodal LLMs, Co-Agent Architectures, and Foundation Models in Manufacturing AI Session Chair: Yun Liu (Sichuan University), Chao Ma (Institute of Information Engineering, CAS; School of Cyber Security,University of Chinese Academy of Sciences) AML-BGNN:Advanced Multi-Level Bayesian Graph Neural Network Framework for Robust Network Security Analysis Chao Ma (Institute of Information Engineering,Chinese Academy of Sciences; School of Cyber Security,University of Chinese Academy of Sciences) and Xiaoyu Kang, Bo Ho, Zhixin Shi, and Weiqing Huang (Institute of Information Engineering,Chinese Academy of Sciences) Abstract Abstract Modern network security demands both accurate threat detection and reliable uncertainty quantification for critical decision-making. Existing Graph Neural Networks (GNNs) fail to capture multi-scale attack patterns and provide confidence estimates essential for operational deployment. We present AML-BGNN, a unified Bayesian framework that addresses these limitations through three key contributions: (i) a hierarchical architecture that jointly learns node, edge, and graph-level representations, enabling comprehensive threat modeling across network scales; (ii) principled uncertainty quantification via variational inference, distinguishing epistemic and aleatoric uncertainty for risk-aware decision support; and (iii) an uncertainty-guided risk propagation mechanism that models threat diffusion dynamics while incorporating prediction confidence. Extensive experiments on synthetic networks (1,000 nodes) and CICIDS2017 demonstrate that AML-BGNN significantly outperforms state-of-the-art baselines, achieving perfect graph-level classification, 5.8\% improvement in edge-level F1-score, and well-calibrated uncertainties (ECE $\textless$ 0.05). The framework identifies critical attack paths with 92.3\% precision while maintaining linear scalability. Our ablation studies reveal that multi-scale modeling and uncertainty-aware propagation contribute synergistically to performance gains. AML-BGNN establishes a new paradigm for interpretable, uncertainty-aware network security analysis with immediate practical applications. SoftFlow‑DDN: Adaptive Flow and Soft Binning for Free‑Form Conditional Density Estimation ChengLong Song (University of Jinan); Mazharul Islam (University of Jinan; Shandong Provincial Key Laboratory of Green and Intelligent Building Materials, Jinan 250022, China.); Lin Wang (University of Jinan, Quan Cheng Laboratory); Bing Chen (University of Waterloo); Shuangrong Liu (Shandong Key Laboratory of Ubiquitous Intelligent Computing, University of Jinan, Jinan 250022; University of Jinan); and Bo Yang (Quan Cheng Laboratory, University of Jinan) Abstract Abstract Free-form conditional density estimation (CDE) models arbitrary distributions without parametric constraints. Deconvolutional Density Network (DDN) achieve this through discretizing the target space, yet they frequently produce spiked and degenerate densities. This failure arises from gradient isolation in hard binning—where only the target bin receives a learning signal—and is worsened by misalignment between the data distribution and the fixed bin grid. We present SoftFlow-DDN, a redesigned DDN that introduces two synergistic components: a flow transformation, which adaptively warps the target space to align with the grid, and a soft binning strategy, which distributes likelihood across adjacent bins via a temperature-controlled kernel. Together, they ensure global grid compatibility and local smoothness. Ablation studies confirm that both proposed components are necessary and complementary. Empirically, SoftFlow-DDN outperforms the original DDN and recent variants across toy and real-world benchmarks, achieving superior log-likelihood and producing visibly smoother, more realistic densities, particularly in low-data regimes. MCIR-RAG: Multi-Agent Chain-of-Thought Iterative Reflection for Multi-modal Document Understanding Chunming Liu (Hebei University of Engineering); sijia Wang, Ziteng Li, Minghao Hu, Quntian Fang, Yinlong Xiao, Jun Zhang, Zhunchen Luo, Wei Luo, and Zibo Yi (Center of Information Research, AMS, Beijing, China); and Yanping Zhang (Hebei University of Engineering) Abstract Abstract Document Visual Question Answering (DocVQA) is a core task in natural language processing, aiming to provide accurate responses to user queries based on complex documents. The main challenges include efficiently processing long texts, multimodal information such as tables and charts, and handling queries with complex logic. However, traditional Retrieval-Augmented Generation (RAG) frameworks face significant limitations: in the retrieval phase, they tend to introduce redundant information and have insufficient recall accuracy for key content; in the generation phase, they struggle with deep collaborative understanding of multimodal information and cannot support cross-modal reasoning or complex logical inference. To address these issues, this paper proposes the Multi-Agent Chain-of-Thought Iterative Reflection Retrieval-Augmented Generation (MCIR-RAG) framework. In the retrieval phase, the framework performs initial document page screening, generates focused summaries for selected pages, and dynamically reorders them based on these summaries to improve retrieval accuracy. In the generation phase, the framework enhances collaborative understanding of multimodal information and complex logical inference through a multi-agent chain-of-thought and iterative reflection mechanism. R-MALS: Role-Specialized Multi-Agent Scheduling with LLMs under Dynamic Disruptions Yun Liu and Yanan Sun (Sichuan University) Abstract Abstract Dynamic Flexible Job-shop Scheduling plays a critical role in intelligent manufacturing, where scheduling decisions must be continuously adapted to dynamic disruptions. However, existing scheduling algorithms typically focus on either global rescheduling or local decision-making under dynamic disruptions. As a result, these algorithms often lack the ability to explicitly reason about the structural impact of disruptions on schedules, which may limit their effectiveness in dynamic environments. To address this limitation, we propose R-MALS, a role-specialized multi-agent framework that typically integrates large language models (LLMs) into the entire scheduling process to enable explicit reasoning about disruption propagation. Specifically, R-MALS decomposes scheduling into role-specialized agents for initial schedule generation, disturbance analysis, and localized rescheduling, allowing LLMs to reason about disruption propagation and adjust affected operations while preserving stable parts of the schedule. Furthermore, a two-stage training strategy that combines role-specific supervised fine-tuning with constraint-aware direct preference optimization is adopted to better align LLMs reasoning with scheduling constraints. We conduct extensive experiments on the Brandimarte and Fattahi benchmark against state-of-the-art algorithms. The results demonstrate that R-MALS achieves competitive makespan performance in static settings and maintains reliable rescheduling under machine breakdowns, highlighting the effectiveness of integrating LLMs into the entire scheduling process. Tuesday Virtual Room 3 IJCNN Paper SS22 Self-organizing Clustering for Continual Learning and its Applications / SS27 Reservoir Computing for Scalable and Energy-Efficient AI: Theory, Dynamics, and Implementations Session Chair: Shenghao Lyu (Kyoto University), Yi Gong (University Of Electronic Science And Technology Of China) GAL: Global-Anchored Learning for Memory-Free Continual Domain Shift Learning Shenghao Lyu, Xiaofeng Lin, and Kashima Hisashi (Kyoto University) Abstract Abstract Deep learning models deployed in dynamic environments often face a challenge of acquiring knowledge from data in continually shifting unlabeled target domains, which is known as Unsupervised Continual Domain Shift Learning. Existing methods typically require a replay buffer that stores past domain data to assist learning of the current target domain. However, this makes their models risky and unsuitable for privacy-sensitive and scalable environments. In this paper, we consider a more practical setting where data from target domains cannot be retained or shared. To address this challenge, we propose a simple yet effective framework named Global-Anchored Learning (GAL). GAL consists of a core module named Anchored Pre-Learning (APL), and a downstream module named Clean Sample Mining (CSM). APL introduces a pre-learning mechanism that captures target-specific knowledge while staying consistent with existing global knowledge. The acquired knowledge is then aggregated to benefit global optimum. CSM further boosts local adaptation through confident sample selection guided by the knowledge from pre-learning. Experimental results on multiple benchmark datasets demonstrate that GAL achieves state-of-the-art (SOTA) or highly competitive performance compared to replay-based methods. Dancing Between Phases: Self-Organized Learning Transitions in Competing Associative Networks Saleh Sargolzaei and Luis Rueda (University of Windsor) Abstract Abstract We investigate emergent learning dynamics through DANCE (Dynamic Associative Network for Cluster Emergence), a clustering algorithm based on competing Hopfield networks. DANCE exhibits a natural transition from exploration to refinement phases without explicit scheduling. This transition emerges from the adaptive repulsion between networks, which scales inversely with the number of prototypes. By tracking energy evolution, cluster formation rates, and network competition, we demonstrate how the system self-organizes its learning behavior through distinct regimes. The dynamics show how adaptive systems can naturally develop a curriculum, first exploring the feature space broadly before shifting to refining established knowledge, offering insights into emergent learning regimes in neural systems. Quantum Reservoir Computing via Raman-Driven Dynamics of a Single Trapped Ion Jia-Yang Gao (Centre for Quantum Technologies, National University of Singapore) Abstract Abstract Reservoir computing leverages the rich nonlinear dynamics of a physical system to transform input streams into high-dimensional features, enabling simple ridge-regression readout training. While many physical platforms have been explored, the use of a single trapped-ion vibrational mode under monochromatic driving remains largely unexplored. Such a system can exhibit rich dynamical behavior—including non- linear resonances, resonance-cell (localization-like) structures, and chaos-like dynamics—depending on the drive strength and detuning, suggesting potential for reservoir computing. Here, we present a reservoir-computing framework based on the vibrational dynamics of a single trapped ion under Raman- type laser driving. Using numerical simulations, we evaluate the framework on standard benchmarks, including time-series tasks, NARMA, the homogeneous SU(2) Yang–Mills field equation, Mackey–Glass, and Lorenz-63 prediction. Our results show that this system can serve as an effective physical reservoir, achieving strong performance across several benchmarks using only a linear ridge-regression readout. These findings demonstrate a minimal, experimentally accessible route to quantum reservoir computing using a single trapped-ion motional mode. APLoRA: Efficient Multi-Task Fine-Tuning for Long-Context LLMs hongbo chen, yi gong, yueming chen, ruiming wen, Zhaokun Wang, xunlei chen, and wenhong Tian (University of Electronic Science and Technology of China) Abstract Abstract Current large language models (LLMs) deliver strong task generalization, yet adapting them to long-context scenarios is hindered by the quadratic complexity of self-attention and hardware memory constraints. We target single-device fine-tuning and deployment, where extending context length from 4k to 12–20k tokens is difficult but critical for document-centric applications in resource-limited environments. We first present Aggregated Shifted Sparse (ASS) attention, a grouped sparse mechanism that grants global visibility to the first k sink tokens and leverages shifted local windows to retain pre-trained attention patterns. It lowers attention costs, maintains perplexity comparable to dense attention on long sequences, and enables extrapolation beyond training windows via position interpolation. Based on this, we propose APLoRA, a high-throughput multi-task LoRA fine-tuning framework for long-context LLMs. It integrates NF4 quantization, shared-base-model batch fusion, and an adaptive scheduler with Padding-Aware Sequence Alignment (PASA) to optimize memory, padding overhead, and task turnaround time under fixed hardware budgets. Evaluations on Llama2-7B and Llama3-8B across PG19, ProofPile, LongBench, and an internal long-context corpus demonstrate that our method stably extends context length from 4k to 12–20k tokens on a single Ascend 910B NPU. Compared with dense-attention and sequential LoRA baselines, APLoRA boosts effective token throughput by up to 23.9% and cuts end-to-end task time by 19%, while achieving comparable or better performance in perplexity and downstream accuracy than existing long-context adaptation methods. Tuesday Virtual Room 4 IJCNN Paper SS37 AI in Healthcare: Harnessing Emerging, Generative and Agentic Technologies for Responsible Innovation / SS45 Computational Intelligence in Transactive Energy Management and Smart Energy Network (CITESEN) Session Chair: Dongjiao Ge (City University of Macau), Zhiqiang He (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications) A Morphology-Aware Contrastive Learning Network for Detecting a Potential Subtype of Knee Osteoarthritis Zixuan Wang (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications); Hu Li and Hui Li (Arthritis Clinic and Research Center,Peking University People's Hospital); Kai Niu (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications); Jianhao Lin (Arthritis Clinic and Research Center,Peking University People's Hospital); and Zhiqiang He (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications) Abstract Abstract During a recent knee osteoarthritis (KOA) screening in the Tibetan population, clinicians identified a rare subtype characterized by subtle yet distinctive morphological abnormalities. In addition to proximal tibial and distal femoral enlargement with diaphyseal narrowing, this subtype exhibits fine-grained edge irregularities and micro-deformations along bone contours. These morphological cues are critical for clinical diagnosis but are difficult to capture using conventional convolutional neural networks (CNNs), which tend to overfit texture patterns and struggle under extreme data scarcity. In clinical practice, diagnosis relies primarily on careful inspection of bone shape and contour rather than internal texture appearance. Motivated by this clinical observation, we propose a Morphology-Aware Contrastive Learning Network (MACLNet) for automated detection of this KOA subtype. Specifically, MACLNet introduces a morphology-aware enhancement mechanism that employs a wide-edge stream as explicit shape guidance, enforcing the network to prioritize structural morphology over texture-dominant cues. Furthermore, a momentum-based supervised contrastive learning strategy is incorporated to enlarge subtle inter-class differences in the feature space, effectively alleviating overfitting and representation collapse under limited and imbalanced data conditions. Experimental results indicate promising detection performance, suggesting potential utility for broader screening and characterization of this subtype. StructDiffSeg: A Structure-Aware Diffusion Framework for Anatomically Consistent Medical Image Segmentation Xin Yu, Yehang Li, and Hui Zhao (Xinjiang University) Abstract Abstract Diffusion-based segmentation models have demonstrated strong capability in refining object boundaries through stochastic denoising, but they often suffer from structural instability in anatomically constrained scenarios, especially for multi-organ segmentation. In this paper, we propose TopoDiff, a structure-aware diffusion framework that integrates learnable structural priors into the diffusion refinement process to improve anatomical consistency. TopoDiff consists of a deterministic base segmentation branch that maintains class-wise structural priors and a conditional diffusion branch that iteratively refines noisy segmentation masks under image guidance and structural conditioning. A Class-wise Structural Prior Interaction (CSPI) module is designed to model category-specific structural information across encoder scales, while a Topology-Aware Affinity Fusion (TAAF) mechanism aligns diffusion features with structural representations at the bottleneck stage to stabilize denoising trajectories. Unlike purely appearance-driven or boundary-focused diffusion refinement strategies, TopoDiff emphasizes global structural alignment while preserving fine-grained segmentation accuracy. Extensive experiments on the Synapse multi-organ dataset and the ACDC cardiac dataset demonstrate that TopoDiff consistently outperforms diffusion-based and shape-prior-based baselines in terms of Dice score and Hausdorff distance. Qualitative results further show that TopoDiff produces more anatomically plausible segmentations with reduced fragmentation and structural artifacts. Enhancing Asset Visibility in Smart Energy Systems: A Computational Intelligence Framework for Heterogeneous Infrastructure Inspection Li Tan and Yibo Li (Beijing Technology and Business University), Dongjiao Ge (City University of Macau), and Xiaofeng Lian and Yixin Peng (Beijing Technology and Business University) Abstract Abstract The fragmentation of energy infrastructure creates an Asset Visibility Gap, constraining digital twin construction and smart grid scheduling. While UAV imagery bridges this gap, precise perception faces challenges from feature dissipation of tiny assets and boundary aliasing in dense arrays. To address this, we propose HFA-YOLO, a lightweight detection framework based on YOLOv11 tailored for heterogeneous energy assets. The model reconstructs the feature pathway via three mechanisms: First, Space-to-Depth Convolution (SPD-Conv) establishes a lossless down-sampling channel to mitigate feature dissipation. Second, a Detail-Enhanced Unit (DEU) leverages gradient priors to sharpen boundaries and decouple dense components. Finally, a Hybrid Feature Aggregation (HFA) module employs frequency-spatial filtering to enhance weak features and suppress environmental noise. Experiments across photovoltaic, transmission tower, wind turbine, and VisDrone2019 scenarios demonstrate superior generalization. With only 3.06M parameters, the method achieves an optimal accuracy-efficiency trade-off, realizing mAP50 gains of 2.12%, 4.91%, 1.40%, and 4.18% in photovoltaic, transmission tower, wind turbine, and VisDrone2019 scenarios, respectively, compared to the baseline. Ablation studies verify the effectiveness of each component. HFA-YOLO offers a robust solution for the visibility bottleneck of energy assets. The source code is publicly available at https://github.com/HermitKin/HFA-YOLO. Energy-Aware 3D UAV Path Planning in Mountainous Terrain via TDAMSO Li Tan and Jiaqin Chai (Beijing Technology and Business University), Dongjiao Ge (City University of Macau), and Ruimeng Tian and Xinshi Zhang (Beijing Technology and Business University) Abstract Abstract Unmanned Aerial Vehicles (UAVs) are increasingly deployed in energy-intensive transportation tasks. While UAVs enable rapid access in mountainous regions, rugged topography and dense obstacles make 3D path planning a high-dimensional, heavily constrained problem. Under strict safety and feasibility requirements, many existing methods converge slowly and yield high-cost trajectories, increasing energy consumption under battery-limited operations. To address these challenges, this paper proposes Thermal-conduction Dynamic Adaptive Mirage Search Optimization (TDAMSO) for energy-aware planning in complex mountainous environments. TDAMSO integrates Sobol low-discrepancy sequence initialization, a differential thermal-conduction search strategy, nonlinear dynamic step-size adjustment, and an adaptive Cauchy-Gaussian dual-layer local search to accelerate convergence and reduce the overall path cost. Moreover, an energy-oriented 3D path cost model is constructed to unify flight distance, climbing, turning, and smoothness costs, along with terrain and obstacle-related safety constraints. The planned trajectories provide a reusable primitive for mission energy budgeting and green aerial logistics under safety constraints. Multi-scenario simulations show that TDAMSO outperforms other representative metaheuristics and advanced optimization algorithms in terms of total energy-related cost, obstacle-avoidance feasibility, and convergence speed. Overall, TDAMSO provides an effective and reliable approach for energy-efficient UAV transportation in complex terrain. Tuesday Virtual Room 5 IJCNN Paper SS40 Design Challenges, Methods and Applications in Sustainable Neuromorphic Hardware / SS42 Games / SS46 Computationally Intelligent Techniques in Early Prediction and Detection of Brain Disorders Session Chair: Wannian Xia (School of artificial intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences), Tao Hong (Shenyang Aerospace University) A Heterogeneous Feature Fusion Network for EEG Seizure Detection Integrating Spatio-Temporal Channel Networks and Vision Transformer Zhiqiang Chang, Qiaoli Zhou, Bin Guo, and Tao Hong (Shenyang Aerospace University) Abstract Abstract Epileptic seizures pose a significant challenge to patient quality of life, highlighting the need for efficient and automated detection systems to support clinical decision-making. Electroencephalogram (EEG) signals, as a critical diagnostic modality, require models that can effectively capture temporal, spatial, and frequency-domain features, along with their complex interdependencies. However, existing methods often fall short in comprehensively modeling these multi-dimensional characteristics. To address this limitation, we propose a Heterogeneous Feature Fusion Network (HFFNet), which integrates a Spatio-Temporal Channel Network (STCN) and a Vision Transformer (ViT) pre-trained model for robust epilepsy detection. Specifically, raw EEG signals and their corresponding STFT-transformed spectrograms are used as dual-path inputs to extract complementary features. STCN utilizes Distributed Residual LSTM and Graph Neural Network layers to model spatio-temporal dependencies and inter-channel relationships, while ViT captures global frequency-domain representations from the spectrograms. The two subnetworks are independently trained to minimize mutual interference and are subsequently integrated through an attention-based feature fusion module. Experimental results on three public datasets—CHB-MIT, TUSZ, and Helsinki—demonstrate that HFFNet achieves state-of-the-art performance and significantly outperforms existing approaches.Notably, we adopt leave-one-subject-out (LOSO) validation for subject-independent evaluation, confirming the model's strong generalization capability in real-world clinical scenarios. MIDINET: Multi-Slice Inter-regional DistanceBased Individual Structural Brain Network Xiaoye Bin and Chuang Liang (Nanjing University of Aeronautics and Astronautics); Tulay Adali (University of Maryland); Vince Calhoun (Georgia State University, Georgia Institute of Technology, Emory University); and Shile Qi (Nanjing University of Aeronautics and Astronautics) Abstract Abstract Unlike group-level networks, individualized structural brain networks are constructed based on region-to-region structural magnetic resonance imaging (sMRI) connectivity at individual-level to capture individual heterogeneity. However, most construction methods overlook the hierarchical distribution patterns of the brain, limiting their ability to capture subtle microstructural variations. To address this, we propose a novel framework: Multi-slice Inter-regional Distance-based Individual Structural Brain Network (MIDINet), which divides each brain region into anatomically ordered slices and extracts slice-wise multi-features to construct inter-regional similarity networks, thereby preserving spatial patterns and better capturing intra-regional heterogeneity. Results show that MIDINet generally outperforms the alternative approaches across diagnostic tasks and classifiers. The identified interpretable diagnostic biomarkers indicate that structural abnormalities associated with psychiatric disorders are more likely concentrated in higher-order cognition and emotion regions rather than in lower-level sensory or motor areas. In summary, MIDINet provides new insights into individualized structural brain networks and facilitates the precise diagnosis and mechanistic understanding of psychiatric disorders. Neuromorphic visual attention for Sign-language recognition on SpiNNaker Šárka Lísková and Olha Vedmedenko (Czech Technical University in Prague), Mazdak Fatahi (Université de Lille), Matěj Hoffmann (Czech Technical University in Prague), P. Michael Furlong (University of Waterloo), and Giulia D'Angelo (Czech Technical University in Prague) Abstract Abstract Sign-language recognition has achieved substantial gains in classification accuracy in recent years; however, the latency and power requirements of most existing methods limit their suitability for real-time deployment. Neuromorphic sensing and processing offer an alternative paradigm based on sparse, event-driven computation that supports low-latency and energy-efficient perception. Language-Guided Value with Latent Dynamics for Textual Card Game Agent Wannian Xia (School of Artificial Intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences) and Shuang Xu and Bo Xu (Institute of Automation, Chinese Academy of Sciences) Abstract Abstract Textual card games, such as Hearthstone, offer a rich environment for exploring decision-making integrated with language understanding, yet achieving efficient policy learning across various game scenarios remains a significant challenge. In this paper, we tackle this issue by introducing a new framework that integrates large language models (LLMs) with reinforcement learning (RL) agents to improve training efficiency. Our approach leverages a fine-tuned T5 model to encode and interpret card strategies expressed in natural language, yielding a language-guided value function that computes Q-values with language embeddings. We further augment self-play RL training of the agent with an auxiliary transition loss in the latent space. The agent learns from scratch to capture the complex dynamics of the game environment, facilitating efficient policy learning across a large set of cards and decks. Experimental results demonstrate that our method significantly enhances learning efficiency, enabling the agent to effectively learn from a massive dataset of over 5,000 different decks and to generalize on novel cards. These findings highlight the promise of incorporating pretrained LLMs in RL agents for complex decision-making problems. Tuesday Virtual Room 6 IJCNN Paper Advances in Computer Vision I Session Chair: Liang Fan (Loughborough University, ai.io), Wei Shi (Hangzhou Dianzi University) DACAM: Depth-Aware Cross-Attention for Defocus-Deblurred Small Object Detection Liang Fan (Loughborough University, ai.io); Zhi Chen (Ulster University); Baihua Li (Loughborough Univeristy); Yan Liu (Southwest Jiaotong University); and Luyang Zhang (Independent researcher) Abstract Abstract Object detection in precision agriculture is often severely degraded by defocus blur resulting from unconstrained camera placement and challenging outdoor conditions, which particularly affects small or distant livestock targets, such as calves. To address this, we propose a novel transformer-based detection framework that integrates a geometry- and depthguided deblurring network. Our core innovation is a deblurring module that leverages monocular depth estimation to generate geometric priors, which guide a cross-attention mechanism to selectively enhance blurred regions. This targeted feature refinement improves edge and texture clarity in defocused areas, producing detection-friendly features that enable accurate object localization and classification by the subsequent detector. Importantly, our approach allows reliable calf detection using lowcost, fixed-focus cameras in real farm environments, facilitating scalable livestock health monitoring and automated precision agriculture without specialized imaging hardware. We validate our framework on a custom livestock dataset as well as the public DPDD benchmark. Experiments show our integrated model outperforms the RT-DETR baseline by a significant margin, achieving a mean Average Precision (mAP) of 95.7%, particularly with the largest gains observed on small or distant calves, where we observe an improvement of over 10%. IPD3D: Camera-Agnostic Lightweight Monocular 3D Detection via Geometric Decoupling Zihui Gao (Zhejiang University), Xinhai Zhao (Huawei Technologies), and Hao Chen (Zhejiang University) Abstract Abstract Monocular 3D object detection is a critical component of autonomous driving systems, offering a low-cost solution for understanding complex 3D environments. While recent Transformer-based approaches have achieved high accuracy, their excessive computational overhead limits their scalability to real-world, resource-constrained edge devices. Conversely, lightweight CNN-based detectors are efficient but suffer from intrinsic bias—they tend to overfit the specific camera focal lengths of the training set, leading to inferior object localization when applied to cameras with different intrinsic parameters. In this work, we propose a lightweight and robust framework, IPD3D, to bridge this gap. We first demonstrate that standard CNN detectors implicitly couple depth estimation with camera intrinsics, limiting their generalization. To address this, we propose an Intrinsic Parameters Decoupling (IPD) module that learns feature representations in a unified virtual camera space via affine transformations, effectively decoupling the geometric reasoning from specific camera setups. Furthermore, we leverage an Object Reallocation Confidence (ORC) loss to adaptively reallocate regression weights based on the information adequacy of objects. Extensive evaluations on the nuScenes dataset show that our method achieves competitive performance among lightweight models while maintaining a compact size (~200MB) and real-time speed. Most importantly, IPD3D demonstrates excellent cross-camera generalization: it achieves state-of-the-art results in a challenging domain adaptation setting (training on nuScenes and testing on KITTI) without any fine-tuning, validating its effectiveness for generic 3D perception. Frequency-Aware RGB Guidance for 3D Object Detection with Extremely Sparse LiDAR Ruixiao Xu, Lei Tang, and Junzhe Zhang (Chang'an University) Abstract Abstract 3D object detection is a critical task in autonomous driving. However, LiDAR-based detectors suffer severe performance degradation when point clouds become extremely sparse due to cost-effective hardware constraints, long-range attenuation, or environmental signal degradation. While fusing monocular images with sparse LiDAR data is a viable solution to bridge this resolution gap, existing methods often confuse high-frequency textures with low-frequency geometric structures, leading to blurred object boundaries and reduced detection accuracy. We propose a 3D detection framework specifically designed for extremely sparse scenarios (e.g., 1% density). Its core is the Dynamic Frequency Guidance (DFG) module. This module decouples and selectively fuses high- and low-frequency information through content-adaptive filtering, enabling more accurate depth reconstruction even with minimal point cloud data. The resulting dense depth map is converted into pseudo-LiDAR data, which can be directly fed into existing detectors. On the KITTI benchmark, when only 1% of LiDAR points are available, our method achieves AP3D scores of 41.92%, 23.70%, and 18.73% for easy, moderate, and hard settings. Notably, this represents a substantial gain of 12.03% AP3D in the moderate setting over the baseline PV-RCNN, verifying our method’s capability to recover detailed geometry from extremely sparse inputs. DFM: Steering Marigold with DINOv2 Priors for Monocular Depth Estimation Wei Shi, Zixuan Wu, Yucheng Zhu, Chenglan Huang, Tingting Han, and Min Tan (Hangzhou Dianzi University) Abstract Abstract Monocular depth estimation from a single image remains challenging, especially for sharp boundaries and fine-grained textures. To address this, we propose DFM (DINOv2-Fused Marigold), a framework that integrates DINOv2 features into the Stable-Diffusion-based Marigold pipeline via a hierarchical ViT encoder. Unlike conventional feature fusion, our encoder selectively amplifies informative channels while suppressing redundant ones, yielding a discriminative, geometry-aware representation that integrates global structural priors with precise local geometry. This design results in sharper boundaries, richer textures, and improved depth accuracy. Analysis of internal activations indicates that channel-wise co-variation in select DINOv2 channels sparsifies noisy responses while amplifying salient geometric structures, providing insight into the improved robustness and stability observed in our evaluations. Extensive experiments demonstrate that DFM fully outperforms the Marigold baseline on challenging outdoor benchmarks, while exhibiting competitive performance on indoor benchmarks, and simultaneously achieving improved training stability. Tuesday Virtual Room 7 IJCNN Paper Advances in Computer Vision II Session Chair: Xujian Fang (Hangzhou Dianzi University), Pei Zhou (SiChuanUniversity) AMU-Net: Physics-Aware Adaptive Mamba U-Net for Efficient Multi-Lead Time Precipitation Nowcasting Baitian Liu, Haiping Zhang, and Dongjing Wang (Hangzhou Dianzi University); Luan Xu, Ying Li, and Feng Chen (Zhejiang Institute of Meteorological Sciences); and Xujian Fang (Hangzhou Dianzi University) Abstract Abstract Precipitation nowcasting is critical for disaster mitigation but struggles to capture the rapid, nonlinear changes in convective systems. Conventional deep learning methods face two main issues: error propagation from autoregressive prediction, which reduces long-term accuracy, and the tendency to blur and underestimate intensity due to deterministic loss functions that fail to capture extremes. To address this, we propose the Adaptive Mamba U-Net (AMU-Net), a physics-aware framework for multi-lead-time direct forecasting of extreme precipitation. AMU-Net adopts a Unified Direct Forecasting Paradigm with a Physics-Aware Time Embedding and Adaptive Layer Normalization to inject temporal information into the network, enabling a single model to dynamically modulate its feature statistics and adapt to physical laws across lead times. It integrates the Mamba module—a linear-complexity State Space Model—into a U-Net to capture long-range convective advection with global receptive fields without the prohibitive quadratic costs of Transformers. To reduce blurring, a Hybrid Generative Training Strategy combining spectral consistency and adversarial learning recovers high-fidelity textures. Extensive experiments on the SEVIR dataset show that AMU-Net significantly outperforms baselines across both meteorological skill scores and perceptual quality. Ablation studies further validate the necessity of each proposed component, confirming that the unified direct forecasting strategy and hybrid training are essential for reconciling computational efficiency with accurate detection of extreme weather events. Extending Precipitation Nowcasting Horizons via Spectral Fusion of Radar Observations and Foundation Model Priors Yuze Qin, Qingyong Li, Zhiqing Guo, and Wen Wang (Beijing Jiaotong University); Yan Liu (Chinese Academy of Meteorological Sciences); and Yangli-ao Geng (Beijing Jiaotong University) Abstract Abstract Precipitation nowcasting is critical for disaster mitigation and aviation safety. However, radar-only models frequently suffer from a lack of large-scale atmospheric context, leading to performance degradation at longer lead times. While integrating meteorological variables predicted by weather foundation models offers a potential remedy, existing architectures fail to reconcile the profound representational heterogeneities between radar imagery and meteorological data. To bridge this gap, we propose PW-FouCast, a novel frequency-domain fusion framework that leverages Pangu-Weather forecasts as spectral priors within a Fourier-based backbone. Our architecture introduces three key innovations: (i) Pangu-Weather-guided Frequency Modulation to align spectral magnitudes and phases with meteorological priors; (ii) Frequency Memory to correct phase discrepancies and preserve temporal evolution; and (iii) Inverted Frequency Attention to reconstruct high-frequency details typically lost in spectral filtering. Extensive experiments on the SEVIR and MeteoNet benchmarks demonstrate that PW-FouCast achieves state-of-the-art performance, effectively extending the reliable forecast horizon while maintaining structural fidelity. Our code is available at https://anonymous.4open.science/r/PW-FouCast. HoloEv-Net: Efficient Event-based Action Recognition via Holographic Spatial Embedding and Global Spectral Gating Weidong Hao, Rui Fan, Juntao Guan, and Rui Lai (Xidian University) Abstract Abstract Event-based Action Recognition (EAR) has attracted significant attention due to the high temporal resolution and high dynamic range of event cameras. However, existing methods typically suffer from (i) the computational redundancy of dense voxel representations, (ii) structural redundancy inherent in multi-branch architectures, and (iii) the under-utilization of spectral information in capturing global motion patterns. To address these challenges, we propose an efficient EAR framework named HoloEv-Net. First, to simultaneously tackle representation and structural redundancies, we introduce a Compact Holographic Spatiotemporal Representation (CHSR). Departing from computationally expensive voxel grids, CHSR implicitly embeds horizontal spatial cues into the Time-Height (T-H) view, effectively preserving 3D spatiotemporal contexts within a 2D representation. Second, to exploit the neglected spectral cues, we design a Global Spectral Gating (GSG) module. By leveraging the Fast Fourier Transform (FFT) for global token mixing in the frequency domain, GSG enhances the representation capability with negligible parameter overhead. Extensive experiments demonstrate the scalability and effectiveness of our framework. Specifically, HoloEv-Net-Base achieves state-of-the-art performance on THU-EACT-50-CHL, HARDVS and DailyDVS-200, outperforming existing methods by 10.29%, 1.71% and 6.25%, respectively. Furthermore, our lightweight variant, HoloEv-Net-Small, delivers highly competitive accuracy while offering extreme efficiency, reducing parameters by 5.4×, FLOPs by 300×, and latency by 2.4× compared to heavy baselines, demonstrating its potential for edge deployment. Learning-Based Modulation Enhancement for High-Dynamic-Range Fringe Projection Profilometry Kejiang Chen, Pei Zhou, and Jiangping Zhu (Sichuan University) Abstract Abstract Fringe Projection Profilometry (FPP) is a key technique for optical 3D surface measurement, yet it remains highly sensitive to high dynamic range (HDR) reflective surfaces, such as metallic objects. Strong specular reflections often lead to overexposure, while insufficient illumination causes underexposure, resulting in degraded fringe modulation, phase distortions, and incomplete 3D reconstructions. These issues severely limit the accuracy and reliability of FPP in industrial scenarios. Existing solutions, including surface treatment and multi-exposure HDR imaging, are either invasive or computationally expensive, restricting their practical deployment. To overcome these challenges, we propose a learning-based modulation enhancement network that explicitly models and corrects modulation distortions in HDR fringe patterns. The proposed framework introduces a Modulation Offset Estimation (MOE) module to accurately estimate modulation offsets in over- and under-exposed regions, followed by a Modulation Intensity Fusion (MIF) module that adaptively integrates the corrected modulation features to generate enhanced fringe images. This design enables effective restoration of fringe modulation while preserving spatial consistency across HDR regions. Tuesday Virtual Room 8 IJCNN Paper Advances in Computer Vision III Session Chair: Shuo Zhang (Shanghai Dianji University; School of Electronic and Information,), Zheng zhihao (Institute of Information Engineering, Chinese Academy of Sciences; University of Chinese Academy of Sciences) CDSegNet++: Enhancing Diffusion-Based Point Cloud Segmentation with Adaptive Scheduling and Consistency Fusion Dongdong An (Intelligent Society and Scientific Computing Laboratory, Shanghai Normal University); MingYu Shou (College of Information Mechanical and Electrical Engineering, Shanghai Normal University); Shuo Zhang (School of Electronic Information Engineering, Shanghai DianJi University); and Wenbing Tang (Northwest A&.F University) Abstract Abstract Point cloud semantic segmentation is challenging due to irregular sampling, sparsity, and frequent class-boundary ambiguity. Diffusion-based segmentors are promising for their robustness to noise, yet existing designs often (i) rely on fixed noise schedules that ignore scene difficulty, (ii) use a single prediction stream that tends to blur boundaries, and (iii) fuse conditional features in a largely indiscriminate manner. We propose an enhanced Conditional-Noise Framework with three components. First, an Adaptive Diffusion Scheduling Strategy (ADSS) adjusts the injected noise level conditioned on scene complexity, improving denoising stability across easy and hard regions. Second, a Dual-Stream Semantic Consistency (DSSC) module introduces an auxiliary prediction head and enforces semantic consistency between streams to sharpen object boundaries. Third, a Lightweight Adaptive Fusion Module (LAFM) employs gated fusion to selectively integrate multi-level conditional features while suppressing noisy cues. Extensive experiments on ScanNet and SemanticKITTI demonstrate consistent gains over strong baselines, achieving 67.50% mIoU (+4.23%) and 41.94% mIoU (+1.47%), respectively, while retaining efficient single-step inference. DuDoST: A Dual-Domain Self-Denoising Transformer for Robust Semantic Segmentation Jinming Li, Wenzhu Yang, and Shengbo Wang (Hebei University) Abstract Abstract Noise interference and semantic instability are key factors that limit the generalization ability and prediction consistency of Transformer-based semantic segmentation models. Although the global modeling capability of Transformers helps alleviate local ambiguity, unconstrained global attention often introduces noise during the encoding stage and amplifies semantic instability. In addition, the lack of explicit modeling of predicted mask structures causes errors to continuously accumulate during decoding and upsampling, ultimately affecting the overall consistency of segmentation results. To address these issues, we propose a Dual-Domain Self-Denoising Transformer (DuDoST), which effectively suppresses background noise and enhances foreground region representations while exhibiting stronger structural awareness and self-repair capability. The Hierarchical Semantic Prior Gating (HSPG) embedded in DuDoST effectively alleviates the interference of background noise in multi-scale features on the Transformer decoding process and enhances the model's attention to semantically relevant regions. The Mask-domain Diffusion Denoiser (MaDD) introduces conditional diffusion modeling in the mask domain, significantly improving the structural consistency and robustness of segmentation results without requiring additional post-processing. Extensive experiments on the Cityscapes and ADE20K datasets demonstrate that the proposed method achieves state-of-the-art performance, obtaining 82.9% mIoU on Cityscapes and 50.5% mIoU on ADE20K. Personalized Diffusion-based Outfit Recommendation with Diffusion-Aware Stochastic Perturbation Hangjie Xia, Yifan Huo, Junhong Zheng, and Lili He (Zhejiang Sci-Tech University) Abstract Abstract Personalized outfit recommendation requires modeling both item compatibility and user preferences. However, given a specific context, multiple valid completion options often exist, and traditional point estimation models struggle to capture this ``one-to-many'' multimodal uncertainty. Existing diffusion models typically rely on static conditional guidance, overlooking noise variations during the denoising process and lacking domain-specific knowledge constraints. This paper proposes the Personalized Diffusion-based Outfit Recommendation (PDOR) framework, which formulates outfit completion as conditional probability distribution learning within a latent space. PDOR injects dynamic perturbations via a diffusion-aware stochastic perturbation mechanism during denoising to learn robust semantic dependencies. It utilizes compatibility reward guidance to provide explicit stylistic constraints during training and employs adaptive uncertainty regularization to prevent feature convergence. The framework integrates a ``generation-retrieval'' paradigm, which achieves a balance between recommendation precision and inference efficiency by generating an ``ideal prototype'' in a continuous latent space followed by nearest neighbor retrieval. Experimental results demonstrate that PDOR significantly outperforms existing methods in compatibility prediction (CP AUC: 0.9074), outfit completion (FITB: 0.7117), and retrieval (HR@10: 0.3264) tasks on the iFashion and Polyvore-630 datasets. Ablation studies confirm the necessity of each module and the critical role of the perturbation mechanism. Drift-Aware Conditional Diffusion with Compact Feature Representations for Malicious Traffic Detection Zheng zhihao, Xia Wei, and Ma Shituo (Institute of Information Engineering, Chinese Academy of Sciences; School of Cybersecurity, University of Chinese Academy of Sciences); Luo Xiang (Institute of Information Engineering, Chinese Academy of Sciences; School of Cybersecurity,University of Chinese Academy of Sciences); and Xiong Gang and Gou Gaopeng (Institute of Information Engineering, Chinese Academy of Sciences; School of Cybersecurity, University of Chinese Academy of Sciences) Abstract Abstract The rapid evolution of cyber-attack techniques poses significant challenges to malicious traffic identification, particularly under concept drift. Existing drift-aware approaches largely rely on autoencoder-based reconstruction, whose limited expressive capacity and generalization hinder the detection of subtle distributional shifts among diverse malicious traffic types. To address this limitation, we propose FlowDiff, a drift-aware framework that leverages class-conditional diffusion-based reconstruction in a compact feature space. FlowDiff first learns stable and discriminative compact representations through a Compact Feature Module (CFM). On top of these representations, a Conditional Diffusion Model (CDM) captures class-conditional feature distributions and enables drift detection via a loss-adaptive diffusion mechanism. In addition, a lightweight LightGBM classifier is employed for accurate in-distribution traffic classification. Extensive experiments on real-world malicious traffic datasets demonstrate that FlowDiff consistently outperforms state-of-the-art methods in closed-world classification, drift detection, and open-world recognition, with particularly significant improvements in detecting previously unseen malicious traffic. Tuesday Virtual Room 1 IJCNN Paper Advances in Computer Vision IV Session Chair: Tian Lan (Renmin University of China, Center for Applied Statistics and School of Statistics), ZhiJian Fang (Zhejiang Sci-Tech University) Diff-Feat: Diffusion-Based Cross-Modal Feature Extraction for Multi-Label Classification Tian Lan, Yiming Zheng, and Jianxin Yin (Renmin University of China, Center for Applied Statistics and School of Statistics) Abstract Abstract Multi-label classification has broad applications and depends on powerful representations capable of capturing multiple interactions. We introduce \textit{Diff-Feat}, a simple yet effective framework that extracts intermediate features from pre-trained diffusion-Transformer models for both images and text, and fuses them for downstream tasks. We observe distinct patterns in optimal feature selection: for vision tasks, the most discriminative intermediate feature along the diffusion process occurs at the middle step and is located at the middle block in Transformer. In contrast, for language tasks, the best feature occurs at the noise-free step and is located at the deepest block. Notably, a consistent phenomenon across diverse datasets: a fixed "Middle Layer" yields the best performance on various downstream classification tasks for images (under DiT-XL/2-256$\times$256 and various U-ViT architectures). We further analyze the underlying mechanism and its scaling behavior across model scales. To efficiently select representations, we devise a heuristic local-search algorithm that pinpoints the locally optimal "image-text"$\times$"block-timestep" pair among a few candidates, avoiding an exhaustive grid search. A simple fusion-linear projection followed by addition-of the selected representations achieves 46.1\% mAP on Visual Genome 500 (with 500 categories) and 80.2\% mAP on Pix2Cap-COCO (with 80 categories), consistently outperforming strong baseline methods. We will make our code available upon acceptance. FaceOff: Preventing Unauthorized Text-to-Image Identity Customization Mingwang Hu and Chenyu Zhang (School of New Media and Communication, Tianjin University); Zili Yi (School of Intelligence Science and Technology, Nanjing University); and Lanjun Wang (School of New Media and Communication, Tianjin University) Abstract Abstract Recently, tuning-free techniques such as PhotoMaker have advanced text-to-image identity (ID) customization. Unlike tuning-based approaches, these techniques employ powerful facial encoders to extract ID information from a single portrait, enabling efficient customization in a single inference pass. However, their misuse can exacerbate the generation of misleading and harmful content, endangering individuals and society. To mitigate this risk, existing methods introduce protective perturbations into user portrait images to distort ID information in the customized images. We identify two limitations of these methods: 1) copyright infringement risk caused by identified faces in customized images; 2) lack of robustness against image transformations. To address these limitations, we propose FaceOff, a framework that protects portrait images from the threat of tuning-free ID customization by disturbing the ensemble of image encoders. Specifically, to reduce the recognized faces in customized images, we design a contrastive loss to shift the image semantics to a target ID and deviate the image semantics away from the original ID. In addition, we introduce a Gaussian augmentation module to mitigate the protection degradation under image transformations. FaceOff outperforms the SOTA method by 12.65% of FDFR and 4.70% of ISM across three customization models and two facial benchmarks on average. The code is available at https://github.com/huzimun/FaceOff. STEDiff: Strengthening Text Embedding for Text-to-Image Alignment in Diffusion Models Hailan Zhang and Haipeng Liu (Hefei University of Technology), Bo Fu (Liaoning Normal University), and Yang Wang (Hefei University of Technology) Abstract Abstract Although pretrained text-to-image (T2I) generation models can produce high-quality images, they often fail to faithfully reflect the semantic intent of complex prompts due to stochastic noise and inherent model limitations. This issue frequently manifests as the model overlooking specific objects or failing to correctly bind attributes to their corresponding entities—a challenge referred to as semantic alignment. Unlike existing approaches that rely on computationally expensive fine-tuning or labor-intensive layout priors, we propose STEDiff, a training-free method designed to enhance semantic representations directly within the text-embedding space. Specifically, we introduce a method that primarily leverages [EOT] tokens to strengthen the relevant semantics of sub-sentences, and then replaces the corresponding tokens in the original prompt. Furthermore, a novel semantic enhancement loss is incorporated to enforce spatial constraints, ensuring that the semantics of each entity are precisely mapped to their respective image regions. Extensive quantitative and qualitative evaluations on the T2I-CompBench demonstrate that our method significantly improves semantic consistency and generation integrity in complex scenarios. Contrastive Learning and Co-Attention for Hierarchical Multi-Label Patent Classification Xin Hu, Huaxiong Zhang, and Zhijian Fang (Zhejiang Sci-Tech University) and Lei Zhou (Zhejiang Talent Development Group) Abstract Abstract Patent classification aims to automatically assign one or more IPC codes to patent documents. To address representation challenges caused by the highly technical nature of patent language and the prevalence of self-defined terminology, as well as the insufficient modeling of label semantics–text associations, we propose a hierarchical multi-label classification framework. Specifically, we weight contrastive constraints by hierarchical label-path similarity to avoid treating semantically close samples as negatives; we further design a co-attention mechanism to fuse label semantics with patent text and strengthen fine-grained matching signals, thereby improving the model’s ability to characterize category decision boundaries. Experiments on USPTO-200K show statistically significant gains over PatCLS with up to 2.63% relative improvement in NDCG@3. Tuesday Virtual Room 2 IJCNN Paper Advances in Computer Vision V Session Chair: pengcheng lu (Sichuan University), Mu Yu (Tongji University) Enhancing Multi-View 3D Object Detection with 2D Auxiliary Information and Feature Fusion Mu Yu and Chaojie Zhang (Tongji University) Abstract Abstract Multi-view 3D object detection has gained increasing attention for its ability to construct a comprehensive 3D scene representation in complex scenarios. However, existing sparse query-based detectors often exhibit performance degradation in crowded scenes or long-range scenarios, primarily due to the reliance on static queries. In this paper, we propose EAFF, an enhanced multi-view 3D object detection framework that leverages 2D auxiliary supervision and adaptive feature fusion to address these limitations. Specifically, EAFF employs a multi-task 2D neural network to extract object-aware semantic representations. These representations are further leveraged to construct adaptive 3D queries, providing more informative and flexible queries than fixed sparse ones. Moreover, we design an attribute-aware feature fusion module that explicitly aligns and integrates attribute-specific features within a unified embedding space, enabling more consistent feature alignment and improved discrimination. Extensive experiments on the nuScenes benchmark demonstrate that EAFF consistently improves detection performance and achieves a 1.3% improvement in mAP compared to the state-of-the-art RayDN method, while maintaining robust detection performance in complex environments. Enhanced Oriented Object Detection in Remote Sensing Using Deformable Quadrilateral Attention and PCA Li-Dan Kuang, Junwu Xie, Yan Gui, Jianming Zhang, and Xiaoyong Tang (Changsha University of Science & Technology) Abstract Abstract Current oriented object detection methods often underutilize the distinctive properties of rotated bounding boxes. To bridge this gap, we introduce a novel approach centered on a learnable Deformable Quadrilateral Attention (DQA) mechanism, explicitly tailored for oriented detection. DQA generalizes window-based attention by dynamically adapting to a flexible quadrilateral formulation. An end-to-end regression module predicts a transformation matrix that deforms a default window into an oriented quadrilateral, enabling adaptive token sampling and enriched context modeling for variably shaped and rotated objects. Furthermore, we integrate Principal Component Analysis (PCA) into the detection head to efficiently deduce rotated bounding boxes from the predicted quadrilateral, enhancing both accuracy and efficiency. Combined, our DQA head with PCA achieves state-of-the-art results: 98.90% mAP on HRSC2016, 79.38% mAP on DOTA-v1.0, and 67.24% mAP on DIOR-R, demonstrating its strong capability for oriented object detection. Multi-Oriented Open-Set Text Recognition via Direction-Aware Channel-Adaptive Fusion Mengdie Han, Fang Yang, and Xiufen Miao (Hebei university) Abstract Abstract The open-set text recognition task aims to recognize unknown text categories in complex scenes, but large variations in character sets, text lengths, and orientations increase difficulty. Existing methods, such as MOoSE, use identical-structure experts for different orientations, but still lack sufficient directional modeling and effective multi-scale feature fusion. To address this, we propose a Direction-Aware and Channel-Adaptive (DACA) framework, integrating a Direction-Aware Attention (DAA) module to enhance directional feature extraction for multi-oriented text. Each expert is specialized for a different text orientation pattern. We also design a Channel-Adaptive BiFPN (CA-BiFPN) for improved multi-scale feature fusion. Experiments show that DACA improves line-level accuracy by 4.99% over MOoSE on a multi-directional open-set benchmark, increases F-score for rejected samples by 3%–7%, and achieves 0.9\% higher accuracy on horizontal text, demonstrating significant advantages in multi-directional text recognition. Enhancing Semi-Supervised Change Detection via Fractional Differential Perturbation Pengcheng Lu and Yifei Pu (Sichuan University) Abstract Abstract Semi-supervised change detection (SSCD) aims to learn robust pixel-level change maps from limited annotated bi-temporal images and abundant unlabeled pairs. Recent consistency-regularization frameworks such as FixMatch and UniMatch rely on strong augmentation to enforce weak-tostrong prediction consistency, but hard-occlusion methods such as Cutout and CutMix may disrupt the topology of ground objects in remote sensing scenes. We therefore propose Fractional Differential Perturbation (FDP), a structure-preserving strong augmentation strategy based on fractional calculus. FDP injects controlled high-frequency perturbations while preserving low-frequency geometric structure, thereby increasing texturelevel difficulty without destroying object boundaries. Extensive experiments on the WHU-CD and LEVIR-CD datasets show that FDP consistently improves semi-supervised baselines, especially in low-label regimes (e.g., a 5.3% IoU gain on WHU-CD with 5% labels). Tuesday Virtual Room 3 IJCNN Paper Advances in Machine Learning I Session Chair: 欣 黎 (华中师范大学), Bin Fang (Chongqing university) Cross-Modal Prior Knowledge Guided Dual-Path Video Captioning Langping Wang, Bin Fang, Mengdi Li, and Ningkai Zhong (Chongqing university) Abstract Abstract Video captioning aims to generate accurate textual descriptions from videos, bridging visual and linguistic understanding. While existing methods have advanced multimodal representation learning during training, they often lack explicit cross-modal semantic guidance during inference, leading to inaccurate or generic descriptions in complex scenes. To address this, we propose CMG-DP, a novel framework that introduces textual prior knowledge for real-time guidance at inference. CMG-DP consists of three key components: a Prior Knowledge Synergy mechanism to retrieve relevant textual descriptions, a Semantic Enhancement module to refine visual semantics, and a Dual-Path Cross-Attention decoder for adaptive fusion. Experiments on MSR-VTT and MSVD benchmarks show that CMG-DP achieves state-of-the-art performance, with CIDEr scores of 58.3 and 124.6, respectively, demonstrating its effectiveness in enhancing caption accuracy and robustness. Event-Anchored Semantic Augmentation for Video Captioning Langping Wang, Bin Fang, Mengdi Li, and Ningkai Zhong (Chongqing university) Abstract Abstract To address insufficient fine-grained perception and noise interference in video captioning, this paper proposes EAS-Cap (Event-Anchored Semantic Captioning), an event-aware semantic-enhanced model. EAS-Cap decomposes videos into atomic events and aligns visual action features with textual event representations via cross-modal knowledge distillation, enhancing semantic perception. A Visual-Text Alignment (VTA) module suppresses noise and improves coherence by mapping visual features into a textual semantic space with contrastive learning. A Visual-to-Event (V2E) constraint further ensures comprehensive event capture while filtering irrelevant interference. Experiments on MSVD and MSR-VTT show that EAS-Cap outperforms existing methods, achieving CIDEr scores of 120.7 and 58.9 respectively, and generates accurate, semantically rich captions. ClusRefineRec: Cluster-Guided Embedding Refinement for Sequential Recommendation Jiangnan Gu, Xinru Liu, and Shengjun Liu (Central South University, School of Mathematics and Statistics) Abstract Abstract While deep learning has significantly advanced sequential recommendation, current methods typically rely solely on back-propagation to learn item embeddings. Consequently, under sparse interaction data, these models often struggle to fully capture the latent relational structures among items and their inherent semantic consistency, resulting in insufficiently expressive embedding representations. To address these challenges, we propose ClusRefineRec, a clustering-based embedding refinement framework. Specifically, we introduce a clustering loss that is decoupled from the main recommendation objective to promote the coherence among positive item embeddings. After each training epoch, we further refine the entire embedding space based on the learned centroids. Furthermore, we design a dynamic state fusion module that adaptively integrates hidden states from all time steps to enhance the final sequential representation. Extensive experiments demonstrate that ClusRefineRec improves both accuracy and robustness while maintaining a plug-and-play property, enabling seamless integration into a wide range of existing sequential recommendation models. FREQ: Frequency-Domain Enhanced Cross-Domain Sequential Recommendation via Item Popularity Dynamics Xin Li, Yong Zhang, Jing Wang, and Shuo Li (Central China Normal University) Abstract Abstract Sequential recommendation predicts users’ next in terests from historical interactions, while cross-domain recom mendation transfers this capability across domains but is often hindered by data sparsity and temporal-rhythm discrepancy. To address this, we propose FREQ, a frequency-domain en hancement module built on PREPREC. FREQ constructs multi scale item popularity windows and applies the discrete Fourier transform to obtain spectral representations, enabling better modeling of long- and short-term dynamics across domains. It further introduces residual frequency-domain injection with lightweight gating and a cross-scale consistency regularization to improve representation stability and transferability. Experiments show that FREQ typically outperforms strong baselines, with larger gains when the source and target domains exhibit greater periodic discrepancy. Tuesday Virtual Room 4 IJCNN Paper Advances in Machine Learning II Session Chair: Jiahui Zhang (Lancaster University), Hongbo Zhao (East China Normal University) Revisting Out-of-distribution Detection in Multi-instance Learning Jiahui Zhang and Yukun Cui (Lancaster University), Mei Yang (Southwest Petroleum University), Zhen Pan (Chengdu Normal University), Xuemei Cao (Southwestern University of Finance and Economics), and Yu-Xuan Zhang (Southwest Jiaotong University) Abstract Abstract Out-of-distribution (OOD) detection is a critical step in ensuring the security and reliability of machine learning systems. However, with the surge of large-scale datasets and complex applications, the efficacy of traditional OOD methods is increasingly limited under weak supervision, and their adaptation to the popular multi-instance learning (MIL) framework proves challenging. Existing MIL-OOD methods often either directly adapt traditional OOD techniques or only perform transfer learning for OOD samples (bags), thereby overlooking the inherent hierarchical structure of instances within each bag in the MIL context. As a result, they struggle to distinguish in-distribution (ID) and OOD bags effectively. To address this challenge, we propose a clustering-aware MIL-OOD detection method (CluMILOOD). This method leverages the MIL attention mechanism to jointly model both the fused intra-bag feature (FIBF) and the clustered inter-bag feature (CIBF), capturing both local and global structure. We design four OOD scoring strategies, i.e., maximum distance, minimum distance, mean distance, and weighted average distance, to evaluate the unknown bags. We conducted experiments on three OOD detection tasks and compared them with 10 state-of-the-art OOD detection methods. The results show that CluMILOOD significantly outperforms rivalss in overall performance, demonstrating stronger detection capabilities and stability. Beyond Energy Smoothing Propagation in Graph Out-of-Distribution Detection Xiaolong Fan, Yuxuan Liu, Maoguo Gong, Mingyang Zhang, Xiangming Jiang, Jianzhao Li, and Shanfeng Wang (Xidian University) Abstract Abstract Recently, energy propagation based methods have shown considerable success in graph out-of-distribution (OOD) detection. However, existing approaches predominantly adhere to an energy smoothing propagation framework. A key limitation arises when nodes from two distinct distributions are directly connected, as this can cause the propagation mechanism to fail, thereby compromising detection reliability. To mitigate the limitations of employing energy smoothing, we propose augmenting the propagation framework with an energy difference propagation. By integrating energy smoothing and difference propagation, we propose SDGSafe, a smoothing and difference graph out-of-distribution detection method, for graph out-of-distribution detection. This results in a hybrid propagation scheme where the operation of each central node is determined by its neighbors' distribution heterophily, i.e., energy smoothing propagates across edges linking nodes of the same distribution, while energy difference operates across edges that bridge different distributions. Extensive experiments on benchmark datasets demonstrate that our method substantially outperforms the baseline that relies solely on smooth energy propagation, confirming the effectiveness of the integrated approach. Adaptive Anchor Clustering and Lightweight Bidirectional Fusion for UAV-Based Fire Detection Xuepeng Huang and Zhiqiang Liu (Inner Mongolia University of Technology), Xu Zhang (Inner Mongolia Technical University of Construction), and Wenjing Li (Inner Mongolia University of Technology) Abstract Abstract Early fire detection is critical for reducing casualties and property losses; however, existing vision-based methods often exhibit limited robustness in complex scenarios involving small targets and heterogeneous data distributions. To address these challenges, this paper proposes an adaptive fire detection framework based on structural optimization of YOLOv11, which enhances scale representation and multi-scale feature fusion. Specifically, a K-means–based clustering strategy is introduced to better capture the scale distribution of flame and smoke targets, providing prior information to improve the matching process, while a lightweight bidirectional feature pyramid network (BiFPN) incorporating depthwise separable convolutions is employed to strengthen fine-grained feature representation for small and sparse fire targets with limited computational overhead. Extensive experiments conducted on a multi-source fire dataset demonstrate that the proposed method achieves 78.443\% $mAP_{50}$, 56.327\% $mAP_{50–95}$, and 70.019\% Recall, outperforming the baseline YOLOv11 by 4.131\%, 3.646\%, and 5.402\%, respectively. Further ablation and cross-dataset generalization experiments verify the effectiveness and complementarity of the proposed components, particularly in small-target and multi-scale scenarios. These results indicate that the proposed approach provides a robust and efficient solution for real-world fire detection applications. Scale-Aware Positional Reparameterization for Long-Context LLMs Yihong Huang and Hongbo Zhao (East China Normal University) Abstract Abstract Large language models (LLMs) suffer from rapidly deteriorating performance once the input sequence exceeds the pre-training context window. Existing methods mitigate this issue through interpolation or extrapolation. Despite being effective, these strategies either incur additional computational overhead or lack flexibility across varying context lengths. In this work, we propose Scale-Aware Reparameterization for long-context eXtrapolation (SAReX), a training-free scale-aware method to extend LLMs' extrapolation capability. SAReX formulates positional reuse as an optimization problem, minimizing a distance-based objective to adaptively determine group sizes, thereby enabling more effective utilization of pre-trained positional indices. In addition, we introduce an anchor-guided group scheme selection mechanism to adapt to inputs of varying lengths, avoiding per-length optimization and incurring no additional inference-time computational overhead. By integrating adaptive positional reuse with anchor-guided scheme selection, SAReX provides a flexible framework for long-context extrapolation that adapts seamlessly to varying lengths. Extensive experiments demonstrate that SAReX consistently achieves up to 4.35% absolute gains over existing training-free extrapolation methods across diverse base models, further validating its robustness and effectiveness. Tuesday Virtual Room 5 IJCNN Paper Advances in Machine Learning III Session Chair: Ruifan Li (Beijing University of Posts and Telecommunications), Qicong Wang (Xiamen Univercity) Multi-Modal Large Model for Robotic Grasping with Multi-Manifold Priors Haibo Li, Qin Lai, Yaming Yang, Yaoxin Chen, and Qicong Wang (Xiamen Univercity) Abstract Abstract Executing instruction-conditioned grasping in cluttered scenes remains challenging because semantic grounding and physical feasibility must be handled jointly. We propose a multimodal 3D perception framework guided by multi-manifold priors that impose geometric constraints for robust grasp prediction. To mitigate error accumulation in cascaded pipelines, we adopt a parallel dual-branch architecture: a semantic branch grounds language in vision, and a geometric branch generates physically plausible grasp candidates from point clouds. Our key idea is to encode grasp representations in non-Euclidean spaces: hyperbolic space captures hierarchical organization of grasp regions, while spherical space preserves directional consistency for grasp orientation. Building on this principle, we introduce (i) a shared-attention point cloud encoder to strengthen cross-layer feature consistency, (ii) a scene-aware and geometry-enhanced grasp pose generation network, and (iii) multi-manifold consistency learning with geodesic constraints. Experiments demonstrate improved accuracy and robustness in complex environments. SCOPE: Bridging Global Form and Local Geometry for Urban Street Network Generation Wenxuan Guo, Yaohui Jin, and Yanyan Xu (Shanghai Jiao Tong University) Abstract Abstract Generative modeling of urban street networks poses a distinct challenge in geometric deep learning. It requires synthesizing large-scale graphs that respect physical constraints while capturing statistical patterns shared across real-world cities. Existing approaches often struggle to balance global spatial coherence with precise topological connectivity. Furthermore, many are limited to modifying reference templates rather than generating diverse, novel structures from scratch. In this work, we propose a scalable two-stage framework that effectively decouples these modalities. We employ an autoregressive Transformer to learn diverse node distributions from coordinate sequences, and introduce SCOPE, a specialized edge predictor that conditions on global spatial context while respecting local geometric constraints. Trained on over 35,000 cities, our approach outperforms baselines with 0.91 average precision in edge prediction and faithfully reproduces high-order graph statistics. Qualitatively, it synthesizes physically plausible networks, demonstrating the capacity to hypothesize novel urban forms beyond the training distribution. Additionally, we show that the learned geometric embeddings intrinsically capture spatial laws, offering interpretability alongside generative performance. CRoSS: A Continual Robotic Simulation Suite for Scalable Reinforcement Learning with High Task Diversity and Realistic Physics Simulation Yannick Denker and Alexander Gepperth (Fulda University of Applied Sciences) Abstract Abstract Continual reinforcement learning (CRL) requires agents to learn from a sequence of tasks without forgetting previously acquired policies. In this work, we introduce a novel benchmark suite for CRL based on realistically simulated robots in the Gazebo simulator. Our Continual Robotic Simulation Suite (CRoSS) benchmarks rely on two robotic platforms: a two-wheeled differential-drive robot with lidar, camera and bumper sensor, and a robotic arm with seven joints. The former represent an agent in line-following and object-pushing scenarios, where variation of visual and structural parameters yields a large number of distinct tasks, whereas the latter is used in two goal-reaching scenarios with high-level cartesian hand position control (modeled after the Continual World benchmark), and low-level control based on joint angles. For the robotic arm benchmarks, we provide additional kinematics-only variants that bypass the need for physical simulation (as long as no sensor readings are required), and which can be run two orders of magnitude faster. CRoSS is designed to be easily extensible and enables controlled studies of continual reinforcement learning in robotic settings with high physical realism, and in particular allow the use of almost arbitrary simulated sensors. To ensure reproducibility and ease of use, we provide a containerized setup (Apptainer) that runs out-of-the-box, and report performances of standard RL algorithms, including Deep Q-Networks (DQN) and policy gradient methods. This highlights the suitability as a scalable and reproducible benchmark for CRL research. MMT-VAE: Controllable Multi-Modal VAE for Malicious Encrypted Traffic Generation in 5G Networks Yongjun Huang (Beijing University of Posts and Telecommunications), Pengfei Du and Shiting Xu (Shandong University of Political Science and Law), and Zehong He and Ruifan Li (Beijing University of Posts and Telecommunications) Abstract Abstract Encrypted traffic dominates modern 5G mobile core and edge networks and increasingly carries sophisticated attacks hidden behind TLS tunnels and protocol obfuscation. Training robust encrypted-traffic IDS models is hindered by the scarcity of labeled malicious traces and by privacy constraints that limit data sharing. We propose MMT-VAE, a controllable multi-modal generative framework that models what defenders can reliably observe under modern TLS: record/packet sizes and timing plus coarse protocol metadata, rather than assuming ciphertext byte values are informative. Specifically, MMT-VAE encodes (i) a window-level TLS record/packet length histogram (ciphertext lengths), (ii) an inter-packet timing graph, and (iii) protocol metadata into modality-specific latent codes. A condition-dependent hyper-prior aligns these latents under requested conditions (attack type, protocol family, and deployment context), and a cross-attention fusion block enables condition-aware generation. We further introduce a continuous obfuscation-strength control at sampling time to generate progressively harder variants for ``what-if'' evaluation. On a mixed real and controlled 5G TLS corpus, augmenting downstream detectors with our generated windows improves macro-F1 by up to 6.3 under cross-scenario evaluation. Tuesday Virtual Room 6 IJCNN Paper Advances in Machine Learning IV Session Chair: Hai-Tao Zheng (Tsinghua University), Jian-Yu Li (Nankai University) Revisiting Classification Taxonomy for Grammatical Errors Deqing Zou and Jingheng Ye (Tsinghua University); Yulu Liu (University of Electronic Science and Technology of China); Yu Wu, Zishan Xu, Zihua Lan, Yinghui Li, and Hai-Tao Zheng (Tsinghua University); Lan Zhou (Shenzhen Giiso Information Technology Co., Ltd); and Hong-Gee Kim (Seoul National University) Abstract Abstract Grammatical error classification plays a crucial role in language learning systems, but existing classification taxonomies are often adopted without rigorous validation, which can lead to inconsistencies and unreliable feedback. In this paper, we revisit previous classification taxonomies for grammatical errors by introducing a systematic and quantitative evaluation framework. Our approach examines four complementary aspects of taxonomy quality, namely exclusivity, coverage, consistency, and usability. To support this evaluation, we construct a highquality grammatical error classification dataset annotated with multiple classification taxonomies and evaluate them under our proposed evaluation framework. Experimental results reveal clear drawbacks and limitations of existing taxonomies. Overall, our contributions aim to improve the precision and effectiveness of error analysis, providing more understandable and actionable feedback for language learners. CogNote: Generating Structured Client-Centric Counseling Notes from CBT Dialogues Tiantian Chen, Xuri Chen, and Ying Shen (Tongji University) Abstract Abstract Mental counseling plays a crucial role in the prevention and mitigation of mental health disorders. Counseling notes—concise summaries of key session elements—are valuable to both counselors and clients for later review and for consolidating what was discussed. However, prior work on counseling note generation is largely counselor-centric, focusing on documenting clients’ symptoms, problems, and diagnoses, while under-serving the client’s need to revisit insights, cognitive shifts, and actionable takeaways after the dialogue. To bridge this gap, we introduce client-centric counseling note generation, which aims to produce client-facing notes that support post-session reflection and everyday application. Concretely, we build C-TIND, a dataset of 1,800 cognitive behavioral therapy (CBT) dialogues paired with structured client-centric notes. Building on C-TIND, we develop CogNote, a client-centric counseling note generation model trained on CBT dialogues covering four common CBT techniques, producing structured notes that help clients review key gains and insights from the session. Both automatic and human evaluations show that our approach can generate useful, well-structured client-centric notes. To the best of our knowledge, this is the first work that explicitly formulates and benchmarks counseling note generation from the client’s perspective. Conditional LIME: Building Incremental Insights with Class-Specific Conditions Lige Gan and Guangzhi Qu (Oakland University) Abstract Abstract Explainable Artificial Intelligence (XAI) is crucial for understanding complex models like deep neural networks, thereby enabling trust, debugging, fairness assessment, and regulatory compliance, particularly in high-stakes domains such as healthcare and manufacturing. Local Interpretable Model-agnostic Explanations (LIME) stands out as one of the most prominent techniques in this category. Despite its widespread adoption, standard LIME suffers from poor fidelity and ambiguous locality, making it difficult for users to measure how much trust to place in a given explanation. In this work, we demonstrate these weaknesses, highlighting significant ambiguity and instability issues when applying LIME to multi-class classification problems. This paper proposes Conditional LIME, an extension framework of LIME designed to address these fundamental weaknesses. By introducing class-specific conditions during explanation generation, conditional LIME resolves ambiguity and enhances stability. Our results indicate that conditional LIME provides explanations incrementally with a more precise scope and significantly improved local fidelity relative to standard LIME. Hierarchical RBF Neural Network Ensemble with Adaptive Global-Local Fusion for Data-Driven Expensive Optimization Qian-Xia Jing and Jian-Yu Li (Nankai University) and Jun Zhang (Hanyang University) Abstract Abstract Data-driven evolutionary algorithms (DDEAs) have emerged as effective approaches for solving expensive optimization problems (EOPs), where radial basis function (RBF) neural networks are widely employed as surrogate models due to their excellent function approximation capabilities. However, existing RBF neural network-based DDEAs typically train only global surrogate models, which tend to lose local information for medium-to-high dimensional complex multimodal functions, resulting in suboptimal solution quality. This paper proposes a Hierarchical RBF Neural Network Ensemble with Adaptive Global-Local Fusion prediction framework (HCLS-DDEA). Specifically, we first employ the K-Means clustering algorithm to partition the decision space into multiple local subspaces. Subsequently, we construct a two-layer large-scale RBF neural network ensemble architecture—the first layer trains a global RBF neural network ensemble pool to capture the overall topological structure of the fitness function, while the second layer trains local RBF neural network ensemble pools on each clustered subspace to finely fit high-frequency local features. Furthermore, we adopt the Bagging strategy combined with out-of-bag (OOB) error evaluation to measure neural network model quality, and dynamically select high-quality neural network subsets through a roulette wheel mechanism. Finally, we design an adaptive weighted fusion prediction mechanism based on distance and quality factors, which dynamically adjusts the fusion weights between global and local models according to the distance from the point to be predicted to the cluster center and the neural network model accuracy, effectively addressing the issues of insufficient local information learning and local model boundary failure. Experimental results on 5 benchmark test functions across 3 different dimensions (10, 30, and 50 dimensions) demonstrate that the proposed hierarchical RBF neural network ensemble method significantly outperforms existing methods on medium-dimensional expensive optimization problems, particularly achieving comprehensive leading performance on 30-dimensional problems. Tuesday Virtual Room 7 IJCNN Paper Brain-Computer Interfaces and Neural Decoding Session Chair: Guodong Li (Central China Normal University), huan peng (Soochoow University, School Of Computer Science & Technology) Grounding Global Attention: A Prior-Guided Fusion Framework for Cross-Subject EEG Decoding Xiao Chen and Xin Zhang (Harbin Institute of Technology), Shiwei Guo (University of Chinese Academy of Sciences), Guodong Li (Central China Normal University), and Jixuan Kang (Harbin Institute of Technology) Abstract Abstract Electroencephalography (EEG) offers millisecond-scale temporal resolution for probing brain dynamics. However, accurate decoding is hindered by pronounced cross-subject variability and the inherent non-stationarity of neural oscillations. Existing hybrid methods often treat local feature extraction and global modeling as decoupled stages, failing to ground the attention mechanism in biologically meaningful local structures. To address this, we propose DMTA-PGT, a unified framework that tightly couples adaptive feature weighting with prior-guided context modeling. Specifically, the Dynamic Multi-Scale Temporal Attention (DMTA) module adaptively fuses multi-resolution temporal features to mitigate non-stationarity. Crucially, the Prior-Guided Transformer (PGT) injects convolution-derived inductive biases directly into the self-attention mechanism via a gating function, ensuring global dependencies are anchored to local temporal continuity. Extensive cross-subject evaluations on two benchmarks demonstrate that DMTA-PGT significantly outperforms state-of-the-art CNN and Transformer baselines , offering a superior trade-off between decoding accuracy and stability. FreqTrans-Meta: Subject-Adaptive Cross-Subject SSVEP Decoding via Frequency-Domain Transformer and Meta-Learning Miao Sun, Jian Zhang, and Xiaohong Lan (Chongqing Normal University, College of Computer and Information Science) Abstract Abstract Cross-subject decoding in steady-state visual evoked potential (SSVEP) brain-computer interfaces remains a significant challenge due to pronounced inter-subject variability and the scarcity of calibration data. Traditional spatial filtering methods often experience severe performance degradation in this setting, while existing deep learning models are prone to overfitting or catastrophic forgetting under few-shot calibration conditions. To address these challenges, this paper proposes FreqTrans-Meta, a unified framework designed to reconcile high-dimensional representation learning with data-efficient adaptation. The proposed framework integrates three components, each targeting a specific deficiency in current decoding pipelines. First, to address the inherent spectral nature of SSVEP signals, a Frequency-Domain Channel-wise Transformer backbone is constructed. Unlike time-domain models, this backbone explicitly encodes stimulus frequencies and harmonic structures in the range of 5-45 Hz, capturing robust subject-invariant spectral dependencies. Second, to resolve the conflict between model capacity and data scarcity, a lightweight residual Subject Adapter is introduced. This decouples individual variability from general features, preventing overfitting by freezing the backbone parameters. Third, to enable rapid convergence with minimal data, a Reptile-based meta-learning strategy is employed to optimize the initialization of the adapter. This ensures the model can effectively personalize using fewer than four calibration trials per class. Extensive experiments on the BETA (70 subjects) and Benchmark (35 subjects) datasets demonstrate that FreqTrans-Meta achieves average classification accuracies of 85.04% and 87.51%, respectively. By logically integrating spectral modeling, parameter-efficient adaptation, and meta-initialization, the proposed method significantly outperforms state-of-the-art baselines such as SSVEPformer and TDCA. DSEA-Net: Multi-Dimensional Attention Fusion Framework for EEG-Based Depression Diagnosis Tianyu Zhou, Xinyi Wang, Junhao Chen, Zhihu Zhou, Yan Ling, and Keji Mao (Zhejiang University of Technology) Abstract Abstract Major depressive disorder imposes a severe global health burden, highlighting the critical need for objective diagnostic tools. Electroencephalography (EEG) offers a promising avenue, but automated detection faces key challenges: susceptibility to noise and complex signal dynamics, bias from manual preprocessing engineering, and limited clinical applicability due to poor generalization and low interpretability. To address these limitations, we propose DSEA-Net, a lightweight multidimensional attention network for EEG-based depression diagnosis. DSEA-Net synergistically integrates Channel Attention, Neuro-Spectral Excitation Block (NeuroSE-Block), and Deep Attention mechanisms within an efficient architecture. The channel attention module dynamically weights electrode contributions to enhance spatial feature learning. The NeuroSE-Block explicitly incorporates neurophysiological priors of depression biomarkers for frequency-band-specific and channel-adaptive feature recalibration. The deep attention module refines feature representations across network layers. Extensive experiments on two real-world datasets under rigorous subject-independent protocols demonstrate DSEA-Net's superior performance: achieving state-of-the-art accuracies of 94.6\% on MODMA and 97.33\% on EDRA. Moreover, derived attention maps provide neurophysiologically plausible interpretations, aligning with established biomarkers like prefrontal alpha asymmetry. DSEA-Net presents a significant advancement towards clinically applicable, interpretable, and robust EEG-based depression detection. Hybrid-TSR: A Distributed MAPF Fusion Framework with Global Map Prior Injection huan peng (Soochow University, School of Computer Science & Technology); lingzhi li (Soochow University, School of Future Science and Engineering); nan che (Harbin University of Science and Technology, School of Computer Science); zhijun li (Harbin Institute of Technology, Faculty of Computing); and rui lin (Soochow University, School of Future Science and Engineering) Abstract Abstract Centralized solvers for Multi-Agent Path Finding (MAPF) possess global conflict resolution capabilities but incur significant computational and communication overhead in large-scale scenarios. Conversely, distributed planning under local observation offers excellent scalability yet is prone to erroneous decisions and deadlocks in high-density environments due to insufficient information. To address this, this paper proposes a partially observable MAPF (MI-POMAPF) that incorporates global map prior knowledge and designs a hybrid decision framework, Hybrid-TSR, that integrates deep imitation learning with the distributed heuristic algorithm 1-step collisions (CS-PIBT). This framework trains a Transformer decision model using global prior knowledge and local observations during the offline phase, then embeds it into CS-PIBT during online execution to enhance local decision quality. Experimental results demonstrate that this method significantly improves execution success rates in high-density scenarios while maintaining stable planning overhead, validating its effectiveness and robustness in large-scale multi-robot systems. Tuesday Virtual Room 8 IJCNN Paper Code Generation and Program Synthesis I Session Chair: Bilal FAYE (Sorbonne Paris Nord University), Qiuhong Zhang (Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Dual-Expert Guided Adaptive Framework for Secure Code Generation Boyu Wei, Yurong Wu, and Qiuhong Zhang (Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences Beijing, China) and Shuo Zhang and Zhiming Ding (Institute of Software, Chinese Academy of Sciences) Abstract Abstract Large language models are increasingly applied to code-related tasks, raising growing concerns about the safety of their generated code. Existing methods either directly update the model or adjust generation behavior without changing parameters to improve security. The latter approach is more valuable due to stronger scalability but often improves safety at the cost of functional correctness. Inspired by knowledge transfer, we propose Adaptive Secure Code Generation, a framework leveraging two expert models, respectively focused on security and correctness, to collaboratively guide the code generation of target model. We collect representative code repair and instruction data to train the two experts. To achieve a balanced integration of objectives, we further introduce an entropy-based logits fusion strategy that adaptively adjusts expert guidance strength, thereby enhancing the reliability of the generated code. Specifically, our implementation achieves improvements of 6–13% in overall performance. Smaller, Faster, Less Secure? A Jailbreak Security Analysis of Lightweight Code Judges Tianyu Xu (State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing 210023, China; Nanjing Marine Radar Institute, Nanjing 211153, China); Chao Li (The 8th Research Academy of CSSC, Nanjing 211153, China); Ning Zhao (Nanjing Marine Radar Institute, Nanjing 211153, China); and Tengfei Liu, Jinghong Zhang, and Chongjun Wang (State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing 210023, China) Abstract Abstract As lightweight code judges (e.g., 1.5B parameters) are increasingly deployed in critical scenarios such as CI/CD pipelines and reinforcement learning reward modeling, ensuring their security becomes paramount. While recent distillation techniques achieve GPT-4o-level accuracy at reduced cost, the security implications for code evaluation remain unexplored. In this paper, we conduct a systematic security analysis of distilled code judges, revealing a critical vulnerability: nonEnglish jailbreak prompts at code end achieve 49% attack success versus 6% for English equivalents. We attribute this to two synergistic mechanisms: (1) OOD Safety Gap, where Englishcentric training fails to suppress non-English jailbreaks; and (2) Recency Bias, where instructions near the generation point receive disproportionate attention. To address this, we propose EPGE (Error Pattern-Guided Evaluation), a training-free defense that redirects attention toward evaluation criteria, achieving near-zero attack success while improving accuracy by up to 5%. Our work demonstrates that security must be prioritized alongside accuracy and efficiency in lightweight model distillation for code evaluation. Learning to Verify: Efficient Verilog Code Reranking via Collaborative Knowledge Distillation Yiheng Shen (Nantong Normal College), Guang Yang (Zhejiang University), Wei Zheng (Northwestern Polytechnical University), Yifan Sun (Hainan International College of Minzu University of China), Fengji Zhang (City University of Hong Kong), Xiang Chen (Nantong University), and Fengping Xu and Hongxing Xia (Nantong Normal College) Abstract Abstract LLMs have shown strong performance in software code generation, yet struggle with Verilog due to limited hardware domain knowledge. While sampling techniques improve pass@k by generating multiple candidates, engineers need a single reliable solution rather than uncertain alternatives. Current reranking methods face a capability gap: execution-based approaches (e.g., CodeT) rely on test case quality and incur execution overhead, while code generation models lack the discriminative capability needed for reliable correctness judgment. This paper investigates whether verification reasoning can be distilled from large expert models into a lightweight discriminator. We propose VCD-Rnk, a discriminator model for Verilog code reranking built on three key insights: (1) the capability to judge functional correctness is learnable and distillable; (2) single-teacher distillation suffers from knowledge bias, which collaborative dual-teacher distillation can mitigate through complementary expertise; (3) effective supervision requires explicit reasoning traces explaining why code succeeds or fails, not just binary labels. After collaborative knowledge distillation, we construct VerilogJudge-47K, a dataset with expert reasoning traces, and fine-tune a 4B-parameter discriminator that approximates simulator-level verification without execution overhead. Experiments show that VCD-Rnk improves pass@1 by 10.4-25.8% across multiple LLMs, achieving 93.5% of the theoretical oracle performance while without execution overhead. Prototype-Guided Diffusion: Efficient Visual Conditioning without External Memory Bilal FAYE and Hanane AZZAG (Sorobonne Paris Nord University) and Mustapha LEBBAH (Versailles Saint Quentin en Yvelines University (UVSQ)) Abstract Abstract Diffusion models achieve state-of-the-art image generation but remain computationally costly due to iterative denoising. Latent-space models like Stable Diffusion reduce overhead yet lose fine detail, while retrieval-augmented methods improve efficiency but rely on large memory banks, static similarity models, and rigid infrastructures. We introduce the Prototype Diffusion Model (PDM), which embeds prototype learning into the diffusion process to provide adaptive, memory-free conditioning. Instead of retrieving references, PDM learns compact visual prototypes from clean features via contrastive learning, then aligns noisy representations with semantically relevant patterns during denoising. Experiments demonstrate that PDM sustains high generation quality while lowering computational and storage costs, offering a scalable alternative to retrieval-based conditioning. Tuesday Virtual Room 1 IJCNN Paper Computer Vision and Multimodal: Emerging Topics I Session Chair: Shuxiang Song (Guangxi Normal University), Jing Qi (Hebei University) FGF-Net: Frequency-Guided Fusion for Hybrid Voxel–Point LiDAR-Based Place Recognition Lineng Chen (Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education; Guangxi Normal University); Haiying Xia (Guangxi Normal University); and Shuxiang Song (Key Laboratory of Education Blockchain and Intelligent Technology, Ministry of Education; Guangxi Normal University) Abstract Abstract LiDAR-based place recognition is fundamental to large-scale outdoor simultaneous localization and mapping (SLAM) and long-term autonomous navigation. Voxel-based networks are efficient and robust but suffer from information loss induced from quantization, whereas point-based networks preserve geometric details but are sensitive to sensor noise. Recent hybrid voxel–point approaches typically fuse these two representations in the spatial domain, overlooking their complementary properties in the frequency domain. In this paper, we propose FGF-Net, a lightweight frequency-guided fusion network for LiDAR-based place recognition with hybrid voxel-point representations. Specifically, we propose a Frequency-Guided Fusion Module that applies local discrete cosine transform to decompose point-wise features into low- and high-frequency bands and performs an energy-based dual-stream gating mechanism. In this way, voxel-wise features provide semantic priors that enhance reliable high-frequency details in the point branch, while low-frequency cues from the point branch refine voxel-wise representations for stable structural encoding. Extensive experiments on multiple large-scale benchmarks show that FGF-Net achieves competitive performance with only 0.22M parameters and exhibits strong generalization across diverse environments and adverse weather conditions, validating the effectiveness of frequency-guided hybrid fusion for LiDAR-based place recognition. Unveiling Deception: A Frequency-Semantic Dual-Attention Network for Rumor Detection Yuxiang Chen and Yang Tang (College of Cyber Science and Technology); Shihui Gao, Hui Zhou, and Zheng Qin (College of Cyber Science and Technology, Hunan University); and Lu Ou (School of Journalism and Communication) Abstract Abstract Multimodal rumor detection aims to automatically identify real or fake news, thereby mitigating the adverse effects of misinformation. Although existing methods are effective, they pay insufficient attention to frequency features and often fail to fully integrate multimodal features. To address these challenges, we propose the FS-DARD model. First, we extract frequency features from images using the discrete cosine transform and design a spatial-frequency fusion stream based on an attention mechanism to enable comprehensive interaction between different image features. Next, we develop an image-text fusion stream that incorporates textual semantic features into the discrimination process. Finally, we introduce a gating mechanism during feature fusion to select the most representative features for detection. Experiments on three public datasets, Weibo, Twitter, and GossipCop, demonstrate the superiority of our model. Now, our code is available at https://anonymous.4open.science/r/FS-DARD-76E2. GPR Relative Localization Method Based on Saliency Detection and Intermittent Fusion Huaichao Wang, Xinyu Guo, and Xuanxin Fan (Civil Aviation University of China); Lei Sun (Nankai University); Haifeng Li (Civil Aviation University of China); Kairat Koshekov (Civil Aviation Academy); and Dezhen Song (Mohamed Bin Zayed University of Artificial Intelligence (MBZUAI)) Abstract Abstract When performing robot/vehicle localization using Ground Penetrating Radar (GPR) to handle adverse weather and environmental conditions, existing GPR-based relative localization methods are prone to generating unreliable displacement estimates in feature-sparse regions. Direct implementation of continuous fusion propagates errors induced by these erroneous estimates into the entire system, which in turn degrades the trajectory accuracy. This paper proposes a GPR intermittent fusion localization method based on saliency detection and filtering. Specifically, a three-branch HVS-Net saliency detection network is employed to assess the feature saliency of consecutive B-scan image pairs. Only when the image pair features are deemed salient are they introduced as observations into the factor graph-based fusion framework. Experiments on the public CMU-GPR dataset demonstrate that the proposed saliency detection module can effectively filter out feature-sparse image pairs. Moreover, the intermittent fusion strategy proposed significantly enhances system-level localization accuracy. Relative to continuous fusion, the overall Absolute Trajectory Error (ATE) is reduced by 18.04\% (0.081 m). The experimental results verify that the proposed method can effectively suppress the propagation of errors from low-quality GPR observations, thereby improving localization stability and accuracy in complex subsurface environments. FA‑Net: Dual‑Branch YOLOv5 with Triplet‑Aware Dynamic Filtering and Bilateral Coordinate Fusion for RGB‑D Hand Gestures Lu Yang Wang (Hebei University), Jing Qi (HeBei University), and You Zhou (Hebei University) Abstract Abstract Hand posture recognition using RGB-D cameras offers intuitive interaction but remains challenged by RGB instability (e.g., illumination changes, skin-like backgrounds) and depth noise. Thus, to address these challenges, we propose a Fusion-Attention Network (FA-Net), a multi-stage dynamic attention fusion network built on a dual-branch YOLOv5 backbone. Firstly, our proposed Triplet-Aware Dynamic Filtering (TADF) module generates data-driven convolutional kernels via channel, width and spatial interactions to precisely align and enhance RGB and depth features. Secondly, our developed a Bilateral Coordinate Fusion (BCF) module boosts cross-modal feature exchange and integrates spatial coordinates into channel attention to suppress background clutter and emphasize key gesture regions. FA-Net adaptively adjusts modality weights and performs scale-aware fusion, fully exploiting RGB-depth complementarity. On CUG, NTU and a self-built robot dataset, FA-Net outperforms state-of-the-art fusion methods, achieving notable gains in accuracy and robustness. Deployment on a wheeled robot with an Intel RealSense camera confirms high-precision gesture recognition and reliable motion command generation, demonstrating FA-Net’s practicality for real-world human-robot interaction. Tuesday Virtual Room 2 IJCNN Paper Computer Vision and Multimodal: Emerging Topics II Session Chair: HaiDi Xu (Zhejiang Sci-Tech University), Cong Liang (Changchun University of Science and Technology) Structure-Constrained GAN: Selective Structure Guided Adversarial Network for Lightweight Image Super-Resolution Na Zhang, Haidi Xu, Cheng Xu, Xiaoan Bao, and Hainan Chen (Zhejiang Sci-Tech University) and Qingqi Zhang (Hangzhou Institute of Medicine Chinese Academy of Sciences) Abstract Abstract Image Super-Resolution(SR) faces a long-standing trade-off between perception and distortion. Perceptual quality approaches based on adversarial training enhance visual realism by synthesizing high-frequency details, yet often introduce hallucinated textures and grainy artifacts in regions that are intrinsically smooth. When coupled with global modeling backbones such as state space models, the adversarial signal can be diffused through long-range interactions, making background artifacts more likely and more widespread. To address these issues, Structure-Constrained GAN (SCGAN) is proposed, which is a lightweight SR framework that explicitly injects structure awareness into both generator design and adversarial supervision. The architecture features a Selective Progressive Attention Module (SPAM) to prune redundant interactions in regions with poor texture, complemented by a branch based on Mamba for global context modeling. These streams are integrated via an Adaptive Gating Fusion Module (AGFM), which employs weights that vary spatially to amplify texture synthesis while suppressing artifact leakage in flat areas. Furthermore, a Variance Guided Adversarial Training (VGAT) strategy restricts the focus of the discriminator to regions dominant in texture, aligning perceptual objectives with structural constraints. Quantitative evaluations on five benchmarks show that SCGAN achieves SOTA performance, outperforming MambaIRv2-light by up to 0.26 dB in PSNR and reducing FLOPs by 18.8% with only 835K parameters. PC-AGL: Light-Weighted Prior Constrained Adaptive Graph Learning for fMRI Analysis to Diagnose Autism Spectrum Disorders Wenbo Ning, Shijie Guo, and Yuxiang Guo (School of Software, Taiyuan University of Technology); Yan Niu (School of Computer Science and Technology (Data Science), Taiyuan University of Technology); and Rui Cao and Xin Wen (School of Software, Taiyuan University of Technology) Abstract Abstract Resting-state functional magnetic resonance imaging (rs-fMRI)-derived functional connectivity (FC) has emerged as a potent metric for analyzing neuropsychiatric disorders, with autism spectrum disorder (ASD) being a prime focus. However, existing deep learning-based automated diagnostic methodologies suffer from several limitations, including the neglect of global representations, excessive parameterization, and insufficient interpretability. In this study, we propose a novel prior constrained adaptive graph learning (PC-AGL) method for ASD diagnosis. Firstly, a graph model integrated with a priori constraints is constructed on the basis of discriminative FC. Subsequently, a comprehensive convolutional kernel is devised to amalgamate short-range and long-range representations. The Autism Brain Imaging Data Exchange (ABIDE) dataset is employed for ASD detection. Our results indicate that the PC-AGL approach achieves an accuracy of 73% in ASD identification, while concurrently reducing the training parameters by 85% compared to the state-of-the-art model. The reduction in model parameters not only showcases its potential for deployment in edge computing applications but also the discriminative patterns unveiled may offer data-driven insights into the underlying brain-neurological mechanisms of ASD. Unlocking Potential in Suboptimal Demonstrations via Progressive Trajectory Valuation Yiyang Xu, Yilin Liu, Tao Wang, Shiwei Li, Xiangfeng Luo, and Shaorong Xie (shanghai university) Abstract Abstract Imitation learning (IL) provides a powerful framework for skill acquisition by leveraging expert demonstrations. However, its real-world application is often hindered by the prevalence of suboptimal demonstrations, which violates the common assumption of optimal expert policies. While existing methods handle imperfect data via static weighting or ranking, they typically fail to capture and utilize the progressive nature of learning—the valuable trends of improvement within trajectories that are evolving towards expert-level performance. To bridge this gap, we propose Progressive Confidence-based Adversarial Imitation Learning (PCAIL). Our framework establishes a dual-assessment paradigm, which dynamically reinterprets imperfect demonstrations throughout training via adaptive target margins that shift the focus from static quality to instructive potential. This enables the model to prioritize demonstrations showing positive learning progression, thereby enhancing both sample efficiency and final performance. Extensive evaluations on six MuJoCo continuous control tasks and a custom unmanned surface vehicle (USV) reconnaissance environment demonstrate that PCAIL achieves consistent and significant improvements over a range of established baselines. Our work offers a practical and principled solution for learning from imperfect data, establishing a paradigm that strategically exploits the progressive potential within imperfect data rather than treating it as homogeneous noise. APTA2D: APT Attribution via Attention-Guided Pruning and 2-D Convolutional Reasoning Weiwu Ren and Cong Liang (Changchun University of Science and Technology) and Ying Lei (Huajin Aramco Petrochemical Company Limited) Abstract Abstract Advanced Persistent Threat (APT) attribution is a cornerstone challenge in cybersecurity. Existing approaches such as KGConvE over-rely on probabilistic associations and fail to uncover the deep causal logic underlying multi-stage intrusions, yielding limited interpretability. We propose APTA2D, a “classify-first, reason-second” coupled framework: stage 1 employs multi-head attention to probabilistically prune the full graph, shrinking the candidate space; stage 2 performs 2-D convolutional reasoning on the condensed subgraph, reducing overall complexity from O(N²) to O(α²N²) (α ≤ 0.08) and achieving multiplicative—rather than additive—error decay. This progressive architecture significantly boosts both the accuracy and the explainability of source-IP localization, and experiments demonstrate that APTA2D substantially outperforms existing baselines in attribution precision while completely reconstructing the attack chain. Tuesday Virtual Room 3 IJCNN Paper Continual and Incremental Learning I Session Chair: Di Shang (Institute of Automation,Chinese Academy of Sciences), Hechang Chen (Jilin University) Beyond Naive Replay: SS-RL for Robust and Effective Continual Offline Reinforcement Learning Yang Yu, Jifeng Hu, Sinuo Zhang, Zhejian Yang, Shengjie Wang, and Hechang Chen (Jilin University) Abstract Abstract To address the increasing demands of mastering potential decision-making tasks based on existing well-trained models, continual offline reinforcement learning (CORL) is proposed to effectively utilize offline data for long-term learning and decision-making. Facing the challenges of mitigating catastrophic forgetting when learning sequentially from fixed offline datasets and overcoming performance limitations imposed by sub-optimal data quality within these datasets, we propose an innovative method, Trajectory Stitching and PCA-Selective Replay for RL (SS-RL), to tackle the above challenges and increase the generalization ability. Specifically, a multi-head neural network structure is adopted to promote agents’ knowledge sharing among multiple subtasks or goals, thereby improving the strategy’s stability and expressiveness. Secondly, an experience replay mechanism is introduced to enhance the long-term stability of the strategy by reusing previous tasks’ offline data. Finally, increasing the dataset’s quality by stitching sub-optimal trajectories, then hold- ing the performance or avoiding rehearsal experience overfitting with diverse trajectories selected by PCA. Experimental results show that the proposed method significantly improves learning efficiency and strategy performance in standard CORL bench- mark tasks, and has strong generalization ability. ViTReplay: Continual Learning with Vision Transformers via Conditional GAN Replay Zhicong Zhu (Worcester Polytechnic Institute), Kun Zhou (Linyi University), and Bo Tang (Worcester Polytechnic Institute) Abstract Abstract Continual learning aims to learn a stream of tasks while mitigating catastrophic forgetting. With the rapid progress of Vision Transformers (ViTs), they have achieved strong performance on a wide range of vision tasks and have attracted increasing attention in continual learning. However, most ViT-based continual learning approaches rely on exemplar buffers for rehearsal, and representation diversity often degrades when the buffer is small or unrepresentative. Replacing exemplars with generative replay can reduce storage, but low-fidelity synthesis may introduce distribution shift. To address this, we propose a ViT-based continual learning framework that integrates class-conditional generative adversarial networks (cGANs) with ViTs and employs bidirectional alignment to improve both distributional diversity and semantic alignment. We define forward alignment as using replay samples synthesized by a cGAN from previous tasks to regularize the ViT, thereby mitigating representation drift and promoting sample diversity. Conversely, we define reverse alignment as using the updated ViT to provide semantic guidance to the cGAN by enforcing classification consistency together with feature-level and prediction-level alignment, thereby improving the semantic fidelity of generated samples. Experimental results indicate that, regardless of whether original samples are stored, this approach significantly enhances model robustness and representation diversity in multi-task environments, thus enabling more efficient and effective continual learning. DePER: Decoupled Multi-Scheme Prioritized Experience Replay for Sample-Efficient Deep Reinforcement Learning Diyuan Shi (Zhejiang University, Westlake University) and Zifeng Zhuang and Donglin Wang (Westlake University) Abstract Abstract Prioritized Experience Replay (PER) is a widely adopted technique in Deep Reinforcement Learning (RL) and could significantly improve the sample efficiency and adaptation ability of RL methods. Tremendous research effort has been put into designing novel prioritization schemes, such as learning error, recency and visitation count. Unfortunately, given these promising and effective methods, few works have focused on how to `integrate' multiple prioritization schemes easily and straightforwardly. And the current approach to integrate multiple prioritization schemes is still weighted sum which suffers shortcomings in both practice and theory. As RL methods are becoming more complex, we also need PER method that scales better. Hence in this work, we propose DePER, which performs independent and decoupled PER and avoids the drawbacks in scaling weighted sum PER. We conduct theoretical analysis and extensive experiments (against heavily tuned weight sum PER) to demonstrate DePER requires less tuning effort, offers broader applicability and could obtain better performance in challenging tasks. Dual-Loop Online Meta-Learning with Subspace-Aware Memory Refresh for Online Class Incremental Learning Di Shang, Yanfeng Lu, Zhiyuan Li, Guoqi Li, and Lu Zhang (Institute of Automation,Chinese Academy of Sciences) Abstract Abstract Online class-incremental learning (OCIL) requires a model to learn a growing set of classes from a strictly one-pass, non-stationary stream, where each sample is observed only once, leading to insufficient learning of new classes and unstable representations. Despite the strong advantages of replay-based methods in OCIL, imbalanced replay and sample loss under a finite buffer further suppress plasticity, disrupt the stability-plasticity balance, and make performance highly sensitive to buffer size. Specifically, we introduce a dual-loop online meta-learning framework that meta-learns both model parameters and drift-adaptive update rules at the sample level, promoting more transferable representations and stronger plasticity under insufficient new-sample learning. We further propose Structured Replay with Subspace-Aware Refresh, which uses class-subspace relations to enable uniform replay and cyclic buffer reuse, improving buffer-size robustness and providing a more representative replay signal for meta-learning transferable representations. Extensive evaluations on standard OCIL benchmarks (T-ImageNet, CIFAR100, MNIST-P) and a challenging few-shot OCIL setting (CUB200) demonstrate state-of-the-art performance with consistently lower forgetting and stronger forward transfer. Our method remains robust across a wide range of buffer budgets while incurring runtime comparable to practical replay baselines. Tuesday Virtual Room 4 IJCNN Paper Federated and Distributed Learning I Session Chair: Yixiang Wang (South-Central Minzu University), jingjing fu (clemson university) On the Convergence of Cyclic Hierarchical Federated Learning with Heterogeneous Data jingjing fu (clemson university), Haibo Yang (Rochester Institute of Technology), Xiaonan Zhang (Florida State University), and Linke Guo (Clemson University) Abstract Abstract Hierarchical Federated Learning (HFL) advances the classic Federated Learning (FL) by introducing the multi-layer architecture between clients and the central server, in which edge servers aggregate models from respective clients and further send to the central server. Instead of directly uploading each update from clients for aggregation, the HFL not only reduces the communication and computational overhead but also greatly enhances the scalability of supporting a massive number of clients. When HFL operates for applications having a large-scale clients, edge servers train their models in a cyclic pattern (a ring architecture) as opposed to the star-type of architecture where each edge develops their own models independently.We refer it as Cyclic HFL(CHFL). Driven by its promising feature of handling data heterogeneity and resiliency, CHFL has a great potential to be deployed in practice. Unfortunately, the thorough convergence analysis on CHFL remains lacking, especially considering the widely-existing data heterogeneity issue among clients. To the best of our knowledge, we are the first to provide a theoretical convergence analysis for CHFL in strongly convex, general convex, and non-convex objectives. In particular, CHFL achieves a Õ(1/(MNRKT)) convergence rate for strongly convex objectives and an O(1/sqrt(MNRKT)) rate for general convex objectives. For smooth non-convex objectives, we prove convergence to ε-stationary point, with the dominant term scaling as O(1/sqrt(MNRKT)). Extensive experiments on real-world datasets corroborate our theory and demonstrate that CHFL performs competitively under both inter- and intra-edge heterogeneity. DFedReweighting: A Unified Framework for Objective-Oriented Reweighting in Decentralized Federated Learning Kaichuang Zhang (University of South Florida), Wei Yin and Jinghao Yang (The University of Texas Rio Grande Valley), and Ping Xu (Georgia State University) Abstract Abstract Decentralized federated learning (DFL) has emerged as a promising paradigm that enables multiple clients to collaboratively train machine learning models through iterative rounds of local training, communication, and aggregation without relying on a central server which introduces potential vulnerabilities in conventional federated learning. Nevertheless, DFL systems continue to face a range of challenges, including fairness, robustness, etc. To address these challenges, we propose \textbf{DFedReweighting}, a unified aggregation framework designed to achieve diverse objectives in DFL systems via an objective-oriented reweighting aggregation at the final step of each learning round. Specifically, the framework first computes preliminary weights for all clients based on the target performance metric (TPM) obtained from an auxiliary dataset constructed using their local data. These weights are then refined using a customized reweighting strategy (CRS), resulting in the final aggregation weights. Theoretically, we prove that the appropriate combination of TPM and CRS ensures linear convergence for general $L$-smooth, strongly convex functions. Experimental results consistently show that our proposed framework significantly improves fairness and robustness against Byzantine attacks in diverse settings. Two examples of multi-objective tasks across clients and within clients further demonstrate that our framework can achieve a broad range of desired learning objectives by appropriately designing the TPM and CRS. Our code is available at \url{https://github.com/KaichuangZhang/DFedReweighting}. Enhancing Robustness of Federated Learning via Server Learning Van Sy Mai (NIST), Richard La (UMD), Kushal Chakrabarti (Tata Consultancy Services Research), and Dipankar Maity (University of North Carolina) Abstract Abstract This paper explores the use of server learning for enhancing the robustness of federated learning against malicious attacks even when clients' training data are not independent and identically distributed. We propose a heuristic algorithm that uses server learning and client update filtering in combination with geometric median aggregation. We demonstrate via experiments that this approach can achieve significant improvement in model accuracy even when the fraction of malicious clients is high, even more than 50% in some cases, and the dataset utilized by the server is small and could be synthetic with its distribution not necessarily close to that of the clients' aggregated data. Model Heterogeneous Federated Learning with Personalized Feature Distillation Yixiang Wang (South-Central Minzu University) and Boyi Liu (City University of Hong Kong) Abstract Abstract Model-Heterogeneous Federated Learning (MHFL) enables collaborative training across clients with diverse backbones while keeping private data local, yet it renders parameter aggregation infeasible and complicates knowledge fusion across incompatible feature spaces. Public-data-assisted distillation provides a model-agnostic alternative, but logit alignment can be unreliable under feature mismatch, and naive feature matching may over-regularize heterogeneous clients. We propose PerFed, a representation-centric MHFL framework that performs personalized representation distillation without any aggregatable parameters. PerFed aligns intermediate features into a shared space and constructs client-specific distillation targets via mask-guided composition, selectively inheriting globally transferable cues while retaining complementary client-specific patterns. Empirical results across standard MHFL benchmarks show that PerFed consistently outperforms strong state-of-the-art baselines. Tuesday Virtual Room 5 IJCNN Paper Few-Shot and Meta-Learning Session Chair: Qinghan Wang (Qilu University of Technology (Shandong Academy of Sciences)), Yichao Fu (China) SAN: Style-Adaptive Zero-Shot Coordination via Dynamic Partner Inference Yichao Fu, Chao Zhang, Wensong Bai, Hanbin Zhao, and Hui Qian (Zhejiang University) Abstract Abstract Zero-shot coordination aims to develop agents that can generalize to unseen human partners without relying on human data. To achieve this, Population-Based Training (PBT) has emerged as a prevailing paradigm, which constructs a diverse partner pool via self-play to enhance the agent’s robustness. However, existing PBT frameworks often fail in real-world interactions where human partners exhibit non-stationary strategies and dynamic style switching. To address this challenge, we propose Style Adaptive zero-shot coordinatioN(SAN), the first style-aware PBT framework designed for dynamic coordination. SAN incorporates a latent style inference module to capture the temporal evolution of partner behaviors. Specifically, we employ a Transformer-based selector to dynamically activate specialized policy heads from a diverse repertoire. To ensure high-quality training, we introduce a style-constrained utility function during the population seeding phase, explicitly incentivizing behavioral diversity via distorted quantile fractions. Extensive evaluations in Overcooked environment demonstrate that SAN significantly outperforms state-of-the-art methods in both static and dynamic switching scenarios, manifesting robust zero-shot generalization. CoDe: Synergizing Contrastive Discriminative Retrieval and Confidence-Aware Fusion for Zero-Shot Image Captioning Yvzhe Lu, Zhongjie Zhu, Zhijing Yu, Di Ge, and Renwei Tu (Zhejiang Wanli University) Abstract Abstract Zero-Shot Image Captioning (ZSIC) serves as a pivotal bridge between visual perception and language generation without relying on paired training data. However, existing retrieval-augmented methods are impeded by embedding space anisotropy and static fusion mechanisms. The former introduces hard negatives that are visually similar yet semantically contradictory, whereas the latter lacks awareness of generation uncertainty, leading to severe semantic misalignment and object hallucination. To address these challenges, we propose CoDe: Synergizing Contrastive Discriminative Retrieval and Confidence-Aware Fusion for Zero-Shot Image Captioning. First, the Contrastive Discriminative Retrieval (CDR) module constructs a discriminative semantic manifold to filter topological noise. By employing contrastive calibration, CDR eliminates geometric outliers to ensure that external memory precisely anchors to the core visual content. Second, the Confidence-Aware Fusion (CAF) module incorporates a dynamic gating unit driven by an entropy-based uncertainty metric. This mechanism adaptively modulates the contribution of visual, linguistic, and memory signals, effectively suppressing hallucinations while preserving linguistic fluency. Extensive experiments on MS COCO, Flickr30k, and NoCaps datasets demonstrate that CoDe consistently outperforms existing methods across multiple evaluation metrics, particularly exhibiting superior robustness in cross-domain scenarios. D2SC: Dual-Stage Diffusion-Guided Semantic-Visual Coupling for Zero-Shot Learning Qitong Fang (Jilin Jianzhu University) Abstract Abstract Generative zero-shot learning remains challenging in practice. When conditioning on diffusion timesteps, semantics become step-dependent: the same concept drifts across noise levels, weakening conditional guidance for synthesis. Adversarial alignment alone further leaves a mismatch between synthesized and real features, both in global statistics and class-wise geometry; and in multi-branch setups, discriminator heads can produce inconsistent gradients that destabilize training. To address these issues, we design D2SC (Dual-Stage Diffusion-Guided Semantic-Visual Coupling) with a simple rationale: first stabilize the conditional prior, then align the visual space. Concretely, the Contrastive Prior Regularizer (CPR) enforces timestep consistency to suppress conditional drift, while the Visual Feature Synthesizer (VFS) performs multi-alignment via maximum mean discrepancy and center loss, complemented by lightweight discriminator distillation for coherent feedback. The two modules share semantic mappings and a unified schedule. With a single hyper-parameter setting across datasets, D2SC achieves robust zero-shot learning (ZSL) and generalized zero-shot learning (GZSL) gains. The code is available at https://github.com/btxq-ily/d2sc. SIPC-SQL: Enhancing Zero-Shot Text-to-SQL via Structure-Aware Intent Parsing and Complexity-Adaptive Routing Qinghan Wang, Chuantao Li, Jintao Li, Yang Zhang, Jianglin Ma, Wenhui Yang, and Zhigang Zhao (Qilu University of Technology (Shandong Academy of Sciences)) Abstract Abstract Zero-shot Text-to-SQL methods often struggle with structural hallucinations in complex schemas and incur unnecessary inference cost due to over-reasoning on simple queries. To address these challenges, we propose SIPC-SQL, a structured inference framework that improves both reliability and efficiency without relying on fine-tuning or in-context demonstrations. SIPC-SQL decouples intent understanding from SQL generation through structure-aware intent parsing, producing a schema-grounded intermediate representation that explicitly encodes target entities, logical operations, and join paths. To ensure structural validity, we further employ schema-strict constrained decoding to eliminate hallucinated schema elements. In addition, we introduce a complexity-adaptive routing mechanism that dynamically selects between a lightweight direct generation path and a structure-enhanced reasoning path based on query complexity. Experiments on the Spider and BIRD benchmarks demonstrate that SIPC-SQL achieves state-of-the-art performance among zero-shot methods, reaching 84.8$\%$ execution accuracy on the Spider test set, while reducing token consumption by approximately 40$\%$ with negligible performance loss. Cross-backbone evaluations further show that SIPC-SQL consistently improves performance across different model scales, enabling smaller models to outperform larger zero-shot baselines. These results highlight the effectiveness of structured intent grounding and adaptive inference control for building reliable and cost-efficient Text-to-SQL systems. Tuesday Virtual Room 6 IJCNN Paper Generative Priors and Model Inversion Session Chair: Chun Yuan (Tsinghua University), Zixuan Chen (School of Cyber Security and Computer, Hebei University) Prior-Guided Deep Inversion for Few-Shot Knowledge Distillation Jiahe Wang, Yongxian Wei, Tangyu Jang, and Chun Yuan (Tsinghua University) Abstract Abstract Data-free knowledge distillation (DFKD) enables model compression without access to original training data, but it fundamentally suffers from being an ill-posed inverse problem, as matching Batch Normalization (BN) statistics alone leaves higher-order moments unconstrained. While few-shot scenarios offer valuable priors, effectively leveraging them to regularize this inversion remains an open challenge. To address this, we propose Prior-Guided Deep Inversion (PGDI), a framework that formulates synthesis as a constrained distribution matching problem. We introduce a dual-space guidance mechanism to rigorously shrink the solution space: (i) in the feature space, we minimize the Maximum Mean Discrepancy (MMD) to enforce high-order moment matching in a Reproducing Kernel Hilbert Space (RKHS), ensuring topological consistency; (ii) in the decision space, we employ a hybrid calibration objective that functions as a geometric bias-variance trade-off, balancing class-center estimation (bias reduction) with instance-level diversity (variance preservation). Furthermore, to address the non-stationary nature of the generative process, we design a confidence-aware curriculum that dynamically modulates the distillation trust region, transitioning from high-fidelity priors to the augmented manifold. Extensive experiments on CIFAR-10, CIFAR-100, and Tiny-ImageNet demonstrate that PGDI consistently outperforms state-of-the-art baselines by effectively recovering the target manifold from sparse priors. PGI: A prior-guided inversion model for image inpainting with enhanced textures and less artifacts Zixuan Chen (School of Cyber Security and Computer, Hebei University); Liang Wang (School of Cyber Security and Computer, Hebei University; Key Laboratory on High Trusted Information System in Hebei Province); and Shaokang Zhang (School of Cyber Security and Computer, Hebei University) Abstract Abstract Deep image inpainting has made significant strides due to recent advancements in image generation and processing algorithms. However, effectively handling fine-grained textures and structural details while avoiding noticeable artifacts remains an unsolved problem. The challenge lies in the uncertainty within the image generation process, where mapping degraded inputs to images inevitably leads to degradation without effective prior guidance. In view of this, we propose a prior-guided inversion model (PGI) that systematically integrates both learning-based and optimization-based inversion strategies. Specifically, generative priors are first obtained and refined through a learning-based inversion, and then enhanced their textures with our designed GFFM and W-mix module. The enhanced priors are then further utilized to guide an optimization-based inversion for improved texture generation. Experimental results on widely used public datasets show that PGI effectively alleviates information loss and produces more natural and coherent restored images compared with existing representative methods. Tuesday Virtual Room 7 IJCNN Paper Generative Vision and Media Synthesis I Session Chair: Xingzhe Luo (Chongqing University), Paul Henderson (University of Glasgow) RetinexGS: A 3D Retinex Framework for Gaussian Splatting in Low-Light Xingzhe Luo, Yingbo Wu, Wenxin Li, and Haoran Wang (Chongqing University) Abstract Abstract Reconstructing and rendering high-fidelity, normally-lit novel views from multi-view low-light images via 3D Gaussian Splatting remains a key challenge. Low signal-to-noise ratio and color distortion of input images severely degrade geometric reconstruction accuracy. Conventional 2D enhancement preprocessing breaks multi-view photometric consistency, causing geometric artifacts. Existing advanced 3D decomposition methods, which rely solely on weak unsupervised priors, cannot stably realize visually consistent reflectance-illumination disentanglement. To address these issues, we propose RetinexGS, a novel hybrid-supervised framework. It uses 2D Retinex decomposition as strong priors to guide 3DGS in disentangling scenes into view-consistent 3D reflectance and structure-aware illumination fields, with a jointly trained enhancement network enabling automatic brightness restoration. Notably, it reconstructs normally-lit, 3D-consistent, and photorealistic views from challenging low-light inputs. Experiments on a challenging benchmark show RetinexGS outperforms or matches state-of-the-art methods in both quantitative metrics and visual quality. GCGS: Obtaining Accurate Surface via Geometric-Consistent Gaussian Splatting from Sparse Input Beiqi Chen, Chenxu Li, and Jinhe Su (Jimei University) Abstract Abstract 3D surface reconstruction from sparse views is a challenging task in computer vision. Recently, there have been many improvements for Gaussian Splatting to extract high-quality geometry from multi-view inputs. However, this ability degrades when faced with sparse viewpoints. Lacking robust multi-view geometric constraints, it struggles to reconstruct complete surfaces in textureless regions and overfitting to the limited input views leads to geometric inaccuracies. This paper introduces GCGS, a method for acquiring accurate geometry from sparse views by introducing geometric consistency feature alignment and normal consistency constraints. Specifically, our method applies a multi-view consistency constraint to the projected depth features, with the weighted based on the geometric relationships between views. Building on this, we utilize monocular estimated normals to supervise the rendered normals, facilitating the reconstruction of detailed surfaces. These optimization strategies are built on a dense initial point cloud generated by the MASt3R model. Experimental results on the DTU and BlendedMVS datasets demonstrate that our method can reconstruct complete and detailed surfaces. It outperforms the leading NeRF-based method Neusurf by 0.04 in mean Chamfer Distance. RTGS: Metric-Scale RGB--Thermal Gaussian Splatting Jian Zhang, Peiwei Lin, Runyu Chen, Hui Zheng, Yingying Wang, Huimin Guo, Xiaotong Tu, Yue Huang, and Xinghao Ding (Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, China; School of Informatics, Xiamen University, China) Abstract Abstract Metric-scale RGB-Thermal (RGB-T) 3D reconstruction is critical for applications like robotics and infrastructure inspection, requiring both geometric accuracy and multi-spectral information. However, since thermal SfM is often unreliable, effective fusion is hindered by the scale ambiguity inherent in monocular RGB Structure-from-Motion (SfM), which fails to align with the metric constraints of factory-calibrated sensor rigs. To address this, we propose a robust, differentiable framework for cross-modal scale recovery. Unlike methods relying on discrete search or explicit depth supervision, we formulate scale estimation as a continuous gradient-based optimization using a frozen RGB 3D Gaussian Splatting (3DGS) backbone as a geometric anchor. Specifically, we synthesize RGB images from candidate thermal viewpoints and bridge the modality gap using a structure-oriented objective (Normalized Gradient Fields). This allows us to exploit the structural correlation between RGB gradients and thermal signatures, backpropagating alignment errors directly to the global scale parameter despite drastic radiometric differences. Experiments on synthetic and real-world scenes demonstrate that our approach reliably recovers the metric scale from random initializations, enabling geometrically consistent RGB-T fusion and high-fidelity thermal novel view synthesis without altering the pre-trained RGB geometry. Sampling 3D Gaussian Scenes in Seconds with Latent Diffusion Models Paul Henderson, Melonie de Almeida, and Daniela Ivanova (University of Glasgow) and Titas Anciukevicius (University of Edinburgh) Abstract Abstract We present a latent diffusion model that can generate 3D scenes, yet is trained using only 2D image data. To achieve this, we first design an autoencoder that maps multi-view images to 3D Gaussian splats, and simultaneously builds a compressed latent representation of these splats. Then, we train a multi-view diffusion model over the latent space to learn an efficient generative model. This pipeline does not require object masks nor depths, and is suitable for complex scenes with arbitrary camera positions. We conduct careful experiments on two large-scale datasets of complex real-world scenes -- MVImgNet and RealEstate10K. We show that our approach enables generating 3D scenes in as little as 0.2 seconds, either from scratch, from a single input view, or from sparse views. It produces diverse and high-quality results yet runs an order of magnitude faster than non-latent diffusion models and NeRF-based generative models. Tuesday Virtual Room 8 IJCNN Paper Generative Vision and Media Synthesis II Session Chair: Xin Chen (China Mobile (Hangzhou) Information Technology Co., Ltd.), Shuangquan Lyu (Carnegie Mellong University) Unified Long Video Inpainting and Outpainting via Overlapping High-Order Co-Denoising Shuangquan Lyu (Carnegie Mellon University), Jian Mao (Jilin University), and Yue Ma (Tsinghua University) Abstract Abstract Diffusion-based text-to-video models are increasingly capable, but mask-based editing over hundreds of frames is still unreliable: naïve long-video generation suffers from memory blow-up, window seams, and temporal drift, and existing editors often require specialized modules or heavy fine-tuning. We propose Overlapping High-Order Co-Denoising, a lightweight framework that turns a single pre-trained text-to-video model into a unified inpainting–outpainting editor. We train only LoRA adapters using mixed interior/border masks and a dual-region loss that improves synthesis inside the mask while explicitly preserving known content. At inference, we denoise long latent sequences via overlapping windows and perform second-order Heun sampling per window, then fuse overlaps with Hamming-weighted blending to reduce boundary artifacts and improve temporal coherence. On InpaintBench (30 real videos, 81–300 frames), our method outperforms Wan 2.1 variants and VACE in background faithfulness (SSIM/LPIPS), temporal consistency (tLPIPS), and text alignment (CLIP), and scales to long horizons with memory bounded by the chosen window size. VideoRoot: Rooting Watermarks in Text-to-Video Generation via Spatiotemporal Codec Guidance Yingyue Yan, Weihai Li, and Zikai Xu (University of Science and Technology of China) Abstract Abstract Recent advances in Text-to-Video (T2V) diffusion models have enabled high-quality video generation, but also raised concerns regarding intellectual property protection and unauthorized model adaptation. Current watermarking solutions tend to lose effectiveness during downstream fine-tuning, failing to safeguard model distribution. Moreover, embedding watermarks into the generation process indiscriminately introduces global perturbation, thereby degrading the model's performance. To address these challenges, we propose VideoRoot, a framework that embeds robust watermarks directly into the video diffusion backbone via a trigger-conditioned mechanism. We employ a decoupled two-phase training strategy. In Phase I, we train a Spatiotemporal Watermark Codec (SWC) to embed signals into video latents while preserving visual quality by Temporal Feature Fusion Module (TFFM). In Phase II, using the frozen SWC as a guidance, we fine-tune the diffusion model to align its generated distribution with the watermarked space when specific triggers are present. Furthermore, to leverage the temporal redundancy of video data, we introduce Dual-layer Adaptive Message Preprocessing (DAMP) that improves capacity and ensures robustness against temporal tampering with a message recovery algorithm. Experiments on ZeroScope and LTX-Video demonstrate that VideoRoot sustains high accuracy even under downstream fine-tuning, effectively addressing the fragility of existing methods. Crucially, this robustness is achieved while preserving generation quality and supporting high-capacity payload. IDC-Animator: Efficient Identity-Consistent Human Image Animation mengting xie and yingchi mao (Hohai University) Abstract Abstract Diffusion models have demonstrated superior performance in the field of human image animation. However, the complexity and diversity of human motion present formidable challenges to effective modeling. Particularly in scenarios involving rapid, large-amplitude movements, existing methods often struggle to maintain structural stability, leading to facial feature collapse or limb motion incoordination. Furthermore, the process of long-form video generation incurs substantial memory overhead. This not only constrains the achievable video duration but also frequently degrades generation quality due to computational resource limitations. To address the aforementioned challenges, we propose IDC-Animator, an efficient video generation network designed to ensure high identity consistency throughout the generation process. First, we introduce the Identity-Adaptive Facial Encoder (IAFE), which significantly enhances feature adaptability through deep parsing of the reference facial region. Second, we construct the Identity-Consistency Calibrator (ICC) to dynamically regulate identity injection via multi-modal interaction and alignment. This mechanism ensures high-fidelity synchronization of identity features while maintaining structural stability throughout animation generation. Finally, we incorporate the Mamba architecture as the core for temporal modeling, successfully reducing the computational complexity of long-sequence processing to a linear level, thereby substantially improving the efficiency of long-video generation. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on the TikTok and UBC Fashion public datasets. RIFE-AOV: An Efficient Real-time Video Frame Interpolation Method for Always-on Video Xin Chen, Daqing Chen, and Jun Lei (China Mobile (Hangzhou) Information Technology Co., Ltd.); Liquan Shen (Shanghai University); and Nan Wu (China Mobile (Hangzhou) Information Technology Co., Ltd.) Abstract Abstract In Always-on Video (AOV) systems, switching from a low frame-rate (LFR) mode to a high frame-rate (HFR) mode often leads to temporal discontinuities in video sequences. Existing video frame interpolation (VFI) methods are not specifically designed for this scenario, and therefore tend to suffer from information loss and blurred motion boundaries during frame-rate transitions. To address these issues, this paper proposes RIFE-AOV, an efficient real-time frame interpolation method tailored for AOV applications. Built upon the RIFE framework, RIFE-AOV targets the limitations of the optical flow estimation network IFNet under frame-rate switching conditions by structurally enhancing its core sub-module, IFBlock, with a Global Directional Attention Mechanism (GDAM). Specifically, Global Channel Attention (GCA) is employed to emphasize critical motion features and alleviate information loss, while Directional Spatial Attention (DSA) strengthens the modeling of motion boundaries and structural directionality, thereby reducing motion boundary blurring. Experiments conducted on the Vimeo90K dataset demonstrate that RIFE-AOV achieves superior interpolation performance in dynamic scenes, in terms of PSNR and SSIM, compared with state-of-the-art real-time VFI methods, while maintaining real-time efficiency. These results validate the practical effectiveness of the proposed method for AOV systems. Wednesday Virtual Room 1 IJCNN Paper IJCNN Various Tracks I Session Chair: M. Tanveer (Indian Institute of Technology Indore, India), An Zhao (Institute of Artificial Intelligence (TeleAI), China Telecom) FineCAP: Fine-grained Cyclic Augmentation Prompting with Zoom-in Enhancement for Few-shot Learning Zhongjiang He (Beijing University of Posts and Telecommunications; China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd); Hongbo Sun, Hao Sun, Han Fang, An Zhao, and Ye Yuan (China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd); and Kongming Liang and Zhanyu Ma (Beijing University of Posts and Telecommunications, Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance) Whova Tag: Virtual only Abstract Abstract Vision-Language models (VLMs) have garnered significant attention in Few-Shot Learning (FSL) tasks for powerful generic representation and generalization abilities. Existing FSL methods typically add either learnable prompts or adapters on VLMs, which generally face two key limitations, i.e., an inadequate mechanism for encoding fine-grained features of objects, and misalignment between textual class information and visual objects. Motivated by these observations, this paper proposes a Fine-grained Cyclic Augmentation Prompting (FineCAP) method of Vision-Language Models for few-shot learning, with two core designs: the Fine-grained Visual Enhancement Mechanism (FVEM) and the Dynamic Textual Enhancement Mechanism (DTEM). Specifically, (1) The FVEM module first leverages the hybrid attention to guide the model to focus on and zoom in on visual objects, extracting discriminative features by visual prompting. (2) The DTEM module proposes a learnable textual prompting method to optimize textual class features with objects’ attribute information by cross-modal mapping of fine-grained object features. (3) These two modules fully exploit object information through a cycling augmentation manner, achieving adaptive fine-grained alignment. Extensive experiments on 11 widely-used benchmarks demonstrate that FineCAP achieves new state-of-the-art in few-shot learning. Learning Heuristic Selection Policies with Transformers for Selection Hyper-Heuristics Mironshoh Sobirov, Doniyor Erkinov, and Mustafa Misir (Duke Kunshan University) and Aldy Gunawan (Singapore Management University) Abstract Abstract The present study introduces a Transformer-based approach for learning heuristic selection policies in Selection Hyper-heuristics (SHHs) on combinatorial optimization. The iterative selection of low-level heuristics (LLHs) is modeled as a sequence prediction task. Two attention-based architectures are trained to imitate high-quality heuristic sequences produced by expert SHHs on the 1-dimensional bin packing problem (1D-BPP) within the HyFlex SHH framework. The models include a standard Transformer and the inverted Transformer (iTransformer). iTransformer uses self-attention across the feature (variable) dimension rather than time steps, with temporal information embedded in each token. Both models learn exclusively from successful decision trajectories without incorporating problem-specific features, thereby preserving the domain-agnostic principle central to SHH design. Evaluation is conducted across all 12 1D-BPP instances from HyFlex under a fixed computational budget. The iTransformer consistently outperforms Long Short-Term Memory (LSTM) and Temporal Convolutional Networks (TCN), as well as the standard Transformer variant across all tested instances. These results demonstrate that architectural choices in attention mechanisms significantly influence policy learning for heuristic selection. The findings establish the iTransformer as a particularly effective model for capturing sequential decision patterns in discrete optimization contexts. GBFRVFL: Granular-Ball Computing-Based Fuzzy Random Vector Functional Link Network Abdul Quadir and Abdur Rahaman (IIT Indore); P. N. Suganthan (Qatar University, Qatar); and M. Tanveer (IIT Indore) Abstract Abstract In practical machine learning tasks, data are of- ten contaminated with noise, outliers, and class imbalance, which can degrade the performance of conventional models. While random vector functional link (RVFL) networks offer fast training and strong generalization, they do not explicitly handle uncertainty or exploit local data structure. To address these limitations, we propose a fuzzy granular-ball random vector functional link (GBFRVFL) framework that leverages granular-ball computing to abstract raw samples into adaptive granular balls. Within this framework, we introduce two mem- bership assignment schemes: (i) F-GBRVFL, which incorporates fuzzy membership to quantify the reliability of each granular ball, and (ii) SDAP-GBRVFL, which we propose, incorporates a novel statistical density-adaptive pythagorean membership (SDAPM) scheme that dynamically adjusts membership and non- membership values based on class variance, local sparsity, and granular-ball compactness. These schemes enhance robustness to noise, outliers, class imbalance, and uncertainty in granular- ball distributions, while retaining the computational efficiency of RVFL networks. Extensive experiments on 37 benchmark UCI and KEEL datasets under both clean and noisy conditions demonstrate that the proposed models consistently outperform baseline models, achieving superior accuracy and stability. The results validate the effectiveness of integrating granular-ball com- puting with adaptive membe Robust Dual-Model Collaborative Random Vector Functional Link Network Abdul Quadir, Abdur Rahaman, Mushir Akhtar, and M. Tanveer (IIT Indore) Abstract Abstract Random vector functional link (RVFL) networks are lightweight and fast neural models that offer efficient train- ing and strong generalization through randomized hidden-layer weights and direct input-output connections. However, conven- tional RVFL models are sensitive to noisy labels, outliers, and imbalanced data, which limits their performance in real-world applications. To address these challenges, we propose the kernel risk-sensitive mean p-power based RVFL (KRPRVFL) model, which integrates the computational efficiency of RVFL with the robustness of the kernel risk-sensitive mean p-power (KRP) criterion. By replacing the standard least-squares objective with a KRP-based loss, KRPRVFL adaptively reduces the influence of corrupted or unreliable samples during training, resulting in im- proved stability and generalization. Additionally, a collaborative learning mechanism is introduced to enable adaptive interaction among model components, further enhancing robustness in complex and noisy environments. The proposed framework also leverages kernel-induced feature mapping to capture nonlinear relationships without requiring explicit hidden-layer selection, maintaining both efficiency and scalability. Extensive experi- ments on UCI and KEEL benchmark datasets demonstrate that KRPRVFL consistently outperforms baseline models in terms of accuracy, robustness, and statistical significance, highlighting its effectiveness as a fast, scalable, and reliable solution for challenging classification tasks. Wednesday Virtual Room 2 CEC Paper CEC Various Tracks I Session Chair: Mohamed Abouhodaima (University of New South Wales) Low-Rank Landscape Learning via Tensor Decomposition for Metaheuristic initialization Mohamed Abouhodaima (University of New South Wales) and Saber Elsayed and Ruhul Sarker (University of new south wales) Abstract Abstract Population-based metaheuristics rely heavily on the quality of their solution initialization; however, traditional strategies scatter solutions uniformly across the search space without exploiting the problem structure, which may limit their search ability. Therefore, in this paper, we propose Tensor-Based Initialization (TBI). This technique approximates high-dimensional fitness landscapes through a low-rank approximation tensor, which divides the input tensor into three components: C (sampled columns/slices), U (a small interlinking core), and R (sampled rows) (CUR). Unlike traditional decomposition methods, such as the Singular Value Decomposition (SVD), CUR preserves interpretability by using actual data elements as basis vectors rather than abstract/eigenvector factors. Therefore, TBI breaks down problems into smaller, easier-to-handle groups of dimensions. It then samples the search space at different levels of detail and combines these samples to build a rough picture of where good solutions are likely to be, without evaluating every possible point. By embedding this approach in optimization algorithms, the proposed framework works in three steps: first, a quick exploration finds diverse candidate regions using the optimizer, second, TBI refines those regions, and third, the optimizer starts from those informed regions instead of random points. Experiments on 29 standard test problems show that TBI helps find better solutions, winning on 16 problems while losing on only 1 when comparing the best results found. TBI adds only a small cost (about 3.4% of the total computation) and works best on problems that have underlying patterns it can exploit. Bayesian Optimization Framework for Multi-Objective Charging Strategy Design in Lithium-Ion Batteries Rashi Verma and Kishalay Mitra (IIT Hyderabad) and Suryanarayana Kolluri (University of Washington) Abstract Abstract Lithium-ion batteries (LIBs) underpin modern electrification, enabling applications ranging from portable electronics and electric vehicles to grid-scale energy storage systems. However, their large-scale deployment is constrained by critical safety concerns, including thermal runaway, alongside performance-degrading phenomena such as lithium plating, capacity fade, and irreversible aging. To overcome these limitations, there is a pressing need to move beyond empirical, trial-and-error based charging protocols toward rigorous optimization frameworks that ensure fast, safe charging while extending battery life. In this work, a high-fidelity, physics-based single-particle electrochemical-thermal model with electrolyte and thermal coupling (SPMeT) is used. This mechanistic modelling framework is integrated within a multi-objective Bayesian optimization algorithm. This scheme is designed to systematically explore and optimize charging protocols. Findings show that Multi-Objective Bayesian Optimization (MOBO) requires only 2.06% of the function calls required by NSGA-II, a well-established and popularly used multi-objective evolutionary algorithm, to achieve similar quality of solutions highlighting its lower computational cost. The proposed approach identifies Pareto-optimal charging strategies that balance charging time, thermal stability, and stored energy, thus providing practically relevant solutions for real-world lithium-ion battery operation at lower computational cost. Ensemble Strategies for Objective Subset Selection: A Decision-Maker-Assisted Framework Using Clustering and Correlation-Based Measures Tomás Marques (Institute for Polymers and Composites, University of Minho); Sunith Bandaru (School of Engineering Science, University of Skövde); and António Gaspar-Cunha (IPC - Institute for Polymers and Composites, University of Minho) Abstract Abstract Many-objective optimization in injection molding often involves dozens of quality indicators that are strongly correlated, making optimization inefficient and difficult to interpret. This paper investigates an ensemble objective-selection strategy for conformal cooling channel design, combining three complementary families: dependence-based clustering (representative selection), relevance–redundancy selection (mRMR), and graph-based ranking via Laplacian Score. Candidate subsets are consolidated with a correlation-aware consensus rule and evaluated under three ensemble scenarios (within-family, cross-family, and decision-maker-weighted). A single global multi-output neural surrogate is trained from 400 simulation samples spanning both gate locations and used to run NSGA-III with 11 runs per objective set. Performance is assessed in a common three-objective space by comparing direct three-objective optimization with the projection of four-objective solutions onto the same space using Hypervolume and IGD+. Results show that objective reduction can preserve the main trade-off structure, and that certain ensemble configurations yield more stable behavior when moving from four to three objectives. Distributed Flexible Jobshop Scheduling for Enterprise-Wide Model Lifecycle in Commercial Banks Keyao Wang (Bank of China/BOC Postdoctoral Research Center) and Dehui Wang (Liaoning University/School of Mathematics and Statistics) Abstract Abstract The widely use of models in commercial banks presents significant resource scheduling challenges across model lifecycle, from development and validation to monitoring and exit. These models demand intensive computational and human resources, making consolidated resource allocation critical for effectively control the model risk associated with all models within an institution. These resources are typically distributed across multiple departments and locations for the purpose of disaster recovery redundancies. The expert skills are heterogeneous to maintain different types of models, and the lifecycle workflow varies for models at different risk levels. This paper addresses the optimization of the enterprise-wide model lifecycle by formulating it as a Bilayered Distributed Flexible Jobshop Scheduling Problem (BiDFJSP), capturing the distributed and heterogeneous machine environment, two-stage processing workflows and model maintainance time window constraints, aiming at minimizing the makespan of the first-stage implementation under mandatory second-stage maintenance time windows. Prevailing dispatching rules, constructive heuristics and metaheuristic algorithms are adopted to solve the proposed BiDFJSP. Experimental results demonstrate that the metaheuristics achieve makespan reduction over dispatching rules, offering a scalable solution for model risk management through optimized resource allocation. Wednesday Virtual Room 3 FUZZ-IEEE Paper FUZZ-IEEE Various Tracks Session Chair: Qian Ma (Nanjing University of Science and Technology), Frank Chung Hoon Rhee (Hanyang University) On The Robustness of Dimensionality Reduction Methods in Selecting Type-1 and Type-2 Fuzzy Membership Functions for High-dimensional Data Annway Samal and Arnav Gupta (Indian Institute of Technology Guwahati) and Frank Chung-Hoon Rhee (Hanyang University) Abstract Abstract The Wilcoxon Minimal Bin Sizes algorithm has been recently proposed as a deterministic method for selecting the most suitable fuzzy membership function for representing a multi-dimensional dataset, based on null hypothesis testing. Although the method is theoretically robust, it may require large amounts of computations in practice for high-dimensional data, consisting of a few hundred dimensions. In this paper, we extend the algorithm for such datasets by evaluating several dimensionality reduction techniques to select the most suitable which may be integrated with the method without significantly affecting its accuracy. The motivation behind this approach is that high-dimensional datasets are gaining popularity in various fields such as medicine, image processing, geolocation, biochemistry, and computational linguistics, and for most real datasets, careful selection of a small number of features may localize enough variance such that the data can be effectively represented. We validate the robustness of our method on the 8-dimensional Yeast dataset and demonstrate its applicability on the 500-dimensional Madelon dataset. Adaptive Fuzzy Event-Triggered Quantized Control for Uncertain Nonlinear Systems Qingtan Meng and Qian Ma (Nanjing University of Science and Technology) Abstract Abstract This paper addresses the adaptive control problem for a family of uncertain nonlinear systems. By employing fuzzy logic systems to approximate the unknown nonlinear terms of the systems, we develop an adaptive control scheme that incorporates a novel dynamic event-triggered quantized mechanism. The approach can effectively conserve communication resources without relying on input-to-state stability properties or assumptions. Theoretical analysis proves that all signals in the closed-loop systems are globally bounded and the Zeno phenomenon is excluded. Simulation results are provided to validate the effectiveness of the proposed control scheme. Wednesday Virtual Room 4 CEC Paper CEC Various Tracks II Session Chair: Feby John (Amrita Vishwa Vidyapeetham, Coimbatore) An Evolutionary Algorithm Guided Search of Chained Adversarial Attacks on Generative Adversarial Networks Feby Elizabeth John, Sharath S R, and Ritwik Murali (Amrita Vishwa Vidyapeetham) Abstract Abstract Though Generative Adversarial Networks (GANs) are crucial in medical imaging to address data scarcity, they are vulnerable to adversarial attacks. This work exploits the exploratory capabilities of Evolutionary Algorithms (EAs) to search for optimal white-box adversarial attack chains, applied to the latent space of a medical-image augmentation GAN during its training. As part of the same, three standard white-box attacks are adapted to affect Deep Convolutional GANs (DCGANs) that generate augmented histopathological images. A sensitivity analysis is conducted across six EA configurations, to study the behavior of EA in this experimental setup. Results show that EA behavior depends strongly on hyperparameter choices and the right hyperparameters help establish a balance between exploration and exploitation with EAs. This work highlights the risk of latent-space attack chaining for medical-image GANs and the potential for using EAs to explore the search space and identify optimal chained attack sequences. Quantum-Inspired Evolutionary Multitasking with Two-Level Transfer Learning Techniques Eduardo Nunes Santiago Ramos, Alimed Celecia, and Marley Vellasco (PUC-Rio) Abstract Abstract This paper aims to explore alternative transfer learning techniques within the quantum-inspired evolutionary multitasking (Q-EMA) framework. We implement several novel transfer strategies, highlighting a two-level approach that performs a transfer between two tasks and later applies a mutation within the same task, adding randomness to the evolutionary process. The experimental results show that both the twolevel approach and the mutation approach consistently achieve superior results compared to the baseline framework, with special attention to the gains obtained by the two-level approach on low similarity tasks. Moreover, both strategies converge more rapidly while maintaining competitive computational costs. Wednesday Virtual Room 1 IJCNN Paper IJCNN Various Tracks II Session Chair: M. Tanveer (Indian Institute of Technology Indore, India), Abhishek Varshney (Indian Institute of Technology Indore) Robust Broad Learning System with Wave Loss for Classification under Data Uncertainty Mushir Akhtar, Abhishek Varshney, Abdul Quadir, Abdur Rahaman, Mohd. Arshad, and M. Tanveer (IIT Indore) Abstract Abstract Broad Learning System (BLS) offers an efficient alternative to deep architectures by enabling fast learning through randomized feature mapping and closed-form solutions. However, its reliance on squared error loss makes it highly sensitive to noise, outliers, and corrupted labels, limiting its reliability in real-world scenarios. To address this limitation, we propose Wave-BLS, a robust broad learning framework that integrates the wave loss function, which is asymmetric, bounded, and smooth, enabling controlled penalization of large errors. The proposed formulation replaces the standard least-squares objective with a wave-loss-based optimization problem, solved efficiently using a Nesterov accelerated gradient (NAG)-based scheme without requiring matrix inversion, thereby improving scalability. Extensive experiments on 30 UCI benchmark datasets demonstrate that Wave-BLS consistently outperforms classical BLS and several robust variants. Statistical validation using Friedman and Nemenyi post-hoc tests confirms the significance of the observed improvements. Furthermore, robustness evaluations under controlled noise and outlier injection reveal that Wave-BLS exhibits substantially slower performance degradation compared to BLS, even in challenging contamination settings. These results establish Wave-BLS as a stable and robust alternative to existing broad learning models for learning under data uncertainty. Code and Supplementary are available at \url{https://github.com/mtanveer1/Wave-BLS}. ECA-BLS: An Efficient Complex-Augmented Broad Learning System Abdur Rahaman, Abdul Quadir, M. Sajid, Mushir Akhtar, and M. Tanveer (IIT Indore) Abstract Abstract Broad Learning System (BLS) is an efficient alter- native to deep architectures due to its fast training, analytical learning, and strong generalization under limited data. However, existing BLS variants are confined to real-valued representations, restricting their ability to capture nonlinear interactions and second-order statistical dependencies inherent in real-world data. Notably, no prior BLS model fully exploits the complete second- order statistics that naturally emerge when data are embedded in the complex domain. To address this limitation, this paper intro- duces the first Complex Augmented Broad Learning System (CA- BLS), which transforms real-valued inputs into phase-encoded complex representations and adopts widely linear modeling to jointly leverage covariance and pseudo-covariance information via complex conjugate augmentation. This enables effective mod- eling of latent nonlinearities, coherence structures, and second- order dependencies inaccessible to conventional BLS formula- tions. To mitigate the additional computational cost of complex augmentation, an Efficient Complex Augmented BLS (ECA-BLS) is further developed, reformulating CA-BLS entirely in the real domain while preserving its exact decision function, achieving up to 75% fewer multiplications and over 60% fewer additions. A rigorous theoretical analysis proves the mathematical equivalence between CA-BLS and ECA-BLS, ensuring zero theoretical loss. Extensive experiments on 26 benchmark datasets from the UCI and KEEL repositories demonstrate that ECA-BLS consistently outperforms classical BLS and recent state-of-the-art random- ized neural networks in accuracy, average rank, and statistical significance, establishing augmented second-order modeling as a critical and previously missing dimension of BLS research. The supplementary material is available in an anonymous repository at https://github.com/AnonymousAuthor0011/ECA-BLS. RoBell-RVFL: A Robust Generalized Bell Random Vector Functional Link Network Abdur Rahaman, Abdul Quadir, and M. Tanveer (IIT Indore) Abstract Abstract The dominance of majority classes in real-world datasets poses a fundamental challenge to randomized neural networks, often biasing decision boundaries and overlooking critical minority samples. Existing remedies, such as synthetic minority over-sampling (SMOTE) and class-weighted loss func- tions, primarily address class proportions while neglecting intra- class distribution, making them vulnerable to label noise and outliers. In this paper, we propose RoBell-RVFL, a robust and lightweight quality-aware generalized bell random vector functional link network that redefines how randomized models handle class imbalance and noisy data. RoBell-RVFL employs a dual-strategy, sample-level weighting mechanism that strictly preserves minority class information using unit weights, while adaptively regulating the influence of majority class samples through a probability-weighted generalized bell (gbell) mem- bership function in a kernel-induced feature space. This design effectively suppresses noisy, boundary, and outlier samples within the majority class, enabling the network to learn from infor- mative samples rather than merely abundant ones. By explic- itly incorporating local class probability and class distribution information into the learning process, RoBell-RVFL achieves adaptive control over sample contributions without sacrificing the closed-form learning efficiency of RVFL networks. Extensive evaluations on UCI and KEEL benchmark datasets, along with robustness tests under up to 40% label noise, demonstrate that RoBell-RVFL consistently and significantly outperforms recent state-of-the-art RVFL variants. The results indicate that adaptive, quality-aware sample weighting is essential for robust RVFL learning, rendering conventional global weighting schemes ineffective in noisy and imbalanced environments. The supple- mentary material is available in an anonymous repository at https://github.com/AnonymousAuthor0011/RoBell-RVFL. Metric-Enhanced Hybrid Kernel Probabilistic Neural Networks for Robust Classification Abhishek Varshney, Mushir Akhtar, Mohd Arshad, and M. Tanveer (Indian Institute of Technology Indore) Abstract Abstract Probabilistic Neural Networks (PNNs) offer an interpretable and uncertainty-aware framework for classification by explicitly modeling class-conditional probability densities. However, classical PNNs and their existing variants rely on isotropic Euclidean distance and single-kernel Gaussian density estimation, which limits their ability to capture feature relevance, inter-feature dependencies, and heterogeneous data geometries commonly observed in real-world and imbalanced datasets. To address these limitations, this paper proposes a Metric-Enhanced Hybrid Kernel Probabilistic Neural Network (MEHK-PNN), which augments the classical PNN architecture through two complementary mechanisms. First, a metric-enhanced feature transformation reshapes the input space by encoding feature relevance and correlations, yielding a more discriminative similarity geometry for density estimation. Second, a hybrid-kernel Parzen estimator is introduced, combining a Gaussian kernel with a complementary similarity function to jointly capture local neighborhood structure and broader relational patterns. Two instantiations of the proposed framework are developed, incorporating cosine and sigmoid kernels as complementary components. Extensive experiments conducted on 35 benchmark KEEL datasets demonstrate that the proposed MEHK-PNN variants consistently outperform classical PNNs, recent probabilistic extensions, and strong discriminative baselines in terms of F1- score, stability, and statistical significance. Overall, the proposed framework preserves the probabilistic transparency and non-iterative training advantages of PNNs while significantly enhancing robustness and classification effectiveness under imbalancedand heteroge neous data conditions. Wednesday Virtual Room 2 IJCNN Paper IJCNN Various Tracks III Session Chair: Stavros Ntalampiras (University of Milan), Ali Muhtaroglu (Oslo Metropolitan University) Energy-Delay Product as an Early-Decision Metric for Real-Time SNNs Eeman Fatima and Ali Muhtaroglu (Oslo Metropolitan University) Abstract Abstract The escalating energy demands of deep learning have prioritized the development of energy-efficient neuromorphic hardware for the edge. However, early-stage architectural evaluation is often hindered by fragmented metrics that fail to balance metabolic costs with real-time latency constraints. This paper proposes a comprehensive evaluation framework for Spiking Neural Networks (SNNs) based on the Energy–Delay Product (EDP). Utilizing a custom event-driven C++ emulator, we model energy via spike-activity proxy and delay via total computation steps, enabling “what-if” analysis of critical design features. Validated through a context-dependent reinforcement learning (RL) task, results demonstrate that migrating from 32-bit to 5-bit precision reduces EDP by over an order of magnitude, while structural synaptic pruning achieves significant efficiency gain without degrading classification accuracy for scenarios when unknown real-time input variations may necessitate an initially redundant network structure. Quantitative trends reveal a minimum L=100 input sequence length in a given problem for amortizing learning overhead. These findings align with established neuromorphic literature, confirming the proposed emulator as a viable tool for informed architectural decision-making prior to costly physical implementation. Low-Cost Real-Time Energy-Efficient Edge Image Classifier Architecture Based on Reservoir Computing With Cellular Automata Jawdat Andraous, Tom Glover, and Ali Muhtaroglu (Oslo Metropolitan University) Abstract Abstract Abstract—This paper introduces a real-time, edge-oriented image classification architecture utilizing Reservoir Computing with Elementary Cellular Automata (ReCA) designed for resource-constrained environments. By integrating single-threshold binary inputs, column-wise evolution, and spatial downsampling with a lightweight, 6-bit spiking-inspired readout, the proposed pipeline drastically minimizes memory footprint and logic complexity. Architectural optimizations across the hardware modules yield a 2–4× reduction in resource overhead relative to existing baselines. Furthermore, transitioning from offline software-based classification to an online, perceptron-based winner-takes-all (WTA) neural network with 6-bit integer weights enhances energy efficiency significantly, with potential of orders of magnitude improvement. Validation through an in-house hardware emulator on the MNIST dataset demonstrates that substantial efficiency gains are achieved due to complexity reduction in learning engine at a small accuracy penalty of approximately 2.5%. The results confirm that integer-based ReCA systems with online learning provide a scalable, battery-efficient foundation for real-time adaptive inference at the edge. Training Spiking Neural Networks Using Lessons from State Space Models Andrew Smith and Taylor Kergan (University of California Santa Cruz) Abstract Abstract Anyone who has trained a spiking neural network (SNN) knows they can be difficult to work with: discrete spike events, surrogate gradients, subtle mismatch with deployment hardware, and the list goes on. A more fundamental issue lurks underneath: most SNNs are still trained by unrolling their dynamics sequentially in time. As sequence length grows, wall-clock time grows with it. Despite the energy efficiency of inference, this gain is partially offset by prohibitively slow training runs. How Pleasant is Your Voice? On Automatic Prediction of Voice Attractiveness Anja Bulajić and Stavros Ntalampiras (University of Milan) Abstract Abstract This work describes an approach to automatically compute the attractiveness of a voice based only on speech audio input, while linking such perceptual criteria to the identified spectro-temporal regions. Voice attractiveness plays a crucial role in the effectiveness of human–computer interaction systems, including speech synthesis technologies and virtual assistants. To this end, both traditional machine learning and deep learning methods were employed. More specifically, a Support Vector Regression model and a Convolutional Neural Network combined with bidirectional Long Short-Term Memory (CNN-BiLSTM) layers, were created and compared. Importantly, we adopted a standardized speaker-independent experimental protocol along with a publicly-available corpus facilitating the reproduction of this study. Interestingly, the CNN-BiLSTM model provided lower prediction errors and a positive coefficient of determination $R^{2}$, outperforming the SVR model. The comparison between actual and predicted Mean Opinion Score values indicated that the largest differences were found at the extremities of the perceptual scale. Moreover, we produced Grad-CAM visualizations characterizing the model operation when processing speech coming from different genders and attractiveness levels. There, we observe that the network is focused on high-frequency regions that are related to brightness and harmonic stability. At the same time, non-attractive voices show more scattered activation patterns. These results contribute to the understanding of the relationship between subjective perceptual impressions and objective acoustic modelling mediated by explainable deep learning. Wednesday Virtual Room 3 IJCNN Paper IJCNN Various Tracks IV Session Chair: Manish Pratap Singh (DYSL-CT, DRDO), Shirin Shujaa (RMIT University) XAI-Guided Feature Flow Refinement for Trustworthy HRRP Classification Avinash Rangarajan (DYSL-CT, DRDO); Aryan Pareek (Techno India University); Manish Pratap Singh (DYSL-CT, DRDO); and Debasis Chaudhuri (Techno India University) Abstract Abstract High-Resolution Radar Range Profile (HRRP) classification is a challenging problem in radar target recognition, particularly in mission-critical sensing scenarios, due to strong sensitivity to aspect angles, signal noise, and subtle geometric variations. Although convolutional neural networks (CNNs) have demonstrated competitive performance, their opaque decision processes raise reliability concerns in safety-critical applications where misclassifications can have severe consequences. Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models Shivanshu Agnihotri (Malaviya National Institute of Technology Jaipur); Snehashis Majhi (Côte d’Azur University, France); and Deepak Ranjan Nayak (Malaviya National Institute of Technology Jaipur) Abstract Abstract Automated polyp segmentation is critical for early colorectal cancer detection and its prevention, yet remains challenging due to weak boundaries, large appearance variations, and limited annotated data. Lightweight segmentation models such as U-Net, U-Net++, and PraNet offer practical efficiency for clinical deployment but struggle to capture the rich semantic and structural cues required for accurate delineation of complex polyp regions. In contrast, large Vision Foundation Models (VFMs), including SAM, OneFormer, Mask2Former, and DINOv2, exhibit strong generalization but transfer poorly to polyp segmentation due to domain mismatch, insufficient boundary sensitivity, and high computational cost. To bridge this gap, we propose LiteBounD, a Lightweight Boundary-guided Distillation framework that transfers complementary semantic and structural priors from multiple VFMs into compact segmentation backbones. LiteBounD introduces (i) a dual-path distillation mechanism that disentangles semantic and boundary-aware representations, (ii) a frequency-aware alignment strategy that supervises low-frequency global semantics and high-frequency boundary details separately, and (iii) a boundary-aware decoder that fuses multi-scale encoder features with distilled semantically rich boundary information for precise segmentation. Extensive experiments on both seen (Kvasir-SEG, CVC-ClinicDB) and unseen (ColonDB, CVC-300, ETIS) datasets demonstrate that LiteBounD consistently outperforms its lightweight baselines by a significant margin and achieves performance competitive with state-of-the-art methods, while maintaining the efficiency required for real-time clinical use. VEP-AGUNet3D: Vascular-Enhancement Prior Improves Post-Treatment Glioma Segmentation Fizza Hassan (King Fahd University of Petroleum & Minerals) and Mufti Mahmud (King Fahd University of Petroleum and Minerals) Abstract Abstract Accurate segmentation of residual or recurrent tumour in post-treatment glioma MRI remains challenging due to resection cavities, treatment-related changes, and heterogeneous enhancement patterns. While deep learning methods achieve strong performance on pre-operative benchmarks, their performance often degrades in post-treatment settings, particularly for small and irregular residual disease. In this work, we introduce a vascular enhancement prior (VEP), a lightweight auxiliary channel derived automatically from routine multiparametric MRI, which emphasizes enhancement- and vessel-like patterns while suppressing edema- and fluid-dominated regions. VEP is computed from T1, T1-Gd, T2, and FLAIR images and integrated into an attention-gated 3D U-Net without modifying the backbone architecture, loss function, or inference procedure. We evaluate three input configurations: a 3-channel structural baseline (T1, T1-Gd, FLAIR), a 4-channel structural baseline (T1, T1-Gd, T2, FLAIR), and a 3-channel+VEP configuration in which VEP is used in place of raw T2. Experiments on the BraTS 2024 post-treatment glioma dataset show that the VEP-based model achieves the best overall external-cohort performance among the evaluated configurations, improving both overlap and boundary accuracy relative to the structural baselines. In the corrected external evaluation, the raw 4-channel baseline performs comparably to the 3-channel baseline, while the VEP-based model achieves the strongest overall results. These findings suggest that structured integration of modality-derived information can be beneficial for robust post-treatment glioma segmentation. PrefixGuard: Lightweight LLMs for Hate Speech and Sexism Detection Ginel Dorleon (SogetiLabs, Capgemini, Toulouse, France) and Shirin Shujaa (RMIT University) Abstract Abstract Hate speech and sexism detection on online platforms remains challenging due to their often subtle and context-dependent nature. Large Language Models (LLMs) offer powerful representations for this task, yet full fine-tuning is computationally expensive and may amplify biases. To address these limitations, we propose a parameter-efficient approach, PrefixGuard, based on prefix tuning. Rather than updating all model weights, we learn small trainable prefix vectors at each Transformer layer alongside a lightweight classification head, while keeping the LLM backbone frozen. We formalize our approach for classification and analyze how it influences internal representations of biased or offensive language. Experiments on three benchmark datasets, EDOS (sexism), OLID (offense), and HatEval (hate targeting women or immigrants), demonstrate that our prefix-tuned LLMs method consistently outperform BERT and RoBERTa baselines and achieve results competitive with full fine-tuning while training less than 1\% of parameters. Further analysis shows that our approach yields balanced group-level performance on HatEval, robustness to adversarial text obfuscation, and improved calibration of predicted probabilities. These findings highlight PrefixGuard as a practical and effective alternative to full fine-tuning for hate speech and sexism detection in real-world moderation settings. Wednesday Virtual Room 4 IJCNN Paper IJCNN Various Tracks V Session Chair: Chang Wang (National University of Defense Technology), Karim Ali (KING Fahd University of Petroleum & Minerals) Shapley-Enhanced Mean Field Multi-Agent Reinforcement Learning for Multi-UAV Pursuit-Evasion Maneuvering Decision-Making Yunxiao Guo (National University of Defense Technology, Sun-Yat-Sen University); Dan Xu (Sun-Yat-Sen University); Long Han and Shaofei Chen (National University of Defense Technology); Chaoyang Chen (Hunan University of Science and Technology); and Chang Wang (National University of Defense Technology) Abstract Abstract With the development of artificial intelligence, Multi-Agent Deep Reinforcement Learning (MADRL)- based UAV autonomous maneuver decision methods have made significant progress, enabling swarm systems to respond effectively to complex pursuit-evasion tasks. However, the current MADRL-based maneuver decision methods still face three main challenges: the curse of dimensionality, credit assignment, and a non-stationary environment. To address these issues, this paper proposes the Shapley Mean Field Multi-Agent Reinforcement Learning (SMFMARL) algorithm for UAV swarm pursuit-evasion maneuver decisions. Specifically, we use Approximate Marginal Contribution (AMC) to estimate the Shapley-Q value for evaluating the real contribution of each UAV in the swarm. We also introduce a mean-field representation into the SQV to mitigate the non-stationarity caused by the AMC estimator and reduce the dimension of the input action. To improve the sample efficiency, this paper designs a training framework that introduces the self-play and experience replay mechanism, which can enhance the UAV's reuse of high-quality samples. Furthermore, we present a comprehensive pursuit-evasion reward that comprises catch, defensive, and advantage functions, guiding the UAV in deciding whether to pursuit or evade in different situations. In numerical simulations, we compare the proposed method with state-of-the-art MADRL algorithms in a 15-UAV swarm air-pursuit evasion scenario. The results show that SMFMARL can emerge with complex strategies. Through-Wall Respiration State Detection Using Wi-Fi CSI and a Lightweight CNN–LSTM for Edge Deployment Juan Ariel Mora Carrion, Diego Andres Andrade-Segarra, and Deyvi Rolando Totoy Ponce (Escuela Superior Politécnica de Chimborazo) Abstract Abstract The growing demand for non-invasive remote healthcare has highlighted Wi-Fi Channel State Information (CSI) as a strong alternative to wearable and vision-based monitoring. This work proposes a behind-obstacle vital-signs detection system using a hybrid CNN–LSTM model. Unlike RSSI, CSI provides fine-grained subcarrier measurements that capture subtle respiratory dynamics in Non-Line-of-Sight (NLoS) settings; the same granularity has also enabled CSI-based indoor localization in multipath environments, indicating that CSI preserves informative spatial cues beyond conventional RSSI Experimental validation was conducted using low-cost commercial off-the-shelf (COTS) hardware, specifically a Raspberry Pi 4 and a TP-Link router, with signals penetrating a 20 cm brick wall at a 5-meter distance. Our results demonstrate a global detection accuracy of 96\% and an Area Under the Curve (AUC) of 0.987 in NLoS scenarios. Comparative analysis shows that the hybrid model significantly surpasses individual 1D-CNN and BiLSTM architectures in classification performance and reliability. The system's efficiency, optimized via TensorFlow Lite for edge computing, ensures low-latency inference while maintaining user privacy. This research underscores the potential of existing Wi-Fi infrastructure for ubiquitous, contact-free medical supervision in complex domestic environments. Graph-Based Consistency Verification for A Reliable Knowledge-Augmented Question-Answering System Han Yan (Shanghai Jiao Tong University); Dan Liang (Sichuan Gas Turbine Establishment, Aero Engine Corporation of China); and Shenghe Li, Hongxin Yan, Han Yu, and Hongming Cai (Shanghai Jiao Tong University) Abstract Abstract Knowledge-augmented question answering (KAQA) based on large language models is promising for industrial decision support, yet safety-critical applications impose zero tolerance for factual and logical errors, while large language models remain prone to hallucinations. RAG improves answers by using external context, but it only helps before generation and lacks a way to check the results afterward. This paper proposes a graph-based consistency verification framework for reliable KAQA in industrial environments. The framework introduces a closed-loop process consisting of three main stages. First, the system structures generated answers into knowledge subgraphs. Second, it verifies these answers by checking logic against a domain knowledge graph and finding evidence in technical documents. Third, when errors are found, the system uses structured feedback to trigger a regeneration loop. This process repeats until the answer is reliable, resulting in a final standardized output that includes a reliability status and verified references. Experiments on an aero-engine testing benchmark show that the proposed method achieves higher reliability than representative baseline approaches, including LightRAG and GraphRAG, improving reliability-related metrics by 27\% to 66\%. Surrogate-Based Multi-Objective Optimization of HDH Desalination Systems Using Artificial Neural Networks Karim Ali, Omar Khater, and Mohammad Abido (KING Fahd University of Petroleum & Minerals) Abstract Abstract Humidification–dehumidification (HDH) desalination is a promising low-temperature option for freshwater production, but accurate performance prediction across different extraction layouts is computationally expensive due to strongly coupled heat and mass transfer effects. To enable rapid design space exploration, this work develops artificial neural network (ANN) surrogate models for three HDH configurations: zero extraction, single extraction, and double extraction. The ANN surrogates are then integrated with the Non-Dominated Sorting Genetic Algorithm II (NSGA-II) to solve a bi-objective optimization problem that maximizes the gain output ratio (GOR) while minimizing a representative dehumidifier area metric. Across the three configurations, the trained ANN models achieve high predictive performance, with R² values close to unity. Pareto front analysis shows a clear improvement in the GOR area trade-off as the number of extraction stages increases. Overall, double extraction provides the most favorable balance between water productivity and dehumidifier size among the studied layouts. Wednesday Virtual Room 5 IJCNN Paper IJCNN Various Tracks VI Session Chair: Xiuwen Liu (Florida State University), Shiva Ahir (Stony Brook University, 934-221-1221) Mechanistic Detection of Patch-Based Backdoors in CNNs via Dispersion-Guided Localization Md. Masum Al Masba, Mao Nishino, and Xiuwen Liu (Florida State University) Abstract Abstract Backdoor attacks pose a serious threat to the re- liability of deep neural networks by embedding hidden trig- gers that induce targeted misclassification at inference time. We propose a gradient-free, post-hoc framework for localizing patch-based backdoor triggers based on latent representation variance. Our key observation is that effective triggers induce background-invariant internal representations, leading to system- atic variance collapse across heterogeneous inputs. Exploiting this phenomenon, our method localizes the spatial position of a trigger using only a single poisoned input and a small clean verification set, without access to gradients, retraining, or clean training data. We evaluate our approach on standard architectures (ResNet-18, WideResNet) and datasets (CIFAR-10, GTSRB), demonstrating accurate and robust localization across different trigger sizes. Empirical results reveal consistent variance collapse in deeper layers of poisoned models, aligning with recent theoretical analyses of backdoor learning dynamics. Residual-Guided Randomized Neural Networks Mushir Akhtar, M. Tanveer, and Mohd. Arshad (IIT Indore) Abstract Abstract Randomized neural networks enable fast and analytically tractable training by fixing the input to hidden layer parameters at random and learning the output weights in closed form; however, their performance critically depends on a single uninformed draw of hidden units. This one shot and task uninformed feature construction often leads to redundant representations and suboptimal utilization of model capacity. To address this limitation, we propose a simple and broadly applicable residual guided procedure that greedily constructs the hidden layer using a closed form residual decrease criterion. At each stage, we (i) generate a pool of random candidate units, (ii) score each candidate by the exact reduction it induces in the ridge regularized objective, (iii) select the top $k$ units, and (iv) refit the readout in closed form using the standard design with direct input links. This procedure yields a progressive training process with a guaranteed monotonic decrease of the training objective. The method is model agnostic: only the candidate generation is architecture specific, while the scoring selection refitting loop is shared across models. Extensive experiments on 71 benchmark datasets from the UCI repository, covering both binary and multiclass classification tasks, demonstrate that the proposed residual-guided models consistently outperform their baseline counterparts in terms of accuracy, stability, and overall ranking performance. Code and Supplementary are available at \url{https://github.com/mtanveer1/RG-RNN}. Context-Aware Backpropagation Using Neuron Activation History Pranjala Kolapwar and Jaishri Waghmare (SGGS Institute of Engineering and Technology) Abstract Abstract Backpropagation is the cornerstone of neural network training, yet its reliance on instantaneous neuron activations makes gradient updates sensitive to noise and short-term fluctuations. This motivates the need for incorporating neuron-level temporal context into the learning process. This paper proposes context-aware backpropagation (CABP), a variant of the standard backpropagation algorithm that augments gradient computation with neuron activation history, enabling temporally contextualized updates while preserving the original forward pass, loss function, and optimization framework. The proposed method maintains a lightweight activation history for each neuron using an exponential moving average (EMA), capturing long-term activation behavior across training iterations. During the backward pass, gradients are adaptively modulated based on this activation history, reinforcing neurons that consistently contribute to learning while attenuating updates for neurons that contribute weakly or are noisy. As a result, CABP achieves lower training loss in fewer iterations, indicating faster convergence and improved optimization efficiency, particularly for large-scale and high-dimensional datasets. Importantly, CABP preserves the same asymptotic computational complexity as standard backpropagation and requires only minimal additional memory for storing activation history. Unlike adaptive moment-based optimizers, CABP avoids parameter-wise moment estimation while delivering stable, robust, and well-generalized learning dynamics. HYPERHEURIST: A Simulated Annealing-Based Control Framework for LLM-Driven Code Generation in Optimized Hardware Design Shiva Ahir, Prajna Bhat, and Alex Doboli (Stony Brook University) Abstract Abstract Large Language Models (LLMs) have shown promising progress for generating Register Transfer Level (RTL) hardware designs, largely because they can rapidly propose alternative architectural realizations. However, single-shot LLM generation struggles to consistently produce designs that are both functionally correct and power-efficient. This paper proposes HYPERHEURIST, a simulated-annealing–based control frame work that treats LLM-generated RTL as intermediate candidates rather than final designs. The suggested system not only focuses on functionality correctness but also on Power-Performance-Area (PPA) optimization. In the first phase, RTL candidates are filtered through compilation, structural checks, and simulation to identify functionally valid designs. PPA optimization is restricted to RTL designs that have already passed compilation and simulation. Evaluated across eight RTL benchmarks, this staged approach yields more stable and repeatable optimization behavior than single-pass LLM-generated RTL. Wednesday Virtual Room 6 IJCNN Paper IJCNN Various Tracks VII Session Chair: Yu Liu (University of International Relations), Shuvra Neel Roy (TCS) Local Sensing to Global Hotspots: Cross-Attention Graph Fusion for Congestion-Aware Multi-Agent Path Finding Shuvra Neel Roy, Arup Kumar Sadhu, and Ranjan Dasgupta (TCS Research) Abstract Abstract How can decentralized agents (here robots) predict and mitigate global traffic jams? This paper introduces a unified framework that forecasts warehouse-wide congestion costmaps using only locally-sensed data and utilizes these predictions to optimize multi-agent path finding (MAPF). We model the environment using two graph representations: a dynamic RoboGraph, derived from local robot interactions, and a static NavGraph, representing the warehouse topology. A Graph Fusion module, implemented via cross-attention, projects sparse, learned robot dynamics onto the global map to predict congestion hotspots. We then propose the congestion-aware MAPF (CA-MAPF) algorithm that leverages these predictions to steer agents through low-density zones. Simulation results demonstrate that our model scales linearly with fleet size and, when integrated into the planner, significantly improves system throughput compared to standard baselines. Tracking Evolving Anomalies with Encrypted Federated Continual Learning Zarka Bashir (IIT Hyderabad, IDRBT); Mridula Verma (IDRBT); and C. Krishna Mohan (IIT Hyderabad) Abstract Abstract Industrial Anomaly Detection (IAD) identifies deviations from normal operations in industrial systems, commonly using visual inspection data captured by camera sensors to monitor product quality. In practice, inspection data is generated across multiple industrial sites with evolving anomaly streams posing challenges in data centralization due to privacy constraints, cross-site generalization due to siloed deployments, and continual adaptation due to catastrophic forgetting. Federated continual learning (FCL) enables IAD systems to learn from distributed industrial data silos and adapt to evolving anomaly patterns without sharing raw inspection data. However, plaintext update exchanges remain vulnerable to gradient inversion and model extraction attacks. We propose Federated Approximate Gradient Matching (FedAGM), a fully homomorphic encryption (FHE)-native FCL framework for unsupervised IAD. FedAGM performs server-side aggregation of encrypted client gradients and aligns them via an FHE-compatible mechanism that approximates non-polynomial federated gradient matching operations with low-degree polynomials. We provide a provable error bound for this approximation, ensuring directional gradient alignment while preserving confidentiality across heterogeneous sites. Experiments on the benchmark MVTec-AD demonstrate that FedAGM achieves performance comparable to SOTA FCL methods, with enhanced privacy. GCLQR: Global Context-Aware Logical Query Reasoning over Knowledge Graphs Pengwei Pan (Soochow University); Jun Ma (University of International Relations); Mingzhe Wang, Xinyi Xia, and Wenxuan Xie (Soochow University); and Yanmei Kang (University of International Relations) Abstract Abstract Logical query reasoning over incomplete knowledge graphs remains challenging for complex queries involving long reasoning chains and semantic constraints. Existing stepwise reasoning methods often fail to capture global logical dependencies, particularly across sequential operations and negation. To address these issues, we propose GCLQR, a global context-aware logical query reasoning framework with three key components: (1) a tree-based serialization strategy that combines depth-first and breadth-first traversals to preserve global structural information; (2) an instruction-context fusion module based on pretrained language models to generate contextualized semantic embeddings; and (3) a transformer-based long-chain query encoder with neural intersection operators that supports full first-order logical reasoning. Experiments on FB15k, FB15k-237, and NELL995 demonstrate that GCLQR consistently outperforms existing baselines, achieving average MRRs of 17.5, 35.9, and 20.0, respectively. Fusion of Structural-Aware Prompts and Contrastive Representation Learning for Commonsense Knowledge Graph Completion Yu Liu, Jun Ma, Xuan Xie, Xun Liu, and Yanmei Kang (University of International Relations) Abstract Abstract Commonsense Knowledge Graph Completion (CKGC) aims to infer missing relationships between entities in a commonsense knowledge graph (CKG). However, existing methods face two significant challenges: the loss of topological information when mapping structured triples into linear text for pre-trained language models (PLMs), and the semantic ambiguity of similar entities that can lead to representation confusion. To address these challenges, we propose a novel inductive commonsense knowledge graph completion method, Fusion of Structural-Aware Prompts and Contrastive Representation Learning for Commonsense Knowledge Graph Completion. Our method introduces two key innovations: (1) structure-aware prompt generation, which encodes local subgraph structures to guide PLMs in generating structure-sensitive representations; (2) contrastive learning, which uses an adversarial loss to better distinguish semantically similar but structurally different entities. Extensive experiments on benchmark datasets demonstrate that our method significantly improves performance, verifying the effective synergy of structural prompts and contrastive representation learning. Wednesday Virtual Room 7 IJCNN Paper IJCNN Various Tracks VIII Session Chair: Liusha Yang (Shenzhen Technology University), Mengyu Yang (Beijing University of Posts and Telecommunications) TemAPT: Efficient Temporal Adaptation of Vision-Language Models for Video Recognition Mengyu Yang and Ye Tian (Beijing University of Posts and Telecommunications), Jianwei Li (Baidu Inc.), Lanshan Zhang and JiaKai Wu (Beijing University of Posts and Telecommunications), YaYa Wei and Ziyan Zhong (Omni-Channel Operation Center China Telecommunications Corporation), and Wendong Wang (Beijing University of Posts and Telecommunications) Abstract Abstract Video recognition aims to interpret dynamic visual content, a task increasingly addressed by transferring knowledge from large-scale pre-trained Vision-Language Models (VLMs). However, existing parameter-efficient fine-tuning (PEFT) methods often face a trade-off: they either ignore complex temporal dynamics to maintain efficiency or incur high computational costs through heavy temporal modules. To bridge this gap, we propose TemAPT, a novel framework that enables robust temporal modeling of vision-language models with minimal overhead. Unlike methods relying on expensive attention mechanisms, TemAPT introduces a token-level temporal shift strategy via zero-cost channel re-indexing. This allows features to exchange information across frames and progressively enlarges the effective temporal receptive field throughout the encoder layers. Furthermore, to improve cross-modal alignment, we devise a semantic-conditioned adaptive prompt generator that dynamically aligns video representations with class-specific textual semantics. Extensive experiments on three widely used benchmarks demonstrate that TemAPT achieves a superior balance between accuracy and efficiency, outperforming state-of-the-art adaptation methods. LSG‑Rec: A Unified LLM‑Enhanced Framework for Cold‑Start and Cross‑Domain Recommendation Zhihao Wang, Ziyi Hu, Pengcheng Zhuo, Wei Ding, and Saijie Ni (Trip.com) Abstract Abstract Item cold-start and long-tail recommendation remain critical challenges in travel platforms due to the sparse interaction data. Although Large Language Models (LLMs) possess rich world knowledge and can extract item semantic information, this is interaction-agnostic and disconnected from collaborative signals. This often leads to popularity bias and "blurry" embeddings that fail to capture nuanced user preferences. To bridge this gap, we propose a unified framework that aligns LLM-enhanced multimodal abstractions with interaction-based dynamics. Our approach incorporates: (1) a LLM-powered encoder to distill high-fidelity semantics through commonsense reasoning; (2) a Minimax-driven alignment module that filters popularity-biased noise to reconstruct precise collaborative representations; and (3) a cross-domain heterogeneous GNN that facilitates robust preference transfer across diverse categories, such as UGC (User Generated Content) and PGC (Professionally Generated Content). Extensive experiments conducted on the homepage feed of the Our App demonstrate that our framework effectively warms up cold items and significantly outperforms state-of-the-art baselines in both long-tail and cross-domain scenarios. Neural Nonlinear Shrinkage of Covariance Matrices for Minimum Variance Portfolio Optimization Liusha Yang and Siqi Zhao (Shenzhen Technology University) and shuqi Chai (Shenzhen Research Institute of Big Data) Abstract Abstract This paper introduces a neural network–based nonlinear shrinkage estimator of covariance matrices for the purpose of minimum variance portfolio optimization. It is a hybrid approach that integrates statistical estimation with machine learning. Starting from the Ledoit–Wolf (LW) shrinkage estimator, we decompose the LW covariance matrix into its eigenvalues and eigenvectors, and apply a lightweight Transformer-based neural network to learn a nonlinear eigenvalue shrinkage function. Trained with portfolio risk as the loss function, the resulting precision matrix (the inverse covariance matrix) estimator directly targets portfolio risk minimization. By conditioning on the concentration ratio (the asset-to-sample size ratio), the approach remains scalable across different sample sizes and asset universes. Empirical results on synthetic data and stock daily returns from Standard & Poor's 500 Index (S&P500) demonstrate that the proposed method consistently achieves lower out-of-sample realized risk than benchmark approaches. This highlights the promise of integrating structural statistical models with data-driven learning. Wednesday Virtual Room 8 IJCNN Paper IJCNN Various Tracks IX Session Chair: Linqi Ye (Shanghai University), Junying Chen (South China University of Technology) High-Coverage LLM-based Unit Test Code Generation through Control Flow Graph Guidance Xinran Luo (South China University of Technology); Cheng Liu (Institute of Computing Technology, Chinese Academy of Sciences); Mengchen Zhao (South China University of Technology); Zhibin Wu and Tao Lu (Commercial Aircraft Corporation of China, Ltd.); and Junying Chen (South China University of Technology) Abstract Abstract Unit testing is a critical yet labor intensive phase in software development. Although large language models (LLMs) have shown promise in unit test code generation, existing LLM-based methods often suffer from limited code understanding, high error rate, and low test coverage. We propose an approach that leverages control flow graphs (CFGs) to improve test quality and coverage. Our method comprises three phases. In the data pre-processing phase, we extract focal methods and dependencies from Java projects and construct CFGs from the abstract syntax tree, which allows LLMs to better focus on the core logic. In the code generation phase, we use LLM to decompose each focal method into CFG-based code slices and generate a coherent unit test code for each slice. In the post-processing optimization phase, we apply a two-stage repair strategy for failing tests to ensure greater test executability, combining template-based repair with LLM-based repair. To further boost coverage, we analyze uncovered control paths via the CFG and guide the LLM to optimize the unit test code test cases accordingly. Experimental results show that our approach significantly outperforms baselines in terms of test execution pass rate and code coverage, with ablation studies highlighting strong inter-module synergy. All experimental data and source code are available at https://github.com/EllaDrCarrot/LUGC. DiffAFGA-Net: A Diffusion-Based Framework for Cervical Precancerous Cell Classification with Adaptive Fine-Grained Attention Tianyu Shi and Xu Ma (Shenyang Ligong University), Zhimin Liu (Dalian Medical University), Xiaoyun Liu and Qingkun Guo (Shenyang Ligong University), and Meng Wang (Liaoning Technical University) Abstract Abstract Cervical cancer remains a significant global health challenge, with early detection critically dependent on accurate classification of precancerous lesions in cytological images. However, conventional deep learning approaches often struggle with robustness issues under complex noise interference, such as staining variations and cellular overlaps, and exhibit limited sensitivity to fine-grained morphological features essential for distinguishing subtle pathological changes. To address these gaps, this paper proposes DiffAFGA-Net, a novel diffusion-based framework that integrates adaptive fine-grained channel attention for enhanced feature representation in cervical cell classification. The model leverages a denoising diffusion probabilistic process to implicitly learn robust latent representations, while an adaptive attention mechanism dynamically captures both global and local discriminative features at multiple granularities, enabling precise emphasis on critical cellular structures like nuclear-cytoplasmic ratios and chromatin distribution. Experiments on multiple cervical cell datasets demonstrate that DiffAFGA-Net achieves superior performance compared to state-of-the-art methods, with notable improvements in classification accuracy, F1-score, and robustness under noisy conditions. The results highlight the potential of combining diffusion models with fine-grained attention for reliable automated screening, offering a scalable solution for clinical applications where label noise and domain shifts are prevalent. This work underscores the value of generative priors and adaptive attention in medical image analysis, paving the way for more generalizable and interpretable diagnostic tools. Position Paper: Compressed Models Should Be Robust and Publicly Shared: A Call for Responsible Model Optimization Rahma Fourati (University of sfax), Jihene Tmamna and Fadoua Drira (University of Sfax), and Javier Sanchez-Medina (University of Las Palmas de Gran Canaria) Abstract Abstract This position paper argues that compressed deep learning models obtained through pruning, quantization, or knowledge distillation should be both robustly evaluated and publicly shared to support responsible model optimization. While model compression has emerged as a vital strategy for enabling efficient deployment and reducing carbon footprint, the machine learning community often shares only the techniques, not the optimized models themselves. Furthermore, these compressed models are rarely evaluated for robustness in terms of uncertainty, adversarial resilience, and sensitivity factors critical for safe deployment. We call for a shift in research culture: robustness evaluation must precede sharing, and sharing should be standard practice for models trained on public benchmarks. This approach will promote reproducibility, democratize access, Position Paper: A New Paradigm for Robot Multimodal Understanding and Decision-Making with Large Language Models as the Cognitive Core Yulai Zhang, Yinrong Zhang, Ting Wu, and Linqi Ye (Shanghai University) Abstract Abstract This paper presents a systematic exploration of the application of large language models (LLMs) as cognitive cores in robotics, focusing on multimodal understanding and intelligent decision-making. While traditional robotic architectures face inherent limitations in environmental modeling precision and task-specific data dependency—struggling with open-ended instruction comprehension, dynamic environment adaptation, and cross-task knowledge transfer—the emergence of LLMs offers a transformative solution. By positioning LLMs as the central cognitive core, robotic systems can achieve deeply integrated perception, reasoning, and decision-making capabilities. This paradigm empowers robots to interpret multimodal inputs more effectively, perform commonsense reasoning, and generate executable action sequences. The paper delineates key implementation pathways, including unified semantic interfaces, commonsense reasoning engines, and metacognitive coordination mechanisms. Furthermore, it examines advanced techniques for enhancing multimodal understanding through vision-language-action models, cross-modal commonsense comprehension, and open-vocabulary semantic construction. The discussion extends to how LLMs facilitate efficient intelligent decision-making processes. Finally, the paper outlines future research directions and proposes mechanisms for achieving long-term adaptability and continuous learning within this new paradigm. Wednesday Virtual Room 1 IJCNN Paper Medical Image Analysis II Session Chair: Yunfeng Liu (Beijing University of Chemical Technology), Song Liu (Qilu University of Technology (Shandong Academy of Sciences)) TREAM: A Temporal Reasoning and Evolutionary Analysis Model for Radiology Report Generation Zhen Zhang, Zhaoyang Wang, Jing Zhao, and Song Liu (Qilu University of Technology (Shandong Academy of Sciences)) Abstract Abstract Radiology report generation requires advanced medical image analysis, effective temporal reasoning, and accurate use of radiological terminology. Although multimodal large language models (MLLMs) align with pre-trained vision encoders to enhance visual–language understanding, most existing methods rely on single-image analysis to process images, failing to fully leverage temporal information in multimodal medical datasets. In addition, insufficient cross-modal alignment can lead to imprecise use of radiological terminology, and the lack of intermediate diagnostic reasoning and incomplete utilization of reference report content further impair interpretability, potentially resulting in hallucinations. These limitations ultimately prevent the models from generating radiology reports that are accurate in clinical terminology and capable of capturing disease progression. To address these limitations, we propose TREAM (Temporal Reasoning and Evolutionary Analysis Model), a temporal-aware MLLM that explicitly captures disease progression from longitudinal imaging features, enhances image-text alignment through an additional contrastive learning loss, and employs a five-step Chain-of-Medical-Thought to generate reports, thereby improving interpretability and reducing hallucinations. We experimented and evaluated our model on the MIMIC-CXR dataset, and the results demonstrate that TREAM achieves superior performance across both natural language generation and clinical effectiveness metrics, highlighting its applicability to longitudinal radiology report generation. Progressive Visual Grounding for Medical Reasoning via Automated Chain-of-Thought Data SOO YONG KIM (Seoul National University), Su in Cho (Boston University), Vincent-Daniel Yun (University of Southern California), and Gyeongyeon Hwang (Jeonbuk University) Abstract Abstract Bridging clinical diagnostic reasoning with AI remains a central challenge in medical imaging. We introduce MedCLM, an automated pipeline that converts detection datasets into large-scale medical visual question answering (VQA) data with Chain-of-Thought reasoning by linking lesion boxes to organ segmentation and structured rationales. Starting from 156k raw generations, quality control mechanisms yield 125k high-fidelity CoT-VQA samples with anatomically grounded rationales. Human evaluation by we authors with help of certified radiologist on 500 randomly sampled instances confirms 91.2\% clinical correctness in our generated data, validating pipeline reliability before model training. To utilize this data effectively, we propose an Integrated CoT-Curriculum Strategy that progressively transitions from explicit visual grounding to implicit localization. On VQA-RAD, SLAKE, PMC-VQA, and challenging benchmarks including OmniMedVQA, MedXpertQA, and BESTMVQA,MedCLM achieves state-of-the-art performance, attaining 72.6\% LLM-judge win-rate over SOTA baseline on radiology report generation when using GPT-5 as judge. Toward Reliable Dermatological Medical Report Generation via Knowledge-Enhanced Structured Multi-modal Reasoning Feng Tan and Yuexian Zou (Peking University Shenzhen Graduate School) Abstract Abstract Accurate and reliable dermatological medical report generation (DMRG) is essential for clinical decision support and improved healthcare accessibility. However, most existing models rely on direct neural mapping from images to text, leading to opaque reasoning and reports that often violate established medical knowledge. To address these limitations, we propose a Knowledge-Enhanced Structured Multi-Modal Reasoning (KE-SMMR) framework that replaces opaque black-box mappings with a structured collaborative paradigm between Multi-Modal Large Language Models (MLLMs) and Large Language Models (LLMs). This collaborative process is governed by three complementary constraints: (1) Perceptual Constraint for auditable and structured observations; (2) Dermatological Knowledge Constraint to restrict reasoning within medically valid boundaries; and (3) Inference Constraint to align multi-stage reasoning with clinical practice. Furthermore, we introduce Triple-Chain Reasoning (TCR), which aligns visual evidence with morphological, anatomical, and pathophysiological knowledge for multi-granularity verification. Extensive experiments demonstrate that our KE-SMMR framework significantly improves report accuracy and medical consistency, offering a knowledge-enhanced reasoning paradigm for reliable DMRG, particularly in data-scarce settings. OphthaVL-Ground: Interpretable Ophthalmic Diagnosis via Attention-Driven Weakly-Supervised Grounding Yunfeng Liu, Ruirui Li, and Yuxin Xiao (Beijing University of Chemical Technology) Abstract Abstract Ophthalmic Multimodal Large Language Models (MLLMs) face a critical dilemma: generic models lack the fine-grained perception required for minute lesions, while specialized models are constrained by the scarcity of high-quality medical reports and bounding box annotations. To resolve this, we present OphthaVL-Ground, a novel framework designed for precise, interpretable, and scalable ophthalmic diagnosis. Our contributions are threefold: (1) We introduce the OphthaVL-W Data Engine, an automated multi-agent pipeline that synthesizes large-scale, pathology-aware diagnostic reports across Color Fundus Photography (CFP), Optical Coherence Tomography (OCT), and Lens Photography (LP), effectively bypassing the bottleneck of manual annotation. (2) We propose a Semantic-Structural Dual-Encoder that synergizes SigLIP’s global semantic alignment with SAM2’s fine-grained spatial features, ensuring the capture of subtle pathological structures often missed by standard encoders. (3) Crucially, we pioneer an Attention-Driven Self-Grounding strategy. By mining internal cross-attention maps to generate pseudo-bounding boxes, we enable the model to learn explicit lesion localization from weak supervision. This facilitates a clinical "localize-then-diagnose'' inference paradigm that significantly mitigates hallucinations. Extensive experiments demonstrate that OphthaVL-Ground achieves state-of-the-art performance in both diagnostic accuracy and visual grounding, setting a new benchmark for interpretable ophthalmic AI. Wednesday Virtual Room 2 IJCNN Paper Medical Image Analysis III Session Chair: Haifeng Zhao (Anhui University, School of Computer Science and Technology), Junyuan Huang (Guangxi University) MGNet: Multi-Axis Similarity Matching and Grouping Contextual Attention Pyramid Network for Deformable Medical Image Registration Junyuan Huang, Mengxiao Yin, Jiachao Li, and Tao Luo (Guangxi University) Abstract Abstract Accurately modeling complex deformations remains challenging in medical image registration tasks, as actual deformations often involve a mixture of large global displacements and subtle local distortions. Although various advanced registration models have been proposed, they still fall short in achieving a balance between long-range similarity matching and multi-scale representation. To this end, this paper proposes a novel Multi-Axis Similarity Matching and Grouping Contextual Attention Pyramid Network for Deformable Medical Image Registration (MGNet). To overcome the limitations of existing explicit matching methods that are constrained by local windows, we design a Multi-Axis Similarity Matching Module (MASM), which explicitly establishes feature correspondences by jointly processing dense local windows and sparse global grids. In addition, to enhance feature representation at each network layer, we introduce a Grouping Contextual Attention Module (GCA) that aggregates grouped features with different receptive fields, enabling refined modeling of complex deformation patterns. Extensive experiments are conducted on three publicly available 3D brain MRI datasets, including LPBA, Mindboggle, and OASIS. The results demonstrate that MGNet achieves state-of-the-art performance in medical image registration. FAReg: Unsupervised Deformable MRI Registration Using Fourier Analysis Yuhao Liu, Boyue Song, *Jinping Tang, and *Ge Zhu (Heilongjiang University) Abstract Abstract Deformable brain MRI registration remains challenging because of complex anatomy and prevalent imaging artifacts. Most existing methods estimate dense displacement fields only in the spatial domain and lack explicit frequency-domain inductive bias, which limits modeling of periodic structures. Many boundary-aware designs also rely on fixed weight maps and are easily misled by pseudo-boundaries caused by partial volume effects and Gibbs ringing, reducing robustness under spatially heterogeneous artifacts. To address these issues, we propose FAReg, an unsupervised registration framework that integrates explicit frequency-domain modeling with adaptive boundary optimization. FAReg introduces a Frequency-domain Feature Fusion Module (FFFM), which explicitly decomposes features with learnable sine–cosine bases to capture quasi-periodic anatomy, and an Edge-Gradient Dynamic Reweighting (EGDR) mechanism, which uses a spatially adaptive gradient-based map to suppress artifact-induced gradients while reinforcing anatomically consistent boundaries. Extensive experiments on two public brain MRI datasets demonstrate that FAReg consistently outperforms existing registration methods in both structural alignment accuracy and boundary fidelity. DAEM: Spectral-Deformable Entropy Modeling for Scalable Learned Image Compression yiming Ding and jianguo Wei (Tianjin University) Abstract Abstract Context modeling is a fundamental component distinguishing modern learned image compression (LIC) from classical transform codecs. However, state-of-the-art entropy models increasingly rely on global self-attention, the quadratic complexity of which becomes prohibitive for high-resolution images. Simultaneously, local context modules are often constrained by fixed grid sampling, which is hindered by misaligned structures and window boundary artifacts. To overcome these limitations, we propose Dual-Aware Entropy Modeling (DAEM), a plug-and-play framework that integrates spectral global context with offset-adapted local context. For global modeling, we employ content-adaptive spectral filtering in the frequency domain via FFT on causal side information (hyperprior features and previously decoded slices), achieving an efficient global receptive field with $O(n\log n)$ complexity. For local modeling, we introduce an offset-adapted context module that predicts content-conditioned sampling offsets atop overlapped window attention. This mechanism improves structural alignment for non-anchor symbols while preserving parallel decoding capabilities. Extensive evaluations on Kodak, Tecnick, and CLIC Professional Validation show that DAEM achieves PSNR-based BD-rate reductions of 12.67\% (Tecnick) and 10.28\% (CLIC) relative to VTM-17.0, and attains the best MS-SSIM BD-rate on Tecnick (-47.55\%). SDRF-Net: Synergistic Dual-Resolution Fusion for Semi-Supervised Medical Image Segmentation Xu Tang (Anhui University, School of Computer Science and Technology); Shiwei Zhou and Dengdi Sun (Anhui University, School of Artificial Intelligence); and Haifeng Zhao (Anhui University, School of Computer Science and Technology) Abstract Abstract Semi-supervised learning (SSL) improves medical image segmentation by exploiting limited labeled and abundant unlabeled data. However, the imbalance between large background and sparse foreground regions hinders the learning of rare classes, while inevitable errors in pseudo-labels cause class confusion and reduce segmentation accuracy. To address these issues, we propose a novel Synergistic Dual-Resolution Fusion Network (SDRF-Net). Our framework introduces a grid-based foreground reorganization (GFR) module to reconstruct and diversify foreground regions, improving rare class recognition. A dual-segmentation head architecture (DSHA) then synergistically integrates global contextual information and local fine-grained details. Furthermore, an entropy-aware weighting (EAW) mechanism is employed to reduce noise during pseudo-label fusion by suppressing unreliable predictions. Extensive experiments on three public medical image datasets show that SDRF-Net effectively alleviates class imbalance and pseudo-label inaccuracy, achieving state-of-the-art performance. Wednesday Virtual Room 3 IJCNN Paper Mixture-of-Experts and Modular Networks Session Chair: Gaosheng Sun (AnHui University), Jinyuan Feng (Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences) FSGMap: Frequency-Space Priors Gated Mixture-of-Experts for HD Map Construction Gaosheng Sun, Luyang Ye, and Yanxu Su (AnHui University) Abstract Abstract High-definition (HD) maps are crucial for autonomous driving. While recent online map construction methods have gained popularity, they often suffer from critical failures in complex scenarios, including directional distortion, missing map elements, and inaccurate keypoint localization. These issues stem from a uniform processing approach that fails to distinguish between heterogeneous map characteristics, such as sharp geometric boundaries and smooth semantic regions. In this paper, we propose FSGMap, a novel framework based on Frequency-Space Priors Gated Mixture-of-Experts. Our approach introduces Frequency-Space Dual Prior Modulation (FSDM) to generate orientation-aware priors and semantic-aware spectral composition priors. To account for prediction reliability under ambiguous frequency characteristics, an Uncertainty-Aware Guidance (UAG) mechanism is integrated to adaptively guide map queries based on entropy-based reliability hints. Furthermore, we design a Prior Refinement Mixture-of-Experts (PR-MoE) that employs hierarchical routing to enable coarse-to-fine expert collaboration. These structured priors and adaptive refinement significantly enhance the model's ability to capture precise topology and reliable geometry. FSGMap achieves exceptional results on the challenging nuScenes dataset, outperforming several widely recognized approaches and improving the baseline model MapTRv2 by 4.5 mAP. Tackling Mask Imbalance: Prototypical Mixture-of-Experts with Hardness-Aware Mining for Medical Segmentation Junjie Jiao, Guoqian Liu, Yanhao Chen, and Shuhao Hu (Xiamen University); Wei Liu, Yin Zhang, and Lei Wang (Academy of Military Medical Sciences, Academy of Military Sciences); and Qingqiang Wu (Xiamen University) Abstract Abstract Medical image segmentation often suffers from severe mask and difficulty imbalance with ambiguous boundaries, where sparse foreground regions are overwhelmed by dominant background pixels. To address these challenges, we propose a Prototypical Mixture-of-Experts framework with a Collaborative Expert Module (CEM) that encourages complementary specialization while preserving global consistency, using a harmonic-mean objective to balance expert competence and inter-expert distillation to enforce consensus. Furthermore, we introduce a Prototypical Contrastive Strategy (PCS) to enhance pixel-level discriminability. PCS leverages consensus predictions to mine hard pixels and contrast their representations against global prototypes, sharpening decision boundaries. Extensive experiments on MoNuSeg, GlaS, COVID-19, and ISIC 2018 demonstrate competitive segmentation performance and improved boundary robustness over strong baselines. PRISM: PRinciple Induced Spectral Mixture of Experts with Test-Time Calibration for Time Series Forecasting Mufeng Liu, Junhui Wang, Xidong Xi, and Guitao Cao (East China Normal University) Abstract Abstract Time series forecasting remains constrained by a fundamental assumption: all samples within a dataset can be effectively modeled by unified parameters. This assumption fails when samples exhibit heterogeneous spectral characteristics—trending series concentrate energy in low frequencies while volatile signals show high-frequency dominance. Existing methods apply dataset-agnostic transformations, unable to exploit sample-specific spectral structure. We propose PRISM, a spectral mixture-of-experts framework addressing this limitation through three principled innovations. First, data-driven expert construction derives band-specific orthogonal transformations from training set correlation matrices, enabling each expert to operate in decorrelated feature spaces tailored to distinct frequency bands. Second, lightweight FFT-based routing assigns samples to experts according to spectral energy distribution, achieving O(NT log T) complexity while providing physical interpretability. Third, test time calibration through low-rank orthogonal parameterization refines transformations for individual samples using self supervised objectives, requiring only O(T r) parameters with r ≪ T. Experiments on standard benchmarks demonstrate that PRISM achieves competitive performance with state-of the-art methods, while maintaining computational efficiency and interpretability. Mixture-of-Experts Policies with Imitation Guidance for Multitask Multi-Agent Reinforcement Learning Yuan Wang, Zhiqiang Pu, Jinyuan Feng, Xiaolin Ai, Tenghai Qiu, and Xiwen Ma (Chinese Academy of Sciences, Institute of Automation) Abstract Abstract Multitask multi-agent reinforcement learning (MARL) aims to train general-purpose cooperative agents across diverse scenarios, but faces fundamental challenges stemming from the multi-modal nature of joint behavior distributions. Different tasks require distinct coordination patterns, posing dual challenges. First, the expressiveness challenge arises because a single shared policy lacks capacity to capture all modes, leading to negative transfer. Second, the mode coverage challenge emerges as standard mode-seeking optimization of reinforcement learning struggles to discover and maintain diverse coordination modes. Combining sparse Mixture-of-Experts (MoE) with generative adversarial imitation learning (GAIL), GEMS (GAIL-Enhanced MoE for multitask Specialization) is developed to address both challenges synergistically. MoE provides structured capacity by decomposing the policy into specialized experts, while GAIL delivers mode-covering guidance through expert demonstrations that embody diverse task-specific coordination patterns. MoE offers the expressiveness to represent multiple modes, while GAIL ensures effective mode coverage by accelerating expert differentiation and preventing mode collapse. Experiments on StarCraft Multi-Agent Challenge (SMAC) and Multi-Agent Particle Environment (MPE) demonstrate that GEMS significantly outperforms state-of-the-art methods in sample efficiency, final performance, and zero-shot generalization. Wednesday Virtual Room 4 IJCNN Paper Model Compression and Quantization III Session Chair: Dat-Thinh Nguyen (University College Dublin), Yunhan Xing (Beijing Jiaotong University) Targeted Bit-Width Quantization-Conditioned Backdoor Quoc-Anh Mai and Van-Thuc Le (Vietnam National University Ho Chi Minh City), Dat-Thinh Nguyen and Nhien-An Le-Khac (University College Dublin), and Kim-Hung Le (Vietnam National University Ho Chi Minh City) Abstract Abstract Quantization-aware training (QAT) is a popular approach to reducing the footprint of large image classifiers while maintaining accuracy, enabling deployments of these models on resource-constrained devices. However, QAT introduces a new attack surface, where adversaries can exploit rounding errors in quantization operators to hide backdoor behaviors that are activated only during quantization; thereby, these backdoors are called quantization-conditioned backdoors (QCBs). Despite recent advances, existing QCB methods exhibit limited stealthiness, as they activate the backdoor upon any quantization precision. In this paper, we propose a targeted bit-width QCB attack that activates only at a specific quantization precision while remaining dormant at all other precisions. Our method modifies the QAT objective to jointly preserve clean accuracy, suppress backdoor behavior at non-targeted bit-widths, while enforcing backdoor activation at the targeted bit-width by leveraging bit-width-dependent rounding discrepancies. We evaluate the proposed attack on CIFAR-10 and CIFAR-100 using MobileNetV2, ResNet34, and VGG16, and show that our method achieves high attack success at the targeted precision while keeping the attack success rate insignificant at other precisions without compromising clean accuracy. SDA-SAM: Semantic-Driven Adaptive Mixed-Precision Quantization for Segment Anything Model Wenbo Tan and Huimin Lu (Changchun University of Technology, Jilin Province Science and Technology Innovation Center for Multimodal Cognitive Computing and Analysis of Medical Biometrics); Jianwei Guo (Changchun University of Technology); and Zexing Zhang (National University of Defense Technology) Abstract Abstract The Segment Anything Model (SAM) exhibits strong zero-shot segmentation capability, yet its parameter count and compute cost hinder deployment on edge devices. Post-training quantization (PTQ) is attractive for compressing SAM with only a small unlabeled calibration set, but existing PTQ pipelines typically adopt uniform bit-widths and are semantic-blindness: layers serving different semantic roles show markedly different quantization sensitivity, while reconstruction metrics (e.g., MSE) correlate poorly with task degradation.To address this, we propose SDA-SAM, a semantic-driven adaptive mixed-precision PTQ framework.Our key idea is to build a Semantic Function Profile (SFP) for each quantizable layer to guide bit allocation before quantization.SFP combines (i) a task-oriented Gradient Sensitivity Score computed from normalized gradient norms using pseudo-label self-training signals, and (ii) Attention Entropy Analysis that categorizes attention layers into localization- versus aggregation-dominant types.Guided by SFP, our Semantic-Driven Bit Allocation assigns higher precision to semantically critical layers and lower precision to robust layers under a global bit budget.On COCO instance segmentation with SAM-B, SDA-SAM matches the accuracy of uniform W6A6 while reducing the effective average weight bit-width to 5.1, and improves W4A4 accuracy by up to 8.5% compared to PTQ4SAM-S, without iterative reconstruction. A Theoretical Analysis of Migration Strength Selection in SmoothQuant Zhihang Liu, Chiu-Wing Sham, Jiale Li, and Sean Longyu Ma (University of Auckland) and Chong Fu (Northeastern University) Abstract Abstract Post-training quantization (PTQ) is essential for efficient inference of large language models, yet remains challenging due to activation outliers with large magnitudes. SmoothQuant mitigates this issue by migrating quantization difficulty from activations to weights using a global migration strength α. However, the choice of α in prior work is empirically motivated and lacks theoretical grounding, while a global α ignores layer heterogeneity. In this paper, we provide a theoretical analysis of SmoothQuant by formulating it as a quantization-noise balancing problem. Under standard quantization noise assumptions, we derive an explicit objective that characterizes the trade-off between activation and weight quantization errors, revealing that the error-minimizing migration strength depends on layer-specific activation and weight statistics. Motivated by this analysis, we propose a layer-wise smoothing strategy in which each layer adopts its own theory-guided α. With minimal modifications to the original method and no additional runtime overhead, the proposed approach consistently reduces layer output error and improves accuracy on LLaMA models under W8A8 quantization, demonstrating the practical effectiveness of the theoretical insights. Quantized DiT with Non-structured Positional Embedding for Signal Synthesis Yunhan Xing (Beijing Jiaotong University), Zikai Zhang and Khaled Harras (Carnegie Mellon University), and Yidong Li (Beijing Jiaotong University) Abstract Abstract Synthesizing high-fidelity wireless signals is essential for addressing data scarcity, yet conventional generative models often fail to account for the non-structured spatial distribution of signal samples and suffer from excessive computational demands. This paper presents SDiT, a quantized Diffusion Transformer for signal synthesis. To address the lack of local spatial correlation in signal sequences, we modify the standard DiT embedding layer by concatenating learnable position embedding into samples and incorporating Rotary Position Encoding to embed continuous spatial coordinates into condition vector. By aligning feature representations with the underlying physical laws of signal propagation, this approach effectively captures the non-structured spatial dependencies of signal samples, successfully transcending the limitations inherent in grid-based image priors. Furthermore, to facilitate efficient deployment, we develop a block-wise quantization algorithm featuring an inter-channel covariance compensation mechanism. By minimizing the Mean Squared Error of layer outputs, this approach effectively suppresses quantization noise accumulation across denoising timesteps. Evaluations on WiFi, FMCW, and MIMO datasets demonstrate that our method significantly reduces model bit-width while maintaining superior generation metrics, including FID and SNR, proving its effectiveness for resource-constrained signal simulation tasks. Wednesday Virtual Room 5 IJCNN Paper Model Compression and Quantization IV Session Chair: Longsheng Zhou (University of Science and Technology of China), Zequan Wang (Northeastern University, China) Reconsidering Discarded Parameters: Aiding Sparse Subnetwork Fine-tuning Man Yuan, Bin Wang, Haodong Bian, Xiaochun Yang, and Zequan Wang (Northeastern University) Abstract Abstract While Pruning-at-Initialization (PaI) algorithms have evolved rapidly to identify efficient sparse structures, their potential is often bottlenecked by a primitive "hard pruning" training protocol. This paradigm strictly freezes unselected parameters, introducing optimization discontinuities and forcing the model to traverse a rugged loss landscape. To address this, we propose ReFIT, a universal training paradigm that is strictly orthogonal to the mask selection criteria. Rather than proposing a new pruning metric, ReFIT serves as a "plug-and-play" booster that revitalizes existing masks. By replacing rigid truncation with Score-Guided Adaptive Regularization, ReFIT utilizes redundant weights as temporary "optimization scaffolds," dynamically suppressing them to stabilize the optimization trajectory during critical early phases. Theoretically, we prove via Schur Complement analysis that ReFIT strictly reduces the maximum eigenvalue of the effective Hessian, guaranteeing a smoother optimization landscape. Extensive experiments on CIFAR-10/100 and ImageNet across diverse architectures (ResNet, VGG, MobileNet) demonstrate that ReFIT consistently boosts the performance of diverse sparse masks—ranging from classical magnitude pruning to random-weight methods like Edge-Popup. By effectively decoupling topology discovery from optimization dynamics, ReFIT validates that "how we train" is just as critical as "what we keep." POST: Continuous Probabilistic Optimization for Structured Dynamic Sparse Training Zequan Wang, Bin Wang, Haodong Bian, Xiaochun Yang, and Man Yuan (Northeastern University) Abstract Abstract Dynamic Sparse Training (DST) serves as a promising paradigm for training efficient deep neural networks from scratch by dynamically adjusting the sparse topology. Although unstructured DST methods can achieve high sparsity, they fail to realize practical acceleration on general-purpose hardware. Conversely, while existing structured DST methods can improve inference speed, they typically rely on sub-optimal heuristic masking criteria or rigid coarse-grained constraints, leading to poor convergence and limited model expressivity. To address these challenges, we propose POST, a novel framework that reformulates the discrete mask search into a continuous probabilistic optimization problem. By introducing learnable structural scores, we establish a gradient pathway directly from the task loss to network topology. This enables the network to leverage topological gradients for structure learning, effectively resolving the misalignment between heuristic indicators and the actual training objective. Furthermore, to balance flexibility with hardware efficiency, we adopt a hybrid topology evolution strategy that combines fine-grained Constant Fan-in constraints with importance-aware neuron pruning, enabling practical acceleration via N:M sparsity. Experiments on CIFAR-10/100 and ImageNet-1K demonstrate that POST outperforms existing structured DST approaches in accuracy. Perplexity-Spectrum Calibration for Post-Training Pruning of Large Language Models Qizhu Wang (Hohai University; Institute of Computing Technology, University of Chinese Academy of Sciences); Xingzhou Zhang (Institute of Computing Technology, University of Chinese Academy of Sciences); Zhihao Qu and Wenfeng Xu (Hohai University); and Wentao Wang (Nanjing University of Posts and Telecommunications) Abstract Abstract Post-training pruning enables efficient deployment of large language models (LLMs) without costly retraining, but it relies on calibration data for weight importance estimation. Related studies analyze the impact of calibration data along two dimensions: data sources and input formats. However, once the dataset is fixed, calibration samples are typically selected by random sampling. Therefore, sample selection within the same data source remains underexplored. We propose a perplexity-spectrum calibration sample selection strategy. This strategy uses the target model’s perplexity to measure sample difficulty. It sorts the candidate pool by perplexity and selects a contiguous difficulty window as the calibration set. We evaluate this strategy on three 7B models: Qwen2-7B, Mistral-7B, and Llama-2-7B. Both Wanda and SparseGPT are tested under 2:4 sparsity. Experiments show that the medium-difficulty window centered at the median achieves the best pruning results across all settings. Compared to random sampling, this strategy reduces perplexity degradation by 20%–26% on WikiText-2. On four zero-shot downstream reasoning tasks, it reduces the standard deviation by 3.4–6.1×. Further analysis reveals that medium-difficulty samples induce stronger and more selective neuron activations. This improves pruning mask consistency (Jaccard similarity) from 0.74 to 0.82. The strategy requires one forward pass to score the candidate pool and a sorting step. It serves as an efficient, plug-and-play calibration optimization step. Prune-Quantize-Distill: An Ordered Pipeline for Efficient Neural Network Compression Longsheng Zhou and Yu Shen (University of Science and Technology of China) Abstract Abstract Modern deployment often requires trading accuracy for efficiency under tight CPU and memory constraints, yet common compression proxies such as parameter count or FLOPs do not reliably predict wall-clock inference time. In particular, unstructured sparsity can reduce model storage while failing to accelerate (and sometimes slightly slowing down) standard CPU execution due to irregular memory access and sparsekernel overhead. Motivated by this gap between compression and acceleration, we study a practical, ordered pipeline that targets measured latency by combining three widely used techniques: unstructured pruning, INT8 quantization-aware training (QAT), and knowledge distillation (KD). Empirically, INT8 QAT provides the dominant runtime benefit, while pruning mainly acts as a capacity-reduction pre-conditioner that improves the robustness of subsequent low-precision optimization; KD, applied last, recovers accuracy within the already constrained sparse INT8 regime without changing the deployment form. We evaluate on CIFAR- 10/100 using three backbones (ResNet-18, WRN-28-10, and VGG- 16-BN). Across all settings, the ordered pipeline achieves a stronger accuracy–size–latency frontier than any single technique alone, reaching 0.99–1.42 ms CPU latency with competitive accuracy and compact checkpoints. Controlled ordering ablations with a fixed 20/40/40 epoch allocation further confirm that stage order is consequential, with the proposed ordering generally performing best among the tested permutations. Overall, our results provide a simple guideline for edge deployment: evaluate compression choices in the joint accuracy–size–latency space using measured runtime, rather than proxy metrics alone. Wednesday Virtual Room 6 IJCNN Paper Multimodal Representation Learning III Session Chair: shenao peng (Hunan University of Technology), Fengjing Song (Qilu University of Technology) Consistency-Aware Cross-Modal Noisy Correspondence Rectification via Perturbation-Invariant Modeling zhongmei wang, shenao peng, jianhua liu, and wenxiu ao (Hunan University of Technology) Abstract Abstract Large-scale image-text data has advanced cross-modal retrieval but inevitably contains noisy correspondences. Most robust training methods rely on the small-loss principle for sample selection, which often discards hard positives semantically matched pairs with large losses caused by semantic complexity. In addition, existing label rectification approaches mainly depend on the model's own predictions, making them vulnerable to error accumulation due to confirmation bias. To mitigate these issues, we propose a Consistency-Aware Noisy Correspondence Rectification (CA-NCR) framework. First, we introduce consistency under perturbations as a second criterion and construct a 2D-Gaussian Mixture Model (2D-GMM) in the loss-consistency space, which effectively separates hard positives from true noise. Second, we design an anchor-guided rectification scheme that utilizes a high-precision clean set as an external semantic anchor to reconstruct robust soft labels through bi-directional geometric consistency, thereby breaking the self-confirmation loop. Extensive experiments on Flickr30K, MS-COCO, and CC152K demonstrate that CA-NCR significantly outperforms state-of-the-art methods under both varying noise ratios and real world noise scenarios. GCC: Global Consistency Correction for Robust Cross-Modal Retrieval Yufan Wen, Yifan Wang, Xinyao Zhang, and Chun Yuan (Tsinghua University) Abstract Abstract Large-scale cross-modal training data collected from the web often contain ambiguously labeled or mismatched pairs, which severely degrade the performance of existing cross-modal retrieval methods. The core challenge in noisy correspondence learning lies in effectively balancing semantic information utilization against noisy interference. However, existing approaches suffer from fundamental limitations: mixture model-based methods waste valuable data through explicit partitioning, while iterative approaches introduce training instability issues. To address these challenges, we propose \textit{Global Consistency Correction} (\textbf{GCC}), a robust cross-modal retrieval framework that corrects sample correspondences by leveraging global similarity relationships across all instances. GCC formulates correspondence correction as a global optimal transport problem and ensures training stability through a two-stage strategy with a unified multi-level alignment (instance-level, global structural, and local selective), thereby enabling full utilization of all available data pairs without explicit partitioning. Extensive experiments demonstrate that GCC achieves state-of-the-art performance under both synthetic and real-world noise conditions, while enabling seamless integration into existing training pipelines without requiring architectural modifications. LG-Cog: Language-Grounded Representation Learning with Cross-Modal Cognitive Priors for Image-Text Retrieval Xiang Liu and Yu Bai (School of Computer Science, Shenyang Aerospace University); Peng Lian (Shenyang Northern Software College of Information Technology); and Haifeng Chi and Haoxiang Li (School of Computer Science, Shenyang Aerospace University) Abstract Abstract Contrastive learning has become a standard approach for image-text retrieval, aligning vision and language within a shared semantic space. However, in real-world scenarios, the assumed semantic consistency of positive pairs is frequently violated by noisy or partially matched data. To address this issue, we propose Language-Grounded Representation Learning with Cross-Modal Cognitive Priors (LG-Cog). LG-Cog leverages a frozen multimodal large language model with carefully designed prompts to generate fine-grained and semantically rich textual descriptions for each image, which serve as cross-modal cognitive priors complementary to the original paired texts. The joint alignment of image representations with the representations of both the original and generated texts drives visual representations toward a shared language semantic space, establishing a language-grounded representation learning paradigm. By integrating language-grounded representation learning with cross-modal cognitive priors, LG-Cog effectively mitigates semantic bias. Extensive experiments demonstrate the superiority of LG-Cog over state-of-the-art methods, achieving 3.7\%-4.0\% RSUM improvements on Flickr30K and MS-COCO benchmarks. Multi-Scale Hypergraph Convolutional Network for Low-Quality Text-Based Scene Retrieval Fengjing Song, Shengrong Zhao, and Hu Liang (Qilu University of Technology) Abstract Abstract Image-text retrieval is a core task in vision-language. However, current cross-modal methods often rely heavily on high-quality training data, which tends to cause issues such as performance degradation when the model is applied to scenarios with semantic ambiguity or incomplete data. To address this problem, we propose the Modal Multi-Scale Hyperedge Retrieval (MMHR) model. Specifically, its Multi-Scale Hyperedge Module (MHM) introduces a fine-grained hyperedge taxonomy to reinforce cross-modal/cross-scale node connections, while the Multi-Dimensional Matrix Hypergraph Convolution Module (MMHC) designs dual multidimensional matrices to quantify node contribution weights and enhance node-hyperedge interactions, helping the model prioritize nodes with high information density during the modeling process. These designs enable the model to adjust relationships between nodes of different modalities, thereby calibrating cross-modal retrieval. Experiments on Flickr30K and MSCOCO demonstrate that MMHR outperforms existing methods in cross-modal retrieval. Wednesday Virtual Room 7 IJCNN Paper Multimodal Representation Learning IV Session Chair: Gengchen Liu (Qilu University of Technology), He Yuanye (Institute of Information Engineering, Chinese Academy of Sciences) DTPD-MEE: A Dual-path Teacher and Pseudo-label Distillation Framework for Multimodal Event Extraction He Yuanye and Hu Shuhao (Institute of Information Engineering, Chinese Academy of Sciences); Wang Jin (School of Information Science and Engineering, Yunnan University); and Xiang Ji, Guo Xiaobo, and Zhao Qingfei (Institute of Information Engineering, Chinese Academy of Sciences) Abstract Abstract Multimodal Event Extraction (MEE) is critical for holistic semantic understanding but is currently severe bottlenecked by the extreme scarcity of high-quality annotated multimodal datasets. To address this challenge, we propose DTPD-MEE, a novel Dual-path Teacher and Pseudo-label Distillation framework that shifts the focus from architectural complexity to a data-centric knowledge transfer paradigm. Our approach adopts a ``divide-and-conquer, then unify'' strategy: first, two specialist teacher models are trained on large-scale unimodal corpora (text-only and image-only) to capture deep modality-specific expertise. These teachers then function as automated annotators for vast unlabeled multimodal data. Through a rigorous cross-modal consistency check filter, we successfully construct a high-quality synthetic dataset that is 3x larger than the standard M2E2 benchmark. Finally, a unified student model inherits this distilled expertise through sequence-level distillation. Experimental results demonstrate that DTPD-MEE establishes new state-of-the-art performance on the M2E2 benchmark, notably improving the multimedia argument extraction F1-score by 7.3 points. Our framework effectively demonstrates how to unlock the potential of unlabeled multimodal data by bridging the gap between unimodal abundance and multimodal scarcity. PHE2: A Chinese Document-level Dataset for Public Health Event Extraction Qian Chen, Jiaju Ren, Xin Guo, and Suge Wang (Shanxi University) Abstract Abstract Given the critical role of public health in national security and social stability, effective surveillance of emerging threats is essential, which necessitates Document-level Event Extraction (DEE) to derive structured event knowledge from public health news for timely risk assessment and decision-making. However, research on DEE in public health remains limited, primarily due to the prohibitive cost of manual annotation that hampers the construction of large-scale, high-quality datasets. To bridge this gap, we propose a modular and extensible LLM-assisted annotation framework for Event Extraction, with which we construct a Chinese document-level dataset for Public Health Event Extraction(PHE2), a large-scale Chinese dataset comprising 10,028 documents, over 21,000 events, and 90,000 arguments. Our proposed framework achieves F1 scores of 80.68% for argument extraction on a human-annotated test set, verifying its effectiveness. We benchmark several DEE models on PHE2, and results show that existing models achieve substantially lower performance on PHE2, highlighting its unique challenges and strong potential to advancein research DEE. PHE2 is publicly available at https://gitee.com/mike29/PHE2. Hierarchical Event-Relevant Knowledge Injection for Event Argument Extraction Lexin Ding, Shuchao Pang, and Xuxu Ma (School of Cyber Science and Engineering, Nanjing University of Science and Technology, Nanjing, China) and Bing Li (School of Information and Communication Engineering, University of Electronic Science and Technology of China, Chengdu, China) Abstract Abstract Recent Event Argument Extraction (EAE) approaches incorporate auxiliary event-relevant knowledge to improve performance. However, existing methods (1) lack a systematic taxonomy of event-relevant knowledge and (2) exhibit limited knowledge coverage and effectiveness. To address these limitations, we hierarchically decompose event-relevant knowledge into Event Commonsense Knowledge, which is event-type-agnostic, and Inter-Event Correlation Knowledge, which captures transferable relational patterns among related events. A Two-Phase, Parameter-Isolated Training Paradigm inject these knowledge types into distinct parameter subspaces, enabling abstract knowledge acquisition while mitigating catastrophic forgetting. Moreover, a Meta-role Enhanced EAE Prompt and an Event-Correlation Capture Module enable adaptive activation of relevant knowledge conditioned on context, enhancing both the diversity and effectiveness of injected knowledge. Experiments on RAMS, WIKIEVENTS, and MLEE demonstrate consistent improvements over strong baselines, indicating deeper utilization of event-relevant knowledge. Extensive ablation and cross-dataset analyses further confirm the effectiveness and strong scalability of our framework. A Multimedia Event Detection Method based on Dual-Path Iterative Synergistic Fusion for Modal Information Asymmetry Gengchen Liu, Tao Sun, Xiaoyu Wang, Zhi Yang, Jiahui Liu, and Zimeng Xu (Qilu University of Technology) Abstract Abstract Multimedia Event Detection (MED), as a fundamental task in multimodal understanding, plays a vital role in downstream applications such as information retrieval and situational awareness. However, current MED approaches often suffer from two key limitations: (1) insufficient exploitation of the rich semantics embedded in visual data, particularly in social media contexts where textual descriptions are brief, ambiguous, or noisy; and (2) limited cross-modal interaction, typically relying on shallow fusion strategies that fail to capture fine-grained alignments between modalities. To address these challenges, this paper proposes a novel Dual-Path Synergy Multimodal Model (DPSMM), which incorporates two core modules: the Vision-Language Synergy Enhancer (VLSE) and the Dual-Iterative Attentive Fusion (DIAF). VLSE enriches image representations by generating verb-centric captions and extracting local textual cues through OCR, thereby transforming implicit action semantics into explicit textual signals while filtering noise. DIAF then performs bidirectional, multi-round attention to progressively align and refine multimodal features. Together, these modules enhance semantic consistency across modalities and improve robustness under weakly aligned or noisy inputs. Experimental results on the M2E2 benchmark demonstrate that DPSMM achieves superior performance compared to existing state-of-the-art models, with a 3.6\% increase in key evidence recall, and a 1.5\% gain in overall F1 score, validating the effectiveness of the proposed model in capturing deep multimodal semantics for robust event detection. Wednesday Virtual Room 8 IJCNN Paper Multimodal Representation Learning V Session Chair: Jinhui Yi (University of Bonn), Yanwei Yu (Ocean University of China) EPIC: Semantic Inverse Prompting and Evidential Corroboration for Multi-modal Object Re-Identification Xingan Ma (University of Bonn), Yuhao Wang (Dalian University of Technology), and Jinhui Yi and Juergen Gall (University of Bonn) Abstract Abstract Multi-modal object Re-Identification (ReID) aims to match object identities by leveraging complementary information from multiple modalities. However, existing methods primarily focus on fusing heterogeneous features, while overlooking modality-specific noise that can severely degrade feature quality. Moreover, most fusion strategies treat modalities uniformly, failing to adaptively weight them under varying imaging conditions. To address these issues, we propose EPIC, a novel framework that tackles modality-specific noise through semantic governance and performs reliability-aware fusion via evidential corroboration. Specifically, we propose Semantic Inverse Prompting and Consensus (SIPC), which converts textual semantics into effective visual guidance via three cascaded sub-modules. Minimum Consensus Gating (MCG) filters unreliable semantics through cross-modal consensus, Semantically-Anchored Inverse Prompting (SAIP) projects the gated semantics into visual prompt tokens, and Multi-grained GeM Aggregation (MGA) aggregates multi-grained representations. We further introduce Evidential Corroborated Reliability (ECR), which estimates modality reliability from evidential theory and strengthens representations with adaptive residual rectification. Together, SIPC and ECR form a coherent pipeline from semantic denoising to reliability-aware fusion, yielding robust multi-modal representations in complex scenarios. Extensive experiments on three multi-modal object ReID benchmarks demonstrate the effectiveness of EPIC. FreqDF: Frequency-guided Disentanglement Framework for Multi-modal Brain Tumor Segmentation Ruibin Li, Meiyi Wei, Chudi Hu, Zhongling Wu, and Gang Chen (Wuhan University) Abstract Abstract Accurate segmentation of brain tumors from multi-modal Magnetic Resonance Imaging (MRI) is critical for clinical diagnosis and treatment planning. Multi-modal learning aims to capture modality-shared global shapes while preserving modality-specific contrasts and textures. However, existing approaches often extract and integrate features indiscriminately, leading to feature redundancy and blurred boundaries. To address these issues, we introduce a novel Frequency-guided Disentanglement Framework (FreqDF) for multi-modal brain tumor segmentation. Our framework employs a dual-path encoding architecture to explicitly model modality-shared anatomical structures and modality-specific texture details in parallel. Based on these decoupled representations, we design a two-stage Enhancement-Fusion Module to collaboratively integrate features, ensuring low-frequency structural coherence while selectively refining high-frequency fine-grained edges. To guarantee effective disentanglement, we introduce a Frequency-Domain Constraint strategy. During training, we impose consistency on low-frequency components to extract unified anatomical shapes while maximizing the discrepancy between high-frequency components to preserve unique contrasts and textures. Extensive experiments and comprehensive ablation studies demonstrate that our method achieves superior segmentation accuracy. The code will be available soon. DDHMamba: Dual Domains Hilbert Mamba for Multi-Contrast MRI Super-Resolution Yuguang Yan, Yonglin Xu, Zipeng Zhu, Boyan Xu, and Ruichu Cai (Guangdong University of Technology) Abstract Abstract Driven by advancements in deep learning and increasing clinical demands in Magnetic Resonance Imaging (MRI), Multi-Contrast MRI Super-Resolution (MCSR) has been developed. It utilizes an easily obtainable reference contrast MRI as an auxiliary to improve the quality of low-resolution target contrast MRI. While the existing approach shows promise, convolution-based methods cannot capture long-range dependencies, and Transformers suffer from quadratic complexity. Recently, image super-resolution technology based on Mamba has received significant attention due to its effectiveness. However, Mamba's sequential scanning hinders the preservation of local details and the learning of dependencies among multi-modal images. Moreover, the frequency features across different contrast MRIs exhibit inherent relationships, modeling their complex correlations within the Fourier domain remains a significant challenge. To mitigate these challenges, we propose Dual Domains Hilbert Mamba (DDHMamba), a new Mamba-based framework that captures long-range dependencies with linear complexity through dual domain modeling, leveraging multi-modal features for efficient MRI super-resolution. Specifically, we propose a Hilbert Scan Fusion Mamba (HSFM) module with two novel cross-modal Hilbert scanning mechanisms: one operating in the spatial domain and the other in the Fourier domain to better capture features of distinct modalities across their respective domains. In addition, we introduce a Frequency Assisted Local Fusion (FALF) module to enhance the fusion capability of local image information from two complementary modalities. Experiments on BraTS and fastMRI datasets demonstrate that DDHMamba achieves superior MCSR performance across multiple scales. Multi-modal Fine-grained Semantic-preserving Hashing for Cross-modal Retrieval Bohan Zhang, Yaru Gao, Chenlong Song, Yuan Cao, and Yanwei Yu (Ocean University of China) Abstract Abstract With the rapid growth of multimodal data, cross-modal retrieval has become increasingly important in both the natural and remote sensing domains. Cross-modal hashing is appealing for its compact storage and efficient retrieval. However, most existing methods rely primarily on intramodal similarity, overlooking structural and fine-grained information. To address this limitation, we propose a fine-grained attention-based multiple similarity matrix derived from transformer output, which aggregates local concept-level features to capture more representative semantics. In addition, we design a loss framework combining semantic consistency, discrete quantization, and feature reconstruction alignment. By using the semantic feature reconstruction based on the automatic encoder framework, our approach unifies cross-modal similarity. Extensive experiments on two public and one remote sensing datasets demonstrate clear advantages over state-of-the-art methods. Wednesday Virtual Room 1 IJCNN Paper Multimodal Representation Learning VI Session Chair: Yichen Huang (East China Normal University), Purui Bai (CASIA; School of Artificial Intelligence, University of Chinese Academy of Sciences) MVPBench: A Multi-Video Perception Evaluation Benchmark for Multi-Modal Video Understanding Purui Bai (Institute of Automation,Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences); Tao Wu (Shanghai, ShanghaiTech University); Jiayang Sun and Xinyue Liu (Institute of Automation,Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences); and Huaibo Huang and Ran He (Institute of Automation,Chinese Academy of Sciences) Abstract Abstract The rapid progress of Large Language Models (LLMs) has spurred growing interest in Multi-modal LLMs (MLLMs) and motivated the development of benchmarks to evaluate their perceptual and comprehension abilities. Existing benchmarks, however, are limited to static images or single videos, overlooking the complex interactions across multiple videos. To address this gap, we introduce the Multi-Video Perception Evaluation Benchmark (MVPBench), a new benchmark featuring 14 subtasks across diverse visual domains designed to evaluate models on extracting relevant information from video sequences to make informed decisions. MVPBench includes 5K question-answering tests involving 2.7K video clips sourced from existing datasets and manually annotated clips. Extensive evaluations reveal that current models struggle to process multi-video inputs effectively, underscoring substantial limitations in their multi-video comprehension. We anticipate MVPBench will drive advancements in multi-video perception. The evaluation code and dataset will be available at https://github.com/MVPBench/MVPBench. Liva: Adaptive Bitrate Ranking with Dynamic Value Perception for Video Streaming Zhiqiang Yang (Shanghai University); Hang Xiao (Shanghai Minhang Polytechnic); and Chongchong Yang, Chaoqian Liu, and Wenhao Zhu (Shanghai University) Abstract Abstract Adaptive Bitrate (ABR) algorithms are extensively deployed in video streaming to improve Quality of Experience(QoE). Common bitrate decision approaches typically emphasize physical factors such as network bandwidth and latency. We argue that designing truly effective mechanisms necessitates a more holistic perspective: beyond adapting to time-varying network conditions, they must also perform fine-grained resource allocation that explicitly integrates video content characteristics with user viewing preferences. In particular, users’ quality sensitivity and retention willingness vary across video segments. By allocating bandwidth differentially according to segmentlevel user attention and content reuse potential, overall QoE can be improved with limited bandwidth overhead. To this end, we propose Liva, a lifecycle value-aware ABR algorithm. Liva introduces Value-Aware QoE (VA-QoE), a dynamic QoE metric that characterizes content value in real time by fusing collective interaction behaviors with content access characteristics. Based on this metric, Liva predicts the performance outcomes of candidate actions and formulates bitrate decisions to maximize long-term rewards, thereby generating high-quality playback sequences that enhance overall QoE. Extensive experiments demonstrate that, across diverse network conditions, Liva improves VAQoE by approximately 11.3%–26.7% over the strongest baseline and effectively enhances the viewing experience on high-value segments. ESDN: Event-Centric Semantic Differential Denoising for Long-Form Audio-Visual Video Understanding longfei zhang, lixu sun, Mengen Huang, and Nurmemet Yolwas (Xinjiang University) Abstract Abstract Multisensory temporal event localization aims to precisely localize modality-aware events and their temporal boundaries in long-form videos. However, dense event overlaps in long videos cause relevant semantics and irrelevant semantic noise to intertwine during cross-modal interaction, leading to severe semantic entanglement. Consequently, existing methods frequently suffer from missed detections and insufficient recog- nition precision. To address this, the Event-Centric Semantic Differential Denoising Network is proposed. Specifically, the Semantic Differential Pyramid Multimodal Transformer utilizes Cross-modal Semantic Differential Attention to reconstruct cross- modal interactions for enhancing relevant semantics fusion; its driven pyramid strategy ensures a balance between noise suppression and temporal dependency modeling for events of varying durations. Multi-Granularity Semantic Loss Guidance mitigates label noise via fine-grained supervision and utilizes an event-aware joint loss to optimize event relation modeling. Experiments demonstrate significant performance improvements. UniMixer: Unifying Context and Latent Distribution Tokens for Generative Watch Time Prediction Yichen Huang, Jing Liu, Li Han, Guanyu Lin, Erxue Zhou, and Zhengkang Zhou (East China Normal University) Abstract Abstract Accurate watch time prediction is essential for optimizing content delivery in streaming short-video platforms, yet it remains challenging due to the inherently complex and multimodal nature of watch time distributions. Through systematic analysis of real-world industrial data, we identify two critical challenges: (1) unobservability of latent multi-duration distributions, where watch time emerges from the entangled interaction between users' latent preferences and videos' content structures, and (2) inter-distribution cross-dependency, where different distribution components dynamically overlap and modulate each other along the temporal axis. To address these challenges, we assume that watch time follows an Exponential-Log-Normal mixture (ELN) distribution, where the exponential component captures quick-skip behaviors and the log-normal components characterize diverse long-tail engagement patterns. Accordingly, we propose UniMixer, a linear-complexity architecture that parameterizes each distribution component as a learnable token and achieves deep interaction between contextual features and distribution prototypes through a unified token-mixing mechanism. Extensive experiments on two large-scale public datasets, KuaiRec and WeChat, demonstrate that UniMixer achieves state-of-the-art performance on both ranking (XAUC) and regression (MAE) metrics. Remarkably, UniMixer maintains computational efficiency suitable for real-time industrial deployment while exhibiting superior distribution fitting capability across diverse user-video interaction patterns. Our code is available at https://github.com/Lambert-vsziii/UniMixer. Wednesday Virtual Room 2 IJCNN Paper Multimodal Representation Learning VII Session Chair: PEIFENG LI (Soochow Univeristy, Soochow University), PEIQIANG WANG (Tsinghua University, Shenzhen International Graduate School) Knowledge-Based Multimodal and Context Incongrity Modeling For Sarcasm Detection Wenjie Xu, Zhong Qian, Peifeng Li, and Qiaoming Zhu (Soochow university) Abstract Abstract Multimodal sarcasm detection is a challenging problem in sentiment analysis, which requires recognizing mismatches between literal content and intended meaning. While recent methods have achieved promising results, they often emphasize inconsistencies associated with strongly affective words, potentially neglecting informative signals conveyed by other tokens. Moreover, without adequate commonsense knowledge, existing approaches may not fully utilize dialogue context, resulting in degraded performance on cases that require deeper semantic inference. To address these limitations, we propose a framework that jointly models fine-grained cross-modal conflicts and knowledge-enhanced contextual semantics. Specifically, we develop a Token-level Multimodal Inconsistency Modeling module to capture token-wise contradictions between textual and non-verbal cues. In addition, we introduce a Knowledge-based Contextual Semantic Inconsistency Modeling module that leverages large language models (LLMs) to incorporate external commonsense, facilitating the identification of implicit logical inconsistencies within the dialogue context. Experiments on the MUStARD benchmark show that our approach yields consistent improvements over prior work. Our model achieves an F1 score of 77.9 in the speaker-dependent setting and maintains competitive performance in the speaker-independent setting. Intent-Oriented Hierarchical Incongruity Guidance for Multimodal Sarcasm Detection Siyuan Li and Bangjun Wang (Soochow University) Abstract Abstract Multimodal sarcasm detection aims to identify whether an image–text pair conveys sarcastic sentiment. Many existing methods have achieved strong performance by modeling cross-modal incongruity. However, these methods typically ignore the important role of implicit visual intent semantics related to sarcasm, thus limiting the further improvement. To address this limitation, we propose an Intent-oriented Hierarchical In congruity Guidance framework(IHIG). We first leverage a large visual-language model to generate deep visual intent semantics. Then, we design a Hierarchical Inconsistency Guidance Module(HIGM), which first models global relationships in each modality to obtain shallow inconsistency signals between text and image, and deep inconsistency signals between textual and visual intent semantics. These signals further guide a three-stage interaction and fusion process between the text graph and the visual graph, thus sarcasm-related cross-modal incongruity is strengthened in the graph structure, mitigating the dilution of key cues. Finally, in the graph contrastive regularized representation space, an adaptive Gaussian kernel-weighted prediction strategy is employed to perform multimodal sarcasm detection. Extensive experiments on benchmark datasets demonstrate our approach outperforms previous state-of-the-art baselines. Aem45k: a Large-scale Multimodal Dataset and Method for Aesthetic Assessment of Student Artwork in Middle-school Education Shuhao Hu, Yanhao Chen, Yan Liu, Ruoxuan Liang, and Junjie Jiao (Xiamen University); Jing Jin (Zhejiang Normal University, Hangzhou Tianchang Guanchao Primary School); Wei Liu, Yin Zhang, and Lei Wang (Academy of Military Medical Sciences, Academy of Military Sciences); and Qingqiang Wu (Xiamen University) Abstract Abstract In middle-school art education, teachers must grade large volumes of student drawings under strict time constraints, while ensuring rubric consistency and multi-dimensional feedback. To support scalable educational assessment, we present AEM45K, a dataset of 44,573 exam-condition drawings annotated with five rubric-aligned attributes and a final grade. A verified subset of 7,360 samples further includes student-written descriptions that clarify intent and theme, treated as an optional modality. We propose MADENet, a unified framework for rubric-aligned educational aesthetics assessment with or without text. MADENet retrieves attribute-specific evidence via a dimension query bank, grounds queries to image/text tokens through query-to-token cross-attention, and applies dimension-adaptive gated fusion to balance modalities under varying availability. Experiments under image-only and multimodal protocols show consistent improvements over strong baselines, with the largest gains on semantics-heavy dimensions when descriptions are available. BiKV: Bi-granularity KV-Cache Optimization with Adjacent-layer Reuse and Modality-Specific Pruning PEIQIANG WANG (Tsinghua University, Shenzhen International Graduate School) and Wei Tao (Huazhong University of Science and Technology) Abstract Abstract Multimodal large language models require efficient inference under complex and long-context scenarios. KV-cache is widely used to accelerate decoding and reduce memory consumption. However, existing approaches often fail to accurately identify critical tokens, overlook inter-layer redundancy, and ignore modality-specific characteristics. To address these limitations, we propose BiKV, a unified KV-cache optimization framework that jointly operates at inter-layer and intra-layer levels. At the inter-layer level, BiKV reuses key–value pairs across highly similar adjacent layers to eliminate redundant computation. At the intra-layer level, it leverages cross-modal attention to select informative tokens, ensuring the preservation of essential multimodal information. Experiments on representative long-context tasks from the MileBench benchmark demonstrate that BiKV significantly reduces memory usage and inference latency while maintaining competitive reasoning performance, validating its effectiveness for efficient multimodal inference. We will make our data/code available upon acceptance. Wednesday Virtual Room 3 IJCNN Paper Multimodal Representation Learning VIII Session Chair: Jia Wang (Xinjiang University), Hui Zhang (Southwest University Of Science And Technology) Fair Spectral Clustering with Multiple Sensitive Attributes via Shapley Values Tao Wu (Southwest University Of Science And Technology), Zhijing Yang (Southeast University), and Junjie Zheng and Hui Zhang (Southwest University Of Science And Technology) Abstract Abstract Fair spectral clustering has attracted increasing attention in recent years. Although numerous methods have been proposed with promising results, most of them focus on a single sensitive attribute, while fair spectral clustering algorithms ensuring multiple sensitive attributes have received relatively little attention. To solve this issue, this paper introduces an innovative Multi-sensitive Attribute Fair Spectral Clustering algorithm (MFSC) by incorporating Shapley value theory. To verify the effectiveness of the proposed MFSC, we conducted experiments on four benchmark datasets, comparing it with spectral clustering and fair spectral clustering methods in terms of fairness and clustering performance. Experimental results demonstrate that MFSC substantially improves fairness while maintaining clustering quality. Improving Individual Fairness in Gaussian Mixture Clustering Chuan Qian (Southwest University of Science and Technology), Zhijing Yang (Southeast University), and Hui Zhang (Southwest University Of Science And Technology) Abstract Abstract Clustering is an unsupervised machine learning algorithm. Among the many clustering algorithms, the Gaussian Mixture Clustering (GMC) is widely used for its effectiveness and flexibility in modeling complex data distributions. Fair clustering has received a lot of attention in recent years. Fair clustering ensures fairness between data points or clusters and is essential for reducing bias due to inherent distributional differences. Fairness in clustering is usually categorized into individual level and group level. Individual fairness implies that similar individuals should be treated similarly, but traditional GMC usually ignores this principle, which makes the clustering results inevitably unfair. This paper presents an innovative fair version of GMC, Individual Fair Gaussian Mixture Clustering (IFGMC). By accounting for individual fairness during each iteration, IFGMC ensures the final clustering results consistently respect the principle of equity. Experimental results demonstrate that IFGMC achieves over 30% higher fairness performance than GMC, significantly enhancing individual fairness. This makes it a highly promising solution for achieving fair and reliable clustering. ATMSC-Net: Deep Unfolding Network for Anchor-based Tensorial Multi-View Subspace Clustering Tianhao Huang, Xiaojun Wu, Wenhua Dong, and Tianyang Xu (Jiangnan University) Abstract Abstract Anchor-based and tensor-based approaches have emerged as two promising directions for multi-view subspace clustering, offering linear-complexity scalability and high-order correlation modeling, respectively. However, existing deep methods typically incorporate these structures in a heuristic or task-agnostic manner, lacking principled integration with the underlying clustering optimization and failing to combine both paradigms effectively. To bridge this gap, we propose ATMSC-Net, a deep network architecture that systematically unfolds the ADMM iterative solution of anchor-based tensorial multi-view subspace clustering into principled network structures. The proposed model decomposes the clustering process into four interpretable modules, RepresentModule, NoiseModule, TensorModule, and AnchorModule, each derived by unfolding a specific optimization subproblem into a dedicated network component, providing structural clarity and optimization traceability. By introducing TensorModule into the deep unfolding framework, ATMSC-Net extends the joint anchor-tensor paradigm from shallow methods to deep learning. Additionally, hand-crafted hyperparameters are transformed into learnable parameters through end-to-end training, enabling automatic adaptation to diverse data distributions. Extensive experiments on multiple datasets demonstrate that ATMSC-Net achieves competitive clustering performance compared with state-of-the-art methods. A Transformer-Based Image Manipulation Localization Network with Dual-Frequency Attention and Multi-Scale Feature Fusion Feijiang Chen, Jia Wang, and Qianxi Pan (Xinjiang University) Abstract Abstract Image manipulation detection and localization aims to accurately identify and annotate tampered regions within images and represents a critical research direction in multimedia forensics. Although deep learning-based methods have achieved significant progress in image manipulation localization tasks in recent years, existing models still tend to suffer from missed detections, false alarms, and inaccurate boundary localization when dealing with tampered regions of diverse scales, as well as scenarios where manipulation boundaries are highly blended with complex backgrounds and manipulation traces are concealed. To address these issues, this paper proposes an image manipulation localization network based on Dual-Frequency attention and Multi-scale feature fusion (DFM-Net). The proposed method replaces the conventional multi-head self-attention in Transformers with a Dual-Frequency Attention (DFA), where two branches focus on local fine-grained details and global structural information, respectively, thereby enhancing the perception of tampering clues across different scales. Furthermore, a Feature Enhancement Aggregation (FEA) module is designed to apply different forms of enhancement to RGB features and noise features and incorporate a coordinate attention mechanism, enabling more effective utilization of RGB and noise features for precise tampering boundary localization. Experiments on CASIA, NIST16, Columbia, and Coverage demonstrate that DFM-Net achieves the highest average F1 score of 67.6% and AUC of 92.0%, higher than the baseline models by 3.4% and 1.2%, respectively. Wednesday Virtual Room 4 IJCNN Paper Network Security and Fault Diagnosis I Session Chair: Shaily Kabir (University of Nottingham), Chengming Liu (Zhengzhou University, School of Cyber Science and Engineering) Synergizing Laplacian Positional Encodings and Graph Transformers for Robust Network Intrusion Detection Chengming Liu and Yongkang Yuan (Zhengzhou University) Abstract Abstract Network intrusion detection is critical for securing modern cyber infrastructure. However, traditional graph-based methods often struggle with limited receptive fields and over-smoothing issues. This paper proposes the Dual-stream Graph feature learning Network (Dual-Net), a structural-aware framework designed for robust global relational reasoning. The core of our method lies in the integration of Graph Transformers and Laplacian Positional Encodings (LapPE). The proposed framework comprises three main components: adaptive graph construction, dual-stream feature learning, and structural augmentation. First, an adaptive module is designed to extract high-confidence dependencies from raw traffic using k-NN estimation and semantic pruning. Second, we implement a dual-stream architecture to capture multi-scale features. In this architecture, Graph Transformers are employed to model long-range topological relationships. To enhance this process, LapPE is integrated into the Transformer stream. This encoding effectively captures the structural topology of the network and overcomes the permutation-invariance of standard Transformers. Finally, a dual-dimensional graph data augmentation strategy is implemented to improve model resilience against structural noise. Extensive evaluations on four benchmark datasets (KDD Cup '99, NSL-KDD, UNSW-NB15, and CICIDS2017) demonstrate that Dual-Net outperforms state-of-the-art baselines in both detection accuracy and robustness. An Offset-Attention Boosted Double Deep Q-Network for Network Intrusion Detection Cuixia Li, Shuxiao Wang, Zhenghao Yang, Jizhe Zhao, and Shuyan Zhang (Zhengzhou University) Abstract Abstract Network Intrusion Detection (NID) serves as a critical defense against increasingly sophisticated cyber threats. While Deep Reinforcement Learning (DRL) has been widely adopted in NID research, existing methods often suffer from unsatisfactory detection accuracy. Due to the dominance of benign traffic in real-world networks, standard DRL agents tend to exhibit a bias towards the majority class, resulting in sub-optimal overall performance and poor sensitivity to fine-grained, minority attack families. To address this, an Offset-Attention Enhanced Double Deep Q-Network (OA-DDQN) is proposed. A novel Offset-Attention (OA) mechanism is introduced to replace standard global aggregation with a subtractive interaction. By calculating the offset between input features and their attention-weighted representations, this process acts as a learnable Laplacian operator, explicitly sharpening the distinction between benign traffic and subtle malicious activities. Besides, the Double Deep Q-Network (DDQN) is employed to provide precise value estimation, thereby mitigating the bias towards majority classes and ensuring robust decision-making. Extensive comparative experiments on the NSL-KDD benchmark validate the effectiveness of the proposed approach, where OA-DDQN outperforms state-of-the-art baselines and achieves the highest overall detection accuracy. Uncertainty-Aware Heterogeneous Graph Meta-Learning with Hybrid Attention for IoT Intrusion Detection Fengzhao Li, Wei Zhang, and Huiling Shi (Qilu University of Technology (Shandong Academy of Sciences), Shandong Computer Science Center (National Supercomputer Center in Jinan)) Abstract Abstract With the rapid expansion of the Internet of Things (IoT), heterogeneous and noisy traffic enlarges the attack surface, increasing the risks of malicious communication, denial-of-service attacks, and malware propagation. Existing methods often model traffic samples independently, failing to capture the interaction structure between traffic flows and network resources, and they generalize poorly to emerging or few-shot attacks under scarce labels and evolving attack patterns. To address these issues, this paper proposes an uncertainty-aware heterogeneous graph meta-learning framework, termed UA-MetaHGN, for few-shot IoT malicious traffic detection. Specifically, network flows, IPs, and ports are modeled as heterogeneous nodes with multiple relations; a hybrid-attention hierarchical representation learning scheme is designed to capture high-order dependencies and enhance discriminative representations; and task-adaptive initialization and dynamic task weighting are incorporated into meta-learning, leveraging loss discrepancies, heteroscedastic uncertainty, and Bayesian regularization to improve fast adaptation and robust generalization. Experiments on the ToN-IoT dataset show that UA-MetaHGN consistently outperforms state-of-the-art baselines under both few-shot (3/5/10-shot) and standard supervised settings. From ML to LLM: A Comparative Analysis of Context-Aware Zero-Day Attack Detection Moumita Shib (University of Dhaka), Shaily Kabir (University of Nottingham), and Mosarrat Jahan and Upama Kabir (University of Dhaka) Abstract Abstract Detection of zero‑day attacks, those exploiting previously unknown vulnerabilities, is essential for mitigating severe cybersecurity risks and preventing large‑scale damage. Conventional intrusion detection systems (IDSs) remain inadequate for such attack because they rely on known signatures or behavioural models that often lack stability. Although machine learning (ML) and deep learning (DL) approaches improve generalization, ML methods frequently fail to capture unseen behaviours, while DL models are sensitive to distribution shifts and limited in modelling long‑range context. Neural networks can learn discriminative traffic representations, yet under severe distribution shifts or rare attack patterns they may still become overconfident and miss low‑frequency zero‑day events. To address these limitations, hybrid models have been proposed that fuse traditional detection signals with contextual scoring learned from neural networks to enhance robustness. This paper mainly focuses on hybrid designs incorporating IDSs with deep neural networks (DNNs) or large language models (LLMs) and evaluates their zero‑day attack detection performance against state‑of‑the‑art baselines across three real‑world datasets—UNSW‑NB15, TON‑IoT, and CIC‑Collection—differing in label structure, class imbalance, attack diversity, and zero‑day evaluation settings. Results show that all hybrid models maintain comparatively stable performance on UNSW‑NB15 and TON‑IoT, whereas their detection capability deteriorates markedly on CIC‑Collection, reflecting a significant collapse under its more challenging distributional characteristics. Wednesday Virtual Room 5 IJCNN Paper Network Security and Fault Diagnosis II Session Chair: Gang Shi (Xinjiang University, School of Computer Science and Technology), Zhixin Meng (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) TrafficGraphMAE: Self-Supervised Masked Graph Autoencoders for Encrypted Traffic Classification Zhixin Meng, Mengyan Liu, Junzheng Shi, Gang Xiong, Zhen Li, and Gaopeng Gou (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Abstract Abstract Encrypted traffic classification plays an essential role in network management and security. However, the widespread adoption of payload encryption poses significant challenges for traffic analysis. Existing approaches either model traffic as sequential data, ignoring structural relationships among packets, or construct graph-based representations that model temporal and structural information separately, limiting their representational capacity. In this paper, we propose TrafficGraphMAE, a self-supervised masked graph autoencoder for encrypted traffic classification. We first introduce the Temporal-Traffic Interaction Graph (TTIG), a novel graph construction method that jointly models packet-level temporal dependencies and burst-level structural interactions within a unified framework. Based on TTIG, TrafficGraphMAE employs a masked graph autoencoder with multi-task learning objectives, including node feature reconstruction and structural relationship prediction, to learn expressive traffic representations in a self-supervised manner. Experiments on four public datasets demonstrate that TrafficGraphMAE outperforms multiple existing methods. Optimal Classifier Alignment via Neural Collapse for Incremental Encrypted Traffic Classification Qin Zeng, Yaqi Chen, Hao Zhang, Chaolong Hao, Fei Tian, and Dan Qu (Information Engineering University) Abstract Abstract The rapid evolution of internet technologies has intensified network traffic dynamics due to the emergence of novel encryption protocols, posing significant challenges to traffic classification. Incremental learning, which enables continuous adaptation to emerging tasks, has emerged as a promising approach to enhance the sustainability of encrypted traffic classification. However, existing methods fail to address the substantial feature representation disparities across incremental tasks, resulting in suboptimal model adaptability. This paper propose NCIL-ETC, a Neural Collapse-based Incremental Learning framework for Encrypted Traffic Classification. Our approach employs a pre-trained Mamba as the feature extraction backbone, leveraging its linear-complexity computational properties to significantly reduce resource overhead. Simultaneously, we introduce a pre-allocated ETF classifier that establishes an optimal classification structure covering observed classes. Through feature-classifier alignment constraints during incremental learning, our method enforces both new and historical class features to converge toward ETF vertices, thereby preserving globally optimal category relationships. Extensive experimental evaluations on two public benchmarks (CIC-IDS2017 and USTC-TFC) demonstrate that NCIL-ETC achieves state-of-the-art performance, surpassing baseline methods in both classification accuracy and incremental learning capability. EdgeFlow: Statistical Feature Modulation for Efficient Encrypted Traffic Classification Xiuwei Zhou (Sichuan Normal University, College Of Computer Science); Min Li (Sichuan Normal University, College Of Internet); and Yuanfang Lu, Chen Huang, and Yaolin Sun (Sichuan Normal University, College Of Computer Science) Abstract Abstract Encrypted traffic classification is fundamental to network management and security, yet presents an inherent challenge: traffic flows comprise heterogeneous feature modalities with divergent statistical properties. Raw packet bytes exhibit high-dimensional, high-entropy representations encoding application-level patterns, whereas metadata constitute compact dense vectors capturing network-level characteristics. Conventional fusion strategies such as simple concatenation may lead to suboptimal representation learning due to dimensional and distributional discrepancies. We present EdgeFlow, a hierarchical Transformer architecture designed to address this multimodal fusion challenge through Statistical Feature Modulation (Stats- Mod). Rather than directly combining heterogeneous features in a shared space, Stats-Mod employs global statistical context to dynamically calibrate the distribution of local byte-level representations via learned affine transformations. This distribution- level conditioning mechanism facilitates cross-modal information integration while maintaining computational efficiency. We further incorporate Token Compression via vertical pooling to reduce sequence length, thereby improving inference throughput. Experimental evaluation across four benchmark datasets demonstrates that EdgeFlow achieves competitive classification performance while reducing computational overhead—up to 5× higher inference throughput and 1.5× lower GPU memory consumption compared to models of comparable scale. FSR-RTDETR:Frequency-Selective Adaptive Sampling and Structural Re-parameterization for Efficient Traffic Sign Detection Peiyao Ma, Mengge Lu, Bingjun Liu, Gang Shi, and Qing Cui (Xinjiang University) Abstract Abstract In real-time traffic sign detection for complex road scenes, background texture interference and cross-scale fusion inconsistency make it difficult to achieve high accuracy and efficiency simultaneously. To tackle this issue, we propose FSR-RTDETR, a real-time detection network built upon RT-DETR. Specifically, we design Frequency-Selective Adaptive Sampling (FSAS) to decouple feature enhancement into frequency-domain suppression and spatial-domain alignment, which mitigates redundant texture responses and strengthens structural representations. We further introduce Learnable Gated Fusion (LGF) to adaptively balance deep semantic cues and shallow fine-grained details, thereby improving cross-scale representation consistency. In addition, we incorporate a Re-parameterized Multi-Branch Convolution (RMBConv) module, which is folded into an equivalent single-branch operator at deployment to avoid any additional inference overhead. Extensive experiments and ablation studies on large-scale datasets validate the effectiveness of the proposed approach. Wednesday Virtual Room 6 IJCNN Paper Object Detection and Recognition V Session Chair: Shuailiang Song (Dalian University of Technology), Wenzhu Yang (Hebei University; Machine Vision Engineering Research Center, Hebei University, Baoding, China) FA-DETR: Frequency-Aware for Aerial Small Object Detection Shuailiang Song, Fan Wang, Xiaopeng Hu, and Xinrong Wu (Dalian University of Technology) Abstract Abstract Aerial small object detection is a challenging task, primarily attributed to the small size of targets and lack of discriminative texture. Existing methods rely on the preservation of shallow feature maps and global context modeling to boost detection performance. However, these methods still suffer from inefficient shallow detail preservation and high-level semantic redundancy, thus compromising their effectiveness. To enhance detailed information representation and semantic modeling, we propose FA-DETR, an efficient Transformer-based framework with frequency awareness for aerial small object detection. Specifically, we introduce a Wavelet-Transform Mamba Enhancement Module (WTMEM), which applies a wavelet transform to decompose shallow features into low-frequency and high-frequency subbands and models them separately. WTMEM enhances high-frequency detail features while capturing low-frequency global context, thereby preserving key target details under complex background interference. To suppress semantic redundancy, we design a Large-Kernel Convolutional Self-Attention (LKCSA) module to enable more effective semantic interaction. We also propose a Cross-Stage Information Reactivation Module (CSIRM) to alleviate the degradation of small object detail information during feature fusion, thereby improving the quality of features fed into the detection heads. Extensive experiments on VisDrone and SIMD datasets demonstrate that FA-DETR achieves competitive performance (+2.2 AP on VisDrone and +1.9 AP on SIMD) while maintaining low computational complexity. A Frequency-Aware and Entropy-Guided Transformer for Real-Time Tiny Object Detection in Remote Sensing Dongdong An and Jiezhi Bao (Shanghai Normal University); Shuo Zhang (School of Electronic Information Engineering, Shanghai DianJi University, China); Wenbing Tang (Northwest A&.F University); and Qin Zhao (Shanghai Normal University) Abstract Abstract Real-time aerial object detection remains challenging due to tiny targets and cluttered backgrounds. In remotesensing imagery, small objects often occupy only a few pixels, and their weak high-frequency cues are easily overwhelmed by dominant low-frequency background components during feature extraction and fusion. Moreover, common feature-pyramid fusion relies on upsampling, which can introduce spatial misalignment and degrade geometric localization. In densely populated scenes, classification confidence and box regression quality may jointly deteriorate, causing valid targets to be suppressed without an explicit spatial prior. We propose FRE-Det (FrequencyResidual and Entropy Detector), a task-oriented architecture that integrates time–frequency analysis and information-theoretic guidance. First, we design a Dual Frequency Enhanced Hybrid Encoder that performs FFT-based frequency filtering to enhance weak target responses against background noise. Second, we introduce a Wavelet Reconstruction Block (WRB) to replace naive upsampling with wavelet-guided reconstruction, improving spatial consistency during multi-scale interaction. Third, to mitigate missed detections in dense regions, we propose Local Spatial Entropy-guided Uncertainty Minimization (LSE-UM), which uses local information density as a spatial prior to recover true targets based on signal presence. Extensive experiments on VisDrone2019 demonstrate state-of-the-art performance, achieving AP50 = 49.3% and APS = 19.9% while running at 87 FPS. WH-Mamba: Wavelet-Enhanced State Space Model with Dynamic Hypergraph Fusion for Robust Traffic Object Detection huanrong 唐, jiaxiong lu, and jianquan ouyang (Xiangtan University) Abstract Abstract Real-time object detection in autonomous driving faces persistent challenges in balancing computational efficiency with perceptual robustness, particularly when confronting minuscule targets and adverse weather conditions. To address the limitations of existing CNN and Transformer architectures—specifically feature attenuation and quadratic complexity—we propose WH-Mamba, a novel framework synergizing State Space Models (SSMs) with advanced signal processing and graph theory. At the backbone level, we introduce the Wavelet-Cross-Mamba Block(WCMB), which integrates a Quad-Directional Scanning (QDS) strategy to align 1D selective scans with the radial topology of traffic scenes, while leveraging Discrete Wavelet Transforms to preserve high-frequency edge details often submerged in deep layers. To further mitigate environmental noise, a Frequency-Spatial Attention (FSA) module employs Fourier analysis for spectral denoising and spatial saliency sharpening. Finally, a Dynamic Hypergraph Fusion (DHF) module is designed to model complex "multi-to-multi" semantic correlations among diverse traffic agents, transcending the receptive field limits of local convolutions. Extensive evaluations on MS COCO and BDD100K demonstrate that WH-Mamba achieves 45.7% AP on COCO and 43.2% AP on BDD100K. Notably, it outperforms recent models such as YOLOv13 and the Transformer-based DEIM, showing a significant advantage in small object detection with an 30.1% AP for small objects. An Aerial Object Detection Model Based on Dynamic Feature Cropping and Multi-Level Guiding Lanbin Liang (Hebei University School of Cyber Security and Computer) and Wenzhu Yang (Hebei University School of Cyber Security and Computer, Hebei University Machine Version Engineering Research Center) Abstract Abstract To address the challenges in aerial object detection, including complex backgrounds, large number of small targets, and the redundant computations incurred by existing detection models during feature extraction, this paper proposes a novel aerial detection model based on Dynamic Feature Cropping and Multi-Level guiding(DFCMNet). First, a Lightweight Feature Extraction(LFE) module based on dynamic information filtering is designed to enable the model to dynamically distinguish between salient and redundant features during training, thereby performing feature extraction operations with different levels of computational complexity to reduce model redundancy. Second, this paper proposes an improved detection layer, incorporating a shallow-level feature interaction module into the neck’s upsampling network to enhance the model’s capacity to extract detailed information. By using the C2 and C3 layers, the model is specifically tailored to detect small scale objects in aerial images. Third, an efficient backbone network is constructed by integrating the proposed LFE module into the feature learning layers, yielding a compact feature representation while emphasizing the acquisition of shallow-layer information. Furthermore, this study performs multi-scale modeling in the feature space and incorporates deformable convolutions together with a full-dimension, parameter-free attention module, enabling the model to learn complementary representations of multi-scale object information and contextual information. An optimized localization loss is subsequently adopted to reduce computational cost, accelerate convergence, and maintain high detection accuracy. On the VisDrone-DET2019 and AI-TOD datasets, the proposed model achieves mAP50 improvements of 7.6% and 8.9% over the baseline, respectively, while maintaining a parameter count of only 1.4M, demonstrates that our method achieves a favorable balance between detection accuracy and model size, exhibiting a clear advantage over existing aerial object detection models. Wednesday Virtual Room 7 IJCNN Paper Object Detection and Recognition VI Session Chair: Keji Mao (Zhejiang University of Technology), Jiong Yu (Xinjiang University, Department of Computer Science and Technology) PGNet: Prototype-Guided Camouflaged Object Detection Junyong Sun and Yuan Sun (Zhejiang University of Technology); Yongbiao Zhao (Zhejiang University of Technology, Zhijiang College of Zhejiang University of Technology); and Zhihu Zhou, Zhuchenghao Wang, and Keji Mao (Zhejiang University of Technology) Abstract Abstract Camouflaged object detection focuses on detecting objects that blend into complex backgrounds. Compared to general detection tasks, camouflaged targets have stronger similarity to their surrounding background in terms of texture and color, and their boundaries are often blurred, which significantly increases the difficulty of accurate segmentation.To address these issues, we propose a novel network framework: Prototype Guided Network (PGNet). It extracts background information through prototype learning, highlights the foreground, and uses an attention mechanism to make the model focus on the target region of the image. Under the guidance of edge information, it effectively segments the target. Specifically, it consists of three modules: 1) Background Prototype Module (BPM), which can extract a learnable background prototype from the image, suppress camouflaged region features similar to the background, and help distinguish between foreground and background. 2) Prototype Attention Module (PAM), which employs the prototype as a global context and uses overlock attention to make the model pay more attention to the target region. 3) Edge Guided Module (EGM), which predicts the target edges and guides the segmentation of the image, solving the problem of blurred edges. Extensive experiments show that PGNet outperforms the current COD method in terms of segmentation accuracy. HPT-SAM2: Adapting Segment Anything Model 2 for Camouflaged Object Detection via Hierarchical Prototypes and Texture Modulation Yu Chen, Chengliang Wang, and Xing Wu (Chongqing University) and Peng Wang and Hongqian Wang (The First Affiliated Hospital of Army Medical University) Abstract Abstract Camouflaged Object Detection (COD) aims to identify and segment objects that seamlessly blend into their surroundings. While the Segment Anything Model (SAM) excels in generic segmentation, its performance in COD is impeded by the significant domain gap between its general-purpose pre-training data and the ambiguous camouflaged targets. Existing SAM adaptations, categorized into explicit prompting and architecture design, attempt to bridge this gap but still suffer from two critical limitations: they often fail to maintain intra-object consistency, leading to fragmented segmentation, and struggle to delineate low-contrast boundaries. To address these challenges, we propose HPT-SAM2. Specifically, we design a Hierarchical Prototype Refinement (HPR) module that adopts a global-to-local strategy, utilizing coarse and fine-grained prototypes to consolidate scattered features into coherent entities. Additionally, a Texture-Semantic Modulation (TSM) module is developed to sharpen object contours by synergizing semantic feedback and texture priors via a gated mechanism. Experimental results show that HPT-SAM2 achieves a 4.2% improvement in weighted F-measure on COD10K and a 16.7% reduction in Mean Absolute Error on CAMO compared to the baseline SAM2-UNet. Hierarchical Task-Decoupled Representation Learning for Efficient UAV Object Detection Yaqi Liu (School of Cyber Security and Computer, Hebei University, Baoding, China) and Wenzhu Yang (School of Cyber Security and Computer, Hebei University, Baoding, China; Machine Vision Engineering Research Center, Hebei University, Baoding, China) Abstract Abstract Object detection in UAV scenarios remains a significant challenge due to small objects, extreme scale variations, and complex backgrounds. Existing research often adopts uniform architectural designs, overlooking the distinct representational bottlenecks across hierarchical layers. Deep layers handle global semantics but suffer from severe feature homogenization and semantic degradation, while shallow layers preserve fine-grained details but are constrained by limited receptive fields and background interference. Moreover, complex designs increase computational overhead, hindering real-time deployment. To address these issues, we propose a lightweight Hierarchical Task-Decoupled representation learning network (HTDNet), consisting of three core modules: the Diverse Representation Reconstruction (DRR) module reconstructs structured semantics through feature decomposition and heterogeneous recomposition, mitigating feature homogenization in deep layers; the Cross-Layer Prior Modulation (CLPM) module leverages both semantic and spatial priors to guide targeted feature recalibration for small objects, suppressing background interference in shallow layers; and the Saliency-Guided Purification (SGP) module eliminates representational conflicts during multi-scale fusion through nonlinear modulation, enhancing small object saliency. Experimental results on VisDrone show that HTDNet outperforms the baseline by 6.3% mAP50 with only 45% of the parameters, achieving an optimal accuracy-efficiency trade-off, offering distinct advantages for real-time edge deployment. The generalization of HTDNet is further validated across AI-TOD and TinyPerson datasets. SIG-Net: Object Detection Networks Focusing on Minor Defects in Complex Contexts Xin Wang, Xue Li, Ziyang Li, Jiong Yu, and Wenjing Li (xinjiang university) Abstract Abstract Automated detection of steel surface defects is vital for ensuring product quality and safety. In complex textured backgrounds, defects are typically minute and low in contrast, making them easily obscured or confused with surrounding patterns and leading to frequent missed and false detections. To address this challenge, we propose SIG-Net, a lightweight detection network tailored for small defects in complex contexts. The framework integrates three core modules: the Small Object-Aware Refined Feature Pyramid (SOAR-FPN), which injects high-resolution details into P3 via SPDConv and enhances representation through CSP-OmniKernel fusion; the Inter-Layer Sparse Guidance Module (ISGM), which applies cross-scale sparse attention where high-level semantics guide low-level features to suppress noise and highlight defects; and the Global Edge Information Transfer (GEIT), which generates multi-scale edge features from shallow layers and propagates them across the network to improve boundary perception. Experimental results on the NEU-DET and GC10-DET datasets demonstrate that SIG-Net achieves mAP50 values of 82.2% and 83.4%, respectively, while maintaining a lightweight architecture, significantly outperforming existing methods and effectively addressing the challenge of small defect detection under complex backgrounds. Wednesday Virtual Room 8 IJCNN Paper Object Detection and Recognition VII Session Chair: xiao shikang (Hunan University of Technology), Sheng-Chun Yang (Northeast Electric Power University) IR-DETR-Based Algorithm for Detecting Surface Defects on Porcelain Cups Xiao mingyang, Deng xiaojun, and Xiao shikang (Hunan University of Technology) Abstract Abstract To address the limitations of manual inspection, including missed and false detections in porcelain-cup manufacturing, we propose IR-DETR, an improved RT-DETR framework for detecting small-scale surface defects under low computational and parameter budgets. In the backbone, we introduce a wavelet convolution (WTConv) module that enlarges the effective receptive field, thereby improving robustness to defect patterns and promoting shape-oriented rather than texture-oriented responses. In the neck, we incorporate a hierarchical scale-sequence feature fusion (SSFF) module to uniformly fuse multi-scale sequence features, enhancing representation capacity and the detection of diverse defect types. To offset the overhead introduced by SSFF, we further integrate a dynamic upsampling (Dysample) module that reformulates the upsampling process from a point-sampling perspective, preserving feature continuity while reducing complexity and improving sensitivity to tiny defects. Experiments on a self-constructed dataset show competitive performance: compared with representative defect-detection methods, IR-DETR achieves mAP@0.5 = 54.3% with a lightweight design of 14.4M parameters and 48.4 GFLOPs. Overall, IR-DETR maintains high detection accuracy at substantially reduced complexity, achieving a favorable accuracy–efficiency trade-off. WMD: A Lightweight YOLO11 Extension for Efficient Complexity Detection in Mechanical Part Inspection Xiaofeng Yue, Jiaqi Yang, Zeyu Qi, Ruochen Cao, Rui Cao, and Xin Wen (Taiyuan University of Technology) Abstract Abstract A key feature of Industry 4.0 is intelligent manufacturing, which calls for the deep integration of advanced technologies in industrial production. In current manufacturing practices, the complexity detection of mechanical parts still heavily relies on manual inspection, leading to high labor costs and low efficiency. To address this challenge, this study proposes an efficient improved object detection algorithm based on YOLO11, incorporating Wavelet Convolution (WTConv), Manhattan Self-Attention (MaSA), and Distance Intersection over Union (DIoU), collectively referred to as the WMD model. The WMD model integrates WTConv and MaSA into the backbone and neck of the YOLO11 architecture to enhance feature extraction capabilities, while DIoU replaces the original bounding box regression loss function to improve localization accuracy. Experimental results demonstrate that the proposed WMD model achieves significant improvements over the baseline model: precision increases by 3.6%, mAP@0.5 improves by 1.5%, model parameters decrease by 0.042×10^6, and computational cost reduces by 0.1 GFLOPs. Overall, the WMD model effectively balances detection performance and computational efficiency, making it a promising lightweight solution for automated mechanical part complexity detection in smart manufacturing. PR-DETR for Efficient and Lightweight Object Detection via Weight-sharing Softmask Pruning and Spatial Reduction Minjie Zhang, Hairui Ye, Yunfei Tong, and Zhe Wang (East China University of Science and Technology) Abstract Abstract DETR achieves competitive detection accuracy but suffers from high computational cost and large parameter size, limiting deployment on mobile and edge devices. In this paper, we propose PR-DETR, a lightweight detection transformer that combines weight-sharing softmask pruning with spatial reduction token compression to improve efficiency while preserving accuracy. The proposed weight-sharing softmask pruning strategy enables differentiable structural pruning across MLP, attention, and convolution blocks with flexible control of the pruning budget. In addition, the spatial reduction scheme compresses key and value tokens in attention to further reduce computation. Experiments on VOC0712 and MiniCOCO2017 demonstrate that PR-DETR achieves substantial reductions in parameters and FLOPs while maintaining comparable mAP to existing lightweight DETR models. MLD-DETR: A Method for Detecting Small Traffic Signs in Complex Scenarios shengchun yang, lixun zhou, huixiang wang, ruiqi li, chiyuan wang, and yuhuan wang (Northeast Electric Power University) Abstract Abstract Traffic sign detection technology is a crucial component of intelligent driving systems. To address challenges such as small pixel areas, susceptibility to background interference, and low detection accuracy during traffic sign detection, the MLD-DETR traffic sign detection model is proposed. First, the MSGC module was introduced to replace the feature extraction component in the backbone network. By employing a dual-path parallel convolution strategy, it simultaneously aggregates local detail information and global contextual information, thereby enhancing the network's ability to represent features of objects at different scales. Subsequently, to reduce background interference, the LASS model was proposed. Finally, the DACFPN module was introduced to fuse features at different scales, further improving detection accuracy. Experimental results demonstrate that the MLD-DETR model achieves outstanding performance on the TT100K traffic sign dataset. Compared to the baseline model, it achieves an mAP50 of 83.5%, representing a 7.9% improvement, and an mAP50-95 of 61.4%, showing a 7.3% increase. The number of parameters decreases from 19.9M to 18.2M, a reduction of 1.7M, outperforming other state-of-the-art methods. Wednesday Virtual Room 1 IJCNN Paper Object Detection and Recognition VIII Session Chair: Xingpeng Zhang (Southwest Petroleum University), Yunpeng Guo (University of Jinan) Efficient-CANet: Rethinking Cross-Temporal Interaction and Multi-Scale Geometry for Ultra-Lightweight Change Detection Xingpeng Zhang, Dian Qi, Yuru Li, and Xingpan Hu (Southwest Petroleum University); Qiuli Wang (The First Affiliated Hospital of Army Medical University); and Yang Yu (Southwest Petroleum University) Abstract Abstract Building change detection in complex urban environments faces a dilemma: distinguishing genuine changes from pseudo-changes (e.g., illumination, seasonal variations) requires robust global context modeling, yet deployment on resource-constrained edge devices demands extreme efficiency. Existing lightweight methods often compromise discriminative capability or multi-scale adaptability to reduce parameters. To address this trade-off, we propose CANet, an ultra-lightweight network (0.73M parameters) designed for efficient and accurate change detection. CANet introduces two core innovations: (1) A Cross-Temporal Change-Aware Module (CTCAM), which incorporates a Factorized Similarity Kernel (FSK). By factorizing the global correlation matrix, FSK explicitly models inter-phase dependencies with linear complexity $O(N)$, effectively suppressing pseudo-change noise without the computational burden of standard self-attention. (2) An Omni-Dimensional Spatial Pyramid (ODSP), which synergizes dynamic kernel weighting with multi-scale dilated convolutions. This design decouples geometric deformation from scale variations, enabling adaptive feature extraction for buildings of diverse shapes and sizes. Extensive experiments on the LEVIR-CD and S2Looking datasets demonstrate that CANet achieves SOTA accuracy among lightweight models, surpassing the heavier ChangerEx while using 16× fewer parameters. FSFENet: An Efficient Frequency-Spatial Feature Extraction Network for Remote Sensing Change Detection Hao Zhang, Tao Xu, Ruke Shang, and Jiayin Zhang (University of Jinan); Yan Zhang (Zouping Natural Resources and Planning Bureau); and Yuqiang Wang (Disaster Reduction Center of Shandong Province) Abstract Abstract Erroneous detection of pseudo-change targets is one of the main reasons for the performance degradation of change detection algorithms. Most research efforts improve the ability to identify pseudo-change targets by constructing complex network structures, which makes it difficult to balance detection accuracy and lightweight performance. To address this, this paper proposes a Frequency-Spatial Feature Extraction Network (FSFENet). First, a Frequency Filtering Module (FFM) is constructed, which uses Fast Fourier Transform and attention mechanisms for selective filtering, mitigating the impact of domain shift between dual-temporal image pairs and enhancing the effective representation of change targets. Second, a lightweight Dual-temporal Change Feature Extractor (DT-CFE) is designed to adaptively extract change features from dual-temporal features through channel-wise feature interaction, reducing the influence of pseudo changes. Finally, a Multi-level Feature Aggregation Module (MFAM) is designed to enhance the completeness of change regions by fusing low-level and high-level features. Experimental results on three public datasets—WHU-CD, SYSU-CD, and DSIFN-CD—show that with only 0.45M parameters and 2.34G FLOPs, FSFENet achieves F1 scores of 92.75%, 81.99%, and 67.77%, respectively, demonstrating the effectiveness of the proposed method. The code is available at https://github.com/zhang2209/FSFENet. DFINE-BRA: A Lightweight Algorithm for Growth Stage Detection in Brassica Crops Zhimin Bai, Shiyi Liu, and Huijun Yang (Northwest A&F University) Abstract Abstract To address challenges in detecting Brassica growth stages in complex environments, this paper proposes DFINE-BRA, a lightweight detection framework. This study constructed the BRA dataset, encompassing 19 varieties and 4 growth stages, and introduced SaMam, an SSM-based augmentation method, to enhance scene diversity. The architecture integrates an MI-Block for morphological feature extraction, alongside a CME-Encoder and CSAM for multi-scale feature aggregation. To optimize efficiency, knowledge distillation and TensorRT acceleration were employed for deployment on the Jetson Orin NX platform. Experimental results show that DFINE-BRA achieves an mAP50 of 83.1% on the BRA dataset, representing a 4.6% improvement over the baseline. Notably, the model reduces parameters and GFLOPs to 40% and 32% of the original values, respectively, while achieving a real-time inference speed of 36 FPS. This study offers an efficient solution for automated Brassica monitoring on edge devices. The code is available at https://github.com/Jackbai360/DFINE-BRA. MSGeoGrasp: A 6-DoF Grasp Detection Based on Multi-Scale Feature and Geometric Enhancement Yunpeng Guo (University of Jinan); Jing Zhang (University of Jinan, Yantai Institute of Science and Technology); Jie Su (University of Jinan); Junzheng Yang (Yantai Institute of Science and Technology); and Tianchi Zhang (Chongqing Jiaotong University) Abstract Abstract In recent years, with the increasing complexity of robotic manipulation tasks, 6-DoF grasp detection has become a key technology in robotic grasping. EconomicGrasp introduces an efficient economic supervision strategy, significantly reducing computational cost while improving grasp detection accuracy. However, when dealing with large-scale variations and severe occlusions, the method still suffers from limitations in fixed-scale neighborhood modeling and geometric representation. To address these issues, we propose MSGeoGrasp, a 6-DoF grasp detection method based on multi-scale feature modeling and geometric enhancement. We introduce multi-scale cylindrical local feature modeling to enhance scale adaptability, and employ geometric feature enhancement to improve geometric representation. Experimental results on GraspNet-1Billion demonstrate that MSGeoGrasp achieves superior performance while introducing only a negligible increase in model parameters. Wednesday Virtual Room 2 IJCNN Paper Object Tracking and Video Understanding I Session Chair: Hongyu Song (Inner Mongolia University of Technology), Zesen Cai (University of Electronic Science and Technology of China) Multi-Scale Skeleton Action Recognition Based on Self-Supervision Zesen Cai, Lisi Mo, Chenxi Li, and Ruiting Dai (University of Electronic Science and Technology of China) Abstract Abstract Self-supervised learning (SSL) provides a scalable and annotation-efficient paradigm for action recognition. However, existing approaches often struggle to bridge the semantic gap between fragmented local motions and global action contexts, as they predominantly rely on instance-level discrimination while overlooking the multi-level hierarchical nature and long-range spatiotemporal dependencies inherent in human dynamics. To address this, we propose Multi-Scale Self-Distilled Learning (MS-Distill), a self-supervised framework that employs student-teacher networks with a multi-view consistency objective. Unlike traditional approaches, MS-Distill eliminates the reliance on negative samples by aligning representations across various temporal resolutions and spatial granularities. Furthermore, we introduce a Global-Local Spatiotemporal Encoder that models interactions among arbitrary spatiotemporal points, yielding highly discriminative unified embeddings that capture both fine-grained motion dynamics and holistic, long-range action patterns. Extensive experiments on the NTU-RGB+D 60 and 120 datasets demonstrate the superiority of MS-Distill. Notably, it achieves a significant 3.8% gain in KNN accuracy on the NTU-120 xset protocol with minimal pre-training overhead, while exhibiting strong transferability and scalability to downstream tasks. Gated Contrastive Alignment Variational Autoencoder for Zero-Shot Skeleton Action Recognition Jiang Liu, Ming Liu, and Rong Liu (Central China Normal University) Abstract Abstract Skeleton-based action recognition has attracted increasing attention due to its efficiency and robustness to visual variations, and plays an important role in applications such as human–computer interaction and healthcare. To alleviate the reliance on labeled data for rare or hazardous actions, Zero- Shot Skeleton Action Recognition (ZSSAR) exploits semantic knowledge to recognize unseen classes. However, skeleton sequences are inherently noisy and ambiguous, especially for finegrained actions with similar motion patterns. Existing crossmodal methods adopt fixed latent space decompositions and perform coarse-grained skeleton–text alignment, often enforcing hard semantic–style disentanglement. Such rigid designs may discard useful correlations, amplify noise, and limit generalization when semantic and style factors are not strictly separable. To address these challenges, we propose Gated Contrastive Alignment Variational Autoencoder (GCA-VAE), a novel framework for robust ZSSAR. GCA-VAE introduces a learnable gating mechanism that softly modulates latent dimensions, enabling adaptive separation of semantic-relevant and style-related factors without relying on predefined partitions. In addition, we design a fine-grained contrastive alignment strategy to calibrate skeleton and text representations, which suppresses skeleton noise while preserving discriminative semantic cues. Extensive experiments on the NTU RGB+D 60 and NTU RGB+D 120 datasets demonstrate that GCA-VAE consistently outperforms state-of-the-art methods on both ZSSAR and generalized ZSSAR benchmarks. Prototype-Guided Meta-Learning for Skeleton-Based One-Shot Action Recognition Jing Zhou, Cuiwei Liu, and Huaijun Qiu (Shenyang Aerospace University) Abstract Abstract This paper addresses Skeleton-based One-shot Action Recognition (SOAR), aiming to recognize novel actions from only a single reference sample. The core challenge lies in learning a generalizable feature space by leveraging data-rich base actions. Meta-learning has emerged as the prevalent paradigm by constructing numerous meta-tasks over base actions to simulate one-shot recognition scenarios. However, the classifiers built from one-shot support samples often fail to capture the complete action distribution in meta-tasks, leading to unstable optimization and limited generalization. To overcome these issues, we propose a Prototype-Guided SOAR (PGSOAR) framework featuring a teacher-student structure. The teacher model is a pre-trained base-action classifier, and the student model learns to solve the SOAR task. During episode training, the teacher simultaneously extracts action prototypes for each meta-task, using them as support data to construct a robust target classifier. The student model, implemented via a Multi-stream Adaptive Spatio-Temporal Graph Convolutional Network(MAST-GCN), is then supervised by the target classifier through knowledge distillation, thereby achieving more reliable and stable optimization. At test time, an Adaptive Prototype Refinement(APR) strategy alleviates single reference bias by rectifying novel action prototypes with a minimal number of unlabeled query samples. Extensive experiments on two large-scale benchmarks demonstrate that the proposed PGSOAR framework consistently outperforms previous methods, validating its strong generaliza- tion and robustness in one-shot action recognition. vHeat-HMCD: Hierarchical Multi-scale Contrastive Discrimination for Electron Microscopy Pollen Image Classification Hongyu Song, Shi Bao, Yatu Ji, and Min Lu (Inner Mongolia University of Technology) Abstract Abstract Pollen grain identification is pivotal in diverse disciplines ranging from palynology to forensic science, yet traditional manual examination remains time-consuming and labor-intensive. Although automated recognition has emerged as a solution, it remains challenged by subtle inter-class morphological similarities and significant background interference in electron microscopy images. While recent physics-inspired backbones like vHeat offer efficient global modeling, they often struggle to distinguish discriminative foreground details from complex background noise. To address this, we propose integrating a Hierarchical Multi-scale Contrastive Discrimination (HMCD) framework into the vHeat backbone to synergize efficient global modeling with robust noise filtering. Specifically, we design a Hierarchical Multi-scale Background Suppression (HMBS) module equipped with an Uncertainty-Aware Adaptive Thresholding (UAAT) mechanism, which dynamically adjusts suppression intensity based on predictive uncertainty to preserve fine-grained edge details. Furthermore, we introduce a Contrastive Feature Decoupling (CFD) module to enhance feature discriminability. Within CFD, a Dynamic Foreground Discovery (DFD) mechanism adaptively localizes foreground regions via learnable differentiable thresholding, while a Cross-Scale Contrastive Consistency (CSCC) strategy enforces semantic alignment across hierarchical stages. Extensive experiments on a large-scale self-constructed dataset and three public benchmarks demonstrate that our proposed framework significantly outperforms state-of-the-art methods, achieving superior accuracy and robustness in electron microscopy pollen image classification. Wednesday Virtual Room 3 IJCNN Paper Object Tracking and Video Understanding II Session Chair: shiao mi (Tianjin University of Science and Technology), Xun Ruiqi (Innovation Academy for Microsatellites of Chinese Academy of Sciences, University of Chinese Academy of Sciences) A 3D Hand Pose Estimation Framework for Crewed Space Mission under Limited Sensing Conditions Ruiqi Xun (Innovation Academy for Microsatellites of Chinese Academy of Sciences, University of Chinese Academy of Sciences); Xingyu Zhu (School of Artificial Intelligence, Jilin University); and Haoshuai Mu, Ziang Qu, and Liang Chang (Innovation Academy for Microsatellites of Chinese Academy of Sciences, University of Chinese Academy of Sciences) Abstract Abstract During on-orbit missions, astronauts interact with cargo packages, control panels, and experimental payloads primarily through hand motions. 3D hand pose estimation is essential for automated operation monitoring and human–machine interaction. However, spacecraft impose strict constraints on sensing systems, where depth sensors are impractical, and perception must rely on limited RGB cameras. This sensing limitation poses a fundamental challenge to the accurate estimation of 3D hand pose. To address this challenge, we propose SpaceRGB-HandPose Net(SHPN), an RGB-based framework tailored to the sensing constraints of crewed spacecraft. SHPN enhances geometric perception from RGB images by injecting a pseudo-depth modality. It explicitly addresses the mismatch between pseudo-depth and metric depth via DTOR and ADRM alignment modules. Furthermore, SHPN introduces a cross-modal fusion network to integrate appearance and relative depth information for robust pose estimation. Experiments on the DexYCB dataset show that SHPN significantly outperforms RGB-based baselines and achieves performance comparable to RGB-D methods using only RGB input, demonstrating its applicability to realistic crewed spaceflight scenarios. DualAlign: A Heterogeneous Distillation Framework for Extremely Lightweight Pose Estimation Liyang Zheng and Miao Zhang (Harbin Institute of Technology, Shenzhen) and Siyuan Fang and Xing Chen (HeyiSpace) Abstract Abstract Deploying Human Pose Estimation (HPE) on microcontroller units (MCUs) often relies on hardware-aware architecture selection (e.g., XiNet-Pose) to meet strict resource constraints, which inevitably sacrifices performance by excluding heavy components like Spatial Pyramid Pooling (SPP) and multi-scale heads. While knowledge distillation (KD) offers a remedy, it falters in such highly heterogeneous scenarios due to Optimization Misalignment from conflicting gradient directions and Structural Misalignment between multi-head teachers and single-head students. To bridge these gaps, we propose DualAlign, a unified framework for extremely lightweight pose estimation. DualAlign incorporates two complementary alignment mechanisms: Gradient-Projected Alignment (GPA) dynamically gates distillation weights based on gradient directional consistency to automatically shield the student from optimization conflicts, while Transient Structural Alignment (TSA) employs auxiliary heads governed by a "Lead Head Principle" to seamlessly inject multi-scale knowledge into the shared backbone during training. Extensive experiments on MS COCO Keypoints demonstrate that DualAlign significantly recovers accuracy without incurring any inference overhead, consistently outperforming other representative distillation paradigms adapted for this constrained setting. MambaFlow: A Mamba-Centric Architecture for End-to-End Optical Flow Estimation Juntian Du, Yuan Sun, Zhihu Zhou, Pinyi Chen, Runzhe Zhang, and Keji Mao (Zhejiang university of technology) Abstract Abstract Recently, the Mamba architecture has demonstrated significant successes in various computer vision tasks. However, its application to optical flow estimation remains unexplored. In this paper, we introduce MambaFlow, a novel framework centered around the Mamba design tailored specifically for the end-to-end optical flow estimation. It comprises two key components: (1) PolyMamba, which enhances feature representation through a dual-Mamba architecture, combining a Self-Mamba module for intra-token modeling with a Cross-Mamba module for inter-modality interaction; and (2) PulseMamba, which leverages an Attention Guidance Aggregator (AGA) to adaptively integrate features with dynamically learned weights in contrast to naive concatenation, and then leverages Mamba’s recurrent mechanism to iteratively perform autoregressive flow decoding. Extensive experiments demonstrate that MambaFlow achieves remarkable results comparable to mainstream methods on benchmark datasets. Compared to SEA-RAFT, MambaFlow attains higher accuracy on the Sintel benchmark, demonstrating stronger potential for real-time deployment in practical systems. HCFS-GNN: Semantic-Aligned Differentiable Heterogeneous Graph Anomaly Detection for Industrial Systems Shiao Mi (Tianjin University of Science and Technology); Xiangyu Wu (Beijing KeDong Electric Power Control System Co., Ltd.); Suxiang Zhang (state grid information & Telecommunication center (Big Data center)); and Yiying Zhang and Hongxin Zhao (Tianjin University of Science and Technology) Abstract Abstract Abstract:Stealthy semantic decoupling attacks targeting control logic pose a severe threat to the integrity of Industrial Cyber-Physical Systems (ICPS). Existing homogeneous graph-based anomaly detection methods often struggle to capture the semantic conflicts between physical processes and control commands due to functional role confusion between sensors and actuators. To address this challenge, this paper proposes a Heterogeneous Control-Flow Sensitive Graph Neural Network (HCFS-GNN). First, we construct a heterogeneous bipartite control graph to explicitly decouple sensor observations from actuator actions at the topological level. Second, we introduce Symbolic Transfer Entropy (STE) combined with Gumbel-Softmax reparameterization to achieve lightweight real-time causal discovery and differentiable end-to-end structure optimization. Furthermore, a novel Control-Flow Sensitivity (CFS) metric is designed as a semantic high-pass filter to significantly amplify logical violation signals in critical control nodes. Extensive experiments on the SWaT and WADI benchmarks demonstrate that HCFS-GNN achieves F1-scores of 0.952 and 0.749, respectively. Notably, the proposed method achieves complete coverage of attack scenarios while maintaining high precision, and exhibits superior robustness under conditions of high noise and limited training data. Wednesday Virtual Room 4 IJCNN Paper Open-Set and Out-of-Distribution Recognition Session Chair: Rahul Biswas (Indian Institute of Technology Hyderabad, IIT Hyderabad), Cong Hu (Jiangnan University) FedGCR : Federated Open Set Recognition Using Gaussian Noises in Client's Representation Ankita Das, Mrinmay Sen, Rahul Biswas, and C. Krishna Mohan (Indian Institute of Technology Hyderabad) Abstract Abstract Open Set Recognition (OSR) aims to correctly classify known classes while detecting unknown samples, relaxing the closed-set assumption of traditional classification methods, which presume that both training and testing data contain samples from the same set of classes. However, most existing OSR methods operate in centralized settings, which often require sharing sensitive data. Federated Open Set Recognition (FedOSR) addresses this by enabling collaborative learning without exchanging raw data; however, current FedOSR methods still rely on sharing client-level statistics (e.g., mean and variance), which can leak private information. To address this issue, we propose FedGCR, a privacy-preserving FedOSR framework that trains a binary classifier for unknown class detection using noise-added representations of local data, alongside the main model trained for known class prediction. Unlike existing methods, FedGCR avoids sharing any data statistics and transmits only the main known-class classifier and the binary unknown-detection model to the server. Experiments on benchmark image classification datasets show that FedGCR achieves superior unknown-class detection while maintaining strong performance on known classes and preserving client data privacy. CALIBER: Benign-Calibrated Dual-Score Rejection for Label-Scarce Open-World Sequence Classification Jiayu Li and Xiangxue Li (East China Normal University) Abstract Abstract Open-world sequence classification under label scarcity requires not only accurate recognition of known categories but also reliable rejection of previously unseen patterns under an explicit false-alarm budget. This challenge arises prominently in in-vehicle CAN message streams, where attack labels are limited and new attack behaviors may emerge after deployment, causing closed-set detectors to misclassify unknowns as known types. We propose CALIBER, a label-efficient framework that couples lightweight temporal representation learning with calibrated unknown rejection. CALIBER converts raw bus frames into fixed-length token sequences by fusing payload content, identifier cues, and short-horizon timing and rate signals, enabling a compact sequence encoder to learn discriminative representations from limited supervision. At inference time, CALIBER performs calibrated rejection via a dual-score scheme that combines an energy score from classifier logits with a representation-space Mahalanobis distance. Rejection thresholds are determined using benign-only quantile calibration, with optional mode-aware indexing to improve operating-point stability. Experiments on two public in-vehicle datasets with leave-one-attack-out zero-day evaluation show near-saturated recognition on known traffic while consistently rejecting unseen attacks at low benign false-alarm rates. UDLVM: Unknown Object Detection with Large Vision Models Shulin Chen, Wangshu Yao, Shuxin Zhang, and Xiaofeng Jiang (Soochow University) Abstract Abstract Open World Object Detection (OWOD) aims to detect objects of known and unknown categories in the real open scenarios, and can incrementally learn novel category knowledge under the guidance of humans simultaneously. However, due to the lack of unknown class labels, existing OWOD methods suffer from two main issues: semantic bias towards known objects and low recall of unknown objects. To address these issues, we propose a new and effective detection framework called "Unknown Object Detection with Large Vision Models (UDLVM)", which leverages two prevalent large vision models for the OWOD task, namely Visual State Space Model (VMamba) and Segement Anything Model (SAM). For one thing, a new VSSFPN module is designed, which incorporates the core of VMamba in the Feature Pyramid Network (FPN) for the better feature extraction to alleviate the semantic bias towards known objects. For another thing, a new DUSC method is developed, which utilizes the masks generated by SAM to assist in producing high-quality pseudo labels to improve the recall of unknown objects. Extensive experiments on public benchmark datasets show that our UDLVM framework can effectively alleviate the semantic bias issue and improve the unknown recall compared with existing methods. G3C: Graph consistency based on contrastive clustering for semi-supervised classification Wei Liu, Jiangtao Song, Yuanbo Li, Cong Hu, and Xiaojun Wu (Jiangnan University) Abstract Abstract Graph Semi-Supervised Learning exploits topological structures to leverage unlabeled data. However, in open-set scenarios, existing methods typically rely on model predictions to construct adjacency graphs. This dependency makes them susceptible to confirmation bias, where incorrect pseudo-labels from unlabeled samples propagate noise through the graph structure. Furthermore, these approaches often neglect global cluster compactness, focusing primarily on local pairwise consistency. To address these limitations, we propose a novel framework named Graph consistency based on contrastive clustering (G3C). Unlike traditional approaches, G3C integrates contrastive clustering to jointly optimize sample similarity and global cluster structures. Specifically, we perform dynamic neighbor searching within the augmented feature space to construct reliable weight matrices. It also constrains the representation of data in label space and feature space to distinguish known classes from outliers through graph consistency regularization. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet-30 demonstrate that G3C outperforms state-of-the-art approaches in both closed-set and open-set settings, validating the efficacy of combining contrastive clustering with graph consistency. Wednesday Virtual Room 5 IJCNN Paper Point Cloud and 3D Vision Session Chair: Ziqi Xi (TongJi University), Ge Li (Peking University) NPPC: Neural Point-Process Codec for Exchangeable and Parallel Point Cloud Geometry Compression Yihang Ji (Peking University), Shiqiang Long and Weiming Zhang (bohuauhd), and Ge Li (Peking University) Abstract Abstract Point clouds are fundamental 3D representations for robotics, AR/VR, and autonomous driving, yet most learning-based compression methods impose artificial orderings to enable sequential entropy models. This creates a critical bottleneck: decoding must proceed sequentially, limiting throughput on modern parallel hardware and hindering real-time applications. We propose Neural Point-Process Codec (NPPC), which models point clouds as spatial point processes with an exchangeable factorization that respects their inherent unorderedness. The key insight is that by decomposing geometry into conditionally independent cell-wise components, we eliminate the need for any point ordering while enabling natural parallel decoding without complex synchronization. This exchangeable design provides both a principled probabilistic interpretation and practical parallelism. Experiments on ModelNet40 and ScanNet demonstrate that NPPC achieves competitive rate-distortion performance while enabling substantial decoding speedup, making it practical for latency-critical applications Hybrid Position Encoding and Boundary-Aware Deep Supervision for 3D Point Cloud Semantic Segmentation Yongyue Ma (Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Abstract Abstract 3D point cloud semantic segmentation serves as a fundamental task for 3D scene understanding, whose performance directly influences decision-making in critical applications such as autonomous driving. Although advanced methods such as Point Transformer V3 have been proposed, they still face challenges in complex scenes, including low accuracy in fine-grained and small-object segmentation, inadequate position encoding, and limited capability in local-global context modeling. To tackle these issues, this paper presents a novel framework for 3D point cloud semantic segmentation that integrates supervision enhancement and feature enhancement. Specifically, we design a Boundary-Aware Deep Supervision (BADS) module to strengthen fine-grained feature learning and refine the quality of the supervision in representative regions. Meanwhile, we propose a Hybrid Position Encoding (HPE) strategy that effectively mitigates the deficiency of position encoding by fusing multi-scale HashGrid absolute encoding and multi-axis Rotary Position Embedding (RoPE). Experiments on ScanNet V2 and S3DIS demonstrate that our method achieves 78.1\% mIoU on ScanNet V2 and 74.4\% mIoU on S3DIS Area-5, outperforming Point Transformer V3 and other state-of-the-art approaches, especially on fine-grained and small-object categories. Ablation studies validate the effectiveness of both BADS and HPE individually and synergistically, offering a new technical solution for 3D point cloud semantic segmentation in complex scenes. SRCNet: Robust Point Cloud Completion via Semantic-Guided Part Augmentation Fei Hu and Li Lu (Sichuan University) Abstract Abstract Point cloud completion is a fundamental challenge in 3D vision, aiming to recover complete geometries from sparse, incomplete sensor data. However, existing methods, predominantly trained on synthetic datasets with uniform viewpoint-based occlusions, exhibit a critical structural blindness. They tend to overfit to specific projection patterns and fail to generalize to the diverse structural defects common in real-world scenarios. To address this, we propose a synergistic framework that enforces structural invariance through both data and architecture. First, we introduce PartDrop, a novel semantic-guided augmentation strategy. Unlike generic erasing techniques that operate on random spatial regions, PartDrop leverages semantic priors to stochastically remove logical components, compelling the network to learn robust part-to-whole dependencies rather than shallow template matching. Second, to effectively reason about these disjointed parts, we propose the Structural-Relational Completion Network (SRCNet). Designed as a specialized Transformer architecture, SRCNet features a structure-aware dual-stream encoder to disentangle semantic context from local topology, and a Region-Relation Decoder equipped with an Inter-region Relation Learning mechanism. By leveraging the inherent global attention capabilities of Transformers, this architecture explicitly models long-range dependencies between observed and missing regions, enabling high-fidelity reconstruction even under severe structural incompleteness. Extensive experiments on PCN, ShapeNet-55, and real-world KITTI benchmarks demonstrate that our method achieves state-of-the-art performance, showing significantly improved robustness and consistency across diverse incompleteness patterns compared to existing approaches. LML-MAE: A Local-Aware Multi-Branch Lightweight Framework for Masked AutoEncoder Ziqi Xi and Xiaoliang Gong (School of Computer Science and Technology, TongJi University) Abstract Abstract Masked autoencoders (MAEs) have been widely adopted for self-supervised learning in point cloud understanding. However, existing methods often fail to sufficiently capture local structural information. While multi-scale frameworks enhance local geometry at different scales and cross-modal approaches leverage auxiliary data, both introduce significant computational overhead, and cross-modal data are often difficult to obtain. To address these limitations, we propose a Lightweight Multi-branch Local feature encoding Masked Autoencoder (LML-MAE) for point cloud representation learning. Our framework employs a multi-branch local feature extraction module based on depthwise separable convolutions to achieve expressive local encoding with reduced complexity. To prevent information leakage during masked reconstruction, a center position prediction module is introduced to enforce more effective pretraining objectives. Moreover, we design a local structural information embedding module that dynamically integrates point-wise neighborhood structural information into point features at the patch level, enabling more expressive local geometry modeling. Compared with existing multi-scale and cross-modal approaches, LML-MAE achieves lower computational cost and faster pretraining convergence. Extensive experiments on downstream tasks demonstrate that our method consistently outperforms the baseline Point-MAE model, achieving 5.67%, 5.47%, and 5.41% performance gains on three variants of the challenging ScanObjectNN dataset. Wednesday Virtual Room 6 IJCNN Paper Recommendation and Information Retrieval I Session Chair: Ziyang Cai (Guangdong University of Technology), Bo Ren (Shanghai University of International Business and Economics) Balancing User Experience and Monetization: An iQoE-Aware Recommendation Framework for Audio Platforms Jie Zhao, Ziyang Cai, and Daiyang Wu (Guangdong University of Technology) Abstract Abstract Balancing user experience and monetization is a fundamental challenge in recommendation systems, particularly for experience-oriented content platforms with highly dynamic user interactions. This paper addresses this challenge by explicitly modeling Instantaneous Quality of Experience (iQoE) and incorporating it into a multi-objective recommendation framework. We propose an iQoE-aware multi-expert neural architecture that jointly learns payment-related behavior and iQoE through task-specific gated experts, together with a dynamic loss adjustment strategy to mitigate objective conflicts during training. In addition, a multi-objective re-ranking strategy is designed to balance experience and monetization at the list level for Top-K recommendation. Experiments on real-world datasets collected from a large-scale interactive audio platform demonstrate consistent improvements over baselines in both prediction performance and recommendation quality, validating the effectiveness of jointly optimizing user experience and monetization. Mitigating Position Bias in Cascading User Behaviors for Recommendation Yan Liu and Yingpeng Du (Nanyang Technological University); Hongzhi Liu (Peking University); Thanh-Son Nguyen (Agency for Science, Technology and Research); and Erik Cambria (Nanyang Technological University) Abstract Abstract Recommender systems (RSs) aim to learn users’ intrinsic preferences from historical behaviors. However, user behaviors are often strongly influenced by item display positions, which introduces exposure bias and can mislead preference learning. Existing position-aware methods still face two key limitations. (1) Some approaches collect “unbiased” data via randomized shuffling or swapping, yet this may degrade recommendation quality and harm user experience. (2) Methods trained on biased logs usually rely on single-type behaviors, making it difficult to decouple true preference from position effects. Specifically, they rely on target behaviors (e.g., purchases) that are sparse with respect to display positions, limiting their ability to accurately model and mitigate position bias. GRED: Graph Representation Editing for Efficient Dataset Recommendation Chenyang Zhao, Xiaolei Du, Haotian Chen, and Qingqing Long (Computer Network Information Center, Chinese Academy of Sciences); Zhiyuan Ning (Westlake University); and Meng Xiao, Hengshu Zhu, and Yuanchun Zhou (Computer Network Information Center, Chinese Academy of Sciences) Abstract Abstract Graph-based recommendation systems have gained great success in recent years. Yet most methods still fail to fully capture the dynamics of user behavior efficiently, often require updating all parameters in the user-node embedding table during finetuning, making it hard to adjust to frequent changes in user behaviors. The problem is amplified in sparse interaction graphs, such as those observed in scientific dataset platforms, where users interact with only three items on average, far fewer than those in commercial platforms, leading to marked drops in accuracy. To tackle these issues, \underline{\textbf{G}}raph \underline{\textbf{R}}epresentation \underline{\textbf{ED}}itor (\textbf{GRED}) is introduced as a lightweight, inject-able module set that ``edits" node embeddings to enhance task adaptation. GRED utilizes LLM-based embedding with Graph Neural Network (GNN) to extract complementary semantic and structural signals from history logs. The design unifies static item content and relational structure, yielding more dynamic, and more transferable representations. GRED is used to fine-tune the model for recommendation and conduct extensive evaluations under multiple configurations. The results show that GRED achieves over 10\% improvements against state-of-the-art recommendation systems while requiring significantly less fine-tuning time and computational cost. HiTower: Hierarchical Interest Modeling with Adaptive Gating for Personalized Recommendation Nijia Mo, Bo Ren, Yongda Wei, and Hui Liu (Shanghai University of International Business and Economics) Abstract Abstract In recommendation systems, conventional approaches typically learn a fixed representation vector for each item. However, the value of an item is inherently condition-dependent; for example, the meaning and appeal of the same movie may differ significantly for teenagers and middle-aged viewers, just as the same product may hold different attractions for novice versus expert users. Existing solutions, such as incorporating attention mechanisms into item representations to embed user–item relevance, partially alleviate this issue. Yet, they are often constrained by flat modeling and struggle to dynamically adapt to the varying interests across different users. To address these issues, we propose a Hierarchical masked gating dual-tower model called \textbf{HiTower}. By introducing an additional concept pool and a dynamic, masked gating mechanism, our model can adaptively select a variable number of interest concepts based on user interaction sequences. These activated concepts are then used to enrich both the original sequence representations and candidate item embeddings. Furthermore, to alleviate the collapse effect of softmax in interest activation, we propose a layer-wise activation strategy that enables the model to attend to multiple focuses hierarchically. Experimental results on three real-world datasets demonstrate that \textbf{HiTower} consistently outperforms state-of-the-art baseline methods, validating the effectiveness and practical value of the proposed framework in user interest modeling. Code is available at \href{https://anonymous.4open.science/r/HiTower-1CAB/}{https://anonymous.4open.science/r/HiTower-1CAB/}. Wednesday Virtual Room 7 IJCNN Paper Recommendation and Information Retrieval II Session Chair: Jiwei Qin (Xinjiang University), jiayue wu (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) SteerRec: Popularity-Steerable Disentanglement for Long-tail Recommendation Jiayue Wu, Chenxu Niu, Mingzhe Lu, Qihao Wang, and Yue Hu (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Abstract Abstract Popularity bias remains a pervasive challenge in recommendation systems, where models tend to over-recommend a few head items while neglecting a vast number of niche items. Existing debiasing methods are typically computationally intensive, dependent on specific model architectures, and fail to fully disentangle popularity bias from intrinsic collaborative signals. To address this issue, we propose \textit{SteerRec}, a novel two-stage framework designed for balanced and precise popularity-steered recommendation. We employ a plug-and-play \textit{Disentanglement Adapter} to separate popularity bias from core item representations and an \textit{Adaptive Gated Mechanism} to dynamically control debiasing intensity for each item, thereby generating refined embeddings that reconcile recommendation accuracy with long-tail fairness in the final results. Extensive experiments on three real-world datasets demonstrate that \textit{SteerRec} significantly improves the exposure and recommendation performance of tail items while maintaining the overall recommendation quality. DAAM: Dual Alignment with Adaptive Margins Framework for Popularity-Aware Fair Recommendation Wenxuan Xie, Jian Cao, Shiyou Qian, and Ziyi Huang (Shanghai Jiao Tong University) Abstract Abstract Recommendation systems frequently suffer from severe popularity bias, which significantly constrains the exposure of unpopular items and degrades overall recommendation quality. Existing debiasing methods, primarily based on causal inference or graph contrastive learning, have achieved progress in mitigating bias. However, they predominantly rely on imposing global constraints to suppress the influence of popular items. This paradigm overlooks users’ heterogeneous preferences regarding popularity and fails to effectively bridge the representation separation between popular and unpopular items within the embedding space, leaving unpopular items with insufficient semantic learning. To address these challenges, we propose the Dual Alignment with Adaptive Margins (DAAM) framework, which reframes popularity as a personalized user preference. DAAM optimizes recommendations via a dual alignment mechanism, achieving intent alignment by matching user popularity expectations with item reality, and representation alignment by transferring collaborative signals from popular anchors to enhance the learning of unpopular items. Furthermore, a popularitybased dynamic margin adjustment mechanism is incorporated to adaptively regulate training difficulty. Experiments on three datasets demonstrate that DAAM outperforms mainstream baselines in both accuracy and fairness, verifying the rationality and effectiveness of our approach. DTMFRec: Dynamic Temporal Weighted Multi-Path State Space Fusion for Sequential Recommendation Jisheng Tian, Xiaoqiang Ren, Hongpei Ji, and Wenpeng Lu (Key Laboratory of Computing Power Network and Information Security, Ministry of Education; Shandong Computer Science Center (National Supercomputer Center in Jinan); Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing; Shandong Fundamental Research Center for Computer Science) and Xueying He (Shandong University of Traditional Chinese Medicine) Abstract Abstract Sequential Recommendation Systems (SRS) aim to predict users' dynamic preferences from their historical interaction sequences. Although recent state space models such as Mamba exhibit strong potential with linear computational complexity, their inherently unidirectional architecture limits contextual comprehension and short-term interest modeling, especially in sparse data regimes. To address these challenges, we propose a Dynamic Temporal Weighted Multi-Path State Space Fusion (DTMFRec) Framework, a multi-path adaptive fusion framework. DTMFRec employs a multi-path architecture to simultaneously capture dependencies from both the raw and temporally enhanced sequences, thereby constructing more comprehensive contextual representations. Furthermore, an Adaptive Fusion Gating Layer is introduced to dynamically weight multiple information paths for optimal fusion. Experimental results on multiple real-world datasets demonstrate that DTMFRec significantly outperforms existing baselines and demonstrates superior performance in long-tail scenarios. Wednesday Virtual Room 8 IJCNN Paper Recommendation and Information Retrieval III Session Chair: Changhong Li (Huazhong University of Science and Technology), Guosheng Kang (Hunan University of Science and Technology) VQ-DCR: Vector-Quantized Disentangled Item Cold-Start Recommendation Li Zou, Changhong Li, and Guohui Li (Huazhong University of Science and Technology) Abstract Abstract Generating collaborative embeddings for cold-start items relying solely on auxiliary content remains a fundamental challenge. Direct alignment methods predict ID embeddings from content features, but sparse supervision and a clustered target space make the regression unstable and biased toward frequent item clusters. RVQ discretizes content into ID-like codes, but its residual hierarchy concentrates most energy in early codebooks and leaves later codebooks with low-energy residuals, so they contribute little during training and inference. Static code summation ignores cross-attribute interactions and collapses compositional signals into a single fixed rule. We propose VQ-DCR to address both issues. We replace RVQ with Finite Scalar Quantization (FSQ) and align the discrete factors with collaborative signals via behavior-aware distillation. An item-side Transformer combiner then contextualizes factor embeddings with self-attention, and a lightweight continuous branch compensates for quantization loss. Across three datasets, VQ-DCR consistently improves cold-start ranking and yields more balanced code energy and usage. B-STAR: Behavior-aware Sequential Transformer with Adaptive Representations for Multi-behavior Recommendation Kaiwen Zhou and Hongjuan Liu (Northeastern University) Abstract Abstract Multi-behavior sequential recommendation (MBSR) leverages heterogeneous user interactions (e.g., views, purchases) to capture dynamic preferences. However, existing Transformer-based MBSR models still face three limitations: (1) they model different behaviors via simple additive embeddings or separate sequences, failing to adaptively capture the varying influence between behavior pairs; (2) they mainly rely on local sequential patterns and overlook high-order collaborative signals in the global user–item graph; and (3) they rarely exploit explicit negative feedback, missing useful information for delimiting user preference boundaries. We propose Behavior-aware Sequential Transformer with Adaptive Representations (B-STAR), which unifies global graph reasoning with local sequence modeling. A global graph encoder first injects high-order collaborative signals into item representations. A behavior-aware attention mechanism with a learnable interaction matrix then re-weights attention scores according to the affinity between behavior types. Furthermore, a hard negative push strategy, combined with a multi-task learning objective encourages the model to separate preferred and disliked items and yields more robust representations. Experiments on three real-world datasets (Yelp, Taobao, and Tianchi) show that B-STAR consistently outperforms recent state-of-the-art baselines. Visualizations of learned attention weights indicate that B-STAR captures meaningful behavioral patterns (e.g., the e-commerce purchase funnel). MSRec: Multimodal Semantic Learning for Cross-modal Recommendation via Multi-curvature Geometric Spaces Yuxin Dong (Beijing University of Posts and Telecommunications) and Zheming Yang (Institute of Computing Technology, Chinese Academy of Sciences) Abstract Abstract Abstract—The rapid growth of multimedia content has further stimulated research on multimodal recommendation. However, existing methods, when integrating modality embeddings with ID embeddings, often implicitly assume a unified latent ge ometry, lack rigorous theoretical support, and tend to dilute item representations. We argue that heterogeneous signals like text, vision, and behavioral IDs—exhibit distinct geometric reg ularities, including local proximity, hierarchical structure, and cyclic constraints, forcing them into a single Euclidean manifold entangles shared, modality-specific, and spurious components, thereby degrading ranking performance. To address this issue, we propose a Multi-geometry Semantic alignment Recommendation framework (MSRec) that embeds each modality and the ID representation into parallel Euclidean, hyperbolic, and spherical manifolds and performs geometry-aware inference. Specifically, MSRecemploys per-space modality weighting and ID-conditioned gating to extract cross-modality commonalities and preserve discriminative modality-specific cues in parallel manifolds, the fused representations are written back to the ID anchor via native manifold algebra, with cross-space consensus and geometry tailored branches jointly forming the final item embedding for ranking. We further introduce a graph-neighborhood contrastive objective that integrates top-K neighborhood sampling, proto type construction, and near-boundary hard example generation, jointly optimized with a standard BPR loss. Experiments on three public datasets demonstrate that MSRec consistently outperforms strong baselines, achieving up to 9.62% relative improvement in Recall@10 and consistent gains of 2–9% across Recall@K and NDCG@K. Multi-View Denoised Graph Collaborative Filtering for Third-Party Library Recommendation Yuxiang Kuang, Xiaoxiang Liao, Guosheng Kang, Ye Cao, Ziyi Niu, Wen Li, and Jiayan Xiang (Hunan University of Science and Technology) Abstract Abstract Third-party libraries are crucial for accelerating software development, yet identifying suitable candidates within the expanding open-source ecosystem remains challenging. Graph neural network (GNN)-based collaborative filtering emerges as a promising solution due to its ability to effectively model high-order interaction relationships. However, data sparsity leads to biased representations, hindering effective node representation learning. Moreover, standard multi-layer convolutions blend irrelevant node features, causing noise accumulation and suboptimal embeddings. To address these challenges, a Multi-View Denoised Graph Collaborative Filtering (MDGCF) approach is proposed for TPL recommendation. Specifically, MDGCF constructs complementary graph views to alleviate sparsity and employs an implicit supervision module to guide representation learning with auxiliary signals. To mitigate noise, we propose a denoised graph convolution that calculates normalized weights for even-order nodes to suppress irrelevant propagation, alleviating noise accumulation. Furthermore, a sampled softmax loss with graph topology-based negative sampling is carefully designed to replace the conventional BPR loss, strengthening the differentiation between positive and negative samples and enhancing recommendation accuracy. Experiments on a real-world dataset demonstrate that MDGCF outperforms state-of-the-art baselines, validating the model's effectiveness in improving the performance of TPL recommendation. Wednesday Virtual Room 1 IJCNN Paper Recommendation and Information Retrieval IV Session Chair: XiangQin Pang (Sichuan Normal University, College of Computer Science), ZhiJian Fang (Zhejiang Sci-Tech University) MKGRec: A Meta-Knowledge Guided Multi-View Recommendation Framework via Integrating LLMs XiangQin Pang, ZhenDong Wu, Yang Yang, and Hetao Chen (Sichuan Normal University, College of Computer Science) and Jingzhi Zhang (Shenzhen University, College of Computer Science and Software Engineering) Abstract Abstract To address the limitations of traditional recommender systems in dynamic interest modeling, multi-source semantic fusion, and semantic preservation during augmentation, this paper proposes a Meta-Knowledge Guided Multi-View Recommendation Framework (MKGRec) via integrating Large Language Models. The core of our approach is to construct a meta-knowledge guided multi-view contrastive learning framework. First, we utilize Large Language Models to parse user behavior sequences and infer latent dynamic interest evolution paths. By filtering noise through mutual information maximization, we construct a dynamic user interest graph. Next, guided by meta-knowledge, we perform cross-view semantic alignment and conflict pruning on the collaborative, interest, and knowledge views to achieve efficient fusion of multi-source heterogeneous information. Finally, we construct a semantic-preserving data augmentation pipeline via a dynamic masking strategy. Experiments on three public datasets demonstrate that our method improves Recall@50 and NDCG@50 by up to 6.87% and 13.16% compared to the best baselines. Ablation studies validate the key contributions of the LLM interest inference, meta-knowledge guidance, and interest augmentation learning modules. Furthermore, noise injection experiments confirm the superior robustness of our method.Our code is available at https://github.com/pxq1/MKGRec. Decoupling Preference Aggregation and Interest Evolution via Multi-Agent Reasoning for Next POI Recommendation WeiWen Cao (Institute of Systems and Information Engineering University of Tsukuba) Abstract Abstract Point-of-interest (POI) recommendation aims to predict a user's next visit based on historical mobility behaviors and is fundamental to location-based services. While large language models (LLMs) have shown promising reasoning capabilities for this task, existing LLM-based approaches still face several challenges: (i) they often rely on monolithic reasoning processes, making it difficult to explicitly aggregate heterogeneous user preference signals; (ii) they lack effective mechanisms to capture the temporal evolution of user interests, resulting in limited modeling of dynamic mobility intentions; and (iii) they struggle to efficiently filter large-scale POI candidates, which increases computational cost and degrades recommendation precision. TEG-Rec: Scene-driven Thought-Evolution Generative Recommender Zhaohui Zhang and Chun Xu (Donghua University, School of Information and Intelligent Science) Abstract Abstract Recommender systems are evolving towards generative methods. However, they often overlook the fundamental decision-making factors behind user interaction and treat the complex decision-making process as a series of fixed identifiers. To overcome this limitation, we propose TEG-Rec (Thought-Evolution Generative Recommender), an innovative framework that integrates large language models (LLMs) and hypergraph reasoning into the generation-based recommendation process. Specifically, TEG-Rec first uses LLMs to extract the potential "thought scenarios" related to the items, thereby bridging the semantic gap between item attributes and user intentions. Next, a thought evolution hypergraph is constructed to represent the dynamic transformation of these user intentions, and hypergraph neural networks (HGNNs) are used to generate structured thought embeddings. These embeddings as continuous prompts in the Transformer decoder, guiding the autoregressive generation of item's semantic IDs. Comprehensive experiments conducted on real datasets show that TEG-Rec can effectively capture high-order collaborative information and outperform the existing best benchmark models in recommendation performance. Diffusion-guided Contrastive Learning with Large Language Model-enhanced Knowledge Graph Recommendation Yichen Zhang, Huaxiong Zhang, and Zhijian Fang (Zhejiang Sci-Tech University) and Lei Zhou (Zhejiang Talent Development Group) Abstract Abstract Knowledge graphs (KGs) provide rich factual and structural information for recommendation, yet real-world KGs often contain noisy links and long-tail sparsity. Most existing methods learn on a fixed KG, making it difficult to explicitly remove task-irrelevant links from a structural perspective. Meanwhile, inner-product-based pointwise scoring has limited ability to jointly exploit multiple structural cues when candidate structures are highly similar or evidence is scarce, limiting fine-grained ranking at the top positions. To address these issues, we propose DiCoLL-KGRec, a unified framework that integrates diffusion-based KG denoising, diffusion-guided multi-view contrastive learning, and large language model (LLM)-enhanced two-stage reranking. Specifically, we perform diffusion-based generative modeling in the item-entity relation space to denoise noisy KGs, deriving a task-relevant clean KG and multiple diffusion views. We then quantify discrepancies across diffusion views as structural reliability signals, which serve as a consistency prior. Guided by this signal, we conduct biased augmentations on the user-item graph and impose multi-view contrastive constraints on the KG side, enabling robust representation learning via generated views and consistency alignment. At inference time, we rewrite the top candidates and their structural features into structured prompts and use an LLM to rerank them, further refining the final list. Experiments on three public datasets show that our method yields consistent improvements in overall performance, robustness, and decision enhancement. Wednesday Virtual Room 2 IJCNN Paper Reinforcement Learning II Session Chair: Li Chengwei (Institute of Automation Chinese Academy of Sciences; School of Artificial Intelligence,University of Chinese Academy of Sciences), Hongmai Xu (Guangdong University of Technology) Observable Multi-Task Cooperative Multi-Agent Reinforcement Learning for Inventory Management Hongmai Xu (Guangdong University of Technology); An Zeng, Qinghua Zhu, Yuzhu Ji, and Baoyao Yang (Guangdong University of Technology); Dan Pan (Guangdong Polytechnic Normal University); and Jiayu Ye (Guangdong University of Technology) Abstract Abstract Unreasonable Inventory Management (IM) strategies lead to extended delivery times and increased costs in the supply chain. Multi-Agent Reinforcement Learning (MARL) is proposed for the research area of IM. However, existing methods are limited to simple one-to-many supply chain models, which cannot match real-world supply chain. Additionally, in the current use of Cooperative Multi-agent Reinforcement Learning (CMARL) in supply chain model, agents communicate through shared rewards, resulting in insufficient communication efficiency. To address these issues, this paper proposes a many-to-many supply chain model and Observable Multi-task CMARL which leverages a learning paradigm of Centralized Training and Decentralized Execution (CTDE) Based on Shared Rewards and Multi-task CMARL. The learning paradigm that we propose applies shared rewards to each centralized critic, enhancing the communication and collaboration abilities among agents. Furthermore, we propose Multi-task CMARL which improves the performance of CMARL by sharing parameters between tasks. To verify the effectiveness of the algorithm, extensive experiments were conducted on both real-world and simulated datasets. The experimental results show that the Observable Multi-task CMARL algorithm outperforms traditional heuristic algorithms and CMARL in the many-to-many supply chain model, maintaining low inventory levels while improving product sales to an average increase of 131.2%. GSLHD: Global Skill Learning with Heterogeneous Decoding for Multi-Agent Reinforcement Learning Tenghai Qiu (School of Automation, Central South University; Institute of Automation, Chinese Academy of Sciences); Ruoyang Gao (College of Civil Engineering, Tongji University); Jinyuan Feng and Zhiqiang Pu (Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences); and Yuqian Zhao and Biao Luo (School of Automation, Central South University) Abstract Abstract In heterogeneous multi-agent reinforcement learning, each type of agent can naturally emerge as different roles, meaning that diversity of agents is conducive to cooperation within the team. However, adopting a fully heterogeneous policy, where each agent learns an independent policy without any shared components, suffers from sample inefficiency. To achieve a better trade-off of diversity and efficiency, we propose a novel framework named Global Skill Learning with Heterogeneous Decoding (GSLHD), which fosters sample efficiency by collaboratively sharing generalizable skill semantics and promotes diversity by heterogeneously decoding according to agent type. Specifically, we introduce a concept of skill, which is extracted from the Global Skill Assignment (GSA) module, to adaptively encode skill based on the inference of global states. Then, generalizable skill semantics are learned by an attention mechanism and shared across types. Finally, a heterogeneous decoder promotes policy diversity by decoding cross-type shared skills into specific actions tailored for each agent type. Extensive experiments are conducted on SMAC and SMACv2 to verify the remarkable performance of GSLHD. Partially Observable Multi-Type Mean-Field Reinforcement Learning Shuhui Chu and Chengzhong Xu (University of Macau) Abstract Abstract Multi-agent reinforcement learning (MARL) holds significant promise for real-world applications, but scaling it to environments with many agents remains a fundamental challenge. Mean-field theory has emerged as a powerful approach to this problem, simplifying complex multi-agent interactions into tractable two-agent interactions. While this greatly enhances scalability, existing methods rely on strong and unrealistic assumptions—homogeneous agents and global observability—which limit their practical applicability and theoretical generality. In this work, we relax these restrictive assumptions by introducing agent heterogeneity and partial observability into the mean-field MARL framework. We explicitly model diverse agent behaviors through multiple agent types and operate under realistic local observation constraints. To this end, we propose POMTMFQ (Partially Observable Multi-Type Mean-Field Q-learning) under these more general conditions. We provide a comprehensive theoretical analysis, proving that POMTMFQ converges close to the Nash Q-value. Experimental results on three large-scale games in MAgent platform demonstrate the effectiveness and superiority of POMTMFQ. Notably, POMTMFQ achieves 42% higher winning rates and significantly faster convergence compared to prior methods, confirming its ability to learn more efficient policies in complex, many-agent environments. Evolutionary Enhanced Multi-Agent Reinforcement Learning for Cooperative Air Combat Chengwei Li, Yang Gao, and Junlin Liu (Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence,University of Chinese Academy of Sciences) and Hui Chang, Xinchen Zhang, and hao Zhao (Institute of Automation, Chinese Academy of Sciences) Abstract Abstract As modern air combat evolves toward beyond-visual-range (BVR) multi-aircraft cooperative engagements, autonomous decision-making for unmanned combat aerial vehicles (UCAVs) faces significant challenges due to high-dimensional state spaces, discrete action commands, and strongly adversarial dynamic environments. To overcome the limitations of existing multi-agent reinforcement learning (MARL) methods in such settings—namely insufficient exploration efficiency, low sample utilization, and poor policy generalization—we propose Adversarial Curriculum and Evolutionary-enhanced Multi-agent Proximal Policy Optimization (ACE-MAPPO), a hybrid learning framework that integrates evolutionary algorithms with MAPPO. Specifically, a genetic soft update mechanism is introduced to enhance population diversity and mitigate convergence to local optima. An evolutionary-augmented prioritized trajectory replay strategy is further employed to improve the utilization of sparse high-value samples. In addition, an adversarial evolutionary curriculum learning mechanism is designed to enable adaptive training with progressively increasing difficulty. Extensive experimental results demonstrate that the proposed method outperforms MAPPO and other baseline algorithms in terms of training stability, convergence speed, and win rate, validating its effectiveness in multi-aircraft cooperative air combat scenarios. Wednesday Virtual Room 3 IJCNN Paper Reinforcement Learning III Session Chair: Zhilin Zhang (Chongqing Institute of Green and Intelligent Technology,Chinese Academy of Sciences; Chongqing School,University of Chinese Academy Sciences), ziming liu (University of Chinese Academy of Sciences, Institute of Software Chinese Academy of Sciences) Risk-Aware Model-Based Offline Reinforcement Learning via Bellman Risk Propagation Zhilin Zhang, Zihao Li, Xiaoyu Shi, and Yun Lu (Chongqing Institute of Green and Intelligent Technology,Chinese Academy of Sciences; Chongqing School,University of Chinese Academy Sciences) and Yanan Bai and Zepeng Gong (ChongQing University of Technology) Abstract Abstract Offline reinforcement learning (RL) learns policies from fixed datasets without environment interaction, improving safety and sample efficiency. However, distributional shift between behavior and learned policies often causes severe value overestimation, especially for out-of-distribution actions. Existing conservative methods mitigate this by restricting policy updates within dataset support, but their excessive conservatism limits policy improvement. Alternatively, data augmentation and model-based approaches expand the training distribution yet lack principled control over risks from unreliable synthesized samples.In this paper, we propose Model-Based Risk-Aware Policy Optimization (RAPO), a framework that explicitly models long-horizon decision risk to address overestimation in offline RL. RAPO uses a learned dynamics model to generate interactive samples and introduces a dedicated risk value function that propagates extrapolation risk via Bellman backups. This decouples risk estimation from reward learning, enabling transparent and adaptive risk-aware policy optimization. We further design a risk-adaptive weighting mechanism that balances exploration and conservatism based on policy behavior and dataset quality. We theoretically show that RAPO achieves a tighter performance lower bound than existing offline RL methods. Empirical results on diverse D4RL benchmarks demonstrate consistent and significant improvements over state-of-the-art baselines. DOOF: Decoupled Offline to Online Finetuning via Dynamics Model Zifeng Zhuang (Zhejiang University, Westlake University); Xiao He (Westlake Robotics, Westlake University); Diyuan Shi (Zhejiang University, Westlake University); and Ting Wang and Donglin Wang (Westlake University) Abstract Abstract Offline reinforcement learning (RL) agents, constrained by suboptimal datasets, typically require further online finetuning before deployment. However, this process suffers from severe distribution shift due to unbalanced state coverage in offline data and the inherent conservatism of offline algorithms, often destabilizing policy improvement and causing performance degradation. In this work, we propose a novel \textit{decoupled offline-to-online framework} that separates policy improvement from distribution shift mitigation. Our key insight leverages the dynamics model in model-based RL: 1) During online interaction, \textit{only} the dynamics model is fine-tuned to adapt to the real environment, avoiding direct policy updates under shifting distributions. 2) The policy is then refined \textit{offline} using the updated dynamics, preserving the conservatism of offline RL while eliminating sudden performance drops. This approach reduces the online phase to supervised dynamics learning—a simpler task than policy optimization under distribution shift—while maintaining sample efficiency. Extensive experiments on standard offline RL benchmarks demonstrate that our framework effectively eliminates distribution shift and achieves superior performance with limited online interaction. Unraveling Max-Return Sequence Modeling via Return Consistency Zifeng Zhuang (Westlake University, Zhejiang University); Dengyun Peng (Harbin Institute of Technology); Jiacheng Liu (Zhejiang University); Xiao He (Westlake Robotics, Westlake University); and Ting Wang and Donglin Wang (Westlake University) Abstract Abstract Offline reinforcement learning (RL) learns from fixed datasets without interaction with online environment, enabling supervised solutions for offline RL. Decision Transformer (DT) casts offline RL as return-conditioned supervised sequence modeling, thereby sidestepping optimal value fitting and policy gradients. This paradigm overlooks RL’s core objective of return maximization, which yields brittle behavior on suboptimal trajectories and limited stitching ability. Reinformer reorients this objective through max-return sequence modeling: during inference, the model conditions on the predicted maximum achievable returns to generate the optimal actions. To better understand both the SOTA performance of this paradigm and its occasional dramatic failures, we adopt a supervised perspective and introduce the return consistency to assess whether similar state-action pairs have similar returns. Indeed, high return consistency guarantees the maximized return reliably cues the optimal action, while low consistency may lead to suboptimal action selection. Through visualizations, two different consistency modes are exposed and we quantify this via the return standard deviation of the data cluster with highest return mean. Furthermore, we reveal the relationship between this metric and 1) final performance, 2) context lengths, 3) model architectures through a systematic study. Finally, we improve return consistency by explicitly decreasing the return standard deviation, thereby further increasing the performance. Generalized-Risk Constrained Policy Optimization for Cross-Morphology Transfer liu ziming (Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences); meng qingxin, chen linjuan, and jing mingxuan (Institute of Software, Chinese Academy of Sciences); and zheng jin (The Science and Technology on Integrated Information System Laboratory) Abstract Abstract Cross-morphology policy transfer in reinforcement learning remains challenging due to significant differences in dy- namics, stability, and safety characteristics across robot embod- iments. Existing methods often overlook how safety constraints shift under morphology variations, leading to transferred policies that are either unsafe or overly conservative. Constrained reinforcement learning provides a promising per- spective for addressing this issue. However, mainstream meth- ods such as Projection-based Constrained Policy Optimization (PCPO) typically assume fixed cost distributions and constraint thresholds, making them ineffective under the cost distribu- tion shifts induced by morphology changes. To address the limitations of standard constrained reinforcement learning in cross-morphology transfer, we propose a generalized-risk con- strained policy optimization framework that extends PCPO. Our approach uses a dual-layer generalized risk mechanism: raw costs are normalized using reference statistics from the source policy to obtain morphology-invariant relative risks, and then mapped into a percentile-based risk space via a generalized risk network, decoupling constraints from absolute cost magnitudes. An adaptive threshold further balances safety and exploration. This framework resolves key challenges in cross-morphology transfer: it mitigates the effect of morphology-induced cost distribution shifts, and enables stable constraint enforcement during policy adaptation. The framework is integrated into PCPO by projecting policy updates onto the feasible set defined by the generalized risk con- straint. We evaluate our approach on four MuJoCo locomotion benchmarks—Ant, HalfCheetah, Hopper, and Walker2d—with two morphology variants each. Experimental results demonstrate that our method substantially reduces constraint violations and improves task performance compared to direct policy transfer and PCPO trained from scratch, and in some cases even outperforms PPO trained from scratch Wednesday Virtual Room 4 IJCNN Paper Reinforcement Learning IV Session Chair: Zhilin Chen (Fuzhou university), Xing Yu (East China Normal University) RISK-SAC: Risk-Driven Soft Actor-Critic for Safety-Critical Scenario Generation in Autonomous Driving Dehui Du, Ermuyun Li, Xing Yu, and Lili Tian (East China Normal University) Abstract Abstract Simulation-based testing has become an essential approach for validating autonomous driving systems. However, existing scenario generation methods often struggle to efficiently discover rare safety-critical interactions in high-dimensional traffic environments and may produce physically implausible behaviors due to unconstrained or manually designed exploration strategies. To address these limitations, we propose RISK-SAC, a risk-driven scenario generation method built upon Soft Actor-Critic (SAC) to generate diverse and physically realistic safety-critical driving scenarios. The proposed approach leverages stochastic exploration and risk-aware learning signals, together with explicit realism constraints, to guide exploration toward hazardous yet valid interaction patterns. Extensive experiments in diverse scenarios within the CARLA simulator demonstrate that RISK-SAC consistently outperforms representative baselines in terms of collision discovery efficiency, scenario diversity, and physical realism. Ablation studies further confirm the effectiveness of the proposed components in improving exploration efficiency and training stability. These results indicate that RISK-SAC provides a practical solution for efficient and realistic safety-critical scenario generation in autonomous driving. A Survey on Large Model-based Safety-critical Scenario Generation for Autonomous Driving Lingzhong Meng (Institute of Software Chinese Academy of Sciences); Cancan Jiang (Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences); Zhidong Wang (Qiyuan Lab,Beijing); and Guang Yang, Hongyun Yu, and Yuxi Ma (Institute of Software Chinese Academy of Sciences) Abstract Abstract Safety-critical scenario generation for autonomous driving aims to efficiently and authentically construct long-tail, high-risk traffic scenarios for simulation testing. The core of this task lies in leveraging generative technologies to address the scarcity and high acquisition costs of extreme accident data in real-world road testing. This paper provides a comprehensive review of existing safety-critical scenario generation technologies based on Large Language Models (LLMs). Decoupling Representation Robustness and Behavioral Robustness in End-to-End Driving: A Feature-Space Robustness Module for Actor–Critic Policies yuanchen liu (China University of Geosciences) and dongcheng li (Department of Computer Science, California State Polytechnic University - Humboldt) Abstract Abstract Closed-loop adversarial training has emerged as an effective paradigm for improving the safety of end-to-end autonomous driving policies. However, existing approaches primarily operate at the environment or observation level, leaving the robustness of internal representations largely implicit and poorly understood. DSMOP-Drive: A Deep Reinforcement Learning Framework for Autonomous Driving with Decoupled Safety and Multi-Objective Optimization Zhilin Chen, Jingtang Chen, Qisong Guo, Mingjian Fu, and Yuanlong Yu (Fuzhou university) Abstract Abstract Deep Reinforcement Learning (DRL) empowers autonomous vehicles with decision-making capabilities to meet diverse driving demands. However, it struggles to jointly optimize multiple conflicting objectives under strict safety constraints. Existing Lagrangian approaches often suffer from gradient conflicts and limited generalization due to coupled representations. To address this, we propose DSMOP-Drive, a primal-based framework decoupling safety rectification from multi-objective optimization. Specifically, it introduces Driving Safety Rectified Multi-Objective Policy Optimization (DSR-MOPO) to harmonize gradients, complemented by a constraint rectification mechanism for immediate safety. Furthermore, we introduce the Dynamic Consistent Joint Embedding Predictive Architecture (DC-JEPA) to disentangle vehicle dynamics from task preferences in latent space, enabling robust adaptation. Experiments in complex highway merge ramp scenarios confirm that DSMOP-Drive achieves a 96.25% success rate under high traffic density, demonstrating low latency in safety response and superior driving utility thanks to the decoupled design. Wednesday Virtual Room 5 IJCNN Paper Reinforcement Learning V Session Chair: Qixuan Cao (East China Normal University), Tianyou Liu (National University of Defense Technology, College of Intelligence Science and Technology) Verifiability-Aware Training for Efficient Reachability Analysis of Deep Reinforcement Learning Systems Qixuan Cao and Min Zhang (East China Normal University) Abstract Abstract Deep reinforcement learning (DRL) is increasingly used in safety-critical settings, where safety properties need to be formally verified. Reachability analysis can provide sound guarantees, but verifying DRL systems is often expensive due to dual over-approximation of plant dynamics and neural policy outputs. Piece-wise Linear Decision Neural Network (PLDNN) alleviates this difficulty by learning a region-wise linear policy via state abstraction and exporting the closed-loop system as a hybrid automaton amenable to standard reachability tools. Despite this progress, verification can still be slow when the extracted linear control units (LCUs) exhibit (i) action discontinuities across neighboring regions and (ii) high sensitivity within a region, which accelerate reachable-set inflation and increase verification runtime. We propose a lightweight, training-time regularization framework that directly shapes LCUs while keeping the PLDNN abstraction-to-hybrid pipeline unchanged. Specifically, we add a neighbor consistency regularizer to align actions on shared boundaries and a slope penalty to suppress large linear coefficients. Experiments on multiple benchmarks show that our approach reduces verification runtime by up to 37.23% while maintaining comparable control performance to the PLDNN baseline. Certificate-Guided Robust Training for Deep Reinforcement Learning via State Abstraction Qixuan Cao (East China Normal University), Dapeng Zhi (Jiangsu University of Technology), and Min Zhang (East China Normal University) Abstract Abstract Deep reinforcement learning (DRL) systems have achieved unprecedented results in recent years. However, even the best DRL control policies are often vulnerable to state perturbations, which may cause these systems to act unexpectedly, thereby prohibiting their widespread deployment. Multiple methods have therefore been put forth to robustify DRL systems during their training. Albeit promising, these approaches suffer from two main setbacks: (i) they tend to be extremely expensive and (ii) they are typically tailored to specific DRL algorithms. To bridge these gaps, we propose a novel approach to improve the robustness of DRL systems. Our approach employs recent advances in abstraction-based training and exploits the induced finite abstract state space to efficiently compute robustness certificates, which are incorporated into a certificate-guided robust objective. The resulting abstract DRL system can then be easily trained using a novel loss function defined by robustness certificates while still maintaining high performance. Our approach is unified in that it fully applies to any DRL training algorithm. We demonstrate our approach on a wide range of benchmarks. It significantly outperforms the existing methods, improving the accumulated rewards by 92.0% on average under various perturbations and attacks. Robust UAV Hovering Control under Aggressive Conditions via Bidirectional Thrust and Deep Reinforcement Learning Tianyou Liu, Zhihong Liu, Xiaoxin Li, and Xiangke Wang (College of Intelligence Science and Technology, National University of Defense Technology) Abstract Abstract Most quadrotor control methods are designed under the assumption that motors produce only positive thrust. However, these methods struggle to maintain stability and ensure the safety of the quadrotor under aggressive conditions. To overcome this limitation, we propose a robust UAV hovering control method under aggressive conditions via bidirectional thrust and deep reinforcement learning, which directly outputs motor-thrust commands. More specifically, we build a bidirectional motor–propeller model that includes an explicit dead-zone representation during thrust reversal. To compensate for this dead-zone and improve robustness to model uncertainties, we include action history in the observation to handle the delay and apply domain randomization over dynamics parameters and initial states during training. Experimental results show that our method recovers hover from aggressive conditions with an inverted attitude and a large downward velocity. In comparison, a conventional only positive-thrust controll and an idealized bidirectional-thrust control baseline both crash. Ablation studies show that dead-zone modeling, action history, and domain randomization are necessary for robust performance. A Neural-Symbolic Approach for Safe and Smooth Simplex-Based Control Qixuan Cao, Peixin Wang, and Min Zhang (East China Normal University) Abstract Abstract Ensuring learning-enabled systems to be safe remains a great challenge, especially with high-performance but black-box neural controllers in safety-critical settings. The Simplex architecture offers a pragmatic solution by switching at runtime between an unverified but performant controller and a verified but conservative one. However, the decision logic in Simplex relies on binary switching, causing abrupt action altering and consequently degrading system smoothness and performance. We propose HyPlex, a hybrid control architecture for both safe and smooth controlling. Specifically, we design a novel neural-symbolic decision logic, which hybridizes the outputs of the two types of controllers by predicting their control weights based on runtime reachability analysis results. We evaluate HyPlex on three representative benchmarks and show that, in the presence of obstacles, it improves trajectory smoothness by up to 88.27% and reduces performance degradation by up to 48.07%, compared to standard Simplex switching. Wednesday Virtual Room 6 IJCNN Paper Reinforcement Learning VI Session Chair: Sheng Wang (Guangdong Institute of Intelligence Science and Technology), Huajin Tang (Zhejiang University) Randomized Spectral Options: Option Discovery via Random Reward Probing Yiming Fei and Lang Qin (Zhejiang University), Rui Yan (Zhejiang University of Technology), and Huajin Tang (Zhejiang University) Abstract Abstract Discovering diverse and useful options is key to hierarchical reinforcement learning, yet many spectral approaches rely on Laplacian decomposition and eigenfunction learning, making it costly to scale option sets and difficult to reuse the learned representations. We propose Randomized Spectral Options (RSO), an RL-native alternative that generates smooth option-discovery potentials by randomized reward probing through successor representations/features. RSO samples a small batch of orthogonal random reward probes via QR decomposition, learns an successor representation backbone under an exploration policy, and obtains a family of randomized value-function potentials through a single readout from the learned representation. Each potential is then turned into a potential-based shaping reward to train signed options that ascend/descend the induced potential field. On MiniGrid, RSO achieves state coverage and downstream performance competitive with Laplacian-representation-based options, and Atari 2600 games visualizations suggest that the learned potentials emphasize semantically meaningful bottlenecks. COFS: Counterfactual Optimization for Reinforcement Learning Feature Selection Zixing Chen, Xiaohan Huang, and Yuanchun Zhou (Computer Network Information Center, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Abstract Abstract Reinforcement learning (RL) has recently been used for automated feature selection by exploring the exponentially large subset space. However, most RL-based methods update policies with a single subset-level reward, which yields coarse feedback and is insufficient for attributing rewards to individual features under complex interactions. We propose Counterfactual Optimization for Reinforcement Learning Feature Selection (COFS), which improves RL-based feature selection by providing feature-wise counterfactual advantages for policy optimization. Specifically, COFS trains a critic to estimate rewards and computes a counterfactual advantage for each selected feature. This is achieved by contrasting the predicted reward of the current subset with that of the leave-one-out variants formed by removing each feature, effectively distinguishing important features from redundant ones without repeatedly invoking the downstream model. Moreover, we introduce a pretraining stage that learns subset state representations based on reward signals rather than solely on data distribution. This strategy enhances the critic's prediction accuracy and facilitates more effective policy optimization. Finally, extensive experimental results demonstrate the effectiveness of our proposed method, showcasing notable enhancements in subset quality and compactness. Radar Jamming Waveform Design Based on Hierarchical Reinforcement Learning and Cooperative Perception Lijun Gao, Mengdi Zhao, Yanjie Wang, Rui Liu, Xiyan Cao, and Yanhui Liu (Shenyang Aerospace University, School of Computing) Abstract Abstract Reinforcement Learning (RL) has attracted wide attention in many fields due to its dynamic decision-making and continuous optimization capabilities, but its practical application in radar jamming waveform generation remains limited. Traditional RL-based jamming algorithms face two core challenges:first, low sample efficiency makes it difficult to converge to the optimal strategy adapting to radar dynamic changes; second, poor adaptability of jamming waveforms fails to match radar signal characteristics in real time. Additionally, designing a reward function that accurately describes the radar-jammer interaction mechanism is highly challenging. This paper proposes an online jamming waveform generation framework: three time frequency analysis methods (Smooth Pseudo Wigner-Ville Distribution, Fourier Synchrosqueezing Transform, and Variational Mode Decomposition-based Hilbert-Huang Transform) are used to convert complex multi-class radar signals under low signal to-noise ratio (SNR) into three-channel time-frequency feature maps. Combined with Pulse Descriptor Words (PDW) to form a dual-stream input, this framework solves the problem of low radar signal type recognition accuracy under low SNR and serves as interpretable prior knowledge for RL. Under the hierarchical RL architecture, radar Power Spectral Density (PSD) matching degree is used as the core constraint to guide the agent for efficient directed exploration, improve sample efficiency, and enable learning more adaptive dynamic jamming strategies in complex agile environments. To further evaluate jamming effectiveness in non-cooperative scenarios, a composite reward function integrating frequency-domain and time-domain metrics is designed, combined with closed-loop optimization of the radar signal processing chain. DeFi Intent Discovery via Maximum Entropy Inverse Reinforcement Learning Aizierjiang Aiersilan and Jerome Yen (University of Macau) and Sheng Wang (Guangdong Institute of Intelligence Science and Technology) Abstract Abstract Decentralized Finance (DeFi) transaction sequences can obscure a user’s high-level goal behind multi-step smart-contract calls, routing abstractions, and rapidly changing market conditions. Many prior intent-mining pipelines rely on transaction-level semantic labels, which can miss the sequential decision structure of long-horizon strategies. We formulate DeFi intent discovery as sequential reward inference and learn a parametric reward $R_\theta(s,a)$ from expert demonstrations via Maximum Entropy Inverse Reinforcement Learning (MaxEnt IRL). To capture short-horizon market trends under volatility and partial observability, we augment the state with a temporal market-gradient term $\nabla_t\mathbf{m}$, defined as the block-to-block finite difference of market-state variables. We refer to this state design as Chemo-IRL. To enable controlled synthetic evaluation under intent hiding, we introduce Gym-DeFi, a Gymnasium-compatible simulator with a configurable action-obfuscation channel that corrupts observed action identifiers during data generation. We evaluate reward recovery using Reward Recovery Error (RRE) and downstream intent probing using macro-F1 from a fixed post-hoc attribution-based decoder. On this controlled synthetic benchmark under high obfuscation, Chemo-IRL attains the lowest RRE among intent-label-free baselines while remaining competitive on macro-F1 within the same label-free setting. These results should be interpreted as benchmark-level comparisons in a controlled synthetic environment rather than as direct evidence of real-world DeFi deployment performance. Wednesday Virtual Room 7 IJCNN Paper Reinforcement Learning VII Session Chair: Hao Wang (University of Science and Technology of China), Song Liu (Qilu University of Technology (Shandong Academy of Sciences)) RT-UIE: A Reward-Guided and Task-Oriented Framework for Unified Information Extraction Wentong Wan, Yingxue Qiao, Yiwen Zhang, Yu Du, and Song Liu (Qilu University of Technology (Shandong Academy of Sciences)) Abstract Abstract Information extraction (IE) faces challenges from task-specific label spaces and heterogeneous data structures. Recent advances in large language models (LLMs) and retrieval-augmented generation (RAG) have enabled unified information extraction across diverse tasks, but existing methods still suffer from several limitations. Conventional RAG does not adequately account for IE-specific preferences, resulting in the underutilization of the background knowledge obtained. Furthermore, most approaches rely on fixed instruction templates and lack adaptive mechanisms for task-aware prompting. In addition, long instruction sequences often cause attention drift, weakening the model’s focus on core extraction targets. To address these problems, we propose RT-UIE, a reward-guided and task-oriented unified information extraction framework. In our framework, we design a task-oriented RAG retrieval and knowledge reconstruction mechanism that structurally reorganizes retrieved background knowledge, improving the consistency between background knowledge and extraction preferences. Moreover, we design a chain-of-thought reward model that dynamically adjusts the depth and quality of the reasoning according to task characteristics, enabling more effective reasoning control. Additionally, a module for instruction and attention calibration is designed to mitigate attention drift and improve the focus of the model on key task instructions. We evaluated our model on the IE_INSTRUCTION, WNUT-17, DocRED, and MAVEN datasets and compared it with other baseline models. A Reward-Guided Semantic Decoupling Framework for Zero-Shot Multi-Intent Detection Chuanwang Xiong and Hongwei Ge (Jiangnan University) Abstract Abstract Unsupervised disentanglement of high-dimensional semantic spaces remains a critical bottleneck in zero-shot multi-intent detection. While Large Language Models (LLMs) possess vast knowledge, they suffer from attentional dilution and semantic hallucinations when processing dense, intertwined intents. Existing approaches often rely on heuristic prompt engineering, treating models as opaque black boxes and lacking explicit inference verification. To transcend these limitations, this paper proposes Reward-Guided Semantic Decoupling (RGSD), a parameter-efficient inference framework designed to unlock the complex reasoning potential of open-source models without fine-tuning. Departing from standard end-to-end generation, RGSD restructures the task into a transparent "Locate-Decouple-Reason-Verify" cognitive cycle: (1) A lightweight reward model constructs a semantic saliency map to minimize information entropy and lock onto core features; (2) Structured subspace projection enforces semantic orthogonality, decomposing mixed contexts into independent units to physically block feature interference; (3) Parallel reasoning is conducted within isolated subspaces, followed by a closed-loop global consistency check to audit logical self-consistency. Experiments on MixATIS and MixSNIPS demonstrate that RGSD, utilizing only a 14B parameter model, achieves zero-shot performance comparable to GPT-4. This approach validates the "Explicit Decoupling + Guided Reasoning" paradigm, offering a generalizable theoretical solution for enhancing the interpretability and robustness of large models in complex structured prediction. Boosting Triple Extraction and Relation Normalization with Trained Retrieval and Policy Optimization Yu Bai and Jiansheng Bian (School of Computer Science, Shenyang Aerospace University) and Jianjun Chen (Shenyang Northern Software College) Abstract Abstract Triple extraction and relation normalization are essential for constructing usable knowledge graphs from text, yet deployable systems often struggle to meet strict usability constraints such as stable formatting, reliable parseability, and consistent relation naming. While prior lightweight large language models (LLMs) pipelines improve robustness through ensembling and definition-aware normalization, they typically rely on heuristic selection and remain misaligned with set-level evaluation metrics. We propose a method that strengthens deployable extraction pipelines along two complementary axes: first, trained retrieval for schema guidance and second, policy optimization for metric-level alignment. We train a schema retriever to embed texts and relations into a relevance-oriented vector space, enabling efficient retrieval of relations likely to appear in the input, including long-tail and low-salience relations. The retriever is fine-tuned from a strong instruction-tuned embedding model using an InfoNCE objective, with positive text-relation pairs and negative relations. We then formulate structured triple generation as a policy optimization problem and fine-tune a lightweight actor with parameter-efficient adapters using Proximal Policy Optimization (PPO) style updates. A hybrid reward is designed, combining a format/parseability signal, a set-level gold-matching score, and a baseline-relative improvement term to encourage gains without regressions. Experiments on standard triple extraction benchmarks show that our approach consistently improves upon a strong baseline across multiple matching criteria, with especially noticeable gains under stricter settings. These results demonstrate that combining trained retrieval with reward-aligned policy optimization is an effective and practical approach to improving both triple quality and consistent relation normalization. IE as Cache: Information Extraction Enhanced Agentic Reasoning Hang Lv (University of Science and Technology of China); Sheng Liang (Huawei Technologies Co., Ltd); Hongchao Gu (University of Science and Technology of China); Wei Guo (Huawei Technologies Co., Ltd); Defu Lian (University of Science and Technology of China); Yong Liu (Huawei Technologies Co., Ltd); and Hao Wang and Enhong Chen (University of Science and Technology of China) Abstract Abstract Information Extraction (IE) is traditionally treated as a terminal objective for converting unstructured content into structured representations. However, this static endpoint view underutilizes IE's potential in dynamic reasoning scenarios. To bridge this gap, we propose IE-as-Cache, a framework that repurposes IE as a dynamic, read-write cognitive cache to enhance agentic reasoning. Critically noticing that standard agents struggle with noise and information decay by treating raw text as a static read-only block, our approach adopts a hierarchical memory design to decouple working memory from raw inputs. Specifically, we implement a schema-decoupled, query-driven extraction mechanism to initialize a compact cache, and maintain it via a cache-aware reasoning loop that performs on-demand extraction and dynamic updates. Consequently, this mechanism enables the agent to actively filter noise and reason over a high-density representation of salient evidence, avoiding the redundancy of repeatedly scanning raw text. Experiments on challenging benchmarks spanning question answering, planning, and query-focused summarization across diverse LLMs demonstrate significant improvements in reasoning accuracy. This study indicates that IE can be effectively repurposed as a reusable cognitive resource, offering a promising direction for future research in downstream applications of information extraction. Wednesday Virtual Room 8 IJCNN Paper Reinforcement Learning VIII Session Chair: Zhengyu Ma (China Mobile Communication Group Tianjin Co., Ltd.; Inspur Group), Jiang Haiqing (Guangzhou City University of Technology) StabTune: A Three-Phase Collaborative Framework for Reliable Microservice Kernel Parameter Optimization Haiqing Jiang, Han Liao, Jiaying Li, and Jing Chang (Guangzhou City University of Technology) Abstract Abstract Deep reinforcement learning (DRL) is promising for automated kernel-parameter tuning in microservice deployments, yet its training often exhibits high variance across runs, which undermines service-level reliability. We propose StabTune, a three-phase collaborative training framework that explicitly targets the stability–efficiency trade-off. StabTune decomposes the learning process into: (i) broad exploration via Prioritized DQN to cover the state space, (ii) fast workload adaptation via Model-Agnostic Meta-Learning (MAML), and (iii) stability-oriented hyperparameter refinement using Bayesian Optimization (BO) to ensure low-variance convergence. We further introduce a stability-aware reward that penalizes the coefficient of variation (CV) over a fixed evaluation horizon. Experiments on Linux kernel tuning under three microservice workload modes show that StabTune improves throughput by 13.8\%–25.7\% over competitive baselines while reducing cross-run CV from 17.48\% to 2.45\%. Crucially, while standard Bayesian Optimization yields competitive performance, StabTune achieves millisecond-level inference latency suitable for real-time auto-scaling. Generative Model-based Exploration for Sample-Efficient Reinforcement Learning xinyan cai, shuwei wang, and qiang guan (Institute of Automation) Abstract Abstract Sample efficiency is a critical challenge in deep reinforcement learning (DRL), driving the development of numerous exploration strategies. However, efficient exploration in high-dimensional observation spaces remains difficult. To address this, we introduce Generative Model-based Exploration (GME), a novel algorithm grounded in world models. Specifically, we characterize state transitions as transformations among latent probability distributions, upon which we learn a generative world model. Our exploration term is defined using mutual information and epistemic uncertainty derived from this model. Experimental results demonstrate that GME substantially enhances the sample efficiency of baseline methods, yielding superior performance and robust generalization in unseen environments. Furthermore, extensive experiments validate that the environmental dynamics insights provided by the world model significantly facilitate exploration. DynRCA: Dynamic Sampling for Efficient Multi-Agent Reasoning in Microservice RCA Yuke Liu, Yitao Xiao, Guoming Yang, Shaoyong Guo, and Feng Qi (State Key Laboratory of Networking and Switching Technology, Beijing University of Posts and Telecommunications, Beijing, China) Abstract Abstract The intricate invocation dependencies and fault cascades within microservice architectures pose significant challenges for Large Language Model (LLM)-based Root Cause Analysis (RCA). Prevailing methodologies are often constrained by static sampling, which fail to adapt reasoning width to fault complexity, leading to redundant trajectory exploration and accumulation of decisional noise. To address these limitations, we propose DynRCA, a multi-agent collaborative framework featuring adaptive reasoning trajectory optimization tailored for root cause analysis of microservice environments. DynRCA introduces a dynamic width sampling mechanism that leverages high-dimensional embeddings and density-based clustering to identify and consolidate redundant reasoning steps in real-time. This approach effectively prunes the search space while preserving diverse, high-value diagnostic paths. Experimental evaluations on the 2025 AIOps Challenge and RE2-OnlineBoutique datasets demonstrate that DynRCA achieves high accuracy of 64.25% while reducing reasoning overhead by 28.5%, striking an optimal balance between diagnostic depth and computational efficiency. Reinforcement Learning for Retrieval Path Optimization in RAG: Tackling Sparse Feedback and Distributional Shift via Dynamic State Encoding and Meta-Policy Adaptation Zhengyu Ma (China Mobile Communication Group Tianjin Co., Ltd.; Inspur Group); Xiaojia Jin and Dongming Zhao (China Mobile Communication Group Tianjin Co., Ltd.); Bo Wang (Tianjin University); and Yongzong Wang (Tianjin University, Inspur Group) Abstract Abstract Retrieval-Augmented Generation (RAG) combines document retrieval with generation models, and its performance depends on \emph{how} evidence is aggregated across multi-step retrieval rather than a single top-k list. We focus on reinforcement learning (RL)-based retrieval \emph{path} optimization, addressing two key challenges: sparse, delayed long-horizon rewards from downstream QA tasks, and generalization degradation under corpus distribution shift. We propose a framework with two modules: Dynamic Multi-Scale State Encoding (DMSE) learns temporal state representations by hierarchically encoding retrieval trajectories (via BiGRU, self-attention, and descriptive statistics, fused via cross-scale attention for phase-aware decisions); Meta-Policy Guided Path Adaptation (MPGPA) uses a lightweight meta-controller to track corpus signals, adjusting exploration and ranking parameters to stabilize model behavior. Experiments on seven open-domain QA benchmarks show significant improvements: our approach outperforms baselines including RAFT~\cite{ref_raft2024} by 2.9–19.6 EM points (7.9 on average), reduces performance variance by 51%, and enhances exploration diversity. Ablation studies validate DMSE and MPGPA’s efficacy. Wednesday Virtual Room 1 IJCNN Paper Remote Sensing and Earth Observation I Session Chair: Ziming Song (Fuyang Normal University), Wei Wang (Changsha University of Science and Technology) Cascaded Spectral-Spatial Enhancement and Consistency Guided Interaction for Remote Sensing Change Detection Ziming Song, Xueqin Liu, JIngpeng Yu, and Guocheng Yan (Fuyang Normal University) Abstract Abstract Remote sensing change detection (RSCD) aims to identify semantic changes from bi-temporal imagery but is often constrained by seasonal variations and structural misalignments. Existing methods struggle to balance noise suppression with geometric integrity. Their reliance on ``blind pixel-level matching'' conflates semantic invariance with observational reliability, leading to boundary false positives. To address these challenges, this paper proposes a unified framework based on Cascaded Spectral-Spatial Enhancement (CSSE) and Spectral Consistency-Guided Interaction (SCGI). First, we introduce the CSSE module. By integrating dynamic spectral gating with a local-global dual attention mechanism, CSSE explicitly models what land covers are and how they are spatially organized, effectively filtering non-semantic artifacts while maintaining structural coherence. Second, to handle misalignments, we propose the SCGI strategy. This strategy decouples global spectral consistency priors from local observation confidence, enabling robust change identification even under viewpoint shifts or partial occlusion. Finally, a Hierarchical Decoder (HD) is employed to progressively aggregate multi-scale cues for boundary refinement. Extensive experiments on the LEV, Google, and BCDD datasets demonstrate that our method outperforms state-of-the-art approaches, achieving superior robustness and accuracy with IoU scores of 79.65\%, 73.40\%, and 81.02\%, respectively. Bridging Semantic Filtering and Structural Alignment for Remote Sensing Change Detection Ziming Song, Xueqin Liu, JIngpeng Yu, and Guocheng Yan (Fuyang Normal University) Abstract Abstract Remote sensing change detection aims to identify semantic changes by analyzing bi-temporal images. However, precision is constrained by semantic interference from seasonal or lighting variations and pseudo-changes caused by structural misalignments like viewpoint shifts. Existing methods often struggle to balance these issues, either employing expensive attention mechanisms to suppress noise or overlooking boundary false positives resulting from geometric deviations. In this paper, we propose a novel framework bridging semantic filtering and structural alignment. First, to overcome semantic interference, we introduce the Mono-temporal Semantic Purification (MSP) method. It synergizes a Spectral-Aware Spatial Filtering strategy to explicitly suppress background redundancy with a Semantic Filtering-Aware Enhancement (SFAE) mechanism. Unlike standard attention, SFAE leverages a filtering-aware prior to guide efficient global context modeling, enhancing semantic consistency while preserving high-frequency details. Second, to rectify geometric deviations, we propose the Bi-temporal Graph Alignment (BTGA) module. By constructing bidirectional KNN graphs in non-Euclidean space, this module dynamically registers bi-temporal features. Furthermore, a secondary purification strategy is applied to the fused difference representations, effectively eliminating boundary artifacts to ensure precise change delineation. Extensive experiments on the LEV, Google, and BCDD datasets demonstrate that our method outperforms state-of-the-art approaches, achieving superior robustness and accuracy with IoU scores of 79.50\%, 77.12\%, and 81.84\%, respectively. SADE-Net: Strip-Aware Difference Enhancement and Multi-Dilation Fusion for Remote Sensing Change Detection Zhenyu Wang, Guoxia Wang, and Gang Shi (Xinjiang University, College of Computer Science and Technology) Abstract Abstract Remote sensing change detection (CD) is a funda- mental task in Earth observation, serving critical roles in urban planning, disaster damage assessment, and environmental mon- itoring. Despite the remarkable success of deep learning-based methods, accurately detecting changes in linear features (e.g., roads, bridges, and pipelines) remains a significant challenge. Existing state-of-the-art methods often struggle with two primary limitations: Feature Fragmentation: Traditional square-kernel pooling operations fail to capture the anisotropic geometry of narrow, elongated targets, leading to disconnected or broken predictions. The receptive fields of these kernels are often filled with irrelevant background context, diluting the feature representation of the linear targets. Insufficient Context Aggregation: Simple feature fusion strategies lack the multi-scale receptive fields necessary to distinguish real semantic changes from pseudo-changes caused by seasonal variations, illumina- tion differences, or sensor noise. To address these issues, we propose a novel Siamese network framework named SADE-Net. Specifically, we introduce a Strip-Aware Difference Enhancement (SADE) module that integrates horizontal and vertical strip pooling mechanisms to explicitly model long-range anisotropic dependencies, effectively preserving the structural connectivity of narrow, elongated changes. Furthermore, a Multi-Dilation Spatial Fusion (MDSF) module is introduced at the encoder- decoder bottleneck to aggregate multi-scale contextual informa- tion via parallel dilated convolutions, replacing computationally expensive recurrent units (e.g., LSTMs) used in prior Siamese CD networks. Extensive experiments on two benchmark datasets, LEVIR-CD and SYSU-CD, demonstrate that SADE-Net achieves state-of-the-art performance. On the LEVIR-CD dataset, our method achieves an F1-score of 90.85% and an IoU of 83.24%, significantly outperforming recent baselines while maintaining exceptional computational efficiency with only 1.63M parameters. Activated-SAM for Remote Sensing Change Detection Network Wei Wang, Aocheng Shu, and Xin Wang (Changsha University of Science and Technology) Abstract Abstract Due to the complex land cover types typically present in remote sensing images, it is difficult to achieve more precise detection of changed areas. Inspired by the Segment Anything Model (SAM), a change detection network named ASFCD is proposed incorporating an activated SAM, which utilizes a novel SAM encoder that incorporates NewActNet layers in place of the original MLPs. This encoder, termed Activated-SAM, enables dynamic feature filtering with a 44.54% reduction in parameter count compared to the original SAM encoder. Specifically, a dual-encoder architecture integrating Activated-SAM and VGG is proposed to extract features at different resolutions. Through the SAM-assisted Global-Local Fusion Module (SAGL), both dual-encoder feature fusion and bitemporal feature fusion, finally generate the output layer by layer. ASFCD was extensively evaluated on the public change detection datasets LEVIR-CD and WHU-CD, and its generalization capability was further validated on the SYSU-CD dataset. On the LEVIR-CD dataset, ASFCD achieved an F1-score of 92.18% and an IoU of 85.50%, outperforming the second-best model by 0.42% and 0.72%. On the WHU-CD dataset, it attained an F1-score of 94.84% and an IoU of 90.18%, exceeding the second-best model by 1.58% and 2.8%, achieving state-of-the-art performance. Our code can be seen at https://github.com/zcbg123/ASFCD.git. Wednesday Virtual Room 2 IJCNN Paper Remote Sensing and Earth Observation II Session Chair: Edson Tavares (UESTC), Juan Luo (HUNAN UNIVERSITY) MVRNet: A Multi-View Fusion Network for Road Extraction from Satellite and UAV Imagery Wenfei Zhao, Juan Luo, Kexuan Feng, Ying Qiao, and Shuyang Teng (HUNAN UNIVERSITY) Abstract Abstract Accurate road extraction from remote sensing imagery serves as a critical data foundation for urban and rural planning and intelligent transportation systems. However, existing methods based on single view often struggle with missed detections and fragmented road networks caused by occlusions and shadows. To address these challenges, this paper constructs a paired satellite-UAV road extraction dataset called SURoad and proposes a multi-view road extraction network named MVRNet. The proposed method adopts a layered collaborative modeling mechanism, which organically integrates shallow-level multi-view feature alignment for detail enhancement with deep-level satellite topological priors for semantic guidance. Through an adaptive weighting strategy, the model achieves adaptive integration modeling of global structural consistency and local road details. Experiments on the SURoad dataset show that MVRNet outperforms existing methods on key metrics such as F1 score and IoU, surpassing the best baseline by approximately 1.3% in IoU. These results confirm the effectiveness of the layered collaborative modeling strategy in handling complex occlusions and feature misalignments. Laplacian Frequency Interaction Network for Rural Thematic Road Extraction baiyan Chen and Weixin Zhai (China Agricultural University) Abstract Abstract Rural thematic road network construction aims to extract topological road structures from movement trajectory images of agricultural machinery. However, this task faces challenges where downsampling methods commonly used in existing studies tend to blur the sparse high-frequency road structures, and the heavy noise from dense field operations often leads to fragmented or redundant topologies in the extracted networks. To address these challenges, we propose LFINet, a Laplacian Frequency Interaction Network. The network begins with a Laplacian Multi-scale Separator (LMS) to decouple the image into low-frequency semantic contexts and high-frequency structural details. These components are then processed by the Cross-Frequency Interaction Block (CFIB) through a dual-pathway architecture in which a High-Frequency Block (HFB) refines local structures while a Spatial Transformer (ST) captures global semantics. Subsequently, a Frequency Gated Modulation (FGM) mechanism integrates the features from pathways by leveraging semantic contexts to calibrate the structural details. Finally, a Progressive Reconstruction Decoder iteratively fuses multi-scale features to ensure topological consistency. Experiments conducted on a real-world agricultural trajectories dataset from Henan Province, China, show that LFINet establishes a new state-of-the-art. Specifically, it achieves an F1-score of 92.54% and an IoU of 86.12%, surpassing the second-ranked method by 0.64% and 1.1%, respectively. This confirms its capability to effectively construct topological road networks from noisy and sparse field data. Stereoscopic Dual-Stream Pseudo-Siamese Network for Cloud Top Height Retrieval from Heterogeneous Dual-Satellite Data zhaoran fan, ming wu, chuang zhang, and mingrui xu (Beijing University of Post and Telecommunications) Abstract Abstract Accurate retrieval of Cloud Top Height (CTH) is critical for meteorological analysis and climate monitoring. However, traditional retrieval methods relying on single geostationary satellites often suffer from limited viewing angles, lack of stereoscopic information, and sensor-specific spectral limitations. Operational observations indicate that FY-4A exhibits superior performance in retrieving high-altitude clouds, particularly cirrus, whereas FY-4B demonstrates higher accuracy in mid-to-low cloud regimes. To leverage these complementary strengths and address single-view limitations, this paper proposes a novel Stereoscopic Dual-Stream Pseudo-Siamese Network that fuses heterogeneous data from the FY-4A and FY-4B satellites. In terms of network architecture, we design a Self-Generated Prior Guidance (SPG) module to extract global contextual priors, enabling the alignment of heterogeneous features at the early encoding stage. Furthermore, a Cross-Attention Bridge (CAB) is incorporated to explicitly model the geometric relationships between the two satellite views, utilizing attention mechanisms to rectify parallax-induced discrepancies and enhance feature interaction. Extensive experiments, conducted using a dataset of synchronous FY-4A/B and CALIPSO observations collected from June 2022 to April 2023 over the Yellow and Bohai Seas, demonstrate that our dual-satellite fusion approach significantly outperforms single-satellite baselines. The proposed model achieves a Mean Absolute Error (MAE) of 1.41 km and a Pearson correlation coefficient of 0.802, proving the effectiveness of deep learning-based stereoscopic fusion for high-precision atmospheric parameter retrieval. AEGIS-GNN: Verifiable Multi-Relational Graph Learning for Secure LEO Satellite Operations Edson Eliezer da Silva Tavares, Qi Xia, Hu Xia, Jianbin Gao, and Adjei-Arthur Bonsu (University of Electronic Science and Technology of China) Abstract Abstract Determining the ground truth from conflicting satellite reports is a critical challenge in Low Earth Orbit (LEO) mega-constellations. Existing Graph Neural Network (GNN) approaches model satellite networks as homogeneous graphs, discarding the rich multi-modal relationships that govern trust. This paper presents AEGIS-GNN, a multi-relational heterogeneous GNN architecture for satellite truth discovery. Our framework makes three contributions: (1) a multi-edge graph construction capturing four complementary relationship types: communication links, spatial proximity, feature similarity, and orbital plane membership, that encode distinct trust-relevant modalities; (2) a hybrid residual GNN stack that alternates spectral (GCN) and spatial (GraphSAGE) convolutions with a novel learned residual gating mechanism that adaptively blends new and prior representations per feature dimension; and (3) relation-specific multi-head attention that learns differentiated neighbor weighting across edge types. We further integrate a lightweight blockchain anchoring layer for operational verifiability. Experiments on 1,008 satellite nodes with 557,158 multi-type edges across 5 random seeds demonstrate 92.47 pm 0.41% accuracy, 0.9568 pm 0.005 F1-score, and 93.85 pm 0.52% AUC, significantly outperforming six baselines including HGT (+3.5%), R-GCN (+4.4%), and HAN (+3.8%). Comprehensive ablation studies validate each component, with learned residual gating contributing the largest single gain (+8.0%) and feature similarity edges providing +7.1% accuracy improvement. Sensitivity analyses across key hyperparameters confirm the robustness of the proposed architecture. Wednesday Virtual Room 3 IJCNN Paper Remote Sensing and Earth Observation III Session Chair: Chengliang Wang (Chongqing University), Shijie Zhang (Northeastern University) Spectral-Aware Text-to-Time Series Generation with Billion-Scale Multimodal Meteorological Data Shijie Zhang (Northeastern University) Abstract Abstract Text-to-time-series generation is particularly important in meteorology, where natural language offers intuitive control over complex, multi-scale atmospheric dynamics. Existing approaches are constrained by the lack of large-scale, physically grounded multimodal datasets and by architectures that overlook the spectral–temporal structure of weather signals. We address these challenges with a unified framework for text-guided meteorological time-series generation. First, we introduce MeteoCap-3B, a billion-scale weather dataset paired with expert-level captions constructed via a Multi-agent Collaborative Captioning (MACC) pipeline, yielding information-dense and physically consistent annotations. Building on this dataset, we propose MTransformer, a diffusion-based model that enables precise semantic control by mapping textual descriptions into multi-band spectral priors through a Spectral Prompt Generator, which guides generation via frequency-aware attention. Extensive experiments on real-world benchmarks demonstrate state-of-the-art generation quality, accurate cross-modal alignment, strong semantic controllability, and substantial gains in downstream forecasting under data-sparse and zero-shot settings. Additional results on general time-series benchmarks indicate that the proposed framework generalizes beyond meteorology. We will release the MeteoCap-3B dataset and code upon acceptance. A Unified Framework for Meteorologically Physical Parameter-Driven Weather Generation Jiantao yu, Diyue Zhang, Ting Luo, Xiangyu Wu, Jing Hu, Xia Yuan, Xi Wu, and Feihu Huang (Chengdu University of Information Technology) Abstract Abstract Converting continuous meteorological data into controllable visual scenes remains challenging due to scarce supervision and unstable physical control. We propose a collaborative solution bridging simulation and generation. We introduce WeatherImage, a dataset of 20,000 pixel-aligned pairs providing the first continuous physical supervision across five weather conditions. We develop WeatherDiff to address two key bottlenecks: a Meteorological Semantic Adapter overcomes semantic insensitivity to fine-grained parameters, while a Structure-First Dynamic Optimization Strategy resolves the texture-structure conflict. Experiments demonstrate significant improvements in FID and CLIP , critically, the first monotonic response to continuous meteorological inputs. This enables quantitatively controllable, photorealistic weather synthesis for autonomous driving and disaster management. FoumCast: Adaptive Fourier-Meteorological Disentanglement and Diffusion for High-Fidelity Precipitation Nowcasting Yangyang Xu, Xianwei Meng, and Lin Jia (Hefei Institutes of Physical Science, Chinese Academy of Sciences; University of Science and Technology of China) Abstract Abstract Accurate precipitation nowcasting is critical for disaster mitigation and urban resource management. However, existing deep learning models often struggle to balance spatial precision with temporal consistency, typically yielding over-smoothed predictions that lack fine-grained meteorological details. To address these challenges, we propose FoumCast, a novel frequency-domain guided framework that synergizes adaptive spectral-spatial disentanglement with conditional diffusion modeling. In the first stage, a deterministic component is optimized via the Fourier Amplitude and Correlation Loss. By explicitly supervising both spectral magnitude and phase, FoumCast accurately preserves rainfall intensity and global spatial positioning. A learnable spectral mask is introduced to adaptively partition the preliminary forecasts into low-frequency advective envelopes (representing macro-scale evolution) and high-frequency convective textures (capturing micro-scale dynamics), which are then integrated through a Cross-Attention Fusion Module for multi-scale temporal alignment. In the second stage, a conditional diffusion network leverages these spectral-aware features as priors to model residual uncertainties. Extensive evaluations on four benchmark radar datasets, demonstrate that FoumCast outperforms state-of-the-art baselines in terms of predictive accuracy, structural realism, and physical consistency. APF-GS: Adaptive Prior Fusion and Compositional Reconstruction for Dynamic Humans in the Wild Tao Cheng, Chengliang Wang, and Chaoyue Chen (Chongqing University) and Jianbin Chen (Chongqing Changan) Abstract Abstract High-fidelity reconstruction of dynamic pedestrians in 3DGS-based driving scenes remains challenging, as existing approaches often produce results that lack high-frequency geometric details and involve structural conflation of humans with their accessories. To address these issues, we propose APF-GS, a unified framework for high-fidelity, structured reconstruction. The Adaptive Prior Fusion (APF) module recovers high-frequency geometric details through a neural residual deformation field, reconciling macro-level geometric correction with micro-level detail generation via dynamic regularization. The Compositional Reconstruction (CR) module mitigates structural conflation by self-supervised decoupling of supervision signals, enabling cleaner human-accessory separation. It further introduces a parent-child kinematic linkage to ensure physically consistent interactions during motion.Extensive evaluations on challenging dynamic scenarios from the Waymo Open Dataset demonstrate that APF-GS outperforms existing methods in reconstruction fidelity and robustness, yielding sharper geometric details and structurally consistent interactions. Specifically, it achieves a PSNR of 33.46 on human-centric regions, a significant improvement of +6.73 dB over existing baselines. Wednesday Virtual Room 4 IJCNN Paper Retrieval-Augmented Generation III Session Chair: yu pei (Xinjiang University), 宁 张 (ShangHai University) AMRAG: Adaptive Multi-step Retrieval-Augmented Generation for Complex Question Answering Ning Zhang, Xinzhi Wang, Xiangfeng Luo, and Jianqi Gao (ShangHai University) Abstract Abstract Retrieval-augmented generation (RAG) is widely recognized as an effective paradigm for mitigating hallucinations in large language models (LLMs). For complex queries requiring multi-step reasoning, iterative retrieval is commonly employed to progressively gather evidence across reasoning steps. However, indiscriminate retrieval at every step not only introduces irrelevant or misleading information that impairs downstream reasoning, but also incurs substantial computational overhead. We address these challenges by leveraging a well-established observation: LLMs exhibit stronger factual accuracy on high-frequency knowledge from pretraining data, while low-popularity knowledge is more prone to hallucinations. Building on this insight, we propose AMRAG (Adaptive Multi-step Retrieval-Augmented Generation), a framework that selectively determines retrieval necessity at each reasoning step by estimating the popularity of key entities in the current sub-question. Retrieval is invoked only when the model is unlikely to possess the required knowledge. Retrieved documents are further filtered and reranked by a lightweight semantic reranker to retain only evidence closely aligned with the sub-question. Experiments on HotpotQA and other benchmarks demonstrate that AMRAG outperforms strong baselines while reducing computational costs. ClauseRouteRAG: A Clause-Level Retrieval-Augmented Generation Framework for 3GPP Telecommunications Standards Shan Lu (China Electronics Technology Group Corporation, The 54th Research Institute); Xiayi Wang (Xidian University, School of Artificial Intelligence); Zhanying Shao (Hebei Construction Material Vocational and Technical College); Jing Bai (Xidian University, School of Artificial Intelligence); Zhu Xiao (Hunan University, College of Computer Science and Electronic Engineering); and Yasheng Zhang (China Electronics Technology Group Corporation, The 54th Research Institute) Abstract Abstract The 3rd Generation Partnership Project (3GPP) produces technical specifications with long documents, hierarchical clause structures, and many tables, which makes automated question answering difficult. While Retrieval Augmented Generation (RAG) has improved robustness for knowledge intensive tasks, existing systems can still miss the right clauses in long specifications and may retrieve incomplete evidence for standards questions. This paper presents ClauseRouteRAG, a RAG framework for multiple choice question answering over 3GPP specifications. ClauseRouteRAG first selects a small set of relevant clauses by aggregating evidence from an initial retrieval pass. It then retrieves passages within these clauses using dense retrieval and BM25 with rank fusion. Finally, it reranks the candidate passages using overlap of technical identifiers extracted from the question and the retrieved passages, and it prompts the language model to answer using only the provided evidence. Experiments on TeleQnA Release 17 and Release 18 show that ClauseRouteRAG consistently improves over LLM only baselines and a naive RAG pipeline, reaching an overall accuracy of 0.849 on TeleQnA under the default setting. ReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation Jinxin Liu and Zhicheng Lee (Tsinghua University); Shulin Cao (Tsinghua University, z.ai); Jiajie Zhang (Tsinghua University); Weichuan Liu and Xiaoyin Che (Siemens AG); and Lei Hou and Juanzi Li (Tsinghua University) Abstract Abstract Large Reasoning Models (LRMs) exhibit strong reasoning abilities, but their reliance on parametric knowledge limits factual accuracy. Recent works adopt reinforcement learning (RL) training to integrate reasoning with retrieval, but such methods are often complex and resource-heavy. We propose ReaRAG, a factuality-enhanced reasoning model trained via strategic distillation, rather than costly RL training. Our solution includes a novel data construction framework with an upper bound on the reasoning chain length. Specifically, we first leverage a LRM to generate deliberate thinking, then select an action from a predefined action space (Search and Finish). For Search action, a query is executed against the RAG engine, where the result is returned as observation to guide reasoning steps later. This process iterates until a Finish action is chosen. Experimental results show that ReaRAG achieves the best overall performance across six question answering (QA) benchmarks. Further analysis highlights its strong reflective ability to recognize errors and revise its reasoning trajectory. Our study enhances LRMs’ factuality while effectively integrating robust reasoning for Retrieval-Augmented Generation (RAG). DMG-RAG: Dynamic Multi-Grained Retrieval Augmented Generation for Multi-Hop QA Yu Pei, Mieradilijiang Maimaiti, and Yi Chen (Department of Computer Science, Xinjiang University) and Wu Le, Zhuofei Xie, and Jiawei Chen (Company: Integrated Laboratory for Space, Air, and Ground Systems) Abstract Abstract Retrieval-Augmented Generation (RAG) improves multi-hop question answering by grounding large language models (LLMs) in retrieved evidence, yet its effectiveness is highly sensitive to chunk granularity and context construction under a fixed token budget. Sentence-level chunks can be precise but often fragment evidence chains, while paragraph- and document-level chunks provide broader coverage at the risk of introducing redundancy and irrelevant content. We propose DMG-RAG, a query-aware multi-grained RAG framework that coordinates chunking, retrieval, and budgeted context construction with two lightweight modules. We build aligned indices on the same corpus at three simple granularities (sentence, paragraph, and document). A query-conditioned Router dynamically allocates stage-1 retrieval quota set across granularities to shape an adaptive candidate pool for each query. Given the candidate pool, a Budgeter performs budget-constrained evidence selection by jointly considering relevance, chunk length, and redundancy, and assembles a compact evidence set via a deterministic greedy packing procedure guided by a learned scoring function. Experiments on HotpotQA, 2WikiMultiHopQA, and MuSiQue demonstrate that DMG-RAG consistently outperforms strong baselines under EM and F1, yielding substantial average gains (up to 7–8 points). The code is publicly available. Wednesday Virtual Room 5 IJCNN Paper Retrieval-Augmented Generation IV Session Chair: Qingqiang Wu (Xiamen University, School of Informatics), ShuLin Chen (Korea University) Epistemic State Modeling for Adaptive Retrieval-Augmented Generation Nan Liang (Xiamen University, School of Informatics); Cheng Yang (Southwest Jiaotong University, School of Information Science and Technology); Kun Yang (Xiamen University, Institute of Artificial Intelligence); and Jingqi Gao, Qingqiang Wu, and Meihong Wang (Xiamen University, School of Informatics) Abstract Abstract Retrieval-Augmented Generation (RAG) enhances LLMs with external knowledge, but existing methods use indiscriminate retrieval that ignores the model's internal knowledge state. We introduce Epistemic State Modeling (ESM), a framework that explicitly represents an LLM's metacognitive awareness as structured epistemic states over decomposed knowledge goals. Unlike prior adaptive RAG methods, ESM enables pre-generation planning by having the model confess what it knows, what it is uncertain about, and what specific gaps require retrieval. We implement ES-RAG via a comprehensive training methodology that jointly optimizes three metacognitive capabilities: (1) structured self-assessment via confession imitation, (2) confidence calibration via contrastive preference optimization, and (3) gap-driven intent generation for precise retrieval queries. Experiments on single-hop and multi-hop QA benchmarks show that ES-RAG achieves state-of-the-art performance while reducing retrieval overhead and providing transparent reasoning. Multimodal Adaptive Retrieval Augmented Generation through Internal Representation Learning Ruoshuang Du (School of Information Science and Technology, Shanghaitech University); Xin Sun and Qiang Liu (Institute of Automation, Chinese Academy of Sciences); Bowen Song, Zhongqi Chen, and Weiqiang Wang (Ant Group); and Liang Wang (Institute of Automation, Chinese Academy of Sciences) Abstract Abstract Visual Question Answering systems face reliability issues due to hallucinations, where models generate answers misaligned with visual input or factual knowledge. While Retrieval Augmented Generation frameworks mitigate this issue by incorporating external knowledge, static retrieval often introduces irrelevant or conflicting content, particularly in visual RAG settings where visually similar but semantically incorrect evidence may be retrieved. To address this, we propose Multimodal Adaptive RAG (MMA-RAG), which dynamically assesses the confidence in the internal knowledge of the model to decide whether to incorporate the retrieved external information into the generation process. Central to MMA-RAG is a decision classifier trained through a layer-wise analysis, which leverages joint internal visual and textual representations to guide the use of reverse image retrieval. Experiments demonstrated that the model achieves a significant improvement in response performance in three VQA datasets. Meanwhile, ablation studies highlighted the importance of internal representations in adaptive retrieval decisions. In general, the experimental results demonstrated that MMA-RAG effectively balances external knowledge utilization and inference robustness in diverse multimodal scenarios. Rethinking the Necessity of Adaptive Retrieval-Augmented Generation through the Lens of Adaptive Listwise Ranking Jun Feng (School of Computer Science and Information Engineering, Hefei University of Technology); Jiahui Tang, Zhicheng He, Hang Lv, Hongchao Gu, and Hao Wang (University of Science and Technology of China); and Xuezhi Yang and Shuai Fang (School of Software, Hefei University of Technology) Abstract Abstract Adaptive Retrieval-Augmented Generation aims to mitigate the interference of extraneous noise by dynamically determining the necessity of retrieving supplementary passages. However, as Large Language Models evolve with increasing robustness to noise, the necessity of adaptive retrieval warrants re-evaluation. In this paper, we rethink this necessity and propose AdaRankLLM, a novel adaptive retrieval framework. To effectively verify the necessity of adaptive listwise reranking, we first develop an adaptive ranker employing a zero-shot prompt with a passage dropout mechanism, and compare its generation outcomes against static fixed-depth retrieval strategies. Furthermore, to endow smaller open-source LLMs with this precise listwise ranking and adaptive filtering capability, we introduce a two-stage progressive distillation paradigm enhanced by data sampling and augmentation techniques. Extensive experiments across three datasets and eight LLMs demonstrate that AdaRankLLM consistently achieves optimal performance in most scenarios with significantly reduced context overhead. Crucially, our analysis reveals a role shift in adaptive retrieval: it functions as a critical noise filter for weaker models to overcome their limitations, while serving as a cost-effective efficiency optimizer for stronger reasoning models. AdaCache: Efficient and Robust Noise-Resistant Semantic Caching for RAG ShuLin Chen, LinTong Zhang, and Seong-Whan Lee (Korea University) Abstract Abstract Abstract—Despite the success of Retrieval-Augmented Generation (RAG), existing frameworks inherently suffer from redundant reasoning for semantically similar queries and excessive sensitivity to input perturbations. To address these universal limitations, this paper presents AdaCache, a dual-mode semantic caching framework designed to transform stateless RAG pipelines into robust, stateful systems. Demonstrated within the state-ofthe-art Probing-RAG architecture, our approach employs a DualMode Selective Persistence policy. Designed to bypass expensive model re-computation for recurring intents, AdaCache selectively persists internal knowledge when retrieval is deemed unnecessary (Condition 1) and verified outcomes from successful retrieval (Condition 2). Furthermore, to overcome the fragility of static similarity thresholds in noisy environments, we introduce a LocalDensity Adaptive strategy via local background analysis. By assessing the Signal-to-Noise Ratio (SNR) of the top candidate against local background noise, our framework ensures dynamic and robust hit detection. Experimental results on three opendomain QA benchmarks demonstrate that AdaCache reduces end-to-end latency by over 30% and storage overhead by up to 85% while maintaining superior accuracy under input perturbations compared to state-of-the-art baselines. Wednesday Virtual Room 6 IJCNN Paper Retrieval-Augmented Generation V Session Chair: Jing Chai (Yunnan University), Ruobing Shang (Soochow University) DSL-Coder: LLM-based Code Generation for Domain Specific Language Ruobing Shang and Shanping Su (Soochow University), Shaoyi Wang (Harbin Institute of Technology), Wenliang Chen (Soochow University), and Zhijun Li (Harbin Institute of Technology) Abstract Abstract Intelligent cockpit systems in modern vehicles require the integration of various functions, resulting in increasingly complex control logic. Domain Specific Language (DSL) is widely adopted to simplify the development process in such scenarios. Recently, Large Language Model (LLM) has been explored as a promising approach for generating code automatically. However, they still struggle with DSL code generation due to limited prior knowledge and domain-specific syntax variations. In this paper, we introduce CockpitDSL, a dataset for DSL code generation in intelligent cockpit systems, containing 888 annotated user instructions, reference code, and corresponding system states. Based on this dataset, we propose a novel approach that employs Backus–Naur Form (BNF) to define grammar rules, uses a vector database to store the APIs and code snippets, leverages LLMs for code generation, and introduces an evaluation chain for automatically evaluating syntax and grammar. Experimental results demonstrate that the Compile Success Rate exceeds 90% and the Pass@1 score exceeds 75%. A Low-Cost Lightweight Large Model Code Generation Optimization Framework Based on Structured Constraint Generation and External Verification lei liu and Junchao cui (Key Laboratory of Cyberspace Situation Awareness of Henan Province) and xuanzi ma (information engineering university) Abstract Abstract Lightweight large language models (LLMs) in resource-constrained scenarios exhibit high reasoning uncertainty, leading to low code generation reliability. Existing zero-shot methods fail to address this systematically. We propose a two-stage ``constraint-generation, code-generation'' framework that governs process uncertainty. It first extracts structured constraints from problem examples, converting open-domain generation into controlled structured search. Then, code is generated under these constraints, and an external verification mechanism (AST + example execution) creates a model-independent feedback loop to suppress hallucinations. On HumanEval, our framework improves Qwen2.5-1.5B-Instruct's Pass@1 from 40.2% to 80.3%. It is model-agnostic and lightweight, offering a practical solution for edge computing and reliable neural generation. TED-RAG: A Tree-Structured Multi-Expert Framework for Domain-Specific Data Synthesis in Retrieval-Augmented Generation Sheng Li (Shanghai Jiao Tong university) and Lei Pan, Shengjie Sun, Haodong Zhou, and Shuai Fan (AISpeech LTD) Abstract Abstract We propose a framework named TED-RAG for generating high-quality, domain-specific question-answer (QA) data using collaboration among multiple expert large language models (LLMs), aiming to train specialized QA models. The framework follows a tree-structured generation approach and includes three core components: Generator, which produces QA pairs through multi-turn dialogues; Rewarder, which evaluates the quality of generated questions; and Prompter, which provides feedback to guide iterative optimization. We employ models such as Qwen2.5-72B-Instruct along with well-crafted prompt templates to generate diverse QA pairs and automate the refinement of outputs through the Rewarder and Prompter. We conducted experiments on DragonBall and ChatRAG. Results show that TED-RAG significantly outperforms traditional single-model approaches in both diversity and quality of generated QA data, leading to improved performance in downstream tasks. We further validate the effectiveness of our approach through comprehensive ablation studies. XBRLTagRec: Domain-Specific Fine-Tuning and Zero-Shot Re-Ranking with LLMs for Extreme Financial Numeral Labeling Gang Hu, Qun Zhang, Jingyao Luo, Yile Jiang, Jing Chai, and Haiyan Ding (Yunnan University) Abstract Abstract Publicly traded companies must disclose financial information under regulations of the Securities and Exchange Commission (SEC) and the Generally Accepted Accounting Principles (GAAP). The eXtensible Business Reporting Language (XBRL), as an XML-based financial language, enables standardized and machine-readable reporting, but accurate tag selection from large taxonomies remains challenging. Existing fine-tuning-based methods struggle to distinguish highly similar XBRL tags, limiting performance in financial data matching. To address these issues, we introduce XBRLTagRec, an end-to-end framework for automated financial numeral tagging. The framework generates semantic tag documents with a fine-tuned FLAN-T5-Large model, retrieves relevant candidates via semantic similarity, and applies zero-shot re-ranking with ChatGPT-3.5 to select the optimal tag. Experiments on the FNXL dataset show that XBRLTagRec outperforms the state-of-the-art FLAN-FinXC framework, achieving 2.64%–4.47% improvements in Hits@1 and Macro metrics. These results demonstrate its effectiveness in large-scale and semantically complex tag matching scenarios. Wednesday Virtual Room 7 IJCNN Paper Retrieval-Augmented Generation VI Session Chair: SuGe Wang (Shanxi University), Qinghe Li (Zhengzhou University) RL-Semantic: RL-Enhanced Semantic Fuzzing for Bluetooth Low Energy Protocols Qinghe Li, Ruijie Cai, Xiaokang Yin, Yaobin Xie, Jiahan Liu, Shengli Liu, and Lukai Li (Information Engineering University) Abstract Abstract Existing protocol fuzzing tools struggle to achieve comprehensive Bluetooth Low Energy (BLE) coverage due to limited understanding of its complex state machines and inefficient test strategy selection. We propose RL-Semantic, a framework integrating protocol-aware semantic mutation strategies with Reinforcement Learning (RL) for adaptive strategy selection. We formulate the BLE fuzzing process as a Markov Decision Process (MDP) and model the protocol stack as a four-layer hierarchical state machine. Our approach then employs five semantic mutation strategies with DQN-based adaptive selection, using two-phase training that combines offline supervised learning and online RL. Evaluation on 10 simulated BLE targets over 360‑minute runs RL-Semantic is compared against Field, AFL, and Random baselines. RL-Semantic achieves an average of 266 state transitions, substantially outperforming Field with 55 transitions, AFL with 61, and Random with 63. Notably, RL-Semantic demonstrates decisive advantages at security-critical layers, achieving a 6.9 times improvement at GATT layer and transitioning from zero to non-zero coverage at SMP layer where Field fails to trigger any transitions. Ablation experiments validate that semantic strategies constitute the primary performance driver while RL provides incremental optimization through adaptive strategy selection. RL-Enhanced ProRAG: A Proactive Retrieval-Augmented Generation Framework for Multi-Hop Reasoning Xuewei Luo (School of Data Science and Information Engineering, Guizhou Minzu University, China); Yujia Huo (School of Data Science and Information Engineering, Guizhou Minzu University, China; Guizhou Provincial Key Laboratory of Applied Mathematics and Computing Power & Algorithms, Guizhou University, Guiyang, 550025, Guizhou, China); Yongbin Qin (Guizhou Minzu University, China); Fujian Feng and Gan Liu (School of Data Science and Information Engineering, Guizhou Minzu University, China); Dawen Xia (College of Microelectronics and Artificial Intelligence and College of Big Data Engineering, Kaili University, China); and Wong Derek F (The University of Macau) Abstract Abstract Large Language Models often struggle with factual inaccuracies and knowledge limitations in complex, knowledge-intensive tasks. While Retrieval-Augmented Generation (RAG) mitigates these issues, existing dynamic RAG methods suffer from “retrieval myopia” in multi-hop reasoning, leading to redundant retrievals, incomplete queries, and contextual fragmentation. RPLCD: Training-Free Safety Enhancement for Instruction-Tuned LLMs via Reverse Prompt Layer Contrastive Decoding Wei Li, Shezheng Song, Xiaodong Liu, Bin Ji, Jinying Xiao, Jie Wang, Zhenyang Gao, and Jie Yu (National University of Defense Technology) Abstract Abstract Instruction tuning substantially enhances the ability of pre-trained large language models (LLMs) to understand and follow human instructions. However, this progress also introduces potential safety risks. Although existing safety-alignment techniques have proven effective, they often suffer from high computational costs and inefficient training. To address these challenges, we propose Reverse Prompt Layer-wise Contrastive Decoding(RPLCD), an automated, plug-and-play decoding method that incurs no additional training overhead and can directly improve the safety of instruction-tuned LLMs. RPLCD uses a reverse prompt as a trigger and identifies unsafe outputs by comparing the logit differences between the final layer and designated safety layers; it then suppresses unsafe candidates to increase the likelihood of generating safe responses. Experiments show that, under RPLCD, Vicuna-7B achieves up to a +9.3-point improvement in tie-aware winning rate over standard decoding on four safe- generation tasks and up to a +7.1-point gain on two safety classification tasks, while largely preserving general capabilities. Moreover, RPLCD demonstrates strong scalability, substantially enhancing the safety of Chinese-Alpaca-7B, Chinese-Alpaca-13B, and Vicuna-13B. ProBA: Prompt Brittleness Mitigation via Layer-wise Localization and Representation Alignment Kexiang Qiao (Shanxi University); Yang Li (Shanxi University of Finance and Economics); and Suge Wang, Jian Liao, Jianxing Zheng, and Deyu Li (Shanxi University) Abstract Abstract Large Language Models (LLMs) demonstrate strong performance across diverse tasks, yet remain susceptible to prompt brittleness: minor perturbations in input prompts can degrade output quality and undermine model reliability. However, existing approaches rely on data augmentation or input adjustments, failing to address representation misalignment that causes the inference trajectory drift. To address these limitations, we draw inspiration from how humans maintain cognitive consistency: we localize sources of uncertainty within the reasoning process and realign internal representations to ensure stable reasoning across varying contexts. Motivated by this cognitive paradigm, we propose Prompt Brittleness Mitigation via Layer-wise Localization and Representation Alignment (ProBA), a general framework designed to ensure inference trajectory consistency. ProBA employs Layer-wise Brittleness Localization (LBL), guided by cosine divergence-based stability analysis, to localize the brittle layers responsible for inference drift. The representations within these layers are then rectified through Intra-layer Representation Alignment (IRA), which employs selective parameter freezing and MSE geometric constraints. Extensive experiments on LLMs with various architectures and different tasks show that our ProBA outperforms existing methods, achieving improvements in effectiveness and stability. Wednesday Virtual Room 8 IJCNN Paper Robustness and Adversarial Learning III Session Chair: Xu Zhang (Hohai University), chunqi li (Nanjing University of Science and Technology) Gradient Correlation Whitening and Wavelet Packet Decomposition Based Transferable Adversarial Attacks on Vision Transformers Shunhui Ji, Xu Zhang, and Yunhe Li (Hohai University); Hai Dong (RMIT University); and Yingying Shen and Pengcheng Zhang (Hohai University) Abstract Abstract Vision Transformers (ViTs) have demonstrated remarkable performance across various vision tasks, yet they remain vulnerable to adversarial attacks. While transfer-based attacks offer a practical black-box threat, existing methods often suffer from limited cross-architecture transferability due to token-level gradient overfitting and inadequate handling of frequency-sensitive features across models. In this work, we propose a novel method that integrates Gradient Correlation Whitening (GCW) and Wavelet Packet Decomposition (WPD) to address these limitations. GCW mitigates gradient overfitting by applying a global whitening transformation to decorrelate token-wise gradients, thereby reducing structural dependency on the source ViT. Meanwhile, WPD replaces the conventional Discrete Cosine Transform with a multi-level wavelet packet decomposition, enabling finer-grained frequency analysis and more adaptive perturbation enhancement across diverse spectral bands. Experiments on five ViTs and four CNNs demonstrate that our method surpasses the previous state-of-the-art in transferable adversarial attacks, achieving an average attack success rate improvement of 2.1% across models. Code is available at: https://github.com/lingyu700/GCW-WPD. DiMiG: Diverse Mismatch Guidance for Transferable Attacks on Vision-Language Models Tianjie Ni, Ziyu Liu, Shuyuan Zhang, and Yuxue Chen (Sichuan University); Wenbo Fang (Southwest Minzu University); and Junjiang He (Sichuan University) Abstract Abstract With the growing deployment of Vision-Language Pre-training (VLP) models, understanding their vulnerability to adversarial perturbations becomes increasingly important. Many multimodal adversarial attacks are effective in white-box settings yet degrade markedly under black-box transfer. Furthermore, many methods perturb both images and text simultaneously, while in many real-world scenarios, the text input is fixed or beyond the attacker's control. Based on this, we propose DiMiG, a transferable attack under an image-only threat model. First, we design a sample-conditional generator that generates adversarial image perturbations with only one forward pass. Then, a set of mismatched texts is mined for each image in the shared embedding space and used to provide diverse supervisory signals, thereby improving transferability. Finally, we devise a joint objective that pulls the adversarial image representation toward the mined mismatches and suppresses the matched text through weighted competition, while also enlarging the distance between adversarial and clean image representations in the same embedding space. Extensive experiments demonstrate that DiMiG achieves strong intermodel transferability while remaining highly efficient for adversarial evaluation. Under the image-only setting, with ALBEF as the surrogate, DiMiG improves the mean transferred ASR@1 across five victim models on Flickr30K and MSCOCO by 36.8 percentage points compared with representative prior methods. Our code is available at https://github.com/maskisbest/DiMiG. DAIA: Dynamic Ahead Iterative Attack for Highly Transferable Adversarial Examples Yongqing Zhao, Fuqiang Wang, and Wenqing Guo (Qilu University of Technology) Abstract Abstract Adversarial example attacks are considered a serious threat to Deep Neural Network (DNN) models. While generating adversarial examples in a white-box setting has been extensively studied, creating transferable adversarial examples that can successfully attack black-box models remains challenging. This work proposes DAIA (Dynamic Ahead Iterative Attack), a two-stage transferable adversarial attack method based on momentum-adaptive optimization. In the early stage of the attack, DAIA introduces a delayed-start phase, performing pre-iterations with a relatively large step amplification factor to accumulate a stable momentum direction and suppress early gradient oscillations. During the formal iterative stage, DAIA adaptively adjusts the step size by computing the cosine similarity between consecutive momentum vectors, combined with an Exponential Moving Average (EMA) smoothing mechanism, to balance stability and exploration.Empirical evaluations on the ImageNet dataset demonstrate that DAIA achieves superior portability compared to existing input-transformation-based methods under both single-model and defense settings. Furthermore, DAIA can be seamlessly integrated with input-transformation approaches to further enhance portability, highlighting the flexibility and effectiveness of our method. Boosting Adversarial Transferability via Spatial-Perceptual Reconfiguration chunqi li, Zijian Ying, Qianmu Li, and Zhichao Lian (Nanjing University of Science and Technology) Abstract Abstract Transferable adversarial attacks aim to craft small perturbations on a surrogate model to mislead unseen target models. Existing input transformation-based attacks enrich gradient information but often produce perturbations with limited spatial coverage, reducing overlap with discriminative regions of unseen models. To address this, we systematically study how input transformations affect gradient distributions and introduce the Mean Distance to Centroid (MDC) to quantify spatial dispersion of high-magnitude gradients. We categorize transformations into Spatial Reconfiguration (SR), which redistributes existing gradients via geometric operations, and Perceptual Reconfiguration (PR), which induces new high-magnitude gradient regions without substantially altering spatial locations. Guided by this analysis, we propose the \textbf{Spatial-Perceptual Reconfiguration (SPR)} attack, combining SR and PR to generate diverse transformed inputs and aggregate gradients with broader coverage. Experiments on benchmark datasets show that SPR consistently outperforms seven baseline methods in cross-model transferability. Wednesday 0.01 London FUZZ-IEEE Paper FUZZ 7: Industry applications; Emerging related topics Session Chair: Autilia Vitiello (University of Naples Federico II) Reinforcement Learning-Based Fuzzy Control for E-Bike Power Adaptation Jin-Shyan Lee and Zheng-Yan Lee (National Taipei University of Technology) Abstract Abstract While electric-assist bicycles (e-bikes) promote sustainable and convenient mobility, current assist mechanisms often ignore the balance between physical exertion and user fatigue. Excessive assistance diminishes exercise benefits, while insufficient power increases injury risk. This paper proposes a fuzzy logic control system that integrates pedal force, riding habits, environmental factors, and real-time physiological data. Unlike traditional heart-rate-only controllers which suffer from latency and instability, our model employs a reinforcement learning (RL) algorithm to optimize assist delivery dynamically. The system maintains the rider’s heart rate and speed within optimal, safe ranges across diverse user profiles. Simulation and field test results demonstrate that the proposed method significantly improves riding comfort and physiological stability compared to standard control models. FSSM-ResNet: Hair-Aware Skin Lesion Segmentation with Embedded Fuzzy Structural Suppression Adnan Yazici, Ayaulym Raikhankyzy, Amira Toksanbay, Saltanat Zarkhinova, Rimma Kubanova, Anel Amantayeva, Symbat Bayakhmetova, and Nail Fakhrutdinov (Nazarbayev University) Abstract Abstract Skin lesion segmentation is a key step in computer-aided skin cancer analysis, yet dermoscopic images often contain hair that obscures lesion boundaries and degrades segmentation performance. Most existing methods address hair via a separate preprocessing stage (e.g., hair removal or inpainting), which complicates the pipeline and can introduce additional artifacts. In this paper, we propose FSSM-ResNet, a hair-aware segmentation model that embeds a Fuzzy Structural Suppression Module (FSSM) directly within the network. FSSM exploits simple, interpretable fuzzy logic based on typical hair characteristics, dark, thin, and linear, to generate a hair-likelihood map and selectively suppress hair regions during decoding, eliminating the need for external preprocessing. To rigorously assess robustness to hair, we introduce VHI (Very Hairy Images), a manually annotated benchmark of 936 strongly hair-occluded dermoscopic images collected from multiple public datasets. Experiments on PH2 and ISIC 2016–2018 show that this fuzzy logic–driven structural suppression not only improves performance on standard benchmarks, but also maintains superior reliability under severe hair occlusion on VHI. Solving fractional diffusion equation with Hilfer derivative operator using the Fuzzy Transform Thi Minh Tam Pham, Irina Perfilieva, and Zhivorad Tomovski (University of Ostrava-IRAFM) Abstract Abstract In recent decades, fractional diffusion equations have found numerous applications in various scientific fields, particularly in industry, medicine, and finance, and especially in computer science, specifically image processing. This paper focuses on the fractional diffusion model and proposes a numerical method known as Fuzzy transform, to approximate its solution. The stability and convergence of the method are rigorously analyzed. Numerical results show that the method achieves a good balance between computational efficiency and accuracy. Thus, confirming its suitability for hybrid computing environments that combine artificial intelligence and numerical methods. An Explainable Fuzzy Logic-Based Approach to Quantum Backend Selection Allegra Cuzzocrea, Giovanni Acampora, and Autilia Vitiello (University of Naples Federico II) Abstract Abstract Quantum backend selection is emerging as a critical step toward achieving quantum computing utility. Indeed, the accuracy of quantum circuit execution strongly depends on the specific quantum machine used, each characterized by a distinct computational noise pattern. Currently, machine learning models for quantum backend selection are being proposed to address this challenge effectively. However, these approaches still lack two fundamental aspects: the integration of execution queue information, which impacts waiting time and scheduling efficiency, and the explainability of the selected backend, which is crucial to understand and trust the decision-making process. This paper introduces the first fuzzy-based machine learning approach for quantum backend selection capable of addressing both aspects. Experimental results show that the proposed method achieves competitive performance compared to state-of-the-art crisp classifiers, while selecting the backend by incorporating execution waiting times and exposing the circuit features underlying the selection decision. UnimNeuron: Automatic Connective Selection for Interpretable Neuro-Fuzzy Rules Paulo Vitor de Campos Souza (NOVA IMS) Abstract Abstract In evolving data streams, neuro-fuzzy systems must adapt to non-stationary distributions while preserving interpretability and numerical stability. However, most evolving fuzzy models rely on fixed logical connectives or implicitly encode logical behavior, limiting transparency and flexibility. GSO-DPCD: A Fuzzy Density-Peak Approach for Overlapping Community Detection Swetha Balasubramanian and Pranab K. Muhuri (South Asian University) Abstract Abstract Overlapping community detection in large-scale networks faces a constant trade-off between structural accuracy and computational complexity. Although traditional Density Peak Clustering (DPC) presents an intuitive mechanism for identifying community centers (density peaks), its applicability is limited due to O(N^2) complexity. In this paper, we propose a Gravity-driven Structural Overlapping Density Peak-based community detection (GSO-DPCD) that models community detection as a fluid-dynamic process within a Structural Gravity Field. Using a structural compactness measure derived from triadic closure, we detect density peaks that are topologically central in their communities. Further, we present a Heuristic Path Approximation that overcomes the quadratic computational bottleneck, achieving near-linear O(N log N) complexity. Extensive experiments on LFR benchmarks and large-scale real-world networks show that GSO-DPCD consistently outperforms existing methods in terms of modularity and accuracy, offering a robust and efficient framework for uncovering hidden structures in highly complex networks with deep overlaps. Wednesday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 16 - Optimization III Session Chair: Aldy Gunawan (Singapore Management University) Deep Reinforcement Learning for Stochastic Orienteering Problem with Metaheuristic Enhanced Experience Replay Kai Zhuo Lim, James Koh, and Aldy Gunawan (Singapore Management University); Cédric Verbeeck (EDHEC Business School); and Bing Tian Dai (Singapore Management University) Abstract Abstract In this work, we address the Stochastic Orienteering Problem with Time Windows (SOPTW) using an enhanced reinforcement learning (RL) framework. We employ a Quantile Regression Double Deep Q-Network (QR-DDQN) with Set Transformers for graph representation to capture the problem’s stochastic nature. To improve learning efficiency, we introduce a Metaheuristic Enhanced Experience Replay (MEER) mechanism that incorporates metaheuristic operators to generate diverse experiences. Results show that the RL agent achieves higher rewards by adapting in real time to stochastic variations, outperforming benchmark methods and standard RL baselines. We also analyze the underlying mechanisms to explain the observed performance gains. Bi-Objective Electric Vehicle Charging Scheduling Under Stochastic Charging Durations Aimen Khiar (IRIMAS, University of Haute-Alsace; CESI LINEACT, CESI); Mohamed el Amine Brahmia (CESI LINEACT, CESI); and Mahmoud Golabi, Abdennour Azerine, and Lhassane Idoumghar (IRIMAS Lab, University of Haute-Alsace) Abstract Abstract This paper addresses the electric vehicle charging scheduling problem under stochastic charging durations, where uncertainty arises from variations in actual charging times that are typically assumed deterministic in existing literature. We formulate a bi-objective optimization problem minimizing the expected values of peak load and total tardiness. We explicitly enforce non-overlapping charging sessions within the objective function evaluation via a repair mechanism where this latter induces cascading stochastic dependencies. This makes both realized charging start and end times random variables, substantially increasing problem complexity compared to existing approaches. We derive the cumulative distribution functions of realized charging start and end times and obtain an analytical expression for expected total tardiness involving a high-dimensional integral, while showing that expected peak load is intractable to compute exactly. We therefore approximate both objectives using Monte Carlo simulation. To solve the problem, we adapt and compare three multi-objective evolutionary algorithms: NSGA-II, MOPSO, and MOGWO. Comprehensive computational experiments on 20 real-world instances derived from the ACN-Data dataset, involving up to 200 vehicles, show that NSGA-II achieves superior hypervolume performance on 15 out of 20 instances, with statistically significant differences confirmed by Friedman and Mann–Whitney U tests. The proposed framework provides an effective decision-support approach for managing electric vehicle charging operations under uncertainty. EvoGrad: Accelerated Metaheuristics in a Differentiable Wonderland Beatrice Francesca Rosy Citterio (Bocconi University), Daniele Maria Papetti (University of Milano-Bicocca), Giovanna Maria Dimitri (University of Milan), and Andrea Tangherloni (Bocconi University) Abstract Abstract Differentiable programming has revolutionised real-valued optimisation by enabling efficient optimisation of highly complex, differentiable, parameterised functions and of complex models such as Deep Neural Networks. However, traditional Evolutionary Computation (EC) and Swarm Intelligence (SI) algorithms, widely successful in discrete search spaces or noisy and multimodal functions, typically do not leverage gradient information, limiting their efficiency and adaptability during the optimisation. To bridge these approaches, we introduce EvoGrad, a unified differentiable framework that integrates EC and SI with gradient-based optimisation. EvoGrad redefines all method components, ranging from candidate solution selection to evolutionary and swarm operators, in a differentiable manner, thereby facilitating their tuning and enabling small local searches. Extensive experiments on benchmark optimisation functions and on a real-world Parameter Estimation problem of a biochemical system reveal that our differentiable versions of EC and SI metaheuristics consistently outperform and compete with traditional, gradient-agnostic methods, setting a new standard for hybrid optimisation frameworks. Optimization of the Three-Dimensional Search allocation Game for UAVs Florian Delavernhe (Université Bourgogne Europe) Abstract Abstract This paper addresses the optimization of evasive target search using unmanned aerial vehicles (UAVs) through a search allocation game framework. We introduce a novel three-dimensional UAV search model in which the UAV altitude is part of the decision-making and explicitly affects detection range, search efficiency, and visibility to the target, thereby influencing both search performance and target behavior. Unlike existing search allocation models, which often rely on simplified or unrealistic target assumptions, the proposed formulation captures complex target reactions induced by the UAVs visibility. This leads to a larger-scale and more realistic optimization problem that better reflects practical UAV-based search operations. We develop a mathematical formulation of the problem and propose a metaheuristic solution based on simulated annealing. Computational experiments demonstrate that the proposed approach can solve realistically sized instances within operationally relevant computation times, highlighting its potential for real-world UAV search applications. A Hierarchical Cooperative Multi-Robot Path Planning Algorithm Based on TSP Encoding Yuxi Shen and Caitong Yue (Zhengzhou University); Jing Liang (Henan Institute of Technology, Zhengzhou University); Ying Bi, Peng Wang, and Kunjie Yu (Zhengzhou University); and Hui Song (University of Newcastle) Abstract Abstract Multi-robot path planning problems (MRPP) in practical applications usually involve multiple optimization objectives simultaneously, such as path length, task balance, and cooperative efficiency. This problem typically requires task allocation to be performed first, followed by path optimization for each subtask. These two processes together form a mutually coupled nested optimization structure. Most existing algorithms find it difficult to efficiently handle both task allocation and path optimization simultaneously within a unified framework. To address these challenges, this paper establishes a multi-objective MRPP model and proposes a hierarchical cooperative multi-robot path planning algorithm based on TSP encoding (HCMRPP-TE). The proposed algorithm solves the MRPP using a hierarchical optimization framework. At the upper level, the overall task is formulated as a Traveling Salesman Problem (TSP) model, where global tour optimization is performed to obtain a high-quality path structure. At the middle level, a TSP-tour-based task allocation encoding mechanism is designed to decompose the global tour into an encoding suitable for multi-robot path planning, enabling reasonable task allocation and coordinated path assignment among multiple robots. At the lower level, based on the obtained task allocation results, further local optimization is performed on each robot’s route to enhance the multi-objective performance. Experimental results demonstrate that the proposed method consistently outperforms the comparison methods in terms of overall multi-objective performance. Moreover, the proposed model and algorithm effectively improve the global optimization capability of multi-robot path planning and exhibit strong effectiveness and robustness in complex multi-robot scenarios. Wednesday 0.05 Paris IJCNN Paper Industrial Vision and Visual Quality Inspection Session Chair: Erdi Sayar (Paderborn University), Siamak Mehrkanoon (Utrecht University) Getting the Numbers Right—Modelling Multi-Class Object Counting in Dense and Varied Scenes Villanelle O'Reilly and Jonathan Cox (University of Lincoln); Georgios Leontidis (University of Aberdeen, UiT The Arctic University of Norway); Marc Hanheide (University of Lincoln); Petra Bosilj (Maastricht University); and James M. Brown (University of Lincoln) Abstract Abstract Density map estimation enables accurate object counting in heavily occluded, and densely packed scenes where detection-based counting fails. In multi-class density estimation, class awareness can be introduced by modelling classes non-exclusively, better reflecting crowded and visually ambiguous contexts. However, existing multi-class density estimators often degrade in less-dense scenes, while state-of-the-art detectors still struggle in the most congested settings. To bridge this gap, we propose the first vision-transformer-based approach to multi-class density estimation. Our model combines a Twins-SVT pyramid vision transformer backbone with a multiscale CNN decoder that leverages hierarchical features for robust counting across a wide range of densities. Further to that, the method adds an auxiliary segmentation task with the \emph{Category Focus Module} to suppress inter-category interference at training time. The module improves the density estimation head without the need for constraining assumptions added by the application of the auxiliary task at inference time, as required in previous methods. Training and evaluation on the VisDrone and iSAID benchmarks demonstrates a leap in performance versus the previous state-of-the-art multi-class density estimation methods, attaining a 33\%, 43\%, and 64\% reduction to MAE in testing evaluation. The method outperforms YOLO11 in less busy scenes, exceeding it by an order of magnitude in the most crowded testing samples. FG-YOLO: An Efficient Framework for Flexible Glove Surface Defect Detection Yuhang Lu, Shenao Li, and Qiuyan Wang (Tiangong University); Yaheng Ren (Hebei Academy of Sciences); and Yongjiang Xue and Qingzeng Song (Tiangong University) Abstract Abstract Flexible gloves are widely used in industrial production. However, their soft and deformable surfaces make subtle defects difficult to detect under complex backgrounds. To address this challenge, we propose FG-YOLO, an efficient and robust detection framework for real-time industrial glove surface defect detection. The proposed method redesigns the Spatial Fast Connection (SFC) module and the Bi-level Routing Attention (BRA) module, and proposes a novel Multi-Scale Cross Attention Network (MSCANet) to enhance multi-scale feature representation and suppress background interference. This design enables accurate perception of small and fine-grained defects while maintaining real-time performance. In addition, a high-quality industrial glove defect dataset is constructed using real production data to support reliable training and evaluation. Furthermore, a production-oriented automatic inspection system integrating multi-camera acquisition, real-time detection, and automated rejection is developed to validate the practicality of the proposed framework in continuous industrial environments. Experimental results show that FG-YOLO achieves 89.54% mAP50 and 60.59% mAP50:95 while maintaining an inference speed of 280 FPS, outperforming the baseline YOLO11n and demonstrating strong generalization across multiple datasets. These results indicate that the proposed method provides an effective and deployable solution for real-time industrial surface defect detection. 3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing Tobias Preintner (BMW AG, Leiden University); Yunfei Deng, Phillip Müller, Sebastian Illing, and Adrian König (BMW AG); and Thomas Bäck, Elena Raponi, and Niki van Stein (Leiden University) Abstract Abstract Despite recent progress in 3D generation, intuitive editing of existing shapes remains limited. Unlike images, which benefit from well-established inpainting tools, general 3D objects such as meshes still lack simple and effective methods for local shape editing. Existing approaches are often global, domain-specific, require complex user interaction, or focus on appearance (color and texture) rather than geometry. We introduce 3DMorph, a training-free framework for single-image-guided local 3D shape editing and morphing. Given an edited image showing a desired shape modification, our method automatically localizes the relevant 3D region and transfers 2D modifications to 3D while preserving unmodified areas. 3DMorph also enables intermediate shape generation between the original and edited objects, facilitating design exploration. To benchmark editing quality, we introduce Delta3D, an image-guided local 3D editing benchmark with paired ground-truth edits. Experimental results show that 3DMorph translates intuitive 2D edits into 3D, outperforming state-of-the-art generative and editing methods. PRISM: Industrial Point Cloud Completion via Meta-Learned Test-Time Self-Supervised Adaptation Seung Yeop Ha (Korea Institute of Industrial Technology, Korea University); Jae Hun Hwang (Korea Institute of Industrial Technology, Hanyang University); Jong Pil Yun (Korea Institute of Industrial Technology, Chung-Ang University); Seung-Kyum Choi (Georgia Institute of Technology); Jun-Geol Baek (Korea University); and Hong-In Won (Korea Institute of Industrial Technology) Abstract Abstract Low-cost LiDAR is increasingly adopted for 3D perception in smart manufacturing, but sparse returns, occlusions, and varying sensor configurations can degrade point cloud completion and downstream safety envelope estimation. Most completion methods rely on sparse point-set losses such as Chamfer Distance, offering weak surface supervision under severe sparsity and occlusion. We propose PRISM (Projection-Regularized Industrial Self-supervised Meta-adaptation), a LiDAR point cloud completion framework that couples completion with projection-based depth constraints and episodic meta-learning over auxiliary geometric objectives. In meta-training, PRISM learns an initialization for fast test-time adaptation and meta-optimizes adaptive auxiliary loss weights. At test time, self-supervised adaptation enforces projection-based depth consistency with the same geometric regularizers. Experiments on ShapeNet and our industrial RoboArm-Depth dataset demonstrate improved reconstruction and robot geometry recovery under severe sensor-induced sparsity and occlusion, supporting more reliable LiDAR-based safety envelope estimation. GroundCount: Grounding Vision-Language Models with Object Detection for Hallucination-Free Counting Boyuan Chen and Minghao Shao (New York University Abu Dhabi, New York University Tandon School of Engineering); Siddharth Garg and Ramesh Karri (New York University Tandon School of Engineering); and Muhammad Shafique (New York University Abu Dhabi) Abstract Abstract Vision Language Models (VLMs) exhibit persistent hallucinations in counting tasks, achieving substantially lower accuracy than other visual reasoning tasks. This phenomenon persists even in state-of-the-art reasoning-capable VLMs. CNN-based object detection models (ODMs) such as YOLO, by contrast, excel at spatial localization and instance enumeration with negligible computational overhead. We propose GroundCount, a framework that augments VLMs with explicit spatial grounding from ODMs to mitigate counting hallucinations. Our prompt-based augmentation achieves up to 80.6% counting accuracy on Ovis2.5-2B (a 5.9pp improvement), with gains reaching 7.7pp on Molmo2-4B across evaluated architectures, while reducing inference time by up to 24% for models prone to hallucination-driven reasoning loops such as Molmo2-4B, whose inference time drops from 8.0s to 6.1s. Comprehensive ablation studies reveal that positional encoding is a critical component, with its effect being architecture-specific: beneficial for most architectures yet detrimental for others regardless of model scale. Confidence scores, by contrast, offer marginal value across most architectures, improving performance in three of five evaluated models with modest gains of 0.2–1.1pp. We further evaluate feature-level fusion architectures and find that explicit symbolic grounding via structured prompts consistently outperforms implicit feature fusion, despite the latter employing sophisticated cross-attention mechanisms. Overall, our approach yields consistent improvements across four of five evaluated VLM architectures, with per-model gains ranging from 5.9 to 7.7pp; the remaining architecture exhibits comparatively limited gains, likely attributable to its iterative reflection mechanism partially overriding structured prompt inputs. We open-source our code at https://github.com/BoyuanChen99/GroundCount. Wednesday 0.10 Sydney IJCNN Paper Reliable and Governed LLMs Session Chair: M S MEKALA (Robert Gordon University ), Márcio Basgalupp (UNIFESP) Automatic Generation of Safety-compliant Linear Temporal Logic via Large Language Model: A Self-supervised Framework Junle Li (Hong Kong University of Science and Technology (Guangzhou), University of Glasgow) and Siqi Chen, Jiakai Li, Meiqi Tian, and Bingzhuo Zhong (Hong Kong University of Science and Technology (Guangzhou)) Abstract Abstract Converting high‑level natural‑language task descriptions into formal specifications such as Linear Temporal Logic (LTL) is essential for ensuring safety in cyber‑physical systems (CPS). Existing work, however, only optimizes translation quality without explicitly verifying the output against safety constraints. We present AutoSafeLTL, a self‑supervised, cloud–edge–collaborative framework that automates the generation of safety‑compliant LTL specifications while preserving logical consistency and semantic fidelity. A lightweight edge‑side three-stage-fine-tuned LLM offers real‑time conversion from natural language to LTL specifications (NL2LTL) and guarantees safety‑critical latency and data locality. Two larger‑capacity cloud‑side agents then iteratively refine the alignment: 1) LLM‑as‑an‑Aligner matches atomic propositions to safety constraints, and 2) LLM‑as‑a‑Critic interprets counterexamples from Inclusion Check to guide corrective regeneration. This collaborative architecture provides a safety-guaranteed alignment mechanism between high-level user intent and formally verifiable system behavior, demonstrating the potential of our framework to advance AI Alignment in safety-critical domains. Our approach achieves 0% violation rates on multiple benchmarks, enabling trustworthy specification generation and verification for both AI and critical CPS applications. Robustness of Language Models against Portuguese Harmful Prompts Eduardo Amorim (Centro de Informática, Universidade Federal de Pernambuco; TELUS Digital Research Hub) and Cleber Zanchettin (Centro de Informática, Universidade Federal de Pernambuco) Abstract Abstract Language Models (LMs) are vulnerable to jailbreak prompts designed to bypass safety constraints. Despite the growing literature in English, resources and defenses tailored to the Portuguese language remain scarce. We present SecBERT, a specialized Portuguese classifier built on the BERTimbau Base architecture to detect policy-violating and harmful prompts. We construct a 29,432-instance Portuguese dataset by translating a subset of WildJailbreak, explicitly preserving the original four-way taxonomy (Vanilla/Adversarial vs. Benign/Harmful) to enable granular analysis. We evaluate multiple BERT-based backbones in both frozen and fine-tuned settings, using F1- score, AUC, and the Kolmogorov-Smirnov (KS) statistic to characterize operational separability. The fine-tuned BERTimbau Base (SecBERT) achieves 95.6% F1, 99.2% AUC, and 91.2% KS, significantly outperforming non-Portuguese-centric baselines in discriminatory capability. We further analyze threshold behavior and discuss the limitations of translation-based data genera- tion, outlining directions for native Portuguese harmful prompt datasets and adaptive robustness testing. MIRROR: Novelty-Constrained Memory-Guided MCTS Red-Teaming for Agentic RAG Inderjeet Singh and Andrés Murillo (Fujitsu Research of Europe Limited) and Motoyoshi Sekiya, Yuki Unno, and Junichi Suga (Fujitsu Limited) Abstract Abstract Multimodal agentic retrieval-augmented generation (RAG) systems expand the attack surface beyond prompt injection to include text poisoning, image injection, direct-query attacks, and orchestrator-level tool manipulation. Existing red-teaming approaches are typically surface-specific and often recycle known attack templates; on text-poisoning benchmarks we measure 73-84% exact duplication. We present MIRROR, a unified cross-surface framework that performs memory-guided Monte Carlo tree search while conditioning candidate generation on retrieved context under an explicit novelty constraint. A deterministic Novelty Gate rejects any candidate matching the retrieval set under normalized comparison, allowing retrieval to inform search priors without enabling prompt copying. Across four attack surfaces on a multimodal agentic RAG target, MIRROR attains 76% ASR on image poisoning compared with 52% for baselines, 97% ASR on orchestrator attacks at half the query cost, and the lowest cross-surface variance (coefficient of variation 0.47). In contrast, specialized baselines collapse across surfaces: suffix optimization reaches 79% ASR on text poisoning but 1% on direct queries. We release ART-SafeBench with 41,815 in-package records and runtime adapters yielding 41,991+ total records across four surfaces. Learnable Contrastive Decoding for Dehallucination in LLMs Jen-Tzung Chien and Chao-Chi Li (National Yang Ming Chiao Tung University) Abstract Abstract Large language models (LLMs) excel in the tasks such as text generation for question answering and reasoning, but still remain prone to hallucination due to semantic drift, factual errors, and verbalized repetition, which restrict the reliability when building the knowledge-grounded applications. Existing remedies such as fine-tuning, prompt engineering, and retrieval-augmented generation improve the performance but suffer from high computational cost, architectural complexity, and limited transferability. This paper presents the learnable contrastive decoding (LCD), a decoding-time scheme that enhances the semantic fidelity without altering parameters. Inspired by the locate–edit paradigm in model editing, LCD identifies the semantically critical intermediate layers via a trained selector and edits the outputs through an attention-based fusion module that contrasts these layers with the final-layer distribution. This dynamic and learnable reasoning corrects the biases and improves the alignment. Experiments on WritingPrompts, TruthfulQA, and GSM8K show LCD outperforms the other methods in factuality, consistency and diversity, while the ablations confirm the necessity of both layer selection and fusion. Overall, LCD provides a lightweight, semantically aware intervention that mitigates hallucination without fine-tuning, advancing the controllability and reliability of LLM based on the state-space model using RWKV. Wednesday 0.11 Cape Town IJCNN Paper Industrial Inspection, Cybersecurity, and Harmful Content Detection Session Chair: Zarka Bashir (IIT Hyderabad, IDRBT), Roberto Corizzo (American) FeatureFox: Sample-Efficient Panoptic Graph Segmentation for Machining Feature Recognition in B-Rep 3D-CAD Models Bertram Fuchs, Altay Kacan, Aaron Haag, and Oliver Lohse (Siemens AG) Abstract Abstract Automatic feature recognition (AFR) on B-Rep 3D-CAD models is central to CAD/CAM automation, yet most learning-based methods are complex, data-hungry, and evaluate instance grouping and semantic labeling separately. We present FeatureFox, a panoptic AFR pipeline that outputs machining instances with semantic labels: a calibrated binary edge classifier on enriched edge attributes localizes feature boundaries, instances are recovered as connected components in a pruned face-adjacency graph, and a per-instance classifier predicts the machining class from aggregated subgraph attributes. We evaluate on MFInstSeg using Panoptic Quality (PQ), which jointly scores instance separation and semantic correctness. FeatureFox is substantially more sample- and compute-efficient than the deep baseline AAGNet, reaching PQ > 0.9 with ~250 training parts versus ~5,000 for AAGNet, and training on the full MFInstSeg set takes seconds on a GPU. On the full training set, AAGNet surpasses FeatureFox marginally in PQ, while FeatureFox remains slightly ahead in feature-level recognition and localization accuracy. Finally, leveraging its low data requirement, we train FeatureFox on 270 manually labeled industrial CAD parts and show qualitative generalization to an unseen real industrial part, indicating practical real-world applicability. Transformer-Based Intrusion Detection with Feature-Selected Inputs for Cross-Day Generalization and Unseen Attack Detection Syed Aon Ali Naqvi, Amit K. Shukla, and Petri Välisuo (University of Vaasa) and Muhammad Faheem (VTT) Abstract Abstract Most Intrusion Detection Systems (IDS) are evaluated using random data splits, leading to inflated performance that does not reflect the real-world deployment conditions. In practical environments, IDS must generalize temporally separated traffic and previously unseen attack patterns. To address this limitation, we propose Transformer-based intrusion detection framework with two stage feature selection strategy. The proposed method combines Random Forest and permutation importance to identify informative features, reduce input noise and improve representation learning. The Transformer encoder further learns contextual feature representations through multihead self-attention, which improves the modeling of complex dependencies. The proposed framework is evaluated on the CICIDS 2018 dataset under a cross-day experimental setting, where training and testing data contain different attack categories. Extensive experimental results demonstrate that the proposed framework significantly outperforms existing Machine learning and Deep Learning baselines. The proposed framework improves the cross-day F1-score by up to 14% and achieves 97% F1 score, demonstrating its effectiveness in detecting unseen attacks in dynamic environments Misspoken Whispers: Attacking ASR Models to Mistranscribe Toxic Audios Dan Kienast and Jiaojiao Jiang (UNSW) Abstract Abstract Automatic speech recognition (ASR) is increasingly used as the front end of content moderation pipelines, where ASR transcripts are passed to downstream text-based toxicity classifiers. This design implicitly assumes that ASR outputs are semantically faithful and robust to benign acoustic variation. In this paper, we show that this assumption can be systematically violated by small, structured, and speaker-specific acoustic perturbations that preserve transcript coherence while selectively suppressing toxic lexical content. We introduce a compositional attack that applies a fixed sequence of simple acoustic transformations, with parameters selected via a speaker-level search using a small number of utterances. Once learned, the perturbation generalises to unseen utterances from the same speaker, enabling practical misuse without per-sample optimisation. We evaluate our attack on Whisper-base and eight additional ASR models. For toxic speech, the attack alters more than 92% of transcripts and affects more than 93% of speakers, yielding substantial reductions in text-based toxicity scores. On non-toxic speech, the average change in toxicity remains near zero, indicating minimal collateral effects. We further demonstrate partial transferability across ASR architectures while maintaining intelligible, syntactically plausible transcripts. Our findings reveal a systemic weakness in ASR-dependent moderation pipelines and demonstrate that text-only, transcript-based toxicity detection is insufficient to counter structured audio-level manipulation. We argue that effective defences must incorporate audio-aware robustness mechanisms rather than relying solely on downstream text classifiers. Towards Explainable Hate Speech Detection in Roman Urdu Ubaid Azam (University of Southampton), Hammad Rizwan (Dalhousie University), Faizad Ullah and Ali Faheem (Lahore University of Management Sciences), Faisal Kamiran (Information Technology University), and Asim Karim (Lahore University of Management Sciences) Abstract Abstract Automatic hate speech detection in low-resource languages remains limited to basic classification, while fine-grained tasks such as span detection and intensity classification are largely confined to English. Moreover, temporal generalization and model interpretability, both critical for real-world deployment, remain underexplored. This work presents a methodology for fine-grained, robust, and interpretable hate speech detection in Roman Urdu. We introduce RUHSOLD++, an enriched dataset featuring annotations for hate intensity and hateful span detection, enabling interpretability evaluation through plausibility and faithfulness metrics. Using multilingual transformer models in a multitask framework, we assess temporal generalizability via a time-skip dataset and robustness via adversarial word variations. Our results demonstrate that while interpretability metrics enhance prediction reliability, model performance degrades significantly under temporal distribution shifts and linguistic variations, underscoring the need for periodic retraining or improved training paradigms. Wednesday 0.14 Singapore IJCNN Paper Remote Sensing and Environmental Intelligence Session Chair: Polat Goktas (Sabanci University), Amirreza Yousefzadeh (University of Twente) Graph Structural Descriptors for Classification of Soybean Grains Eduardo da Silva Afonso and Jianglong Yan (University of São Paulo), Murillo G. Carneiro (Federal University of Uberlandia), Kaizhou Gao (Macau University Science and Technology), Huaqiang Yuan (Dongguan University of Technology), and Liang Zhao (University of São Paulo) Abstract Abstract The soybean production chain continues to face important technical and operational challenges, among which automatic grain quality classification remains a major concern. Although deep learning techniques have been extensively employed for soybean grain classification and related agricultural tasks, they still exhibit key practical limitations. These include dependence on large annotated datasets, limited sensitivity to subtle damage patterns, reduced interpretability, and considerable computational cost. To mitigate these issues, we propose a graph construction strategy based on images that encodes rich visual cues, namely color, texture, and curvature, from segmented image regions. On top of this representation, we develop a classification approach for soybean grain images using graph-derived features. By modeling structural relationships in the data instead of depending exclusively on pixel-level detail, the proposed method is able to achieve strong classification results even with a limited amount of training data. The experimental findings also indicate that the inclusion of graph-based features extracted from soybean images leads to clear improvements in classification accuracy. Multi-Agent Proximal Policy Optimization for Network-Aware Drone-based Crop Image Analysis Andrew Hellman and Bishwas Wagle (University of Missouri - Columbia); Sean Peppers (Florida Gulf Coast University); Vincent Zheng (Stony Brook University); Alicia Esquivel Morel and Juan Mogollon (University of Missouri - Columbia); Arunava Roy (University of Memphis); and Jianfeng Zhou, Kannappan Palaniappan, and Prasad Calyam (University of Missouri - Columbia) Abstract Abstract Drone-based applications are transforming precision agriculture by enabling advanced sensing and analysis of crops by collecting high-resolution visual data. However, these applications need to handle dynamic edge network conditions, limited onboard resources, and fixed flight paths, making it challenging to decide when computation tasks should be performed locally on the drone, offloaded to nearby edge servers, or sent to the cloud. In multi-drone deployments, these decisions are inherently coupled: multiple drones make concurrent offloading choices while sharing wireless links and compute resources, where the action of one drone directly impacts the performance of others. In this paper, we propose a network-aware multi-drone-based computation task offloading framework viz., “FieldVision” that can continuously learn and adapt based on real-time system (e.g., battery level and compute) and network (e.g., bandwidth and latency) conditions. Our FieldVision features a Multi-Agent Reinforcement Learning (MARL) approach based on the Centralized Training with De-centralized Execution (CTDE) paradigm to address the coupled and partially observable nature of multi-drone offloading, aiming to learn a decentralized policy that minimizes latency, reduces deadline violations, and conserves energy across the drone swarm. We evaluate FieldVision using representative precision agriculture workloads under a time-varying wireless network model with continuous bandwidth and latency fluctuations. Our experimental results show that FieldVision improves average episode reward by up to 28.6% over single-agent PPO and significantly outperforms heuristic and rule-based baselines in terms of reward stability and deadline reliability, while preserving fully decentralized execution without inter-drone communication. SmaAT-QMix-UNet: A Parameter-Efficient Vector-Quantized UNet for Precipitation Nowcasting Nikolas Stavrou and Siamak Mehrkanoon (Utrecht University) Abstract Abstract Weather forecasting supports critical socioeconomic activities and complements environmental protection, yet operational Numerical Weather Prediction (NWP) systems remain computationally intensive, thus being inefficient for certain applications. Meanwhile, recent advances in deep data-driven models have demonstrated promising results in nowcasting tasks. This paper presents SmaAT-QMix-UNet, an enhanced variant of SmaAT-UNet that introduces two key innovations: a vector quantization (VQ) bottleneck at the encoder–decoder bridge, and mixed kernel depth-wise convolutions (MixConv) replacing selected encoder and decoder blocks. These enhancements both reduce the model’s size and improve its nowcasting performance. We train and evaluate SmaAT-QMix-UNet on a Dutch radar precipitation dataset (2016–2019), predicting precipitation 30 minutes ahead. Three configurations are benchmarked: using only VQ, only MixConv, and the full SmaAT-QMix-UNet. Grad-CAM saliency maps highlight the regions influencing each nowcast, while a UMAP embedding of the codewords illustrates how the VQ layer clusters encoder outputs. MSTR-Net: A Multi-Scale Transformer Refinement Network for Underwater Image Enhancement Pranjali Singh and Prithwijit Guha (Indian Institute of Technology Guwahati) Abstract Abstract Underwater images often suffer from severe visual degradation caused by light scattering, absorption, and wavelength-dependent attenuation, leading to reduced visibility, color distortion, and loss of structural details. To address these challenges, this paper proposes MSTR-Net, a Multi-Scale Transformer Refinement Network for underwater image enhancement (UIE). The proposed framework adopts a transformer-based encoder with efficient self-attention to capture long-range contextual information while maintaining computational efficiency. Hierarchical features extracted at multiple spatial resolutions are aggregated through a multi-scale feature fusion decoder, which predicts a residual image to restore fine structural details. To further improve visual quality, a multi-scale structural refinement module enhances edge continuity and suppresses local artifacts, while a luminance-guided refinement module corrects uneven illumination without introducing color distortion. As a result, MSTR-Net produces visually consistent outputs with improved color balance and sharper boundaries, effectively avoiding over-saturation and halo artifacts commonly observed in existing methods. Extensive experiments conducted on four benchmark underwater image datasets demonstrate that the proposed method consistently outperforms state-of-the-art approaches in both full-reference and no-reference evaluation metrics, confirming its effectiveness and robustness across diverse underwater conditions. The codes of MSTR-Net are available at \url{https://github.com/pranjali1996/MSTR-Net}. UIESNN: A Scale-Aware Spiking Network for Underwater Image Enhancement Shuang Chen and Ruochen Li (Durham University), Zihan Zhu (University of Cambridge), Ronald Thenius Thenius (University of Graz), and Farshad Arvin and Amir Atapour-Abarghouei (Durham University) Abstract Abstract Underwater image enhancement (UIE) is a practically important yet underexplored application of spiking neural networks (SNNs), where the dominant degradations are large-scale and low-frequency, such as wavelength-dependent colour casts and scattering-induced veiling. Existing SNN restoration designs rely on locally bounded spiking perception, which can limit global correction and lead to saturated or inconsistent representations. To address these challenges, we propose a scale-aware SNN framework for UIE named UIESNN. At its core is a Multi-scale Pooling LIF Block (MPLB) that injects hierarchical multi-scale pooling responses into membrane dynamics, thereby enlarging the effective receptive field while preserving fine-grained details and inducing heterogeneous scale-dependent activations. Building on MPLB, we design a spiking residual architecture that integrates frequency decomposition and attention-based refinement in a fully spike-driven pipeline. Extensive experiments on the EUVP and LSUI benchmarks demonstrate that UIESNN achieves state-of-the-art performance among SNN-based methods, delivering improved colour fidelity and spatial coherence with competitive energy cost. Wednesday 0.15 Washington IJCNN Paper Efficient and Resource-Constrained AI I Session Chair: Christos Kyrkou (University of Cyprus; Department of Electrical and Computer Engineering, University of Cyprus), Yanick Christian Tchenko (University of Paris Saclay, University of Evry) Complementary Attention Head Pruning for Efficient Transformers Yaniv Livertovsky, Shahar Somin, and Gonen Singer (Bar Ilan University) Abstract Abstract The remarkable success of Transformer-based models in natural language processing stems from architectural scaling, which leads to a large number of parameters and hinders deployment in resource-constrained environments. While structured pruning offers a pathway to compression, existing state-of-the-art methods often rely on gradient-based importance ranking or stochastic gating, which suffer from instability, structural degeneration, and the need for extensive manual hyperparameter tuning. In this paper, we introduce CAHP (Complementary Attention Head Pruning), a novel post-hoc framework that redefines head selection as a global graph-theoretical problem. Rather than evaluating heads in isolation, CAHP utilizes graph-based clustering combined with information-theoretic distance measures to identify and preserve a topologically diverse subset of complementary attention heads. Without requiring a predefined sparsity level or pruning ratio, the framework automatically determines the number of selected attention heads across layers by identifying a diminishing marginal performance curve, where pruning additional heads leads to a sharp degradation in performance, as determined by the chosen polynomial degree. Extensive evaluations on the SST-5 and MNLI benchmarks, across different Transformer model scales, demonstrate that CAHP consistently outperforms competitive baselines, particularly in high-compression regimes. Furthermore, our structural analysis shows that CAHP avoids the “proximity bias” of gradient-based pruning methods, which tend to preserve heads mainly in layers close to the output, and instead retains a functionally critical set of attention heads in the model’s intermediate layers. Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs Michael Rottoli, Subhankar Roy, and Stefano Paraboschi (Università degli Studi di Bergamo) Abstract Abstract Diffusion-based Large Language Models (D-LLMs) represent a promising frontier in generative AI, offering fully parallel token generation that can lead to significant throughput advantages and superior GPU utilization over traditional autoregressive paradigm. However, this parallelism is constrained by the requirement of a fixed-size response length prior to generation. This architectural limitation imposes a severe trade-off: oversized response length results in computational waste on semantically meaningless padding tokens, while undersized response length cause output truncation requiring costly re-computations that introduce unpredictable latency spikes. To tackle this issue, we propose Predict-then-Diffuse, a simple and model-agnostic framework, that enables compute-budgeted inference per input query by first estimating the response length and then using it to run inference with D-LLM. At its core lies a Adaptive Response Length Predictor (AdaRLP) auxiliary predictor that predicts the optimal response length given an input query. As a measure against under-predicting the response length and re-running inference with a higher response length, we introduce a data-driven safety mechanism, which trades a negligible padding overhead. As a whole, our framework limits the significant waste of computation on padding tokens and preserves output quality. Experimental validation on multiple datasets demonstrate that Predict-then-Diffuse significantly reduces computational costs (FLOP) compared to the default D-LLM inference mechanism and baselines based on heuristics, while being robust to skewed data distributions. Quantization-Aware Regularizers for Deep Neural Networks Compression Dario Malchiodi, Mattia Ferraretto, and Marco Frasca (Università degli Studi di Milano) Abstract Abstract Deep Neural Networks reached state‑of‑the‑art performance across numerous domains, but this progress has come at the cost of increasingly large and over‑parameterized models, posing serious challenges for deployment on resource‑constrained devices. As a result, model compression has become essential, and---among compression techniques---weight quantization is largely used and particularly effective, yet it typically introduces a non‑negligible accuracy drop. However, it is usually applied to already trained models, without influencing how the parameter space is explored during the learning phase. In contrast, we introduce per‑layer regularization terms that drive weights to naturally form clusters during training, integrating quantization awareness directly into the optimization process. This reduces the accuracy loss typically associated with quantization methods while preserving their compression potential. Furthermore, in our framework quantization representatives become network parameters, marking, to the best of our knowledge, the first approach to embed quantization parameters directly into the backpropagation procedure. Experiments on CIFAR‑10 with AlexNet and VGG16 models confirm the effectiveness of the proposed strategy. LaTRO: Language-Token Routing for Scalable and Parameter-Efficient Multilingual ASR Ilyes Oukid (LIPN CNRS UMR 7030 / Sorbonne Paris Nord University, Ecole Militaire Polytechnique); Bilal Faye and Hanane Azzag (LIPN CNRS UMR 7030 / Sorbonne Paris Nord University); Mustapha Lebbah (DAVID Lab, UVSQ, Paris-Saclay University; LIPN CNRS UMR 7030 / Sorbonne Paris Nord University); and Said Yacine Boulahia (Ecole Militaire Polytechnique) Abstract Abstract Multilingual automatic speech recognition (ASR) seeks to transcribe speech from multiple languages within a unified framework, yet existing approaches often suffer from poor scalability and high computational cost. Current state-of-the-art models typically rely on language-specific adapters or complex architectures, which increase the number of parameters as new languages are added and remain ineffective for low-resource and linguistically complex languages, particularly African languages. We propose LaTRO (Language-Token Routing Optimizer), a compact multilingual ASR model that addresses these limitations by (i) maintaining a constant number of trainable parameters regardless of the number of languages, (ii) introducing a learnable language token that guides the model through a shared parameter routing mechanism without language-specific modules, and (iii) supporting both simultaneous multilingual training and progressive language integration. The proposed approach is designed to better handle low-resource and under-explored African languages characterized by rich phonological and morphological variability. Experiments conducted on nine African languages show that LaTRO achieves competitive or superior performance compared to strong state-of-the-art multilingual ASR models, while significantly reducing computational and memory requirements. Compute-Efficient Event-Driven Proceed for Online Time Series Forecasting under Concept Drift JUAN CARLOS LA ROSA PAREDES and LILIAN BERTON (UNIFESP) Abstract Abstract Online time series forecasting under concept drift requires models to adapt continuously as data distributions evolve. Recent proactive frameworks such as Proceed improve predictive accuracy by estimating drift and modulating model parameters before prediction, but this benefit comes at the cost of additional computation at every online step. In this work, we study efficiency-oriented variants of Proceed that aim to improve its accuracy–efficiency trade-off without changing the forecasting backbone. We evaluate an error-triggered variant, Proceed-ET, and a retrieval-based variant, Proceed-ER, as well as their combination. Experiments on ETTH2 and WEATHER with an iTransformer backbone show that Proceed-ET provides the clearest practical benefit: on WEATHER, it reduces total runtime from 1556.27 s to 870.15 s (−44.1%) and step latency from 33.30 to 16.20 ms/step (−51.4%), while maintaining accuracy close to the original Proceed; on ETTH2, it reduces runtime from 90.14 s to 52.02 s (−42.3%) and latency from 6.37 to 2.93 ms/step (−54.0%), although with a larger loss in accuracy. By contrast, Proceed-ER does not show consistent gains under the current configuration, and its benefits appear more plausible as a conditional rather than always-on component. Overall, our results suggest that selectively deciding when to adapt is a more effective lever for efficient proactive online forecasting than adding retrieval on top of constant adaptation, which is especially relevant in real-world settings with strict latency and memory constraints. Wednesday 2.18 Mekong IJCNN Paper Biomedical Signal Processing Session Chair: Roseline Mary Rozario (University of Wollongong), Susana Vieira (Instituto Superior Técnico) Don't Sweat It: Transforming Electrodermal Activity for Stress Detection Mjellma Çitaku, Jae-Eun Nam, Larissa Zott, and Silja Meyer-Nieberg (Universität der Bundeswehr München) Abstract Abstract With stress being a commonly occurring emotion in society as well as a major component in causing serious physical and mental heath issues, studying and monitoring this phenomenon is an important task in the modern world. Recent studies have shown the success of using deep learning approaches in stress classification. A promising way to proceed is to transform the one-dimensional biophysiological time series into two-dimensional representations by image encoding. However, the question arises which encoding type is suitable for which kind of time series. Here, several often used image encodings are analyzed in regards to their ability to extract and display signal characteristics relevant for stress detection. In this paper, the focus lies on the electrodermal activity (EDA) which represents a promising modality to detect stress. Due to data scarcity, the utility of transfer learning starting from ImageNet-pretrained convolutional neural networks (CNNs) is also analyzed. Here, several backbone families, established and modern, are evaluated under a locked fine-tuning protocol to isolate architectural effects. Furthermore, the impact of tonic-phasic decomposition is examined by comparing the full EDA signal to its phasic component. Experiments on the WESAD dataset demonstrate that scalograms appear more robust for stress detection. Additionally, the trade-off between lightweight and larger pretrained CNNs is characterized. Evaluating Vision-LSTM for Stress Detection Mjellma Çitaku, Silja Meyer-Nieberg, and Marko Hofmann (Universität der Bundeswehr München) Abstract Abstract Vision-LSTM (ViL) has recently been introduced as a recurrent alternative to self-attention-based architectures, but its suitability for physiological data has not been studied. This paper provides the first systematic evaluation of ViL for stress detection from photoplethysmography-derived blood volume pulse (BVP) data. BVP segments are transformed into two dimensional image encodings, time-frequency scalograms and symmetric projection attractor reconstruction (SPAR), and ViL is compared against a Vision Transformer (ViT) under a strictly matched protocol: ImageNet-1K initialization, identical patchification, the same hyperparameter search space and tuning budget, and leave-one- subject-out (LOSO) evaluation. On WESAD dataset, ViL shows a consistent gain from SPAR across model sizes, while ViT achieves the strongest overall three-class performance and benefits more reliably from scalograms and increased capacity. In the binary stress vs. non-stress setting, ViL becomes competitive and slightly surpasses ViT in in-dataset LOSO. Cross-dataset experiments between the WESAD and UBFC-Phys datasets highlight a non-trivial domain shift and show that limited target-domain exposure can improve transfer. Overall, the results clarify when ViL is advantageous for physiological image encodings and highlight a strong architecture-representation interaction in BVP-based stress detection. Neural Network Cognitive Feedforward for Habit Change Decision-Making: A systems thinking approach in developing medical software for behavior change and quality assurance Pantea Keikhosrokiani (University of Oulu) Abstract Abstract The advancement of artificial intelligence, cognitive modeling, and computational system thinking presents a gap in quality assurance for hybrid intelligence software systems. Therefore, this study aims to develop an artifact as a cognitive model and quality assurance framework to predict habit-change decision-making in AI-driven software systems. Data was collected from 220 respondents prior to the full deployment of the AI habit-support software system. The study utilized cognitive computation to model a neural network cognitive feedforward model to predict the status of habit change. This model integrates the effects of other assistive factors to enhance the relationship between cognitive constructs and habit-change decision-making. The feedforward model is further optimized to identify the best model configuration. The results indicate that cognitive emotions, lifestyle habits, persuasive cognition, and literacy significantly affect the decision-making process in human-AI software systems. However, assistive factors may help reduce the impact of controlled cognition and behavioral change, leading to resistance instead of adherence to change. Agent-based modeling and systems thinking approaches facilitate the simulation of cognitive computation in the software development life cycle, thereby enhancing decision-making. In addition to the predictive modeling results, several strategies such as consistency, trustworthiness, transparency, follow-through, maintenance, and adherence to new routines, ensure quality assurance in AI-driven software systems while emphasizing computational thinking. The study's findings assist scientists in enhancing maintenance adherence, supporting pre- and post-habit change decisions. The computational intelligence cognitive feed-forward predictive model reinforces new behaviors within the human-AI software development life cycle. Contrastive Pretraining for Ensemble Graph Neural Networks in Alzheimer’s Disease Classification Walter Mangione, Alberto Gaetano Valerio, and Gabriella Casalino (University of Bari Aldo Moro); Miguel Angel Ferrer (University of Las Palmas de Gran Canaria); Giovanna Castellano (University of Bari Aldo Moro); Moises Diaz (University of Las Palmas de Gran Canaria); and Gennaro Vessio (University of Bari Aldo Moro) Abstract Abstract Alzheimer’s disease (AD) is a progressive neurodegenerative disorder that necessitates accurate early diagnosis. We propose a novel training strategy using an ensemble graph neural network (GNN) with contrastive pretraining for AD classification, utilizing atlas-based graphs from the AAL atlas to ensure anatomically meaningful node semantics. This ensemble combines various GNN architectures to capture diverse connectivity patterns, while contrastive pretraining enhances robust latent representation before supervised training. Evaluated on the publicly available ADNI dataset, our approach outperforms state-of-the-art methods across multiple metrics. Additionally, integrating PGExplainer improves model transparency by identifying disease-relevant subgraphs and contributions from specific brain regions. Our framework enhances both classification accuracy and interpretability in graph-based AD diagnosis. Tackling brain signal inter-subject variability with adaptive neural architectures Sébastien VELUT (ISAE SUPAERO, Université Paris Saclay); Stella Douka and Theo Rudkiewicz (Université Paris Saclay); Alex Davey and Stephane Rivaud (INRIA Saclay); Francois Landes (Université Paris Saclay); Julien Mille (INSA Centre-Val de Loire); Guillaume Charpiat (INRIA Saclay); Sylvain Chevallier (Université Paris Saclay); Marie-Constance Corsi (Paris Brain Institute); and Frederic Dehais (ISAE SUPAERO) Abstract Abstract Decoding electroencephalography (EEG) signals with deep learning remains challenging due to strong inter- and intra-individual variability, low signal-to-noise ratio, and the absence of an optimal architecture. While deep neural networks have demonstrated promising performance for EEG decoding, selecting an appropriate model architecture remains a costly and subject-specific process. Neural architecture search approaches partially address this issue but are computationally prohibitive in practice. In this work, we investigate a growing neural network strategy that adapts model expressivity during training, only expanding the architecture when required by the data. Building upon a lightweight convolutional neural network (CNN) architecture commonly used for a recent Brain-Computer Interface paradigm, code modulated visual evoked potential, decoding, we propose to grow only the classification head with two different growing modes, directed acyclic graph (DAG) and growing linear layers (GLL). This design explicitly targets subject-specific variability while preserving architectural priors learned from the literature. We evaluate our approaches on a five-class EEG dataset comprising 24 participants and compare them against fixed CNN, Riemannian TS-LDA, and GREEN architectures. Results show that the growing CNN significantly outperforms fixed CNN by increasing the accuracy by 1% while having 80% less neurons in the last layers. Our findings suggest that selectively growing the final layers of a deep network provides the capacity to adapt to subject-specific signatures, thus addressing the inter-subject variability. Moreover, we find an interesting trend where participants with better accuracy need fewer additional neurons to optimize their performance. Wednesday 0.02 Berlin FUZZ-IEEE Paper FUZZ 9 : Fuzzy control and robotics, sensors, fuzzy hardware, fuzzy architectures Session Chair: Uiliam Nelson Lendzion Tomaz Alves (Federal Institute of Education, Science and Technology of Paraná - IFPR) Robust Switched Control with Guaranteed Cost of a Furuta Pendulum Prototype using T-S Fuzzy model Uiliam Nelson Lendzion Tomaz Alves and Ricardo Breganon (Federal Institute of Education, Science and Technology of Paraná - IFPR) and Hugo Fernando Yamanaka, Dante Javier Solis Oncoy, Flávio Andrade Faria, and Marcelo Carvalho Minhoto Teixeira (São Paulo State University - UNESP) Abstract Abstract Takagi-Sugeno (T-S) fuzzy models provide an exact representation for a broad class nonlinear systems within defined operating regions. This work addresses the complexity of such models by representing selected nonlinearities through polytopic vertices and treating the remaining ones as norm-bounded terms. We propose a switched control strategy with guaranteed cost performance, where the controller gains are synthesized based on local models, offering a balanced alternative: it is less conservative than a single state-feedback gain, yet more computationally efficient than Parallel Distributed Compensation (PDC), as it eliminates the need for online membership function calculations in the control signal. Design conditions based on Linear Matrix Inequalities (LMIs) are presented. The method is experimentally validated on a Furuta pendulum prototype, with results confirming reduced conservatism and effective real-time operation under disturbances. PDC Output Feedback Control of T–S Fuzzy Systems with Premise Selection and Performance Constraints Hugo Fernando Yamanaka and Dante Javier Solis Oncoy (São Paulo State University - UNESP); Uiliam Nelson Lendzion Tomaz Alves (Federal Institute of Education, Science and Technology of Paraná - IFPR); Flávio Andrade Faria (São Paulo State University - UNESP); Márcia Luciana da Costa Peixoto (Universite Polytechnique Hauts-de-France); and Marcelo Carvalho Minhoto Teixeira (São Paulo State University - UNESP) Abstract Abstract This paper addresses the PDC output feedback control design for nonlinear systems represented by a Takagi–Sugeno (T–S) fuzzy model. Sufficient conditions are derived in terms of linear matrix inequalities (LMIs) to guarantee that the closed-loop eigenvalues of the local models lie in a prescribed $\mathfrak{D}$-region in the complex plane, ensuring stability and desired transient performance. To handle the inherent non-convexity of the output feedback design, a two-stage matrix decomposition strategy is adopted to improve numerical feasibility. Moreover, beyond the matrix decomposition itself, a systematic selection of premise variables for the feedback controller is proposed, providing additional design flexibility and leading to less conservative results. A numerical example illustrates the effectiveness of the proposed approach. Wednesday 0.01 London FUZZ-IEEE Position Paper, FUZZ-IEEE Paper FUZZ 8 : FUZZ-IEEE SS05 Fuzzy Federated Learning: Theoretical advances and novel applications (FL-A) & FUZZ Position paper Session Chair: Asier Urio-Larrea (Universidad Pública de Navarra, Spain) Fuzzy-Driven Weighting Policy for Adaptive Federated Learning on non-IID Data Antoni Jaszcz, Katarzyna Prokop, Dawid Połap, Marcin Woźniak, and Rafał Brociek (Silesian University of Technology) Abstract Abstract Federated Learning (FL) describes a model training technique in which the process is distributed among separate clients, holding private training data. During an FL round, locally-trained models from clients are aggregated, forming a new global model, which is then redistributed amongst the clients. During this process, the global model indirectly utilizes information from clients' private data. However, accurate and adaptive measurement of the client's contribution to the global model is an important aspect of the FL system, especially in na on-IID scenario. In this paper, we propose a fuzzy logic-based aggregation policy that dynamically estimates the learning demand of the global model and weights the importance of clients accordingly. The server evaluates class-wise precision and recall and infers per-class learning demand that reflects whether the model needs more training exposure to the samples of each class. Client-side class distribution and dataset size are then used to determine client's contribution in the aggregation process. Unlike learned scheduling policies, the proposed approach provides transparent and easily interpretable decision rules using Sugeno-type inference. The proposed methodology was tested on MNIST, EMNIST and CIFAR-10 benchmark datasets, showing promising results in a non-IID setting. FedFuzzyGA: A Federated Evolutionary Framework for Learning Compact Fuzzy Rule-Based Classifiers Asier Urio-Larrea and Humberto Bustince (Universidad Pública de Navarra), Javier Andreu-Perez (University of Essex), and Graçaliz Pereira Dimuro (Universidade Federal do Rio Grande do Sur) Abstract Abstract This paper presents FedFuzzyGA, a new federated learning framework designed to train interpretable Fuzzy Rule-Based Classifiers using decentralized and non-IID data. The primary challenge we address is that existing methods for federated fuzzy learning often result in excessively complex rule bases that are difficult to interpret. To address this problem, we adapted a centralized Genetic Algorithm into a privacy-preserving collaborative process through a method we call cyclical population synchronization. In this setup, clients perform local evolutionary searches on their private data and periodically exchange only their elite rule sets with a central server. This allows the system to share knowledge through a client-driven fusion strategy without ever compromising raw data privacy. Our experimental results show that FedFuzzyGA not only enables effective collaborative learning but often yields accuracy that meets or exceeds what a client could achieve through training in isolation. Most significantly, the framework ensures model transparency by generating compact, readable rule sets. This makes federated evolutionary learning of fuzzy rule-based classifiers a viable method for high-stakes applications that require human-understandable AI. A Practical Framework for Federated Fuzzy Clusterwise Regression, Preprocessing and Hyperparameter Tuning Morris Stallmann, Anna Wilbik, and Gerhard Weiss (Maastricht University) Abstract Abstract Federated learning has the potential to enable novel use cases in environments with data sharing constraints. While practitioners acknowledge the potential and researchers show applicability to many different use case domains, this new paradigm has not been widely adopted in practice yet. Various studies identify impact factors on adoption, including the communication overhead of the FL approach, non-identically or independently distributed (non-IID) training data, data quality problems, complexity of implementing best practices from (centralized) machine learning operations (MLOps), just to name a few. Moreover, federated regression problems remain an underexplored area as recent research activity has shown. Empirical Bernstein Certified Bandit-KL Mixing with Tree Stacking for Explainable Federated Intrusion Detection Md. Zubayer Ahmad Shibly and Md. Jahidul Islam (Department of computer science and engineering, Southeast University); Md.Mahbubur Rahman Sakib and FNU Al-Amain (Department of computer science and engineering Southeast University); Maybin K. Muyeba (DSAI Hub, School of Engineering and Environment, The University of Salford); and Mo Saraee (University of Salford) Abstract Abstract Network intrusion detection under federated learning (FL) must learn from distributed, heterogeneous non-independent and identically distributed (non-IID) client data where label or feature distributions differ across sites, while privacy constraints limit centralized aggregation. This setting suffers from severe label skew, unreliable client contributions, and the lack of trustworthiness guarantees in standard FedAvg. We propose TreeFed-CERT+ (Treebased Federated Learning with Certified Robustness & Trustworthiness Plus), a tree-based FL framework that combines nested cross-validation certificates with empirical Bernstein Lower Confidence Bounds (LCB) to estimate teacher reliability, shift-aware Kullback–Leibler (KL) regularisation for distributional stability and bandit-style adaptive weighting for round-wise mixing. Four clients train Extra Trees Classifier (ETC) teachers on severely skewed local splits; the server performs CERT+ mixing and trains a compact Hist Gradient Boosting (HGB) student via stacking on encoded features concatenated with teacher and mixture probabilities. On the IFPE Palmares Network Threat Detection Dataset, TreeFed-CERT+ achieves 99.80% accuracy, 99.96% ROC- AUC, and 0.0089 log-loss, and thus outperform uniform aggregation by 0.12%, stacking removal by 0.14%, and traditional baselines up to 2.4% on accuracy alone. Further detailed studies confirm nested cross-validation (CV) provides the largest contribution by 0.18% more, while stacking reduces log-loss by 53%. The framework ensures leakage-free evaluation via three tier partitioning and achieves ten times communication efficiency versus neural federated learning (FL). Shapley Additive exPlanations (SHAP), Local Interpretable Model-agnostic Explanations (LIME), Partial Dependence Plots (PDP) explainability analyses indicate that teacher probability features are primary drivers of the student’s predictions. Among the compared federated intrusion detection studies, TreeFed-CERT+ uniquely includes both nested cross-validation and XAI, supporting interpretable federated defence for decentralised security operations. Wednesday 0.01 London FUZZ-IEEE Paper FUZZ 10 : FUZZ-IEEE SS02 Fuzzy Machine Learning Session Chair: Jie Lu (University of Technology Sydney) A Fast Interpretable Fuzzy Tree Learner Javier Fumanal-Idocin (University of Essex), Raquel Fernandez-Peralta (Slovak Academy of Sciences), and Javier Andreu-Perez (University of Essex) Abstract Abstract Fuzzy rule-based systems have been mostly used in interpretable decision-making because of their interpretable linguistic rules. However, interpretability requires both sensible linguistic partitions and small rule base sizes, which are not guaranteed by many existing fuzzy rule-mining algorithms. Evolutionary approaches can produce high-quality models but suffer from prohibitive computational costs, while neural-based methods, such as ANFIS, have problems retaining linguistic interpretations. In this work, we propose an adaptation of classical tree-based splitting algorithms from crisp rules to fuzzy trees, combining the computational efficiency of greedy algorithms with the interpretability advantages of fuzzy logic. This approach achieves interpretable linguistic partitions and substantially improves running time compared to evolutionary-based approaches while maintaining competitive predictive performance. Our experiments on tabular classification benchmarks prove that our method achieves comparable accuracy to state-of-the-art fuzzy classifiers, while incurring significantly lower computational costs and producing more interpretable rule bases with constrained complexity. Code will be available upon publication. A Novel Fuzzy Adaptive Density-Based Clustering Algorithm Mohamed-Ali Belloum, Valentin Fouillard, and Jean-Philippe Poli (CEA) Abstract Abstract Density-based clustering methods are of paramount importance in data science. They present the advantage to automatically determine the number of clusters, relying on a definition of density. Among these methods, DBSCAN also detects outliers. However, it suffers from hyperparameters that may be difficult to set up, in particular by end-users. Different fuzzy approaches of DBSCAN have been proposed to fuzzify the frontiers between clusters and also the hyperparameters. In this paper, we present FA-DBC (Fuzzy Adaptive Density-Based Clustering), a fuzzy density-based clustering algorithm with self-adaptive thresholding, which significantly reduces manual parameter tuning. We compare the performances of this algorithm on different datasets. Adaptive Level-Set Fuzzy Heterogeneous Autoregressive Model for Realized Volatility Forecasting Leandro Maciel, Gabriel Navarro, and João Pedro Canabrava (University of Sao Paulo) and Fernando Gomide (University of Campinas) Abstract Abstract This paper suggests an Adaptive Level Set Fuzzy (ALSM) extension of the Heterogeneous Autoregressive (HAR) model for realized volatility forecasting. The ALSM-HAR model preserves the interpretability of the HAR structure while introducing nonlinear, regime-dependent, and adaptive dynamics through level set–based fuzzy inference and recursive estimation. The model is used to forecast the volatility of eight major cryptocurrencies one-step-ahead using high-frequency data. Out-of-sample results show that ALSM-HAR consistently performs better than standard HAR and its main extensions in both forecast accuracy and statistical robustness, particularly under mean squared error criteria. ALSM-HAR advances the current state of the art in the area since it significantly improves the effectiveness of forecasting in highly volatile and uncertain markets. Fuzzy Logic-Driven Reward Optimization for Deep Reinforcement Learning in Carbon Emission Trading Seyed Ali Hosseini (Italy, Politecnico di Milano) and Francesco Grimaccia and Alessandro Niccolai (Politecnico di Milano) Abstract Abstract This study presents a novel algorithmic trading framework for carbon emission markets, integrating fuzzy logic with deep reinforcement learning (DRL) to optimize trading strategies under uncertainty. We employ an Advantage Actor-Critic (A2C) algorithm with a fuzzy logic-enhanced reward function, leveraging triangular membership functions and a predefined rule base to model volatile market dynamics, including price fluctuations and technical indicators. Using five years of tick-by-tick carbon emission data, the framework is validated through a rigorous walk-forward approach, demonstrating robust generalization. Results show significant improvements in cumulative return, Sharpe ratio, and Sortino ratio, with the fuzzy reward function enabling stable and adaptive decision-making in complex market conditions. By synergizing fuzzy systems with DRL, this work advances trading strategy development, offering a scalable approach for dynamic financial environments. Fuzzy Accuracy Compensates for Label Subjectivity in Classification of Skin Tone Using Wearable Photoplethysmography Signals Padmini Krishnadas (National Physical Laboratory), Urs Hackstein (Mittelhessen University of Applied Sciences), Alen Bosnjakovic (Institute of Metrology of Bosnia and Herzegovina), and Philip Aston (National Physical Laboratory) Abstract Abstract We consider the problem of classification of skin tone using photoplethysmography (PPG) signals with labels of the ordinal six-class Fitzpatrick skin tones. A typical accuracy for this task is a poor 40-55 \%. However, the labels are subjectively determined by comparing the skin with a colour chart, and hence contain widespread small-scale inaccuracies. By working with a ``fuzzy accuracy'', which deems a prediction of skin tone class to be correct if its difference from the labelled class is not greater than one, much higher accuracy is obtained which provides more convincing evidence that skin tone can be accurately predicted from PPG signals. Three machine learning approaches were used, namely deep learning or tree-based approaches on raw PPG signals, deep learning on image representations of the signals generated by the Symmetric Projection Attractor Reconstruction (SPAR) method, and machine learning on features extracted from the signals. The first method also employed a fuzzy version of the cross entropy loss function, which gave the best results. Tree-based models on raw signals give accuracies up to 55 \% and higher fuzzy accuracies up to 96 \%, while deep learning models on the SPAR images obtained lower results of 44 \% accuracy and 85~\% fuzzy accuracy. The machine learning on PPG features gave similar results to the SPAR method with accuracy of 42 \% and fuzzy accuracy of 87 \%. We have shown that classification of skin tone using PPG signals is possible with high fuzzy accuracy which implies that our modelling approach enables accurate prediction of skin tone class within at most one class of the observer's choice of class, from which we conclude that PPG signals are affected by skin tone in a discernible way. Detecting Complex Money Laundering Patterns with Incremental and Distributed Graph Modeling Haseeb Tariq (TU Eindhoven, ING Bank); Alen Kaja (Adyen); and Marwan Hassani (TU Eindhoven) Abstract Abstract Money launderers take advantage of limitations in existing detection approaches by hiding their financial footprints in a deceitful manner. They manage this by replicating transaction patterns that the monitoring systems cannot easily distinguish. As a result, criminally gained assets are pushed into legitimate financial channels without drawing attention. Algorithms developed to monitor money flows often struggle with scale and complexity. The difficulty of identifying such activities is further intensified by the (persistent) inability of current solutions to control the excessive number of false positive signals produced by rigid, risk-based rules systems. We propose a framework called ReDiRect (REduce, DIstribute, and RECTify), specifically designed to overcome these challenges. The primary contribution of our work is a novel framing of this problem in an unsupervised setting; where a large transaction graph is fuzzily partitioned into smaller, manageable components to enable fast processing in a distributed manner. In addition, we define a refined evaluation metric that better captures the effectiveness of exposed money laundering patterns. Through comprehensive experimentation, we demonstrate that our framework achieves superior performance compared to existing and state-of-the-art techniques, particularly in terms of efficiency and real-world applicability. For validation, we used the real (open source) Libra dataset and the recently released synthetic datasets by IBM Watson. Our code and datasets are available at https://github.com/mhaseebtariq/redirect. Wednesday 0.02 Berlin IJCNN Paper Neuromorphic and Spiking Neural Networks Session Chair: Toshihisa Tanaka (Tokyo University of Agriculture and Technology), Guangzhi Tang (Maastricht University) Scaling high-dimensional synaptic plasticity on neuromorphic hardware: parameter mapping and vectorized implementation Amani Atoui, Jakob Kaiser, Philipp Spilger, Eric Mueller, and Johannes Schemmel (Heidelberg University) Abstract Abstract We present analog accelerated neuromorphic hardware, specifically BrainScaleS-2, as a reliable and efficient tool for studying high-dimensional biologically-plausible synaptic plasticity. We focus on a calcium-based synaptic plasticity rule that integrates different neuron and synapse variables on multiple timescales. Compared to our previous work, we automate the calibration of the analog circuits which emulate the calcium dynamics and scale the plasticity rule to multiple synapses using the SIMD extension of the embedded digital processor. To validate our implementation, we used standard stimulation protocols. The analog nature of the hardware, reduced-precision arithmetic, and the larger integration time steps of the plasticity rule update compared to the simulation cause numerical differences between the emulation and simulation results. Nevertheless, our results show that the implementation can faithfully emulate the model dynamics for multiple synapses in parallel. The final judgment on the implementation of the plasticity rule depends on its applications, specifically modeling cognitive functions in computational neuroscience and machine learning. Our approach aims at exploiting the energy efficiency and acceleration factor of BrainScaleS-2 as well as its analog nature for the study and application of synaptic plasticity. Automated Model-to-Transistor Design Framework for Memristor Crossbar-Based Spiking Neural Network Architectures Hanwen Xuan, Zalfa Jouni, Mikhail Manokhin, Dimitri Galayko, and Haralampos-G. Stratigopoulos (Sorbonne Université, CNRS, LIP6) Abstract Abstract In-memory computing using memristor crossbar arrays provides an energy- and latency-efficient approach for implementing the matrix–vector multiplications required in neural network layers. In parallel, neuromorphic computing based on spiking neural networks (SNNs) offers a compelling alternative to conventional artificial neural networks (ANNs) in terms of energy efficiency and inference speed. Combining these two paradigms, namely executing SNNs on memristor crossbar arrays, presents a highly promising direction. In this work, we present a transistor-level design of a memristor crossbar array architecture for SNNs incorporating analog spiking neurons. A novel neuron circuit is developed to closely match the behavioral dynamics of neurons used in the widely adopted snnTorch framework for SNN training. The corresponding behavioral neuron model captures all key neuron parameters, enabling hardware-aware training within snnTorch. We further introduce an automated framework that generates SPICE netlists directly from trained models, including the mapping of trained weights to programmable memristor conductance values. The proposed design and framework are validated on two SNN benchmarks for ECG and MNIST classification, demonstrating close agreement between snnTorch and transistor-level inference results. Knowledge Distillation at the Edge: Lightweight Deep Neural Networks Deployed on Neuromorphic Hardware Rashedul Islam, Shahanur Alam, Chris Yakopcic, and Nayim Rahman (University of Dayton); Simon Khan (Air Force Research Laboratory); and Tarek Taha (University of Dayton) Abstract Abstract Neuromorphic processors are highly suitable for edge intelligence due to their ultra-low power consumption and event-driven computation. However, deploying lightweight neural networks on such hardware often results in accuracy degradation caused by model compression, quantization, and ANN-to-SNN conversion. Improving accuracy while maintaining low energy and computational cost therefore remains a key challenge. In this paper, we investigate the impact of knowledge distillation for improving lightweight model performance on the Brainchip Akida. To the best of our knowledge, this is the first neuromorphic hardware-based study that systematically evaluates knowledge distillation. Two distillation strategies are explored: (1) knowledge distillation before quantization and (2) knowledge distillation applied after quantization. Experimental results show that knowledge distillation significantly mitigates quantization-induced accuracy loss, with post-quantization distillation consistently outper-forming pre-quantization distillation. Even with extreme model compression (95% parameter reduction compared to the teacher models), the student networks achieved up to 11.45% accuracy improvement over the baseline, while preserving performance after ANN-to-SNN conversion. The results show that knowledge distillation is effective for both complex and lightweight models, enabling compact networks to achieve higher accuracy with minimal energy overhead. Thus, the proposed method is a key enabler for accurate and energy-efficient deployment of neuromorphic processing at the edge. SpikeDiffusion: A Fully Spiking Structure-Guided Diffusion Framework for Energy-Efficient Image Generation Xue Han (Beijing Normal-Hong Kong Baptist University), Zhiwen Luo (Concordia University), Wenchuan Zhang (Beijing Normal-Hong Kong Baptist University), Nizar Bouguila (Concordia University), and Weifeng Su and Wentao Fan (Beijing Normal-Hong Kong Baptist University) Abstract Abstract Spiking Neural Networks (SNNs) offer energy-efficient and biologically plausible computation, yet their discrete spiking dynamics make high-fidelity image generation challenging, particularly in preserving global structure and fine-grained details. To address this challenge, we propose SpikeDiffusion, a fully spiking structure-guided diffusion framework for stable image generation. Unlike attention-based global interaction designs, SpikeDiffusion employs a spiking variational autoencoder (VAE) to produce a structure-preserving prior, which is generated once and reused throughout the denoising process via a lightweight early-fusion strategy, with the greatest benefit in the high-noise early stages. This design enhances structural consistency while avoiding the substantial computational and memory overhead of attention mechanisms. Furthermore, we develop a fully spiking U-Net with multi-step temporal dynamics and membrane-potential integration, enabling end-to-end spike-based diffusion modeling. Extensive experiments on standard benchmarks demonstrate that SpikeDiffusion achieves competitive image quality among SNN-based generative models with significantly reduced energy consumption, highlighting its suitability for neuromorphic and edge deployment. The code is available at https://github.com/hhx0511/spiking-diffusion. Sparsity-Adaptive Sharpness-Aware Minimization Shiryu Ueno, Yoshikazu Hayashi, and Kunihito Kato (Gifu University) Abstract Abstract Deploying deep neural networks in real-world settings requires models that are both compact and robust to common corruptions. However, at deployment-relevant high sparsity, standard pruning pipelines often degrade corruption robustness, and existing sharpness-aware training/pruning approaches provide limited robustness gains. We address this issue by introducing Sparsity-Adaptive Sharpness-Aware Minimization (SA-SAM), which derives a sparsity-dependent SAM/ASAM perturbation radius by keeping the mean absolute perturbation (an L1-based proxy) approximately invariant as sparsity increases. As a simple complementary option, we evaluate Magnitude-Weighted Hessian (MWH), derived from a second-order removal-path analysis, yielding an importance proportional to Diag(F)_i |w_i|, where Diag(F) is the diagonal empirical Fisher used as a curvature proxy in our implementation. Across CIFAR-10-C, CIFAR-100-C, and ImageNet-100-C, our approach achieved stronger corruption robustness than the considered pruning baselines at 80--90\% sparsity, while preserving clean accuracy. We additionally quantify the robustness--throughput trade-off by reporting measured inference throughput under sparse execution at deployment-relevant sparsity levels. Wednesday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 17 - Related Topics III Session Chair: Carlos Coello Coello (Cinvestav; Faculty of Excellence of the School of Engineering and Sciences, Tecnologico de Monterrey, Monterrey, Mexico) Evolutionary Edge Bundling as a Large-Scale Multi-Objective Optimization Problem: A Comparative Study Chihiro Noda and Ryosuke Saga (Osaka Metropolitan University) Abstract Abstract Evolutionary edge bundling reduces visual clutter in network visualization by optimizing edge geometries under multiple objectives, yet the optimizer’s contribution is still poorly understood. We present a controlled comparison of five representative multi-objective evolutionary algorithms—MOEA/D, SMPSO, NSGA-II, NSGA-III, and SPEA2—within a unified edge-bundling pipeline under an identical computational budget. Fixing the control-point representation, objective functions, datasets, and evaluation protocol allows observed differences to be attributed mainly to optimization behavior. Experiments on one synthetic network and two real-world airline networks highlight optimizer-dependent trade-offs among convergence speed, Pareto-set coverage, and visual diversity, offering practical guidance for LSGO-type edge-bundling tasks. MEASE: Mixed Type Evolutionary Algorithm for Exceptional Survival Model Mining Bruno Fonseca, Déborah Yamamoto, and Renato Vimieiro (UFMG) Abstract Abstract Exceptional survival subgroup discovery aims to identify interpretable descriptions of subpopulations whose time-to-event behavior differs from a reference group under censoring. Most descriptive approaches for survival data still rely on user-defined discretization of numerical variables, which is costly, dataset-dependent, and can strongly affect the discovered patterns. This paper proposes an evolutionary subgroup discovery and exceptional model mining method for survival analysis that induces numerical constraints during the search, enabling mixed categorical and numerical descriptions without manual discretization. The Mixed Type Evolutionary Algorithm for Exceptional Survival Model Mining (MEASE) maintains a redundancy-aware top-K archive and uses adaptive crossover and mutation rates with restart mechanisms to balance exploration and exploitation while promoting diversity. Experiments on six benchmark survival datasets, with thirty runs per dataset, compare the proposed approach against representative baselines and state-of-the-art ant colony optimization methods. The results show that the proposed method achieves competitive exceptionality while typically improving subgroup coverage, without requiring any user preprocessing on the datasets. Evolutionary Online Swarm Path Planning for UAV Formation Transition in Dynamic Environments Jyun-Siang Huang, Jia-He Tee, and Chuan-Kang Ting (National Tsing Hua University) Abstract Abstract Unmanned aerial vehicle (UAV) swarms have attracted increasing attention for cooperative missions such as surveillance, environmental monitoring, and formation flight. In formation transition tasks, a swarm must move from an initial formation to a prescribed target formation while respecting kinematic limits, maintaining safe inter-UAV separation, and avoiding environmental hazards. The transition tasks become more challenging in dynamic environments, where observable obstacles must be avoided proactively, and unobserved disturbances can perturb executed motion and degrade formation accuracy unless timely path planning is performed. This study proposes the evolutionary online swarm path planning (EOSPP) framework, which formulates swarm navigation as a sequence of time-bounded constrained optimization problems solved along the flight. The EOSPP employs evolution strategies (ES) as the backbone optimizer to generate control vectors for each UAV within a fixed computation budget, explicitly enforcing speed limits, minimum inter-UAV separation, and obstacle clearance during optimization. Unobserved wind disturbances are modeled as execution-time perturbations and are compensated through repeated online path planning. Experiments in a bounded two-dimensional workspace show that EOSPP achieves collision-free convergence to the target formation in the presence of uniformly sampled obstacles and maintains convergence under unobserved wind gusts, demonstrating the practicality of evolutionary online path planning for safety-critical swarm formation transitions. Limits of Lamarckian Evolution Under Pressure of Morphological Novelty Jed Muff, Karine Miras, and Agoston Eiben (Vrije Universiteit Amsterdam) Abstract Abstract Lamarckian inheritance has been shown to be a powerful accelerator in systems where the joint evolution of robot morphologies and controllers is enhanced with individual learning. Its defining advantage lies in the offspring inheriting controllers learned by their parents. The efficacy of this option, however, relies on morphological similarity between parent and offspring. In this study, we examine how Lamarckian inheritance performs when the search process is driven toward high morphological variance, potentially straining the requirement for parent-offspring similarity. Using a system of modular robots that can evolve and learn to solve a locomotion task, we compare Darwinian and Lamarckian evolution to determine how they respond to shifting from pure task-based selection to a multi-objective pressure that also rewards morphological novelty. Our results confirm that Lamarckian evolution outperforms Darwinian evolution when optimizing task-performance alone. However, introducing selection pressure for morphological diversity causes a substantial performance drop, which is much greater in the Lamarckian system. Further analyses show that promoting diversity reduces parent-offspring similarity, which in turn reduces the benefits of inheriting controllers learned by parents. These results reveal the limits of Lamarckian evolution by exposing a fundamental trade-off between inheritance-based exploitation and diversity-driven exploration. Graph Models to Forecast Road Traffic Crashes and Resource Optimization on Brazilian Federal Roads Júlio César de Freitas Taveira (University of Pernabuco, University of Highway Police) and Hugo de Andrade Amorim Neto, Andina Alay Lerma, and Fernando Buarque de Lima Neto (University of Pernabuco) Abstract Abstract Road traffic crashes present a critical public health challenge, requiring integrated predictive and operational strategies. This paper proposes an end-to-end framework for Brazilian federal highways that bridges crash risk forecasting with proactive resource allocation. By modeling the highway infrastructure as a complex undirected graph, we evaluate three GNNs architectures (GCN, GAT, and EvolveGCN) for road crash prediction. These predictions inform a multi-objective optimization layer using NSGA-II, MOPSO, and P-ACO to maximize risk coverage while minimizing response times and costs. Results show that GCN and EvolveGCN yield stable predictive performance, while optimization outcomes vary with regional network topology. Notably, the framework yields more substantial improvements in decentralized units, suggesting that automated spatial analysis can provide valuable insights into resource positioning in areas where road connectivity is more dispersed. This work validates the feasibility of a unified, data-driven pipeline that transforms statistical risk into actionable patrolling strategies, providing a decision-support tool for road safety management. Boosting the Performance of Evolutionary Multi-Instance Classification with Transfer Learning Nadia Omri and Maha Elarbi (SMART Lab, Computer Science Department, ISG, University of Tunis, Tunis, Tunisia); Slim Bechikh (ComCOL Lab, School of Computer Science, University of Nottingham, United Kingdom); and Carlos Artemio Coello Coello (Computer Science Department, CINVESTAV-IPN, Av Instituto Politécnico Nacional 2508, San Pedro Zacatenco, Gustavo A. Madero, 07360 Mexico City, Mexico; Faculty of Excellence of the School of Engineering and Sciences, Tecnologico de Monterrey, Monterrey, Mexico) Abstract Abstract Multiple Instance Learning (MIL) has proven effective for learning from weakly labeled data, where labels are provided at the bag level rather than at the instance level. Among existing MIL approaches, MILEIS (Multiple-Instance Learning with Evolutionary Instance Selection) is an evolutionary instance selection–based framework that reduces instance ambiguity by mapping bags into a discriminative feature space using optimized instance prototypes. However, many existing MIL approaches including MILEIS, rely solely on target task data and do not exploit knowledge from related tasks, which can limit their performance in data-scarce scenarios. To address this limitation, transfer learning enables the reuse of knowledge from source tasks to improve target task learning. In this paper, we propose TL-MILEIS, a transfer learning extension of MILEIS that incorporates a task similarity evaluation mechanism to ensure safe knowledge transfer. The proposed approach first performs instance selection following the MILEIS paradigm, and then evaluates the similarity between source and target tasks to determine whether knowledge transfer is beneficial. When task relevance is established, discriminative information from the source task is leveraged to improve target task learning; otherwise, standard MILEIS is applied. Experimental results on multiple benchmark datasets show that TL-MILEIS consistently outperforms classical MIL methods, evolutionary MIL approaches, and existing multiple instance transfer learning techniques, highlighting the effectiveness of incorporating transfer learning into evolutionary MIL frameworks. Wednesday 0.05 Paris IJCNN Paper AI for Cybersecurity, Public Safety, and Surveillance Session Chair: Fabrizio Pittorino (Politecnico di Milano), Israel Efraim de Oliveira (Universidade Federal de Santa Catarina) Hybrid CNN-SSM Model for Robust Multi-Tab Website Fingerprinting over Tor Zulu Okonkwo, Zhe Hou, Ernest Foo, Qinyi Li, and Zahra Jadidi (Griffith University) Abstract Abstract Website fingerprinting (WF) threatens anonymity networks such as Tor by inferring visited sites from encrypted traffic. While recent deep learning attacks report high accuracy, they are often evaluated under simplifying assumptions: single-tab browsing or multitab settings where the number of open tabs is known. In addition, most work treats WF as a pure classification problem and provides limited insight into what the model learns. We present a multitab WF attack that handles an unknown number of concurrently opened tabs and explicitly incorporates explainability into evaluation. Our pipeline segments traces, encodes packet, burst, and timing-level statistics that remain informative under common Tor defences, and learns representations with a dilated residual CNN for hierarchical local structure and a state-space model for long-range temporal dependencies. Experiments on public WF benchmarks, including defended traffic, show that our method outperforms prior state-of-the-art methods in multitab settings while remaining competitive in single-tab classification. Deep Learning for Acoustic Side-Channel Attacks on Keyboards. Sterile vs. Noisy Environments Julia Przybytniowska and Adam Żychowski (Warsaw University of Technology) and Jacek Mańdziuk (Warsaw University of Technology, AGH University of Krakow) Abstract Abstract Acoustic side-channel attacks (ASCAs) on keyboards pose a significant security threat, yet existing research often overestimates their viability by focusing on sterile, noise-free environments. This paper presents a comprehensive study on the robustness and generalization of modern hybrid neural architectures (CoAtNet, MOAT, Swin Transformer) for this task. Under controlled conditions, our CoAtNet-based model establishes a new state-of-the-art accuracy of 95.23% on the most popular benchmark. Furthermore, to address the challenge of ASCA in noisy, realistic conditions, we propose a novel benchmark dataset recorded across four realistic noisy environments, and uncover a crucial phenomenon of asymmetric generalization: while models trained on noisy data generalize remarkably well to clean audio, models trained exclusively on clean data fail catastrophically in the presence of noise. Furthermore, we demonstrate that integrating Large Language Models (LLMs) as a post-processing step creates a robust information-recovery pipeline, effectively "denoising" corrupted sequences and reducing error rates to near-zero. Finally, we show the practical impact of the attack through a probabilistic guessing strategy, which recovers complex 12-character passwords in a mean of only 87.76 guesses. Our findings suggest that the integration of diverse training data and LLM-based semantic correction makes ASCAs a potent threat in real-world, unpredictable environments. Towards Detecting Knife-Carrying Individuals Using YOLO and mmWave Radar EAR Representations Ogonna Okafor, David Ada Adama, and Doratha Vinkemeier (Nottingham Trent University) Abstract Abstract Millimetre-wave (mmWave) radar offers a privacy-preserving sensing modality for indoor security applications, but effective learning from radar data remains challenging due to its high dimensionality and non-visual structure. This paper presents an exploratory study on the adaptation of standard vision-based object detectors to radar-derived representations. Specifically, the study investigates whether Elevation-Azimuth-Range (EAR) tensors obtained from a 77 GHz FMCW radar can be transformed into 2D pseudo-images suitable for training off-the-shelf YOLO detectors without architectural modification. Using a small, controlled indoor dataset with weak cross-modal supervision from synchronised RGB cameras, multiple YOLO variants are evaluated under identical training conditions, with performance analysed using both conventional detection metrics and safety-oriented measures such as Miss Rate and False Alarm Rate. The results demonstrate that modern convolutional detectors can learn discriminative patterns from EAR-based pseudo-images despite the absence of explicit geometric calibration between radar and camera coordinate frames, requiring the detector to implicitly learn a cross-modal correspondence. These findings highlight both the feasibility and limitations of representation transfer from radar to vision models. The study is intended as a methodological exploration rather than a deployment-ready system. Wednesday 0.10 Sydney IJCNN Paper Enterprise LLMs, RAG, and Knowledge-Augmented Systems Session Chair: Thiago Oliveira-Santos (UFES, I2CA), Cleber Zanchettin (Universidade Federal de Pernambuco, Northwestern University) TraceRAG: Targeted Graph Traversal via Adaptive Query Resolution for Multi-Hop Retrieval Giuseppe Trimigno (University of Parma), gianfranco lombardo (Università di Parma), and Stefano Cagnoni (University of Parma) Abstract Abstract Retrieval-Augmented Generation (RAG) has become the standard for extending Large Language Models (LLMs) with non-parametric, up-to-date knowledge. Standard RAG systems face challenges with multi-hop reasoning, which involves combining information from separate and unconnected passages to derive answers. While recent graph-based approaches leverage Knowledge Graphs (KGs) to bridge these gaps, they typically rely on single-shot global query embeddings to initialize retrieval, which often leads to semantic drift and weak modeling of sequential dependencies. We introduce TraceRAG, a novel framework that reframes multi-hop retrieval as an adaptive and sequential resolution problem. Given a complex query, we decompose it into atomic sub-questions, which are resolved sequentially by grounding them in factual triples extracted from the knowledge graph. These grounded triples serve as high-precision structural seeds for a Personalized PageRank (PPR) traversal, enabling controlled and interpretable graph exploration across hops. By explicitly binding intermediate resolution steps to verified structure, our framework mitigates semantic drift and preserves the query intent throughout retrieval. Extensive experiments on MuSiQue, HotpotQA, and 2WikiMultiHopQA demonstrate that TraceRAG consistently outperforms the strongest baselines, achieving gains of up to 3.5 for Recall@5 and 4.6 for F1 score. Moreover, it remains robust even when the underlying graph is derived from smaller, open-weights models, highlighting its effectiveness and enabling scalable, structured RAG systems. Adaptive Targeted Dynamic Chunking for Tokenization-Free Hierarchical Model Thang Dang (Fujitsu Research of America) and Akira Nakagawa, Kenichi Kobayashi, and Koichi Shirahata (Fujitsu Limited) Abstract Abstract Tokenization-free hierarchical models are emerging as a promising alternative to traditional Large Language Models (LLMs), addressing inherent preprocessing issues such as vocabulary design complexity, out-of-vocabulary (OOV) errors, and language-specific constraints. However, a significant challenge in these byte-level methods is the optimization of the compression ratio, a critical factor that dictates model performance for processing bytes data via chunks. In this paper, we propose Adaptive Targeted Dynamic Chunking (ATDC), a novel byte-compression control mechanism designed to enhance the effectiveness of dynamic chunking within hierarchical architectures. Our approach utilizes curriculum learning to progressively adjust the compression ratio during training, transitioning from low to high compression to stabilize the learning process. We provide an analysis establishing the relationship between the target compression ratio and Bytes-Per-Innermost-Chunk (BPIC), allowing for tracking of chunk-size evolution throughout the training phase. Evaluations conducted on the FineWeb-Edu 100B dataset demonstrate that hierarchical models equipped with ATDC achieve competitive Bits-Per-Byte (BPB) performance compared to conventional baselines operating at both byte and token levels. Furthermore, the proposed method exhibits more stable training dynamics and superior final performance across diverse downstream tasks compared to models using fixed compression ratios, while maintaining the inherent robustness and flexibility of byte-level processing. RAG-ME - Retrieval Augmented Generation Multi-database Evaluation: Towards a Fairer Assessment Guilherme Goes Zanetti (Universidade Federal do Espírito Santo); Thiago Paixão (Instituto Federal do Espírito Santo); and Filipe Mutz, Sérgio Mucciaccia, Claudine Badue, Alberto F. De Souza, and Thiago Oliveira-Santos (Universidade Federal do Espírito Santo) Abstract Abstract Traditional Retrieval Augmented Generation (RAG) system evaluation often suffers from optimistic bias when using a single text source for both question generation and retrieval context. This paper introduces RAG-ME, a framework whose core contribution lies in decoupling the text source for question generation from the source used for retrieval context. RAG-ME employs one source to generate evaluation questions and a distinct, similar source to provide the context for answer retrieval. System-generated answers are then evaluated with LLM-as-a-Judge. Our experiments, conducted on scientific papers and veterinary pathology textbooks, indicate that RAG-ME's dual-database approach yields a fairer assessment, significantly reducing bias and more effectively differentiating model capabilities compared to conventional single-database methods. Furthermore, RAG-ME's capacity of analyzing retrieval parameters (e.g., chunk size and overlap) highlights its practical value for system optimization, offering a cost-effective and reliable tool for advancing fairer RAG assessment. LogiPart: Local Large Language Models for Data Exploration at Scale with Logical Partitioning Tiago Fernandes Tavares (Insper) Abstract Abstract The discovery of deep, steerable taxonomies in large text corpora is currently restricted by a trade-off between the surface-level efficiency of topic models and the prohibitive, non-scalable assignment costs of LLM-integrated frameworks. We introduce \textbf{LogiPart}, a scalable, hypothesis-first framework for building interpretable hierarchical partitions that decouples hierarchy growth from expensive full-corpus LLM conditioning. LogiPart utilizes locally hosted LLMs on compact, embedding-aware samples to generate concise natural-language taxonomic predicates. These predicates are then evaluated efficiently across the entire corpus using zero-shot Natural Language Inference (NLI) combined with fast graph-based label propagation, achieving constant $O(1)$ generative token complexity per node relative to corpus size. We evaluate LogiPart across four diverse text corpora (totaling $\approx$140,000 documents). Using structured manifolds for \textbf{calibration}, we identify an empirical reasoning threshold at the 14B-parameter scale required for stable semantic grounding. On complex, high-entropy corpora (Wikipedia, US Bills), where traditional thematic metrics reveal an ``alignment gap,'' inverse logic validation confirms the stability of the induced logic, with individual taxonomic bisections maintaining an average per-node routing accuracy of up to 96\%. A qualitative audit by an independent LLM-as-a-judge confirms the discovery of meaningful functional axes, such as policy intent, that thematic ground-truth labels fail to capture. LogiPart enables frontier-level exploratory analysis on consumer-grade hardware, making hypothesis-driven taxonomic discovery feasible under realistic computational and governance constraints. A Summarization Framework with Self-Improvement for RAG Retrieval in the Legal Domain Yuriy Perezhohin, Flavio Ivo Riedlinger, Victor Costa, and Mauro Castelli (NOVA IMS) Abstract Abstract Retrieval-Augmented Generation (RAG) systems are increasingly used for legal question answering, but their effectiveness depends critically on the quality of the representations indexed during retrieval. Legal documents pose particular challenges due to their length, hierarchical structure, and domain-specific semantics, which can result in semantically weak or incomplete representations. This work proposes a quality-controlled representation framework that targets the retrieval stage of legal RAG systems. Documents are segmented into fixed-size chunks and summarized using a lightweight language model. A second model, acting as a synthetic evaluator following the G-EVAL methodology, assesses the semantic adequacy of each summary and triggers refinement when necessary. Only validated summaries are admitted into the vector database.The acceptance criterion governing this process is analyzed through a sensitivity study that characterizes the trade-off between retrieval effectiveness and computational overhead. Experiments on the LEGALBENCH-RAG benchmark demonstrate that quality-controlled summaries yield substantial gains over untreated indexing. In a direct comparison using the same model, pairing text-ada-large-3 with our validated summaries increases accuracy at k =8 from 36.63% to 72.70%. Because retrieval operates on compact, semantically verified representations rather than raw long chunks, the approach improves computational efficiency and narrows the performance gap between proprietary and lightweight models: MultiQA with validated summaries reaches 64.24% accuracy, outperforming the closed-source text-ada-large-3 baseline without quality control. Interactive Optimization Modeling via Preference Intermediate Representation: Automating Domain Adaptation from Raw Data Li Chen and chun yu (Tsinghua University) Abstract Abstract Existing natural language-to-optimization-model approaches face significant challenges in real-world scenarios, including handling large-scale heterogeneous data and accommodating evolving user preferences. To address these, we propose the Preference Intermediate Representation (PIR), which decouples problem data from preference logic and separates preference models from implementation code. PIR supports the iterative refinement of user preferences within the intermediate layer and enhances the scalability and flexibility of optimization modeling. Furthermore, we introduce an automated framework to generate domain-specific PIR components (Item Selectors and Preference Expressions) from raw problem data and task descriptions. We evaluate our approach in the domain of student course timetabling. Experimental results demonstrate that our framework can effectively generate high-quality domain-specific PIR components. Furthermore, when integrated into an interactive dialogue system, the PIR-based natural language parsing approach achieves 93.50\% exact accuracy, significantly outperforming baselines. Wednesday 0.11 Cape Town IJCNN Paper IJCNN SS31 Systems-Theoretic Approaches to Learning 4.0: From Classical to Quantum Neural Networks Session Chair: Vignesh Narayanan (University of South Carolina), Avimanyu Sahoo (University of Alabama in Huntsville), Krishnan Raghavan (Argonne National Laboratory) Safe Adaptive Output Tracking Controller for Robotic Manipulators Ritirupa Dey (University of South Carolina), Avimanyu Sahoo (University of Alabama in Huntsville), and Vignesh Narayanan (University of South Carolina) Abstract Abstract Accurate tracking control of robotic manipulators often requires reliable velocity measurements and precise knowledge of system dynamics. However, in practice, joint velocity sensors are either unavailable or prone to noise, and exact robot dynamics may not always be known. To address these challenges, especially when the robot trajectories are required to adhere to prescribed safety constraints, in this paper, we propose a neural network (NN)-based adaptive observer and controller framework that relies solely on joint position feedback for an n-link robotic manipulator. To this end, we develop an adaptive observer with a custom online weight-tuning law to estimate the joint velocities of the robotic system. Using the estimated states, we design a tracking controller that learns the robot dynamics online and captures the model nonlinearities. Using the proposed closed-loop observer–controller system, we derive sufficient conditions based on control-barrier function (CBF) that guarantee system trajectories respect prescribed safety constraints, thereby ensuring safe robotic operation. We include numerical simulations on a two-link robotic manipulator system to validate the effectiveness of the proposed controller. Representation Learning of Distribution Flows using Ensemble Control Systems Wei Zhang, Lin Tang, and Jr-Shin Li (Washington University in St. Louis) Abstract Abstract Constructing interpretable representation learning models for flows of probability distributions has become a thriving research area in machine learning and artificial intelligence. This effort not only sparks new research avenues but also provides distinctive insights into established fields such as image processing. For example, the recently developed flow matching (FM) model has demonstrated its effectiveness in addressing image restoration problems. However, a unified, comprehensive, and interpretable framework for learning probability flows from data in a general setting remains underexplored. To fill this gap, we develop an ensemble control system (ECS) model for learning probability flows. Our model is represented as an ensemble of heterogeneous control systems, with the control inputs acting as time-dependent trainable parameters. The heterogeneous dynamics and time-dependent parameters significantly enhance the model's capabilities, making it exceptionally powerful. To further capitalize on these strengths, we introduce a moment kernel transform that creates a reduced kernel representation of the ECS model over a reproducing kernel Hilbert space. This approach enables efficient training without compromising learning performance. We demonstrate the significant advantages of the ECS model through various image restoration tasks and provide a detailed comparison with baseline FM-based image processing models. Resource Constrained Safe Reinforcement Learning for Lane Following of Wheeled Mobile Robots Mahtab Noor Shaan (University of Alabama in Huntsville); Vignesh Narayanan (University of South Carolina, Columbia); and Avimanyu Sahoo (University of Alabama in Huntsville) Abstract Abstract This paper proposes a safe and resource-efficient near-optimal lane following control scheme for wheeled mobile robots (WMRs). The method performs online co-optimization of both the tracking control inputs and their execution instants, while rigorously enforcing lane-keeping safety and actuator constraints. A differential game–inspired event-based adaptive dynamic programming (ADP) and reinforcement learning (RL)-based scheme is developed to synthesize a constrained control policy and adaptively schedule control execution instants. A non-quadratic performance index is formulated that simultaneously penalizes the system states, the constrained control effort, the input error induced by aperiodic (event-like) execution, and a reciprocal control barrier function (CBF) for lane keeping, enforcing safety. The associated Hamilton–Jacobi–Isaacs (HJI) equation is solved approximately using a single critic neural network. Simulation results for state and input safety constraints demonstrate the effectiveness of the proposed approach, yielding up to a 61\% reduction in control updates compared with continuous-feedback baselines. How Embeddings Shape Graph Neural Networks: Classical vs Quantum-Oriented Node Representations Nouhaila Innan (New York University (NYU) Abu Dhabi, UAE); Antonello Rosato (Sapienza University of Rome); and Alberto Marchisio and Muhammad Shafique (New York University (NYU) Abu Dhabi, UAE) Abstract Abstract Node embeddings act as the information interface for graph neural networks, yet their empirical impact is often reported under mismatched backbones, splits, and training budgets. This paper provides a controlled benchmark of embedding choices for graph classification, comparing classical baselines with quantum-oriented node representations under a unified pipeline. We evaluate two classical baselines alongside quantum-oriented alternatives, including a circuit-defined variational embedding and quantum-inspired embeddings computed via graph operators and linear-algebraic constructions. All variants are trained and tested with the same backbone, stratified splits, identical optimization and early stopping, and consistent metrics. Experiments on five different TU datasets and on QM9 converted to classification via target binning show clear dataset dependence: quantum-oriented embeddings yield the most consistent gains on structure-driven benchmarks, while social graphs with limited node attributes remain well served by classical baselines. The study highlights practical trade-offs between inductive bias, trainability, and stability under a fixed training budget, and offers a reproducible reference point for selecting quantum-oriented embeddings in graph learning. Wednesday 0.14 Singapore IJCNN Paper IJCNN SS27 Reservoir Computing for Scalable and Energy-Efficient AI: Theory, Dynamics, and Implementations I Session Chair: Claudio Gallicchio (University of Pisa) PRISM: Physical Reservoir Computing Input Spectral Mapping for Chaotic Time-Series Prediction Forrest Gentry, Alex Ochs, and Sangmin Yoo (Oregon State University) Abstract Abstract Physical Reservoir Computing (PRC) offers a high-efficiency, low-power alternative to traditional recurrent neural networks by leveraging the intrinsic nonlinear dynamics of physical substrates. However, memristive and other physical reservoir devices often suffer from a limited dynamic range, leading to Dynamic Saturation that degrades the reservoir's temporal memory. To address this, we propose Physical Reservoir Computing Input Spectral Mapping (PRISM) that converts time-series dynamics into sets of evenly sparse inputs tailored to device-level constraints. PRISM utilizes unevenly-spaced partitioning based on input distribution and soft linear interpolation to maintain information fidelity while ensuring structured sparsity. Our evaluation on canonical chaotic benchmarks (Mackey-Glass, Rossler, Lorenz, and Santa Fe Laser) demonstrates that PRISM improves memory capacity by up to 7x and reduces autonomous forecasting error by up to 33% in 300-step-ahead predictions. These suggest that encoding strategies designed for device-specific limitations are essential for enhancing the performance of physical reservoir systems in complex, real-world temporal tasks. As Simple as Possible, but Not Simpler: When Reservoir Computing Rivals Deep Learning in Hydrological Prediction Farzad Hosseini (Universidad de Cantabria) and Javier Del Ser (TECNALIA, University of the Basque Country (UPV/EHU)) Abstract Abstract Deep learning models -- especially Long Short-Term Memory networks (LSTMs) -- are strong baselines in research related to rainfall–runoff modeling, but their accuracy typically relies on GPU-intensive training, extensive hyperparameter optimization, and multi-seed ensembling. Echo State Networks (ESNs) offer a lightweight reservoir-computing alternative with efficient training, yet their practical potential for operational hourly hydrology is underexplored. In this work we benchmark ESNs against competitive LSTM baselines across three axis: (i) predictive skill across flow regimes and high-flow events, (ii) robustness to hyperparameter choices and seasonal variability, and (iii) compute cost. Using an hourly dataset from 40 humid, flashy catchments located in the Basque Country (North of Spain), we train one tuned local ESN per catchment and compare them with a highly optimized regional LSTM ensemble and 40 tuned local LSTMs. Results reveal that well-configured ESNs can approach LSTM skill for flood-relevant dynamics (peak timing/magnitude and wet-season regimes). A cost–skill analysis further indicates that ESNs achieve competitive performance at a fraction of the computational budget (CPU-scale training vs cluster GPU-centric LSTMs). Our findings clarify when ESNs are a sufficient operational choice and when the added complexity of LSTMs yields gains of practical value. Non-Dissipative Random Oscillators Networks through Negative Feedback Coupling Gioele Zerini, Andrea Ceni, Andrea Cossu, Davide Bacciu, and Claudio Gallicchio (University of Pisa) Abstract Abstract The careful design of neural architectures is a prominent research area in machine learning. Recurrent neural networks are notoriously tricky to train, making their architectural design often a determining factor for performance. We adopt the reservoir computing approach, where the recurrent component of the network is left untrained, but its architectural design guarantees effective sequential data processing abilities. Inspired by the Random Oscillators Network (RON), we propose the Non-Dissipative RON (ND-RON) and benchmark it against popular time series datasets, showcasing its effectiveness against RON and other popular reservoir computing approaches. We provide mathematical proof of the increased ability of ND-RON to propagate information without dissipation over long time spans, surpassing other reservoir computing approaches and remaining competitive with fully trainable recurrent models while requiring one order of magnitude fewer parameters. Hand Gesture Recognition from Doppler Radar Signals Using Echo State Networks Towa Sano (Nagoya Institute of Technology) and Gouhei Tanaka (Nagoya Institute of Technology, The University of Tokyo) Abstract Abstract Hand gesture recognition (HGR) is a fundamental technology in human computer interaction (HCI). In particular, HGR based on Doppler radar signals is suited for in-vehicle interfaces and robotic systems, necessitating lightweight and computationally efficient recognition techniques. However, conventional deep learning-based methods still suffer from high computational costs. To address this issue, we propose an Echo State Network (ESN) approach for radar-based HGR, using frequency-modulated-continuous-wave (FMCW) radar signals. Raw radar data is first converted into feature maps, such as range-time and Doppler-time maps, which are then fed into one or more recurrent neural network-based reservoirs. The obtained reservoir states are processed by readout classifiers, including ridge regression, support vector machines, and random forests. Comparative experiments demonstrate that our method outperforms existing approaches on an 11-class HGR task using the Soli dataset and surpasses existing deep learning models on a 4-class HGR task using the Dop-NET dataset. The results indicate that parallel processing using multi-reservoir ESNs are effective for recognizing temporal patterns from the multiple different feature maps in the time-space and time-frequency domains. Our ESN approaches achieve high recognition performance with low computational cost in HGR, showing great potential for more advanced HCI technologies, especially in resource-constrained environments. Connectome-Based Reservoir Computing With Coupled Stuart–Landau Oscillators Tenshin Okuma and Makoto Fukushima (Hiroshima University) Abstract Abstract Recent computational neuroscience studies have explored the functional roles of the structural connectome in the human brain using the methodology of reservoir computing. However, existing connectome-based reservoir computing models exhibit a spatial scale mismatch between the reservoir edges, derived from the macroscopic connectome, and the reservoir node dynamics, which are inspired by microscopic neural circuitry. To address this issue, here we develop a new connectome-based reservoir computing model. In this model, the spatial scale mismatch is resolved by modeling node dynamics with coupled Stuart–Landau oscillators, which can simulate macroscopic, metastable brain activity. To examine the model's basic properties, we applied it to a standard nonlinear transformation task using sinusoidal and square waves as the input and desired output, respectively. The results demonstrate that regression performance improves with increasing input magnitude and is maximized when the input frequency is close to the frequency peak of the oscillator ensemble with no input. Metastability, which is a key factor in simulating realistic brain dynamics, does not contribute directly to performance during the task; however, an intermediate level of global coupling for the oscillators that yields metastability in the absence of input improves performance robustness against changes in input frequency. This study clarifies the basic characteristics of the new model and reveals the intriguing effects of metastability on reservoir computing performance. Sculpting Deep Attractors in Reservoir Networks via Oja's Plasticity Yuji Kawai and Hiroshi Atsuta (The University of Osaka) and Minoru Asada (International Professional University of Technology in Osaka, The University of Osaka) Abstract Abstract Recurrent neural networks are powerful tools for learning and generating complex temporal dynamics yet training them to efficiently produce stable and robust patterns remains a significant challenge. In this study, we propose a novel learning method that integrates a local Hebbian-like plasticity mechanism, Oja's rule, with the global feedback provided by first-order reduced and controlled error (FORCE) learning within a reservoir computing model. We demonstrate that this combined approach enhances the network's ability to learn periodic target dynamics, achieving significantly faster convergence from arbitrary initial states than the conventional FORCE model while generating dynamics that are robust to noise. Through eigenmode analysis of the recurrent weight matrix, we show that the learning process systematically forms a low-rank connectivity structure characterized by the emergence of dominant outlying eigenvalues. We further demonstrate that Oja's rule captures essential modes of the network dynamics and embeds them within this low-dimensional structure. This work establishes a clear pathway from local synaptic learning rules, through the formation of a global low-rank architecture, to the emergence of robust computational function in recurrent neural networks. Wednesday 0.15 Washington IJCNN Paper Robotics and Autonomous Agents Session Chair: Erdal Kayacan (Paderborn University, Germany), Leon Reznik (Rochester Institute of Technology) STAMP: Spatio-Temporal Augmented Memory Policy for Robotic Manipulation Zhirun Yue and Mingxin Wang (Tsinghua University, Shenzhen International Graduate School); Tianyi You (Nanjing Tech University); Houde Liu (Tsinghua University, Shenzhen International Graduate School); and Jun Cheng (Chinese Academy of Science, Shenzhen Institutes of Advanced Technology) Abstract Abstract Efficient robotic manipulation via imitation learning aims to distill complex skills from expert demonstrations. Acquiring skills for complex manipulation tasks necessitates a consistent understanding of both spatial configurations and temporal dynamics. While conditional information guides action generation in current models, the absence of explicit 4D world modeling leads to a misalignment between environmental perception and execution, hindering the temporal consistency of generated behaviors. Driven by this, we introduce STAMP (Spatio-Temporal Augmented Memory Policy), a novel architecture that embeds inherent spatiotemporal reasoning directly into the policy's decision-making process. Specifically, STAMP performs temporal-aware feature extraction on both observations and actions to capture their co-evolution. Inspired by human cognitive science, we design a hierarchical memory pyramid that represents historical context across multi-level granularities, enabling the policy to model the evolution of memory over time in a way that mimics biological information consolidation. Extensive evaluations on Adroit and MetaWorld demonstrate that STAMP achieves success rates with a 2.4% improvement, while maintaining high efficiency for real-time control. Reinforcement Learning for Drone Control: Lessons from Hyperparameter Optimization Keiichi Ito, Jed R. Muff, and A. E. Eiben (Vrije Universiteit Amsterdam) Abstract Abstract Reinforcement Learning (RL) can be an effective tool for optimizing robot control policies. However, RL performance is notoriously sensitive to hyperparameter configurations, and the relative importance of these parameters remains poorly understood. In this paper, we investigate this sensitivity by training control policies for a quadrotor drone across four complex navigation tasks: Figure8, Circle, Slalom, and Shuttlerun. Using Optuna, an open-source hyperparameter optimization framework, we systematically optimize a Proximal Policy Optimization (PPO) algorithm and analyze the results. Our analysis reveals two key findings: (1) The discount factor (Gamma) is the dominant hyperparameter, explaining 32--38\% of performance variance. This is more than twice the impact of any other hyperparameter. Surprisingly, the learning rate has low importance (3.9\%), suggesting it is necessary for fine-tuning but not the primary performance driver. (2) Hyperparameter values transfer across related tasks. A generic configuration derived by averaging task-specific optima achieves good task-specific performance across all four tasks, and leave-one-out cross-validation shows no significant transfer gap for 3 of the 4 tasks. These findings provide practitioners with actionable guidance: prioritize gamma tuning and leverage hyperparameter values from related tasks as strong starting points. Neuro-Symbolic Task Routing in Semantic Voxel Environments for Generalist Robotic Manipulation Anindya Jana (TCS Research, Jadavpur University); Snehasis Banerjee (TCS Research); Arup Kumar Sadhu (Tata Consultancy Services Research); and Ranjan Dasgupta (TCS Research) Abstract Abstract Vision-Language-Action (VLA) models have demonstrated significant potential in robotic manipulation but suffer from a performance gap when comparing generalist models to task-specific specialists. Furthermore, deploying these models in unstructured indoor environments requires robust mapping and navigation capabilities. In order to close this gap, this paper presents a unified framework that combines a novel neuro-symbolic task configuration engine with a modular semantic mapping system. In order to make navigation and grounding easier, our method first builds a time-constrained semantic voxel map of the surroundings. By using relaxed waypoint algorithms, we optimize exploration efficiency by 50%. When a task configuration framework arrives at a workspace, it analyzes the spatial arrangement using zero-shot object detection and dynamic scene graph generation. 78–100% of the performance difference between generalist and specialist approaches is recovered by this analysis, which directs the manipulation request to the best specialized model. Empirical findings show that our system maintains real-time computational feasibility appropriate for practical implementation while increasing success rates from 74% to 98% in challenging long-horizon tasks. Affordance-Aware Manipulation using Knowledge Graph and Structured Reasoning Agniprabha Chakraborty (TCS Research), Arup Kumar Sadhu (Tata Consultancy Services Research), and Snehasis Banerjee and Ranjan Dasgupta (TCS Research) Abstract Abstract Robotic manipulation in unstructured environments requires accurate perception, pose estimation, and semantic understanding of object affordances. While vision-language models enable zero-shot detection, they lack structured reasoning for functional grasping. Geometry-based grasp synthesis ignores mass distribution and performs only static collision checking, causing failures in cluttered scenarios. We present the first end-to-end neurosymbolic framework integrating zero-shot perception, dynamic scene graph generation, 6-DoF pose estimation with temporal tracking, and grasp synthesis with symbolic reasoning. Key novelties are: (1) scene graphs encoding collision risk scores for real-time execution monitoring, (2) center-of-mass aware grasp optimization prioritizing mechanically stable grasps, and (3) automatic anchor generation for few-shot pose estimation eliminating manual intervention. In simulation, on a Franka Panda robot with 2-finger parallel-jaw gripper, we achieve 100% success on most reasoning-intensive tasks versus 0-50% for pure neural methods, 17 objects detected versus 5-13 for baselines, 99.97% mean ADD-S versus 23-98.7% for state-of-the-art on Any6D benchmarks, 95% success with 2% collision rate versus 45% success with 38% collision for geometry-only methods, and 98% grasp stability with 42% grip force reduction through center-of-mass optimization, demonstrating the superiority of neurosymbolic integration for robust robotic manipulation. End-to-End Learning of Collaborative Loco-Manipulation for Load Transportation using Quadrupeds with Manipulators Lokesh Kumar, Titas Bera, and Sarvesh Sortee (TCS Research Kolkata, India) Abstract Abstract Collaborative load transportation using quadrupedal robots presents a complex yet highly promising challenge in the field of mobile robotics. Quadrupeds, with their superior mobility over uneven and unstructured terrains, are well-suited for tasks in environments such as disaster zones, off-road logistics, and exploratory missions. However, when tasked with carrying shared payloads, the system becomes significantly more difficult to control due to its underactuated nature where the payload's motion must be regulated indirectly through coordinated whole-body movements of the robots. This demands high levels of synchronization, force distribution, and dynamic balance among agents. To address these challenges, we leverage deep reinforcement learning (DRL) to develop a control policy that enables multiple quadrupedal robots to collaborate in transporting a shared load. The proposed system employs a centralized policy framework, which allows for holistic observation and control during training and deployment. While decentralized policies offer greater scalability and robustness in real-world scenarios, the centralized setup provides a tractable foundation for learning and evaluating fundamental coordination strategies. This paper explores the integration of DRL with centralized control for multi-agent quadrupedal load transportation, highlighting learned behaviors, stability mechanisms, and generalization across payload variations and terrains, laying groundwork for future decentralized extensions. MVB-Grasp: Minimum-Volume-Box Filtering of Diffusion-based Grasps for Frontal Manipulation Bibek Poudel, Abdul Basit, and Muhammad Shafique (New York University Abu Dhabi) Abstract Abstract State-of-the-art 6-DoF grasp generators excel on tabletop benchmarks with overhead cameras but struggle in frontal grasping scenarios on low-cost manipulators with con- strained workspaces, where kinematic limits and approach- direction constraints cause high failure rates. We address this challenge for the Unitree Z1 arm by proposing MVB- Grasp, a novel grasping stack that injects a Minimum Volume Bounding Box (MVBB) geometric prior into diffusion-based grasp generation to dramatically improve success rates in frontal, workspace-constrained settings. Our key scientific contributions are threefold: (i) an MVBB-based geometric filter that exploits oriented bounding-box face normals to reject grasps approaching through the table or misaligned with accessible object faces in O(N ) time; (ii) a combined re-scoring function that blends learned discriminator scores with face-alignment geometry (α = 0.85), specifically calibrated for the Z1’s frontal workspace and kine- matic constraints; and (iii) a systematic MuJoCo evaluation protocol measuring grasp success across object types, distances, lateral positions, and pitch orientations to validate embodiment- specific performance. We implement MVB-Grasp on a Unitree Z1 arm with an Intel RealSense D405 camera, integrating YOLOv8 object detection, GraspGen for candidate generation, PCA- based MVBB fitting, and inverse-kinematics trajectory planning. Experiments across 81 MuJoCo episodes (cylinder, asymmetric box, waterbottle) demonstrate that MVB-Grasp achieves 59.3% success versus 24.7% for vanilla GraspGen, a 2.4× improvement, by filtering geometrically infeasible candidates and prioritizing face-aligned grasps suited to the Z1’s frontal approach constraints. Real-world trials confirm that the MVBB prior substantially improves grasp reliability on constrained, low-cost manipulators without requiring model retraining. Wednesday 0.01 London FUZZ-IEEE Paper FUZZ 11 : FUZZ-IEEE SS02 Fuzzy Machine Learning Session Chair: Jie Lu (University of Technology Sydney) Likelihood Optimization of Probabilistic Fuzzy C-Means based on k-NN and Local Search Davide Cazzorla and Corrado Mencar (University of Bari Aldo Moro) Abstract Abstract Fuzzy C-Means (FCM) is a widely used clustering algorithm that offers robustness, efficiency, and simplicity. However, its parameters (the membership matrix) and hyperparameter (the fuzzification coefficient) can be interpreted in probabilistic terms, thus opening the door for the application of a probabilistic methodology for clustering that is alternative to classical soft clustering methods like Gaussian Mixture Models (GMM). In particular, the Probabilistic FCM (PFCM) is inspired by FCM but adopts a different objective function that is based on negative log-likelihood (NLL). Based on the hypothesis that the optimal values of prototypes are close to data points, this paper describes an optimization algorithm using k-NN and local search that significantly reduces both the computation time and the NLL in comparison to a standard optimization based on the basin-hopping algorithm. Furthermore, a preliminary experimental comparison of PFCM with FCM and GMM shows that prototypes resulting from PFCM are always located in data-dense regions, thus providing more significant clustering of data. Quantum-Enhanced Fuzzy K-Nearest Neighbor Srishak Dash and Q. M. Danish Lohani (South Asian University) Abstract Abstract This paper introduces a quantum-enhanced fuzzy k-nearest neighbor (QE-FKNN) algorithm that integrates quantum feature encoding with fuzzy decision-making to improve classification accuracy. Classical data are mapped into quantum states using angle embedding and entanglement, allowing complex nonlinear relationships to be represented in a high-dimensional Hilbert space. Each training sample is assigned a fuzzy membership value, and the class of a test sample is determined by a weighted combination of its k nearest neighbors. As a result, the proposed method produces smoother decision boundaries and shows greater robustness in regions where class boundaries overlap. Experimental results on ten benchmark datasets demonstrate that QE-FKNN outperforms related classical and quantum KNN-based methods in terms of classification accuracy. Interpretable Soil-Robust Bipedal Locomotion on Tilled Soils with Material Curriculum and a Stance-Gated Slip Reward Anhar Risnumawan, Achmad Fahrul Aji, and Naoyuki Kubota (Tokyo Metropolitan University) Abstract Abstract Autonomous systems for precision agriculture increasingly support crop monitoring, targeted spraying, harvesting, and in-field soil assessment, yet the wheeled and tracked platforms used most often can lose traction on tilled soils and exacerbate soil compaction. We study bipedal locomotion under this soil uncertainty using reinforcement learning with per-episode sampling of friction and compliant-contact parameters. Training follows a multi-stage curriculum that progressively expands the friction and stiffness ranges and employs a slip-averse objective that is gated to loaded stance by a small normal-force threshold. We introduce a fuzzy soil-state model that maps the sampled parameters to traction and support membership degrees and a fuzzy soil difficulty index that summarizes curriculum progression in interpretable terms. We also present a smooth stance-membership relaxation of the hard contact gate used in the slip penalty, yielding a continuous measure of stance confidence. Simulation experiments demonstrate reduced stance-phase slip, improved stability, and better disturbance recovery relative to standard baselines under the tested conditions, while the fuzzy formulation provides an interpretable bridge between soil mechanics, curriculum design, and reward shaping. Improving Generalization Performance of Multi-objective Fuzzy Genetics-Based Machine Learning using Synthetic Training Data Generated from High-Performance Black-Box Models Yuji Shuto, Naoki Masuyama, and Yusuke Nojima (Osaka Metropolitan University) Abstract Abstract Recently, the interpretability of classifiers has become increasingly important, and fuzzy classifiers based on linguistically interpretable fuzzy IF-THEN rules have attracted attention. Multi-objective Fuzzy Genetics-Based Machine Learning (MoFGBML) is a method capable of efficiently generating a set of fuzzy classifiers considering both classification performance and interpretability, using an evolutionary multi-objective optimization algorithm. However, MoFGBML generally exhibits inferior classification performance compared to black-box models. To enhance the classification performance of MoFGBML, in this study, we propose transferring the classification knowledge of LightGBM, which possesses strong generalization ability. Specifically, we augment the training data by adding new patterns with the class label predicted by LightGBM. Computational experiments on real-world datasets demonstrate that MoFGBML trained with the proposed method can achieve higher classification accuracy while maintaining interpretability, compared to the conventional one trained solely on the original dataset. xFODE+: Explainable Type-2 Fuzzy Additive ODEs for Uncertainty Quantification Ertuğrul Keçeci and Tufan Kumbasar (Istanbul Technical University) Abstract Abstract Recent advances in Deep Learning (DL) have boosted data-driven System Identification (SysID), but reliable use requires Uncertainty Quantification (UQ) alongside accurate predictions. Although UQ-capable models such as Fuzzy ODE (FODE) can produce Prediction Intervals (PIs), they offer limited interpretability. We introduce Explainable Type-2 Fuzzy Additive ODEs for UQ (xFODE+), an interpretable SysID model which produces PIs alongside point predictions while retaining physically meaningful incremental states. xFODE+ implements each fuzzy additive model with Interval Type-2 Fuzzy Logic Systems (IT2-FLSs) and constraints membership functions to the activation of two neighboring rules, limiting overlap and keeping inference locally transparent. The type-reduced sets produced by the IT2-FLSs are aggregated to construct the state update together with the PIs. The model is trained in a DL framework via a composite loss that jointly optimizes prediction accuracy and PI quality. Results on benchmark SysID datasets show that xFODE+ matches FODE in PI quality and achieves comparable accuracy, while providing interpretability. Position Paper: Neuro-Fuzzy Fusion for Dataset-Agnostic Deepfake Detection Soumyadeep Chattopadhyay, Bhoomi Priya, and Pranab K. Muhuri (South Asian University) Abstract Abstract The rise of deepfake generation models has made manipulated content increasingly sophisticated and pervasive. While current detectors and domain-specific models perform excellent on intra-dataset evaluations, they often fail to generalize across unseen generators, resulting in a large robustness gap and limited detection reliability. Hence, in this position paper, we argue that fuzzy systems, especially neuro-fuzzy systems, can provide an interpretable and adaptive mechanism for decision fusion across heterogeneous detectors. To support this position, we propose a frequency-based ensemble detection network fused through an Adaptive Neuro-Fuzzy Inference System(ANFIS). By leveraging the frequency domain artifacts left behind by specific generators and utilizing neuro-fuzzy reasoning to mitigate uncertain predictions from domain specific models, the network could effectively learns to generalize to unseen variations across varied deepfake scenarios. Instead of providing a benchmark-driven study, this paper proposes a novel direction for integrating fuzzy logic into effective deepfake detection networks, aiming to achieve data-agnostic, enhanced explainability and robust analysis for high stake media forensics. Wednesday 0.02 Berlin IJCNN Paper Neuromorphic Computing and AI Hardware Systems Session Chair: Guilherme DeSouza (University of Missouri, Vision-Guided and Intelligent Robotics Lab (ViGIR)), Giorgio Morales (University of Caen Normandy) Bridging Theory and Practice in Crafting Robust Spiking Reservoirs Ruggero Freddi, Diana Nigrisoli, and Nicolas Seseri (Manava Plus) and Alessio Basti (“G. d'Annunzio” University of Chieti-Pescara) Abstract Abstract Spiking reservoir computing provides an energy-efficient approach to temporal processing, but reliably tuning reservoirs to operate at the edge-of-chaos is challenging due to experimental uncertainty. This work bridges abstract notions of criticality and practical stability by introducing and exploiting the robustness interval, an operational measure of the hyperparameter range over which a reservoir maintains performance above task-dependent thresholds. Through systematic evaluations of Leaky Integrate-and-Fire (LIF) architectures on both static (MNIST) and temporal (synthetic Ball Trajectories) tasks, we identify consistent monotonic trends in the robustness interval across a broad spectrum of network configurations: the robustness-interval width decreases with presynaptic connection density $\beta$ (i.e., directly with sparsity) and directly with the firing threshold $\theta$. We further identify specific $(\beta, \theta)$ pairs that preserve the analytical mean-field critical point $w_{\text{crit}}$, revealing iso-performance manifolds in the hyperparameter space. Control experiments on Erdős–Rényi graphs show the phenomena persist beyond small-world topologies. Finally, our results show that $w_{\text{crit}}$ consistently falls within empirical high-performance regions, validating $w_{\text{crit}}$ as a robust starting coordinate for parameter search and fine-tuning. To ensure reproducibility, the full Python code is publicly available. Online Adaptive Reinforcement Learning with Echo State Networks for Non-Stationary Dynamics Aoi Yoshimura (Nagoya institute of technology) and Gouhei Tanaka (Nagoya institute of technology, The University of Tokyo) Abstract Abstract Reinforcement learning (RL) policies trained in simulation often suffer from severe performance degradation when deployed in real-world environments due to non-stationary dynamics. While Domain Randomization (DR) and meta-RL have been proposed to address this issue, they typically rely on extensive pretraining, privileged information, or high computational cost, limiting their applicability to real-time and edge systems. In this paper, we propose a lightweight online adaptation framework for RL based on Reservoir Computing. Specifically, we integrate an Echo State Networks (ESNs) as an adaptation module that encodes recent observation histories into a latent context representation, and update its readout weights online using Recursive Least Squares (RLS). This design enables rapid adaptation without backpropagation, pretraining, or access to privileged information. We evaluate the proposed method on CartPole and HalfCheetah tasks with severe and abrupt environment changes, including periodic external disturbances and extreme friction variations. Experimental results demonstrate that the proposed approach significantly outperforms DR and representative adaptive baselines under out-of-distribution dynamics, achieving stable adaptation within a few control steps. Notably, the method successfully handles intra-episode environment changes without resetting the policy. Due to its computational efficiency and stability, the proposed framework provides a practical solution for online adaptation in non-stationary environments and is well suited for real-world robotic control and edge deployment. MARS: A Multi-Agent RTL Synthesis Framework Likith Anaparty (Indian Institute of Technology Palakkad, Jurin AI Inc.); Ananthakrishnan Thulasiraman (Indian Institute of Technology Palakkad); Minghao Shao (New York University Abu Dhabi); Vivek Chaturvedi (Indian Institure of Technology Palakkad); and Muhammad Shafique (New York University Abu Dhabi) Abstract Abstract Automating Register Transfer Level (RTL) designs from natural language specifications remains a long-standing challenge in electronic design automation (EDA), where correctness-by-construction remains critical yet elusive. Despite the rapid progress in Large Language Models (LLMs), existing approaches focus more on synthesizing single module RTLs and under-perform in generating functionally correct multi-module RTL designs. A key limitation lies in treating specifications as flat and monolithic units overlooking the modular structure, hierarchical dependencies and iterative refinement process fundamental to real-world RTL development. In this work, we introduce MARS, a novel multi-agent LLM framework that breaks down a hardware specification into sub-modules enabling specialized agents to focus on ambiguity resolution, complexity classification, design decomposition, per-module RTL generation, and integration. Agents collaborate through a coordinated pipeline producing modular testbench-aligned and verifiable RTL. Experimental results on RTL benchmarks demonstrate that MARS outperforms existing state-of-the-art baseline framework by 5.22% in generating functionally correct RTL designs. FedMTFI: Feature Importance Based Optimized Multi Teacher Knowledge Distillation in Heterogeneous Federated Learning Environment Nazmus Shakib Shadin, Aaron Cummings, Xinyue Zhang, and Bobin Deng (Kennesaw State University) Abstract Abstract Federated learning (FL) is a decentralized approach that enables collaborative model training without exposing raw data. Instead of transferring sensitive data, it allows devices to share only model weights, keeping personal data locally and secure. However, in real world settings, the data held by devices is often not evenly distributed and devices mostly differ in computing power and memory capacity. These differences make FL harder to maintain consistent performance across the system. To address these issues, we propose FedMTFI, a novel architecture that combines multi-teacher knowledge distillation (MTKD) with feature importance to improve the FL process in heterogeneous environments. In FedMTFI, clients are clustered based on similar hardware and model types. Each cluster trains a specific model on not independently and identically distributed (non-IID) data. Within a cluster, every client updates that model using only its own local private data. The server then aggregates the locally trained models in each cluster using FedAvg to form multiple prototype models. Then these prototypes serve as teacher models to train a global generalized student model using MTKD. What makes FedMTFI more unique is the integration of Shapley values (SHAP) to emphasize important features during distillation, which enhances both accuracy and interpretability. Experimental results show that FedMTFI achieves higher accuracy than traditional FL algorithms and performs more effectively under non-IID data conditions. Dynamic Graph-Augmented Transformers for Exogenous-Aware Time Series Forecasting Tianxiao Ren, Zhen Wang, Yunzhi Hao, Yang Gao, Shunyu Liu, Kaixuan Chen, and Mingli Song (Zhejiang University) Abstract Abstract Exogenous-aware time series forecasting requires capturing complex temporal dependencies within endogenous signals and incorporating heterogeneous exogenous variates that influence the target variables.However, existing approaches often use static attention based on a single similarity computation, which makes it hard to capture temporal dependencies, and use simple encodings for exogenous inputs that may miss important covariate patterns. Therefore, we propose Dynamic Graph-Augmented Transformers (DGAFormer), which equips Transformer attention with our Dynamic Graph Attention (DGA) mechanism and incorporates a Gated Multi-scale Exogenous (GME) encoder for covariate representation. Specifically, our DGA mechanism augments attention with dynamic graph-based refinement, alleviating the limitation of static attention.Using this mechanism, DGAFormer applies DGA-based self-attention to learn endogenous dependencies and DGA-based cross-attention, where a global endogenous token queries exogenous tokens, to incorporate exogenous information consistently. In addition, to provide more informative covariate representations, we design a GME encoder that extracts patterns at multiple temporal resolutions using a convolutional bank and adaptively fuses them into compact exogenous tokens.Extensive experiments on multiple real-world benchmarks show that DGAFormer consistently outperforms strong baselines, with particularly pronounced gains in covariate-rich settings. Wednesday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 18 - Algorithms IV Session Chair: Andries Engelbrecht (Stellenbosch University) Social Dissent in Particle Swarm Optimization: Breaking Groupthink via Adaptive Stochastic Noise Swapnoneel Sarkar (Accelequant India Pvt. Ltd. Bengaluru), Somnath Mukhopadhyay (Assam University Silchar), and Andries P. Engelbrecht (Stellenbosch University) Abstract Abstract Adaptive social-dissent particle swarm optimization (PSO) addresses the vulnerability of PSO to groupthink and social lock-in by introducing social dissent, which refers to deliberate disagreement in social learning. The framework utilizes three operators: leader corruption, where particles follow foreign leaders; membership migration, where particles temporarily switch to another sub-swarm; and the scout operator, which triggers stochastic re-initialization and momentum resampling for stagnated particles to relocate them to unvisited areas of the search space. Dissent intensity is modulated by the path length ratio, a straightness measure that scales effective probabilities to encourage exploration in oscillatory particles while allowing steadily moving particles to converge. When combined with topology annealing that linearly reduces the number of sub-swarms to balance parallel exploration and unified exploitation, the algorithm achieves superior expected running time and success rates on COCO/BBOB benchmarks. These improvements are particularly pronounced in high-dimensional tasks where the method counteracts diversity collapse and prevents the social lock-in that traps traditional implementations. Beyond Behavioral Sequences: Leveraging Short-Term Memory to Optimize Decision Policies in Anticipatory Learning Classifier System Mateusz Łabędzki and Olgierd Unold (Politechnika Wrocławska) Abstract Abstract The Behavioral Enhanced Anticipatory Classifier System (BEACS) has demonstrated robust capabilities in developing compact and interpretable representations of non-deterministic environments. However, recent evaluations have highlighted a critical limitation: the system often fails to achieve optimal decision policies, resulting in suboptimal average path lengths to the goal in complex aliased environments. This work introduces an augmented BEACS framework designed to address this deficiency by incorporating a Short-Term Memory (STM) mechanism. Unlike standard Anticipatory Learning Classifier System (ALCS) models that rely solely on immediate sensory input, the proposed memory component enables the agent to maintain an internal representation of recent state-action history, effectively resolving ambiguities in states where optimal behavior depends on temporal dependencies. Furthermore, we extend the classifier enhancement process to evolve memory-dependent rules, allowing the system to differentiate between perceptually identical states based on their historical context. Experimental results across a comprehensive suite of benchmark mazes demonstrate that the integration of Short-Term Memory significantly reduces the average path length to the goal while preserving the interpretability and generalization performance characteristic of the original BEACS framework. These findings suggest that incorporating temporal context is a vital strategy for improving path efficiency and achieving optimal performance in partially observable environments within the ALCS paradigm. A Fitness Landscape Analysis of Multi-Objective Neural Architecture Search Cosijopii Garcia-Garcia (Universidad del Istmo) and Bilel Derbel (Univ. Lille, CNRS, Inria, Centrale Lille) Abstract Abstract The performance of local search in Multi-Objective Neural Architecture Search (MONAS) is strongly influenced by the landscape topology, which is itself directly induced by the choice of a neighborhood operator. In this paper, we investigate how different neighborhood definitions shape the MONAS landscape and, in turn, affect the empirical performance of Pareto Local Search (PLS). Specifically, we conduct a systematic multi-objective landscape analysis of the NATS-BENCH topology search space. Using the Pareto Local Optima Networks (PLOS- nets) framework, we analyze the landscapes induced by six neighborhood operators. Our results show that "disruptive" operators generate highly fragmented, poorly searchable landscapes, whereas incremental, “small-step” operators produce well-connected landscapes that effectively link Pareto local optima to the global Pareto front. Building on these insights, we examine the performance of two PLS variants and analyze their relative effectiveness as a function of both the underlying neighborhood structure and the exploration strategy employed. Center-based Sampling in Optimization and Machine Learning: A Theory and Review Rasa Khosrowshahli, Beatrice Ombuki-Berman, and Shahryar Rahnamayan (Brock University) Abstract Abstract High-dimensional optimization and machine learning (ML) pipelines frequently rely on randomized proposals within box-constrained domains, yet it is often unclear when concentrating proposals near the box center increases the likelihood of generating useful candidates. We study center-based sampling (CBS) in the D-dimensional unit hypercube, where candidates are drawn from a central subcube rather than uniformly from the full box. Under a simple single-draw model, we prove that CBS reduces expected squared distance to an independent target by a fixed per coordinate amount, so the total reduction grows linearly with the dimension. When the target is uniform on the hypercube, this means advantage upgrades to a high probability regime in which the probability that a center draw is closer than a uniform draw converges to one exponentially fast in D, with explicit Hoeffding and Bernstein guarantees and an accurate central limit theorem (CLT) approximation. The expectation comparison also extends beyond the uniform target model to any independent target with a finite second moment, for which the mean advantage is unchanged because the uniform and CBS samplers share the same center while CBS has a smaller variance. We complement the theory with the same budget random search experiments on shifted Sphere, Rastrigin, and Ackley functions over box-constrained domains, where CBS consistently improves the best objective under a fixed sample budget, with the strongest gains when the optimum is centrally located. These results position CBS as a useful geometric baseline for box-constrained randomized search, while clarifying that the concentration theorem applies only to the uniform target cube model. Refined CMA-ES for CEC 2026 Bound-Constrained Single-Objective Optimization Adam Stelmaszczyk, Rafał Biedrzycki, and Jarosław Arabas (Warsaw University of Technology) Abstract Abstract We present a simplified and parameter-adjusted variant of CMA-ES designed for the CEC 2026 Bound-Constrained Single-Objective Optimization competition. Unlike typical algorithmic developments that increase complexity by adding mechanisms, Refined CMA-ES systematically removes components that do not improve performance, such as the flat‑land escape, while retaining standard features that provide measurable benefits. Adjusting key parameters---population size and initial step-size---yields substantial gains, outperforming the baseline CMA-ES and several state-of-the-art competitors. The paper highlights the practical relevance of individual CMA-ES components under the CEC 2026 ranking, which equally weights solution accuracy and convergence speed, demonstrating that simplification combined with improved parameter settings can lead to superior performance. RCMAES: A Robust CMA-ES Variant for CEC2026 Competition Khoirul Faiq Muzakka, Sören Möller, and Martin Finsterbusch (Forschungszentrum Jülich GmbH) Abstract Abstract This paper proposes RCMAES, a novel variant of the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) for CEC benchmark optimization. RCMAES integrates a dimension-dependent nonlinear population-size reduction strategy with an adaptive restart mechanism within a pure CMA-ES framework. RCMAES is evaluated on three benchmark suites (CEC2017, CEC2020, and CEC2022) and compared with state-of-the-art DE algorithms as well as its closely related counterpart, BIPOP-aCMAES. Experimental results show that RCMAES achieves competitive and robust performance across all benchmarks. Wednesday 0.05 Paris IJCNN Paper AI for Industrial Operations and Planning Session Chair: Ali Minai (University of Cincinnati), Francesco Alesiani (NEC Laboratories Europe GmbH) ARE: Adaptive TD-λ Return Estimation for Learning Control in Differentiable Simulation Quang Dung Dinh, Adrian Redder, and Erdal Kayacan (Paderborn University) Abstract Abstract Differentiable simulators have gained traction in robotics because they provide the analytic pathwise gradient of the agent's reward, promising better sample efficiency for model-based learning. However, in continuous control settings, the gradient grows exponentially with the decision horizon, making the pathwise gradient susceptible to exploding/vanishing and high variance. Recent works have addressed these fundamental challenges by introducing first-order actor-critic algorithms which optimize over very short learning horizons. In this work, we propose to further improve such algorithms by introducing the TD-λ return into the optimization of the actor network as a generalized return estimator, which reduces the actor gradient estimator's variance, leading to consistent and robust learning. In addition, we propose to adaptively adjust λ using a value-fitting objective to avoid expensive manual tuning and further enhance the learning stability. Our algorithms demonstrate improvements over challenging locomotion tasks, with an average improvement of roughly 50% over the Ant environment and almost 100% with the simulated Unitree Go2 quadruped environment. Moreover, our design allows exploiting the gradient information over much longer learning horizon, enabling more effective long-term credit assignment for first-order model based reinforcement learning methods. Efficiently Solving the TSP with Non-Autoregressive Self Improvement Learning Debarpan Debnath and Praveen Paruchuri (IIIT Hyderabad) Abstract Abstract Neural approaches that rely on autoregressive (AR) solution construction for solving the Traveling Salesman Problem (TSP) - a cornerstone challenge in combinatorial optimization - can be computationally expensive during both training and inference, since each step of the solution construction requires a forward pass through the neural network. Although non-autoregressive (NAR) methods offer the promise of efficient solution generation in a single network pass, prior NAR approaches have typically depended on Monte Carlo Tree Search (MCTS) to reach near-optimal performance, which substantially slows down the solution generation and demands significant additional computational resources. To address this, we propose NARSIL, a NAR method that generates TSP solutions extremely fast without sacrificing the solution quality. NARSIL is trained in an unsupervised manner using Self-Improvement Learning (SIL) and does not require access to optimal solutions during training. Experimental results demonstrate that NARSIL achieves a 9-11x speed-up over the previous fastest baseline while delivering superior solution quality, and a 71-265x speed-up over high-accuracy baselines while producing comparable solution quality. Metalearning for Enhanced Algorithm Selection in Time Series Decomposition-Based Forecasting José Araújo (FEUP); Carlos Soares (LIACC/FEUP, Fraunhofer AICOS); and Luis Reis and Moisés Santos (LIACC/FEUP) Abstract Abstract Time series forecasting is essential in fields such as finance, healthcare, and supply chain management, where accurate predictions support critical decisions. Traditional methods often fail to model the complex nonlinear patterns and multiple seasonal variations found in real-world data. Although decomposition techniques like MSTL improve forecasting by separating time series components, the selection of an appropriate model for each component remains a challenging, heuristic-driven task. To address this, we introduce a novel metalearning framework that automates model selection in a decomposition-aware manner. Unlike existing methods that apply fixed models or treat series monolithically, our approach uses a granular, component-based strategy. It first decides whether to decompose a series via residual analysis, then recommends the suitable forecasting model from a diverse set of statistical and Deep Learning algorithms for each component. Experiments on 5000 monthly series from the M4 Competition show that our framework achieves statistically significant improvements over both non-decomposed metalearning baselines (23% lower NMAE, 61% lower error variance) and state-of-the-art models such as TCN and DeepAR. Preference-Conditioned Dynamic Attention Model for Multi-objective Capacitated Arc Routing Problem Xiaoyuan Wei, Yang Wang, Ya-Hui Jia, and Feng-Feng Wei (South China University of Technology); Qiang Yang (Nanjing University of Information Science and Technology); and Wei-Neng Chen (South China University of Technology) Abstract Abstract Multi-objective Capacitated Arc Routing Problem (MOCARP) is a complex multi-objective combinatorial optimization problem. Existing Neural Combinatorial Optimization (NCO) methods often treat preference vectors merely as static context, failing to evaluate how dynamic state transitions affect conflicting objectives. To address this, we propose a novel single-model NCO method called the Preference-Conditioned Dynamic Attention Model (PCDAM) to solve MOCARP. First, we introduce a preference-aware attention mechanism that injects preference weights directly into the decoder's query vectors, enabling the model to adaptively adjust decision priorities according to given preferences with minimal computational overhead. Second, we design dynamic objective-aware features by combining objective-specific metrics with preference vectors, enabling the model to evaluate the impact of each action on multiple conflicting objectives in real-time during decoding, which provides a correction to the attention model. Extensive experiments demonstrate that PCDAM significantly outperforms representative evolutionary algorithms and existing NCO methods in terms of hypervolume. PRISM: Deployable Counterfactual Recourse for Metro Maintenance via Physics-Constrained Search and RL Distillation Xing Su (Zhejiang Rail Transit Operation Management Group Co. Ltd., Southwest Jiaotong University); Fei Teng (Southwest Jiaotong University); Yifan Zhuang (Zhejiang Rail Transit Operation Management Group Co. Ltd.); Jiasheng Fang (Zhejiang Rail Transit Operation Management Group Co. Ltd., Southwest Jiaotong University); and Zixuan Xu and Dong Zou (Zhejiang Rail Transit Operation Management Group Co. Ltd.) Abstract Abstract Urban metro safety affects millions of daily journeys. Abnormal contact-wire wear is a major threat to reliable metro current collection. Severe wear can trigger arcing and overheating, accelerate component degradation, and in extreme cases lead to service disruptions and safety hazards. Preventing such events requires more than early warning: operators also need actionable and physically feasible mitigation suggestions for the rare but high-risk severe cases. To address this gap, we propose PRISM, an industrial framework that links reliable ordinal wear grading with physics-constrained mitigation. PRISM first performs imbalance-aware multi-head grading classification, and then formulates mitigation as a constrained counterfactual recourse problem. PRISM instantiates a prototype-anchored search teacher, guided by physical priors derived from a differentiable surrogate to accelerate optimization and enable audit-friendly explanations. To meet online constraints, PRISM distills the teacher into an offline actor--critic student with verification and teacher fallback in reinforcement learning. Experiments on three real-world datasets demonstrate that PRISM consistently improves wear grading and achieves the best severe-wear mitigation success over strong baselines. GridSCAPE: An LLM-Based Multi-Agent Framework for Real-Time Intelligent Supply Chain Planning in Electric Power Enterprises Haoyang He (State Grid Beijing Electric Power Company) and Yan Gu (North China Electric Power University) Abstract Abstract Supply chain planning is a core challenge in supply chain management, orchestrating heterogeneous item categories across multiple vendors, and jointly optimizing end-to-end material flows, from sourcing and inbound logistics, through warehousing, to on-time delivery at demand locations. Electric power enterprises, as large and safety-critical consumers of engineering materials, require a comprehensive supply chain planning system, spanning demand planning, supplier and contract management, inventory control, to ensure reliable and cost-effective material availability. Moreover, recent advances in artificial intelligence (AI), particularly large language models (LLMs), have opened new directions for supply chain management, enabling knowledge-grounded reasoning, tool-augmented integration with enterprise platforms, and multi-agent coordination for closed-loop, adaptive planning. Accordingly, we present GridSCAPE, an LLM-based multi-agent framework, built to orchestrate closed-loop, real-time, intelligent supply chain planning for electric power enterprises. GridSCAPE deploys coordinated LLM agents with tool support, covering the entire supply chain, from demand and sourcing to inventory, warehousing, and logistics. We present implementation details of GridSCAPE and validate its efficacy via a utility case study and controlled experiments, showing reduced end-to-end cycle time, higher plan accuracy, and an increased in-stock rate. Wednesday 0.10 Sydney IJCNN Paper AI for Mobility, Transportation and Infrastructure Systems Session Chair: Yatharth Agarwal (Purdue University), Hanadi Alhamdan (Princess Nourah bint Abdulrahman University, Durham University) CUSP: Cross-domain Uncertainty-aware Driving Style Posterior Xinyue Liu (The University of Manchester), Fanlin Meng (University of Exeter Business School), and Xiao-Jun Zeng (The University of Manchester) Abstract Abstract Inferring driving style from real-world trajectories is challenging under domain shift: datasets collected across regions or sites differ in speed distributions, congestion regimes, sensing noise, and annotation quality, causing learned ``styles'' to absorb domain-specific shortcuts such as absolute speed. We propose CUSP, an uncertainty-aware framework that represents driving style with interpretable mechanism parameters of the Intelligent Driver Model (IDM), such as desired speed and time headway, and infers a posterior from short interaction windows under domain shift. CUSP performs scalable amortised posterior inference and learns the parameters via a physics-grounded likelihood with variational regularisation. To keep style semantics comparable across datasets, we stabilise the inferred parameters by discouraging domain leakage and absolute-speed shortcuts through adversarial training, optionally aided by masking speed-related channels and using simple neighbourhood context. Experiments under cross-location and cross-site evaluations on highD and NGSIM show that CUSP preserves mechanistic fidelity while reducing speed and domain information in the inferred parameters, yielding more stable style semantics and uncertainty signals that support downstream interaction modelling. MamMA: A Mamba-Based Pedestrian Trajectory Prediction Algorithm Considering Occupancy Map and Pedestrian Awareness States Juncen Long, Xiaofeng Jin, Gianluca Bardaro, Simone Mentasti, and Matteo Matteucci (Politecnico di Milano) Abstract Abstract Many pedestrian trajectory prediction algorithms have been proposed to improve the safety of navigation for mobile robots working in human-robot coexistence environments. Some pedestrian trajectory prediction algorithms extract information about obstacles near pedestrians from top-down view images to improve the accuracy of trajectory prediction. However, mobile robots typically create local occupancy maps using LiDAR, rather than top-down view images. Meanwhile, the vision sensors on board robots provide egocentric view images, which contain fine-grained behavioral information about the pedestrians near the robot. To better use the information collected by LiDAR and on-board vision sensors, we propose MamMA, a Mamba-based pedestrian trajectory prediction algorithm considering occupancy maps and pedestrian awareness states. MamMA divides the occupancy map by patches and extracts obstacle features from each patch to create map features. Pedestrian awareness states are divided and considered, as some studies show that awareness states affect the perception and speed of pedestrians. Furthermore, a Mamba-based model is proposed to predict the future trajectories of pedestrians based on different types of features. Experiments on the STCrowd, SiT, JRDB, ETH, and UCY datasets show that MamMA achieves better average displacement error and final displacement error than the state-of-the-art algorithms. CBANet: A Compact Attention-Based CNN–BiLSTM Network for Aggressive Driving Event Detection Hanadi Alhamdan (Princess Nourah bint Abdulrahman University, Durham University); Ghadah Alosaimi (Imam Mohammad Ibn Saud Islamic University (IMSIU), Durham University); and Amir Atapour-Abarghouei and Farshad Arvin (Durham University) Abstract Abstract Aggressive driving is a major cause of traffic accidents and poses a serious threat to road safety. Although deep learning methods have shown promising results in detecting risky driving behaviours from vehicle sensor data, their performance in real-world conditions is often limited by severe data imbalance, large variability between drivers, and the lack of physically interpretable vehicle dynamics representations. In this paper, we propose an enhanced deep learning framework for aggressive driving detection using multivariate vehicle dynamics signals. Instead of relying solely on raw measurements, the proposed approach constructs engineered dynamic features that capture steering, acceleration, and braking behaviour. To address the extreme rarity of aggressive events in naturalistic driving data, we introduce a stable training strategy that combines controlled SMOTE-based oversampling with a class-weighted loss formulation, and evaluates focal loss variants for imbalance handling. Furthermore, a safety-oriented decision strategy based on class-specific threshold calibration is adopted to better reflect the asymmetric risks of missed detections and false alarms in real-world applications. The proposed framework is evaluated on a newly collected naturalistic driving dataset. Extensive experiments show that the proposed method consistently outperforms standard deep learning baselines with significant improvements in minority-class recall and safety-critical F-score metrics while maintaining practical computational efficiency. Mind2Drive: Predicting Driver Intentions from EEG in Real-world On-Road Driving Ghadah Alosaimi (Imam Mohammad Ibn Saud Islamic University (IMSIU), Durham University); Hanadi Alhamdan (Princess Nourah bint Abdulrahman University, Durham University); and Wenke E, Stamos Katsigiannis, Amir Atapour-Abarghouei, and Toby Breckon (Durham University) Abstract Abstract Predicting driver intention from neurophysiological signals offers a promising pathway for enhancing proactive safety in advanced driver assistance systems, yet remains challenging in real-world driving due to EEG signal non-stationarity and the complexity of cognitive–motor preparation. This study proposes and evaluates an EEG-based driver intention prediction framework using a synchronised multi-sensor platform integrated into a real electric vehicle. A real-world on-road dataset was collected across 32 driving sessions, and twelve deep learning architectures were evaluated under consistent experimental conditions. Among the evaluated architectures, TSCeption achieved the highest average accuracy (0.907) and Macro-F1 score (0.901). The proposed framework demonstrates strong temporal stability, maintaining robust decoding performance up to 1000 ms before manoeuvre execution with minimal degradation. Furthermore, additional analyses reveal that minimal EEG preprocessing outperforms artefact-handling pipelines, and prediction performance peaks within a 400–600 ms interval, corresponding to a critical neural preparatory phase preceding driving manoeuvres. Overall, these findings support the feasibility of early and stable EEG-based driver intention decoding under real-world on-road conditions. Code: https://github.com/galosaimi/Mind2Drive LPLCv2: An Expanded Dataset for Fine-Grained License Plate Legibility Classification Lucas Wojcik and Eduardo A. F. Machoski (Federal University of Paraná); Eduil Nascimento Jr. (Paraná Military Police); Rayson Laroca (Pontifical Catholic University of Paraná, Federal University of Paraná); and David Menotti (Federal University of Paraná) Abstract Abstract Modern Automatic License Plate Recognition (ALPR) systems achieve outstanding performance in controlled, well-defined scenarios. However, large-scale real-world usage remains challenging due to low-quality imaging devices, compression artifacts, and suboptimal camera installation. Identifying illegible license plates (LPs) has recently become feasible through a dedicated benchmark; however, its impact has been limited by its small size and annotation errors. In this work, we expand the original benchmark to over three times the size with two extra capture days, revise its annotations and introduce novel labels. LP-level annotations include bounding boxes, text, and legibility level, while vehicle-level annotations comprise make, model, type, and color. Image-level annotations feature camera identity, capture conditions (e.g., rain and faulty cameras), acquisition time and day ID. We present a novel training procedure featuring an Exponential Moving Average-based loss function and a refined learning rate scheduler, addressing common mistakes in testing. These improvements enable a baseline model to achieve a 89.5% F1-score on the test set, considerably surpassing the previous state of the art. We further introduce a novel protocol to explicitly addresses camera contamination between training and evaluation splits, where results show a small impact. Dataset and code will be publicly available at [hidden for review]. Wednesday 0.11 Cape Town IJCNN Paper AI for Healthcare, Medical Imaging, and Biomedical Discovery I Session Chair: Silvia Multari (Ca’Foscari University of Venice), Larissa Zott (Universität der Bundeswehr München) MedGraphGen: A Unified Pipeline for Generating Medical Knowledge Graphs from Medical Images Chengzhi Cao and Min Xu (MBZUAI) Abstract Abstract Medical image analysis has evolved beyond single-organ segmentation towards holistic understanding of anatomical relationships. However, current approaches struggle to explicitly represent the complex structural dependencies between multiple organs, which limits their utility for downstream tasks. In this paper, we propose MedGraphGen, a unified pipeline that directly generates interpretable medical knowledge graphs from CT images in an end-to-end manner, enabling structured understanding beyond pixel-level perception. Specifically, we introduce a Spatial-Semantic Context Aggregator to bridge segmentation features with graph-based reasoning by explicitly modeling spatial relationships between organs. These spatial priors are integrated into a Relation-Aware Hypergraph Module, which leverages hyperedges to represent higher-order interactions among multiple organs. Then, we introduce Semantic-Aware Hyperedge Attention (SA-HEA) to differentiate semantically meaningful multi-organ interactions from trivial co-occurrences during message passing. Crucially, to support this task, we present a new benchmark dataset with dense pixel-wise segmentation masks and medical knowledge graph annotations, enabling training and evaluation of models capable of medical image understanding. We demonstrate that the generated graphs are not only anatomically plausible but also serve as powerful priors to enhance the relational reasoning capabilities of medical vision-language models. Efficient KernelSHAP Explanations for Patch-based 3D Medical Image Segmentation Ricardo Coimbra Brioso and Giulio Sichili (Politecnico di Milano); Damiano Dei (IRCCS Humanitas Research Hospital); Nicola Lambri (IRCCS Humanitas Research Hospital, Università degli Studi di Milano); Pietro Mancosu and Marta Scorsetti (IRCSS Humanitas Research Hospital); and Daniele Loiacono (Politecnico di Milano) Abstract Abstract Perturbation-based explainability methods such as KernelSHAP provide model-agnostic attributions but are typically impractical for patch-based 3D medical image segmentation due to the large number of coalition evaluations and the high cost of sliding-window inference. We present an efficient KernelSHAP framework for volumetric CT segmentation that restricts computation to a user-defined region of interest and its receptive-field support, and accelerates inference via patch logit caching, reusing baseline predictions for unaffected patches while preserving nnU-Net’s fusion scheme. To enable clinically meaningful attributions, we compare three automatically generated feature abstractions within the receptive-field crop: whole-organ units, regular FCC supervoxels, and hybrid organ-aware supervoxels, and we study multiple aggregation/value functions targeting stabilizing evidence (TP/Dice/Soft Dice) or false-positive behavior. Experiments on whole-body CT segmentations show that caching substantially reduces redundant computation (with computational savings ranging from 15% to 30%) and that faithfulness and interpretability exhibit clear trade-offs: regular supervoxels often maximize perturbation-based metrics but lack anatomical alignment, whereas organ-aware units yield more clinically interpretable explanations and are particularly effective for highlighting false-positive drivers under normalized metrics. HGDC-Fuse: Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction yueheng jiang and Peng Zhang (Zhejiang University) Abstract Abstract Accurate multi-disease diagnosis using multi-modal data such as electronic health records and medical imaging is a critical clinical challenge. Although existing deep learning methods have achieved initial success in this area, a significant gap persists for their real-world application. This gap arises because existing methods often overlook unavoidable practical challenges, such as modality missingness, noise, temporal asynchrony, and evidentiary inconsistency across modalities for different diseases. To overcome these limitations, we propose HGDC-Fuse, a novel framework that constructs a patient-centric multi-modal heterogeneous graph to robustly integrate asynchronous and incomplete multi-modal data. Moreover, we design a heterogeneous graph learning module to aggregate multi-source information, featuring a disease correlation-guided attention layer that resolves the modality inconsistency issue by learning disease-specific modality weights based on disease correlations. On the large-scale MIMIC-IV and MIMIC-CXR datasets, HGDC-Fuse significantly outperforms state-of-the-art methods. Hierarchical Deep Learning for De Novo Molecular Structure Prediction from EI-MS Spectra Mohammad Falah (Maastricht University), John Mommers (Envalior), and Anna Wilbik and Marcin Pietrasik (Maastricht University) Abstract Abstract Inferring molecular structures directly from electron-ionization mass spectrometry (EI-MS) spectra remains a challenging problem in chemoinformatics. While recent deep learning approaches have shown promise, most operate over the entire chemical space using a single model, limiting scalability and interpretability. This paper proposes a hierarchical deep learning framework for de novo molecular structure prediction from EI-MS spectra using SELFIES representations. The framework progressively narrows the chemical search space through a sequence of stages comprising group classification, subgroup classification, clustering, and molecular structure generation. The approach is evaluated on 183,516 compounds from the NIST EI-MS dataset using an end-to-end pipeline. Results show that the hierarchical model achieves a Tanimoto similarity of 1.0 for 39% of test molecules overall, with 41% accuracy for non-aromatic and 37% for aromatic compounds, outperforming non-hierarchical baselines on several subsets. Comparative analysis against state-of-the-art methods demonstrates that the proposed framework offers competitive accuracy while enhancing interpretability and robustness for de novo molecular identification. These findings highlight the effectiveness of hierarchical modeling combined with SELFIES encoding for molecular structure elucidation from EI-MS data. MLD-Sup: Multi-scale Latent Deep Supervision for Joint Segmentation and Classification of Breast Ultrasound Images Nimalesh Elangovan, Tushar Shinde, and Hitika Tiwari (IIT Madras Zanzibar) Abstract Abstract Breast ultrasound screening is critical for dense breast tissue populations where mammography sensitivity is lim- ited (40–60%). Accurate joint lesion localization and malignancy classification are essential for screening workflows. However, existing joint-learning frameworks face two challenges: (1) severe class imbalance, and (2) progressive spatial information loss in deep decoders that degrades boundary precision for highly irregular lesion geometries. While deep supervision has proven effective for single-task segmentation (UNet++), its application to joint tasks under severe class imbalance, where gradient interference between segmentation and classification destabilizes training, remains underexplored. To address these challenges, we propose MLD-Sup (Multi-scale Latent Deep Supervision), a joint-learning architecture with three contributions: (i) an EfficientNet-B6 encoder with CBAM attention for noise-robust feature extraction, (ii) class-aware latent augmentation that perturbs bottleneck features of minority classes to mitigate imbalance without corrupting pixel-level spatial labels, and (iii) multi-scale decoder supervision that anchors intermediate fea- tures at resolutions {1/2, 1/4, 1/8, 1/16} to ground-truth masks via MSE loss, preventing spatial degradation while stabilizing joint-task gradients. Comprehensive evaluation on three multi- institutional breast ultrasound datasets under patient-wise 4- fold cross-validation demonstrates that MLD-Sup achieves DSC 0.819 ± 0.021 and F1-score 0.903 ± 0.015, representing 20.4% relative DSC improvement and 12.8% relative F1 improvement over the strongest baseline UNet++, validating the effectiveness of multi-scale supervision for robust joint lesion localization and classification in breast ultrasound screening applications. Fully Automatic Deep Learning Approaches for Cervical Spine Fracture Characterisation in CT Images Elena Goyanes, Carmen Lozano, Joaquim de Moura, Jorge Novo, and Marcos Ortega (Varpa Group, University of A Coruña) Abstract Abstract Cervical spine fractures are a critical clinical condition that requires an accurate diagnosis to prevent severe neurological complications. Although Computed Tomography (CT) is the imaging modality of choice for cervical spine evaluation, manual interpretation takes time and is subject to inter‑observer variability, motivating the development of robust automated solutions. In this work, we present a fully automatic deep learning pipeline for cervical spine fracture analysis on CT scans, the first in the state-of-the-art comprising three complementary tasks: fracture screening, vertebrae segmentation, and fracture localisation. Fracture screening is formulated as a classification problem to identify CT slices containing fractures. Vertebra segmentation is achieved through a transfer‑learning strategy based on brain MRI pretraining, enabling effective transfer of knowledge between modalities to cervical CT. For fracture location, we explore two alternative approaches: coordinate‑based regression and prediction of the bounding‑box, providing precise spatial characterisation of fracture regions. The proposed pipeline is evaluated on large dataset comprised by 711,601 images from 2019 patients. The experimental results demonstrate strong and consistent performance across all three tasks, supporting accurate and efficient fracture characterisation. Overall, our methodology has the potential to assist radiologists by improving diagnostic accuracy while reducing analysis time and workload. Wednesday 0.14 Singapore IJCNN Paper AI for Energy and Resource Analytics Session Chair: Enrico De Santis (University of Rome "La Sapienza"), Jean Senellart (Quandela) Meta-Learning Deep Kernels for Turbulence Forecast Verification Grzegorz Zakrzewski (Warsaw University of Technology, Institute of Meteorology and Water Management – National Research Institute) and Jacek Mańdziuk (Warsaw University of Technology, AGH University of Krakow) Abstract Abstract Accurate turbulence forecasts help prevent aviation accidents, but improving forecasts requires reliable ground truth for evaluation. Constructing the best available estimate of the true atmospheric turbulence state -- called reanalysis -- is not a trivial task, as turbulence encounters are sparse, confined to flight paths, and measured in heterogeneous units. We tackle this problem using Gaussian Processes, which learn spatial and temporal covariances directly from the data and provide uncertainty estimates. To capture complex patterns, we combine standard RBF and linear kernels with deep kernels, where a neural network extracts features before covariance computation. We show that meta-learning of kernel hyperparameters and network weights across many reanalysis construction tasks outperforms per-task optimization. We also demonstrate that our approach is competitive with standard reanalysis construction methods. Assessing Covariate-Informed Grid Load Forecasting with a Time-Series Foundation Model Varsha Pendyala, Yiwei Fu, Weizhong Yan, and Nurali Virani (GE Vernova Advanced Research Center) Abstract Abstract Modern power systems are growing increasingly complex as they integrate diverse generation sources to meet rising demand, making accurate load forecasting challenging. Recent advances in time-series foundation models (TSFMs) resulted in promising performance in zero-shot univariate load forecasting tasks. However, real-world load forecasting often involves multiple target variables and requires the integration of exogenous variables, raising important questions about the utility of TSFMs in realistic settings. In this study, we position Chronos-2, a recently developed model by Amazon, as a representative multi-channel TSFM that supports univariate, multivariate, and covariate-informed forecasting, and conduct a systematic investigation of how such models can be used for real-world load forecasting. While prior work has evaluated Chronos-2 on a limited number of energy-related tasks in a zero-shot setting, its performance relative to established task-specific deep learning models and its behavior when adapted using task-specific historical data remains insufficiently understood. In this work, we evaluate Chronos-2 on two real-world utility datasets, ISO New England and ENTSO-E, and benchmark it against widely used task-specific deep learning models. Our results show that Chronos-2 benefits substantially from task-specific fine-tuning and achieves strong short-horizon forecasting performance, but its zero-shot accuracy lags behind task-specific models and its forecasting error grows more rapidly with increasing forecast steps. Overall, this study provides a detailed characterization of the strengths and limitations of TSFMs such as Chronos-2 in grid load forecasting and offers practical insights into how a pretrained TSFM can be effectively adapted for operational load forecasting applications. A Conditional GAN Framework for Joint Generative Modeling of Geological Facies and Acoustic Impedance Israel Efraim De Oliveira and Rafael De Santiago (Universidade Federal de Santa Catarina); Yineth Viviana Camacho-De Angulo (Universidad Tecnológica de Bolívar, Instituto SENAI de Inovação em Sistemas Embarcados); Tiago Mazzutti (Instituto Federal Catarinense); Bruno B. Rodrigues (PETROBRAS); and Mauro Roisenberg (Universidade Federal de Santa Catarina) Abstract Abstract Facies modeling and the generation of associated rock properties are fundamental tasks in subsurface reservoir characterization, particularly in settings where direct observations are sparse and training data are limited. Traditional geostatistical workflows typically model categorical facies and continuous petrophysical properties in separate steps, while recent deep generative approaches often require large datasets and focus on single-variable synthesis. In this work, we propose a conditional generative adversarial network (cGAN) framework for the simultaneous generation of geological facies and acoustic impedance, explicitly conditioned on hard data from wells. The methodology builds upon a multi-scale SinGAN architecture, employing a hierarchical sequence of generators and discriminators that progressively refine joint facies–property realizations from coarse to fine resolutions. By extending the generator to a multi-channel output, the proposed approach captures the spatial dependencies between discrete facies distributions and continuous rock properties within a unified model. The method is validated using a limited set of two-dimensional realizations derived from the Stanford Earth Science dataset. Quantitative evaluation based on distributional metrics and conditioning accuracy demonstrates that the model effectively honors well data while producing geologically plausible and diverse joint realizations. The results indicate that the proposed framework offers a data-efficient and flexible solution for integrated facies and rock property modeling under constrained data availability. AutoML for Crude Oil Price Forecasting: A Comparative Study Marcus Silva and Márcio Basgalupp (UNIFESP) Abstract Abstract Crude oil price forecasting is a critical task for industrial decision-making in volatile economic environments. This study proposes an automated forecasting system based on historical crude oil price data and investigates the problem using both time-series and tabular data representations. An Automated Machine Learning (AutoML) framework is employed to automate model selection and hyperparameter optimization, reducing reliance on domain-specific expertise. A unified AutoML pipeline is developed to evaluate multiple forecasting models across different temporal granularities and prediction horizons, using time-aware validation to preserve data chronology. Historical prices are collected via an application programming interface, ensuring reproducibility and enabling continuous data updates. Experimental results demonstrate that AutoML-based approaches deliver robust, scalable forecasting performance, supporting their use as effective decision-support tools in industrial contexts. Comparisons with traditional methods, including autoregressive models and recurrent neural networks, are conducted to assess the empirical benefits of adopting AutoML. Multi-Adapter PPO: A Cross-Attention Enhanced Wavelength Selection Framework for LIBS Quantitative Analysis Hao Li and Man fung Zhuo (University of Arizona) Abstract Abstract Laser-induced breakdown spectroscopy (LIBS) quantitative analysis faces critical challenges in wavelength selection due to high-dimensional spectral data and the fundamental trade-off between prediction accuracy and feature efficiency. This paper presents a novel Multi-Adapter PPO framework that transforms wavelength selection into a reinforcement learning problem, leveraging cross-attention mechanisms and multiple specialized adapters to capture complex spectral relationships. Our approach outperforms traditional Particle Swarm Optimization (PSO) by an average of 28.4\% in comprehensive score and 45.2\% in prediction accuracy across steel and coal datasets. The proposed method demonstrates superior performance in balancing prediction accuracy with feature efficiency, achieving state-of-the-art results in LIBS quantitative analysis while maintaining interpretability and computational efficiency. Wednesday 0.15 Washington IJCNN Paper AI for Finance and Economics Session Chair: Aleksei Liuliakov (Bielefeld University), Isaac Amankona Obiri (Quanzhou University of Information Engineering) Graphical Dual Axial Transformer for Risk-adjusted Portfolio Optimization Youjia Liu and Yasumasa Matsuda (Tohoku University) Abstract Abstract Portfolio optimization with financial time series requires jointly modeling temporal dynamics and cross-sectional asset interactions under noisy and non-stationary market conditions. Existing deep learning approaches often struggle to disentangle these heterogeneous dependencies or rely on reinforcement learning frameworks that lead to unstable training. We propose the Graphical Dual Axial Transformer (GDAT), which models financial data as a unified 3D tensor across assets, time, and features. GDAT employs a dual axial attention mechanism to disentangle temporal dependencies from cross-sectional correlations, while enabling structured information flow between axes. To stabilize relationship selection and filter market noise, the model incorporates a graph-structured inductive bias derived from sparse inverse covariance estimation, guiding attention toward economically meaningful interactions. By risk-adjusted return optimization, the proposed framework allows stable end-to-end training. Experiments on real-world datasets, including S&P 500, Russell 1000, and FTSE 100, demonstrate that GDAT consistently outperforms classic online learning and deep learning methods in terms of risk-adjusted performance and robustness. These results demonstrate the benefits of structured attention and cross-sectional modeling for portfolio optimization. AlphaTree: Structure-Aware Reinforcement Learning for Formulaic Alpha Mining Xiaolong Yang, Yang Wang, and Ya-Hui Jia (South China University of Technology); Qiang Yang (Nanjing University of Information Science and Technology); and An Song (GF Asset Management) Abstract Abstract Formulaic alpha discovery is crucial in quantitative finance but remains challenging due to the vast search space for expressions. Existing reinforcement learning methods treat alpha formulas as flat token sequences, failing to capture their inherent tree structure. In this paper, we propose \textbf{AlphaTree}, which uses a Tree-structured Long Short-Term Memory (Tree-LSTM) network to encode the hierarchical structure of mathematical expressions during generation. Unlike sequential models that process tokens left-to-right, Tree-LSTM captures operator-operand relationships through bottom-up information propagation over the expression tree. Furthermore, we integrate Tree-LSTM with distributional reinforcement learning to address non-stationarity and reward sparsity. Experiments on the CSI300 and CSI500 datasets demonstrate that AlphaTree significantly outperforms existing methods, achieving higher information coefficients and producing substantial improvements in risk-adjusted returns. Our code is available at \url{https://github.com/ftwangyang/AlphaTree} Fundamental vs. Technical Features for Stock Trend Prediction: A Multi-Horizon Data Stream Learning Analysis José Júnior de Oliveira Silva (Universidade Federal de Pernambuco, Instituto Federal de Alagoas); Roberto Souto Maior de Barros (Universidade Federal de Pernambuco); and Silas Garrido Teixeira de Carvalho Santos (Universidade Federal de Pernambuco, SiDi Institute) Abstract Abstract This paper systematically compares the predictive power of fundamental versus technical features to forecast stock price movements in a data stream learning setting. To enable this comparison, we construct the B3 Technical and Fundamental (B3TF) dataset, which provides daily fundamental ratios rigorously aligned to avoid look-ahead bias, as well as adjusted market data for companies listed on the Brazilian Stock Exchange (B3). This novel contribution solves the temporal misalignment problem. A comparative analysis was conducted using Hoeffding Trees and a delayed Prequential protocol across nine forecast horizons. To the best of our knowledge, this is the first work to make this comparison using a strictly incremental learning protocol. Our results reveal a structural divergence. While technical models maintain stable accuracy around the random baseline, their discriminative power degrades, with global kappa turning negative (reaching κ = −0.0246 at h = 240) in the long term. In contrast, fundamental features demonstrate greater robustness. Although the predictive signal remains marginal, it persists over longer horizons, reaching a global kappa peak (κ = 0.0469) and accuracy = 52.39% at h = 150 trading days. Validating these findings on 10 highly liquid assets confirms the trend: the fundamental model retained stability (κ = 0.0129), whereas the technical model collapsed (κ = −0.0537). These findings delineate a boundary of predictability: fundamental ratios provide persistent signals driven by value convergence, distinguishing them from noise even when price-based models fail. Random Convolution Kernel–Augmented Temporal Convolutional Networks for Volatility-Dominated Financial Time Series Akanksha Sharma (Maulana Azad National Institute of technology, Bhopal) and Chandan Kumar Verma (Maulana Azad National Institute of Technology, Bhopal) Abstract Abstract Due to the nonstationary, noise-dominated, and nonlinear nature of volatility time series, predicting financial market volatility is challenging. Although the accuracy of forecasts has been enhanced by deep learning models in recent times, a lot of these models still depend on trainable components that are vulnerable to noise and sudden spikes in volatility. This research presents an RCK-TCN, an enhanced version of a standard TCN that incorporates fixed random convolution kernels for robust local feature extraction, to overcome these constraints. A TCN backbone models long-range temporal relationships through dilated causal convolutions, while random kernels capture short-term volatility patterns without increasing model complexity. Several equities and commodities implied volatility indexes, such as the VIX, India VIX, VXN, OVX, and GVZ, are used to assess the suggested methodology. Through the use of several error metrics, cross-validation, and SHAP-based interpretability, the experimental results show that RCK-TCN reliably produces better volatility forecasts than current deep learning baselines. An Enhanced Temporal Graph Network with Retrieval-Augmented Graph for Dynamic Link Prediction in Cryptocurrency Transactions Isaac Amankona Obiri (Quanzhou University of Information Engineering), Ansu Badjie (University of Electronic Science and Technology of China), and Abigail Akosua Addobea (Quanzhou University of Information Engineering) Abstract Abstract Link prediction in cryptocurrency transaction networks is inherently challenging due to their large scale, temporal dynamics, and evolving interaction patterns between entities. Traditional static graph-based approaches often fail to capture these temporal dependencies and are limited in modeling missing or incomplete transactional information. This work presents an Enhanced Temporal Graph Network (ETGN) with a Retrieval-Augmented Graph (RAG) model for dynamic link prediction in cryptocurrency transactions. The proposed framework integrates a Multi-Graph LSTM (MGLSTM) with graph attention mechanisms to jointly model structural and temporal dependencies while incorporating a missing information prediction module to mitigate data sparsity and incompleteness. Node representations are enriched through deterministic feature extraction and centrality-driven labeling using affinity and anti-affinity scores derived from transaction graphs. To enhance generalization across time slices, a graph memory module is introduced, enabling the retrieval of historical edge embeddings and contextual fusion for improved prediction accuracy. The model performs edge-level classification to predict the likelihood of future transactional links. Experiments conducted on real-world Bitcoin and Ethereum transaction network data, organized into temporal slices, demonstrate that the proposed ETGN-RAG framework achieves strong performance in terms of AUC-ROC, precision, recall, and F1-score. The results highlight the effectiveness of combining temporal graph learning, attention-based aggregation, and retrieval-augmented memory for robust link prediction in dynamic cryptocurrency transaction networks. Direct Log-Utility Optimization for Portfolio Allocation under Exogenous Market Dynamics Ivan Kurnosau and Abdulrahman Altahhan (University of Leeds) Abstract Abstract We propose an efficient learning framework for long--short portfolio allocation that combines an objective-driven optimisation criterion with the simplicity of supervised training. Under \emph{exogenous market dynamics}, where market transitions are independent of the agent's actions under the price-taker assumption, maximising \emph{log-utility} reduces to a \emph{single-step} objective; hence, trajectory rollouts and multi-step credit assignment used by policy-gradient trading agents (e.g., \textit{AlphaStock}, \textit{DeepTrader}) often provide limited benefit. We therefore train the allocation policy end-to-end by direct log-utility optimisation of a \emph{regularised} objective using standard mini-batch backpropagation on observed returns. We evaluate against both trajectory-based reinforcement-learning agents and supervised portfolio-allocation baselines, including \textit{Decision by Supervised Learning} (DSL), and find that our training setup consistently delivers stronger performance in our evaluation. Empirically, our method achieves substantially higher average annual return and Sharpe ratio than the evaluated baselines, while converging in fewer training iterations and with improved stability. Overall, direct log-utility optimisation under exogenous dynamics provides a practical middle ground between reinforcement learning and supervised portfolio allocation. Wednesday Virtual Room 1 IJCNN Paper LLM Adaptation and Fine-Tuning III Session Chair: Di Wu (Hebei University of Engineering), Yongqing Wang (School of Computer Science and Technology, Tongji University) Dual-Mechanism Balanced Graph Partitioning: Contrastive Learning for Coarse Partition Generation and Reinforcement Learning for Fine-Tuning Yongqing Wang and Guosun Zeng (School of Computer Science and Technology, Tongji University) Abstract Abstract Traditional balanced graph partitioning methods are mainly heuristic and experience-driven. Recent advances adopt deep learning, but existing approaches suffer from loss function mismatch and modeling redundancy. To address these issues, this paper proposes a Dual-Mechanism Balanced Graph Partitioning method (DMBGP), combining contrastive learning for coarse partition generation and reinforcement learning(RL) for fine-tuning. Specifically, positive and negative samples are leveraged to train a graph neural network via contrastive learning, with a tailored partition contrastive loss to extract label-agnostic features and generate initial partitions. RL then performs pairwise fine-tuning of adjacent partitions, and the final global result is obtained through iterative optimization over all adjacent pairs. The pairwise tuning paradigm alleviates modeling redundancy, decouples the state-action space from partition cardinality, and supports arbitrary k-partitioning tasks. Extensive experiments on diverse datasets demonstrate that DMBGP outperforms traditional heuristic methods and state-of-the-art deep learning methods, providing an efficient solution for deep learning-enabled balanced graph partitioning. Region Partitioning and Prototype-Query Calibration for Few-Shot Fine-Grained Image Classification Zhiqiang Guo, Longgang Xiao, and Qiwen Jin (Wuhan University of Technology) Abstract Abstract Few-shot fine-grained image classification aims to distinguish sub-classes within the same parent class using only a few labeled support samples. The main challenges arise from severe background interference and significant pose variations. Existing methods often rely on external auxiliary models and fail to mitigate the bias between support and query samples. In this paper, we propose an end-to-end framework that overcomes these challenges by reducing background interference and pose instability. Our approach decouples feature extraction into global and local branches to enhance foreground regions and suppress irrelevant background noise. Query features and prototypes are then calibrated through similarity weighting and cross-attention to ensure robust matching despite distribution shifts. Moreover, we introduce a Region Partition Module (RPM) to better address background noise and a Prototype-Query Calibration Module (PQCM) that stabilizes class prototypes and aligns query features, improving matching accuracy. Extensive experiments on three fine-grained benchmark datasets demonstrate that our model significantly outperforms state-of-the-art methods. Labeled TrustSet Guided: Batch Active Learning with Reinforcement Learning Guofeng Cui (Nvidia); Yang Liu (Amazon); Pichao Wang (Nvidia); and Hankai Hsu, Xiaohang Sun, Xiang Hao, and Zhu Liu (Amazon) Abstract Abstract Batch active learning (BAL) is a crucial technique for reducing labeling costs and improving data efficiency in training large-scale deep learning models. Traditional BAL methods often rely on metrics like Mahalanobis Distance to balance uncertainty and diversity when selecting data for annotation. However, these methods predominantly focus on the distribution of unlabeled data and fail to leverage feedback from labeled data or the model’s performance. To address these limitations, we introduce TrustSet, a novel approach that selects the most informative data from the labeled dataset, ensuring a balanced class distribution to mitigate the long-tail problem. Unlike CoreSet, which focuses on maintaining the overall data distribution, TrustSet optimizes the model’s performance by pruning redundant data and using label information to refine the selection process. To extend the benefits of TrustSet to the unlabeled pool, we propose a reinforcement learning (RL)-based sampling policy that approximates the selection of high-quality TrustSet candidates from the unlabeled data. Combining TrustSet and RL, we introduce the Batch Reinforcement Active Learning with TrustSet (BRAL-T) framework. BRAL-T achieves state-of-the-art results across 10 image classification benchmarks and 2 active fine-tuning tasks, demonstrating its effectiveness and efficiency in various domains. Inverse Prompt Response Reconstruction and Contextual Momentum Contrast for Few-Shot Dialogue Generation Di Wu, Fanming Meng, and Yao Peng (Hebei University of Engineering) Abstract Abstract Dialogue generation is an essential component of task-oriented dialogue systems, aiming to assist users in completing domain-specific tasks via natural language interaction. Existing few-shot dialogue generation models face challenges in perceiving domain dialogue state and generating responses that are inconsistent with the dialogue context. To address these challenges, we propose the IPCMC-TOD model, a Few-shot Dialogue Generation Model with Inverse Prompt Response Reconstruction and Contextual Momentum Contrast. Specifically, considering the critical role of the model’s adaptability to domain dialogue state, we design an inverse prompt response reconstruction. An inverse-guided dual-prompt template is constructed. An inverse-prompt-trained slot extractor is employed to reconstruct the original system responses. To enhance the logical consistency of dialogue responses, we design a Contextual Momentum Contrast learning. Based on the posterior inference response, the complete set of samples is categorised into positive and negative sets. Momentum-based pseudo-samples are generated by the momentum encoder. We define a contrast loss function with margin constraints for positive and negative samples. This function significantly enhances the fine-grained discrimination ability of the response selector. Experimental results on the MultiWOZ 2.0 and MultiWOZ 2.1 datasets demonstrate the effectiveness of IPCMC-TOD over all baseline models. Wednesday Virtual Room 2 IJCNN Paper LLM Adaptation and Fine-Tuning IV Session Chair: Ken Zhong (Shanghai Jiao Tong University), Long Chen (China West Normal University) A Syntactic Feature Interaction and Cross-Table Relation Search Model for Threat Intelligence Triple Extraction Long Chen, Chong Zhao, Jingjun Deng, and Dongming Zuo (China West Normal University) and Jianjiang Zheng (Chongqing Vocational Institute of Tourism) Abstract Abstract Joint relational triple extraction models based on table filling are the key methods for Cyber Threat Intelligence (CTI) analysis. However, traditional table models typically encode the subject and object of CTI independently and excessively focus on the local features of a single table, resulting in overlooking subject-object semantic interaction and the global associations of relations and entities. To address these issues, we propose a threat intelligence triple extraction model based on Syntactic Feature interaction and Cross-table Search strategy (SyCro). First, we design a Syntactic-aware Graph Convolutional Network (SGCN), which calculates a syntactic feature matrix by integrating syntactic dependency, thereby establishing syntactic associations between subjects and objects. Second, we construct a table feature for each relation of triple and design a cross-table relation search strategy, which utilizes FourierKAN-Attention to capture the interaction relationships between single tables, thereby extracting global relation features. Finally, we integrate the syntactic feature matrix into the global relation features and iteratively apply the above process multiple times to obtain the final table features, thereby further exploring the global associations of relations and entities. Experimental results demonstrate that SyCro significantly outperforms existing models, achieving SOTA results on two intelligence datasets. Our code is available at https://github.com/xiaochen-xihua/SyCro. Spatial Relation Classification on Few-shot In-context Learning Guoqi Yang, Peifeng Li, and Qiaoming Zhu (Soochow University) Abstract Abstract The goal of the spatial relation classification task is to identify the type of spatial relations between entities extracted from text. Although large language models (LLMs) have demonstrated the potential to achieve excellent performance across various tasks through In-Context Learning (ICL), there remains a lack of exploration in the domain of spatial relation classification. The gap between ICL and fully supervised training models primarily stems from two aspects: current example retrieval methods fall short in capturing the semantic relevance of spatial elements and spatial relations; and the absence of explicit explanations for the mapping between inputs and spatial relation labels in examples hinders the model’s ability to effectively leverage contextual information.To address these challenges, we propose the Spatial-ICL method, which specifically optimizes the application of ICL in spatial relation classification. Experimental results demonstrate that Spatial-ICL outperforms existing ICL baselines based on GPT-4 in terms of performance. DUAL-RP: Dual-Source Structural and Textual Re-ranking for Relation Prediction Siyan Wu, Chenghua Zhu, and Jieyu Zhan (South China Normal University) and Lihua Cai (South China Normal University; Xiamen Rekey Medical Technology Co., LTD, Xiamen, China) Abstract Abstract Knowledge graph completion (KGC) is a fundamental task in database and AI systems, aiming to infer missing facts in knowledge graphs and thereby improve the completeness and usability of structured knowledge. However, existing methods often struggle to jointly capture structural reasoning and semantic precision within a unified framework. We present DUAL-RP, a relation prediction framework that bridges knowledge graph embedding (KGE) models and large language models (LLMs) through a semantic re-ranking paradigm. DUAL-RP first employs lightweight KGE models to generate a compact set of candidate relations for a given entity pair. A token translator module then aligns continuous structural embeddings with the token embedding space of LLMs, enabling effective interaction between structural priors and semantic reasoning. In addition, multi-hop relational paths are incorporated as natural-language context to provide complementary semantic cues for relation ranking and interpretability. Experiments on FB15k-237, CoDEx-S, and DBpedia50 show that DUAL-RP consistently delivers strong overall ranking performance across diverse baselines, achieving the most consistent improvements in Mean Reciprocal Rank and Hits@1. MetaRTL: Meta-path Attention Enhanced Relational Table Learning Ken Zhong, Weichen Li, and Zheng Wang (Shanghai Jiao Tong University) Abstract Abstract Relational table learning has gained increasing attention with the widespread use of relational databases. Existing methods typically rely on deep GNN or HGNN stacks, leading to high computational costs and limited performance on large real-world databases. We propose MetaRTL, a two-stage framework for scalable and expressive relational table learning. In the first stage, MetaRTL obtains initial table embeddings via lightweight pre-training. In the second stage, it performs non-parametric message passing to derive meta-path features, which are then aggregated by an attention module, MetaAttn. By shifting computation from deep message passing to efficient meta-path aggregation, MetaRTL captures rich relational semantics while maintaining high efficiency. Experiments on 10 real-world datasets across 24 tasks demonstrate the effectiveness of the proposed method. Wednesday Virtual Room 3 IJCNN Paper LLM Agents and Tool Use I Session Chair: Junkai Zhang (Institute of Automation, Chinese Academy of Sciences) Cooperative Multi-Agent Reinforcement Learning for Heterogeneous Vehicle Routing Junkai Zhang, Yifan Zhang, and Jinmin He (Institute of Automation, Chinese Academy of Sciences); Yifan Zang (Beijing Institute of Astronautical Systems Engineering); and Jian Cheng and Yang Wu (Institute of Automation, Chinese Academy of Sciences) Abstract Abstract Combinatorial Optimization (CO) is fundamental to numerous fields with Heterogeneous Capacitated Vehicle Routing Problem (HCVRP) standing out as a critical application, which optimizes delivery fleet operations. Existing deep learning based (DRL-based) methods show promising performance but suffer from high action space. In this work, we propose Cooperative Multi-Agent Reinforcement Learning for Heterogeneous Vehicle Routing (MAVR) to enhance the performance. We decouple the routing problem into a two-stage learning process: (1) Node Allocation, where customer nodes are assigned to vehicles; (2) Node Traversal, where each vehicle searches for its shortest path. For Task (2), the reduced number of assigned nodes reduces the search space, making the task can be effectively solved using existing DRL-based methods. Consequently, our focus shifts to the Node Allocation task, which can be naturally seen as a cooperative multi-agent task. Specifically, we model each vehicle as an agent within a cooperative multi-agent framework. The process sequentially relays node information to each vehicle agent, allowing them to decide whether to select it. We further implement an adaptive process that dynamically adjusts the vehicle trip number based on a theoretical analysis which narrows down the practical vehicle trip number. We conducted experiments on different routing problems and the results showcase our improved performance. Adaptive Search in Cooperative Multi-Agent Reinforcement Learning with Parameter Sharing Yurui Li, Li Zhang, and Shijian Li (Zhejiang University) Abstract Abstract Parameter sharing is a widely used approach to tackling the search challenges in the joint observation-action space within Multi-Agent Reinforcement Learning (MARL). Despite its popularity, the effectiveness of parameter sharing varies across different tasks. While previous efforts have sought to enhance the applicability of parameter sharing, none have provided a theoretical analysis of its underlying mechanism. In this study, we demonstrate that the effectiveness of parameter sharing stems from its ability to solidify the search space of MARL methods in the joint observation-action space. However, this solidification isn't applicable universally, rendering it ineffective in certain scenarios. To address this limitation, we propose a novel framework that enables parameter sharing to adaptively determine the search subspace based on specific task. We implement this framework with two representative MARL methods and evaluate their performance across two widely used MARL testbeds. The experimental results confirm that our framework can dynamically adjust the search subspace across tasks, thereby enhancing the performance of the original methods. HDCF: Bi-level Diversity Coordination for Hierarchical Cooperative Multi-Agent Reinforcement Learning zhou wang, mengke wang, xiangfeng luo, and shaorong xie (Shanghai University, School of Computer Engineering and Science) Abstract Abstract Hierarchical multi-agent reinforcement learning is a promising approach for complex cooperative tasks, since temporally extended macro-actions can support diverse behaviours over long horizons. However, high-level macro-action learning is often weakly grounded in behavioural outcomes, which may cause a few macro-actions to dominate while many become redundant, thereby limiting the behavioural diversity available for cooperation. To address this, we propose HDCF, a novel framework that formulates diversity learning as a bi-level optimisation process coordinating high-level macro-action regulation with low-level execution. At the high level, HDCF introduces counterfactual similarity-based shaping to regulate macro-action learning using behaviour-level signals. At the low level, HDCF incorporates an advantage-based inter-level incentive to strengthen hierarchical alignment, so that different macro-actions translate into reliably differentiated executions. Experiments on SMAC and GRF show that HDCF outperforms representative baselines, and ablation studies validate that both components contribute to the overall improvement. Multi-Agent Reinforcement Learning for Non-Stationary Edge Task Scheduling David Li (IEEE) Abstract Abstract Task scheduling in edge-enabled Internet-of-Things systems becomes fundamentally harder when device capability is non-stationary: repeated execution can degrade near-term processing efficiency, while idleness enables recovery. Existing learning-based schedulers typically optimize latency or utilization under stationary or exogenous service models, which limits their ability to learn when temporary deferral is beneficial. We formulate this problem as a multi-agent stochastic game in which each device carries a time-varying processing-efficiency state governed by controlled mean-reverting dynamics. Building on this model, we develop a degradation-recovery reinforcement-learning scheduler whose main instantiation is a tabular Q-learning policy over backlog, efficiency, and shared-context variables, and we discuss a policy-gradient extension for continuous-state settings. Experiments on heterogeneous edge IoT systems show that the learned policy acquires recovery-aware behavior, executing aggressively in high-efficiency states and deferring selectively when devices are degraded. Across low, medium, and high load regimes, the proposed method reduces deadline violations by 8–31\% and improves throughput by 12–19\% relative to fatigue-unaware reinforcement learning, with gains increasing as degradation strength grows. The method also maintains stable backlog as the number of devices increases and achieves sub-millisecond inference with constant per-device model size. These results show that explicitly modeling endogenous degradation can substantially improve reinforcement-learning-based scheduling in non-stationary resource-constrained environments. Wednesday Virtual Room 4 IJCNN Paper LLM Agents and Tool Use II Session Chair: Linqi Ye (Shanghai University), Yibo Chen (PLA Academy of Military Science) Multi-Panzer: a LLM-based Multi-agent System for Tank Detachment Combat Simulation Yibo Chen and Yang Ping (PLA Academy of Military Science) Abstract Abstract To further expand the adversarial intelligence capability of tank detachment combat simulation, this study designs a large language model-based multi-agent system for tank detachment combat, Multi-Panzer, using a modular design approach. The Multi-Panzer architecture comprises key components including a scenario information converter, combat unit agents, symbolic planner, communication module, and command agent, wherein combat unit agents possess planning and memory structures. Simulation results demonstrate that various large language models (LLMs) can be applied within the Multi-Panzer architecture; Multi-Panzer exhibits superior combat capabilities compared to traditional agents and demonstrates effectiveness across various scenario terrains; the communication module and command agent enhance Multi-Panzer’s collaborative abilities; the prompt chaining method and memory bank configuration of combat unit agents ensure logical decision-making and behavioral coherence. The simulation results indicate that enhancing cooperation among combat units in tank detachment combat simulations is more important than strengthening individual cognition. This study provides a novel framework for LLM-based multi-agent systems in tank detachment combat simulation, with its architectural design offering guidance for developing LLM-based multi-agent systems in combat simulation. Naval Missile Strike Layout: A Hybrid LLM Agent Architecture for Naval Firepower Planning Yibo Chen and Yang Ping (PLA Academy of Military Science) and Shuhang Zhou (North University of China) Abstract Abstract To address the challenge of traditional naval firepower planning methods in balancing scheme rationality, diversity, and robustness under complex battlefield uncertainties, this study proposes a Naval Missile Strike Layout (NMSL) framework, integrating Large Language Models (LLMs) with enhanced conventional planning methods, which achieves the generation, evaluation, and optimal selection of naval firepower planning schemes during pre-combat preparation through multi-LLM collaboration. The architecture consists of four core modules: a situation analysis module, decoupled generation module, route planning module, and scheme evaluation module. Experimental results demonstrate that in two typical naval combat scenarios, NMSL achieves comprehensive performance improvements ranging from 39.1% to 74.7% compared with baseline methods, achieving state-of-the-art performance while validating the necessity and effectiveness of modular design. This study presents the first planning-oriented LLM agent in the domain of naval firepower strike capable of generating a complete strike scheme solely from initial conditions. This research provides a novel framework for firepower planning agents and offers universal architectural design implications for intelligent planning systems with various adversarial styles. HMem: A Hawkes Process Memory Framework for Long-Horizon Agents Zixuan Yan (Zhejiang University) and Xuanyi Wu and Jinling Wei (School of Computer and Computing Science,Hangzhou City University) Abstract Abstract In recent years, large language models (LLMs) have demonstrated increasingly remarkable capabilities, but they still have limitations in dealing with long context. Therefore, the LLMs-powered agents still have deficiencies in long-horizon scenarios. Recent studies on the memory of agents can alleviate this problem. In the context-aware agent memory, how to effectively organize and retrieve memories is a problem. Inspired by the Hawkes process, its self-excitation property can simulate the suddenness and associative mechanism of human reasoning, and its time-decay property can simulate human memory management, we proposed HMem, an agent memory framework, which combine Hawkes process and human cognitive mechanism. It consists of Dual Intensity Kernel Memory(DIKM), which distinguishes between working memory and factual memory and applies different intensity kernel functions, and a Hawkes-Enhanced Scoring Engine(HESE), which combines semantic similarity and temporal logic for better retrieval. In this way, HMem enables agents effectively filter out noisy memories and maintain behavioral consistency in long-horizon scenarios. Experiments on the LoCoMo dataset show that HMem outperforms baselines expecially in Single Hop and Multi Hop tasks, which demonstrated the effectiveness and practical value of our method. HIM: A Human-Inspired Memory Loop for LLM Agents via Encoding, Consolidation, and Retrieval Haoze Tang, Linqi Ye, and Shaorong Xie (Shanghai University) Abstract Abstract Large language model (LLM) agents often degrade during long-horizon, multi-session, and multi-speaker interactions due to accumulating noise, loss of source attribution, and interference from stale memory traces. Existing memory systems frequently emphasize capacity or retrieval effectiveness, while leaving reliability under-specified, which can lead to attribution drift and suboptimal retrieval in complex dialogues. We introduce HIM, a human-inspired memory framework that enforces a reliability-oriented lifecycle loop across Encoding, Consolidation, and Retrieval. During Encoding, interactions are stored as structured notes with explicit speaker attribution and contextual cues, and low-value entries are regulated through an importance-gated policy into stratified retention levels. During Consolidation, the system performs periodic rehearsal-like maintenance, strengthening repeatedly useful traces via usage signals and stabilizing memory associations over time. During Retrieval, HIM combines similarity-based search with source-attributed evidence presentation and an activation-style score for interpretability, improving robustness under interference. Evaluated on the LoCoMo benchmark, HIM achieves stronger performance than representative baselines including ReadAgent, MemoryBank, MemGPT, and A-MEM in F1 and BLEU-1, with notable gains on adversarial, multi-hop, and temporal queries, indicating improved attribution fidelity and interference resistance in prolonged interactions. Wednesday Virtual Room 5 IJCNN Paper LLM Agents and Tool Use III Session Chair: Lisan Al Amin (University of Maryland, Baltimore County), Zhuoyi Huang (National University of Defense Technology) Deploying a ML Agent for OS: A Comparison between User-Space and Kernel-Space Zhuoyi Huang, Long Peng, Zhuo Li, Xiaodong Liu, and Jie Yu (National University of Defense Technology) Abstract Abstract Machine learning (ML) has shown great potential for improving operating system (OS) performance, yet it is currently deployed primarily in user space. This leaves the performance, overhead, and effectiveness of kernel-space ML largely unquantified. To address this gap, we present a controlled empirical study that compares user-space and kernel-space deployment of the same ML agent. Our approach utilizes a dual-mode ML framework extended to support reinforcement learning (RL), allowing an identical Deep Q-Network (DQN) model to run in both spaces, thereby isolating deployment location as the only variable. We evaluate this using proactive page reclaim, a kernel-intensive task requiring deep access to system state. Experiments show that our kernel-space DQN agent reduces page refaults by 15.9% compared to the conventional LRU policy. More critically, it reduces total decision latency by 98.9% and achieves a 75.8% reduction in context switches, a 62.4 times higher inference throughput and a 16.7 times reduction in runtime memory overhead compared to user-space deployment. These results provide a quantitative foundation for integrating ML algorithms into OS, confirming that kernel-space integration is essential for realizing low-latency, high-throughput, and efficient learning-driven OS optimizations. Simple Multiple Kernel k-Means with Heat Kernel Diffusion Zhiwei Xia, Xiaohong Jia, and Xuejun Zhang (School of Electronic and Information Engineering, Lanzhou Jiaotong University,); Yao Zhao (Institute of Information Science Beijing Jiaotong University,); and Wenqian Yu (Lanzhou No. 5 High school) Abstract Abstract Multiple kernel clustering (MKC) aims to capture complex nonlinear structures by integrating multiple predefined kernels. Simple multiple kernel 𝑘-means (SMKKM) jointly learns kernel weights and clustering assignments within a min–max optimization framework, achieving efficient and competitive performance. However, its direct linear fusion of original kernel matrices fails to exploit the intrinsic geometric structure of data, making it sensitive to noise and unreliable similarities. To address this issue, we propose SMKKM with heat kernel diffusion (SMKKM-HK). For each base kernel, a kernel-induced graph Laplacian is constructed, and a heat kernel operator is applied to perform bilateral spectral smoothing on kernel matrices, enhancing geometric consistency and suppressing high-frequency noise. Furthermore, a dynamic graph update mechanism guided by current clustering assignments is introduced, enabling the diffused kernels to adapt to evolving cluster structures. Extensive experiments demonstrate that SMKKM-HK achieves improved clustering performance and robustness compared with state-of-the-art MKC methods. Structure-Aware Operator Fusion and Dynamic Materialization for Efficient Transformer Inference Ruiting Sun, Honglu He, Guanwen Zhang, and Wei Zhou (Northwestern Polytechnical University) Abstract Abstract Deploying Transformer models in latency-critical applications is frequently hindered by memory bandwidth bottlenecks caused by fragmented operators and redundant global memory access. Conventional deep learning compilers often fail to optimize these heavy workloads effectively because their rule-based fusion heuristics remain overly conservative regarding complex graph dependencies. In this paper, we propose a novel optimization framework that combines Structure-Aware Operator Fusion with a Dynamic Materialization strategy. We introduce specialized kernels for Multi-Head Attention (MHA) and Feed-Forward Networks (FFN) that leverage shared memory tiling and streaming Softmax to rigorously minimize off-chip data traffic. Additionally, our approach resolves fusion conflicts at residual boundaries by dynamically searching for optimal materialization configurations instead of enforcing static safety barriers. Extensive evaluations on benchmarks including Vision Transformers (ViT) and Swin Transformer demonstrate significant improvements over industry standards. Our method achieves a 1.52 times speedup over PyTorch on ViT-Base and consistently reduces inference latency compared to an auto-tuned baseline, while producing smaller binary artifacts than established frameworks in cross-compiler comparisons. Quantum Kernels for Audio Deepfake Detection Using Spectrogram Patch Features Lisan Al Amin and Rakib Hossain (Potomac Quantum, USA); Mahbubul Islam (United International University, Dhaka, Bangladesh); Faisal Quader (University of Maryland, College Park, MD); and Thanh Thi Nguyen (Monash University: Melbourne, Victoria, AU) Abstract Abstract Quantum machine learning has emerged as a promising tool for pattern recognition, yet many existing approaches for audio treat spectrograms as generic images and do not explicitly leverage their time-frequency structure. We propose Q-Patch, a quantum feature map tailored to audio that encodes local time-frequency patches from mel-spectrograms into quantum states using shallow, hardware-efficient circuits with adjacency-reflecting entanglement. Q-Patch summarizes each selected patch with a compact four-dimensional acoustic descriptor and maps it to a four-qubit circuit with a depth of at most three, enabling practical quantum kernel construction under near-term constraints. We evaluate Q-Patch on an audio spoofing detection task in a controlled, balanced protocol and compare it against size-matched classical baselines. Q-Patch improves discrimination between bona fide and spoofed samples, achieving an AUROC of 0.87 compared to 0.82 for an RBF-SVM trained on the same patch-level features. Kernel-space analysis further shows a clear class structure, with cross-class similarity around 0.615 and within-class self-similarity of 1.000, indicating that the proposed feature map induces a separable similarity geometry. Overall, Q-Patch provides a practical framework for incorporating time-frequency-aware representations into quantum kernel learning for audio authenticity assessment in low-resource settings. Wednesday Virtual Room 6 IJCNN Paper LLM Agents and Tool Use IV Session Chair: Gangao Liu (Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences), Yuanjian Zhao (Sichuan University) Mitigating Context Divergence in Multi-Turn Embodied Reasoning via Hierarchical Decay Gangao Liu, Mengna Wang, and Peng Li (Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences) Abstract Abstract The integration of Large Language Models (LLMs) into embodied agents has enabled impressive capabilities in zero-shot task planning. While Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) have proven effective for aligning LLMs in dialogue tasks, applying them to multi-turn embodied reasoning introduces a critical challenge: Context Divergence. In sequential decision-making, a divergent action alters all subsequent observations, rendering the comparison between chosen and rejected trajectories mathematically ill-posed due to state drift. In this paper, we propose H-D2PO, a novel alignment framework specifically designed for embodied agents. H-D2PO introduces a hierarchical temporal decay mechanism that dynamically calibrates the optimization objective, prioritizing the causal branching decision while robustly filtering the noise from the divergent trajectory tail. Furthermore, we propose a Priority-based Preference Construction strategy that explicitly prioritizes long-horizon failures over short ones, encouraging the agent to learn reasoning stability even from unsuccessful attempts. Extensive experiments on ALFWorld and LOGICWorld demonstrate that H-D2PO significantly outperforms SFT and vanilla DPO baselines. Notably, on ALFWorld, our Qwen3-4B based agent achieves a success rate of 89.78%. Our analysis reveals that this performance gain stems from the agent's enhanced capability to persist through and solve complex, long-horizon tasks that baselines fail to complete. Exploitation Is All You Need... for Exploration Micah Rentschler (Vanderbilt University) and Jesse Roberts (Tennessee Technological University) Abstract Abstract Exploration is a central challenge when learning in novel situations. Conventional solutions to the exploration-exploitation dilemma inject explicit incentives such as randomization, uncertainty bonuses, or intrinsic rewards to encourage exploration. In this work, we hypothesize that an agent trained with reinforcement learning (RL) to maximize a purely greedy exploitation objective can exhibit exploratory behavior, provided three conditions are met: (1) Recurring Environmental Structure, where the environment features repeatable regularities that allow past experience to inform future choices; (2) Agent Memory, enabling the agent to retain and utilize historical interaction data; and (3) Long-Horizon Credit Assignment, where learning propagates returns over a time frame sufficient for the delayed benefits of exploration to impact current decisions. Through experiments in stochastic multi-armed bandits and temporally extended gridworlds, we verify that, when both structured tasks and agent memory are present, a policy trained on a strictly greedy objective exhibits information-seeking exploratory behavior. These findings suggest that, under the right prerequisites, exploration and exploitation need not be treated as orthogonal objectives but can emerge from a unified reward-maximization process. Graph-RHO: Critical-path-aware Heterogeneous Graph Network for Long-Horizon Flexible Job-Shop Scheduling Yujie Li (Beijing University of Posts and Telecommunications), Jiuniu Wang (City University of Hong Kong), Mugen Peng (Beijing University of Posts and Telecommunications), Guangzuo Li (Chinese Academy of Sciences), and Wenjia Xu (Beijing University of Posts and Telecommunications) Abstract Abstract Long-horizon Flexible Job-Shop Scheduling~(FJSP) presents a formidable combinatorial challenge due to complex, interdependent decisions spanning extended time horizons. While learning-based Rolling Horizon Optimization~(RHO) has emerged as a promising paradigm to accelerate solving by identifying and fixing invariant operations, its effectiveness is hindered by the structural complexity of FJSP. Existing methods often fail to capture intricate graph-structured dependencies and ignore the asymmetric costs of prediction errors, in which misclassifying critical-path operations is significantly more detrimental than misclassifying non-critical ones. Furthermore, dynamic shifts in predictive confidence during the rolling process make static pruning thresholds inadequate. To address these limitations, we propose Graph-RHO, a novel critical-path-aware graph-based RHO framework. First, we introduce a topology-aware heterogeneous graph network that encodes subproblems as operation-machine graphs with multi-relational edges, leveraging edge-feature-aware message passing to predict operation stability. Second, we incorporate a critical-path-aware mechanism that injects inductive biases during training to distinguish highly sensitive bottleneck operations from robust ones. Third, we devise an adaptive thresholding strategy that dynamically calibrates decision boundaries based on online uncertainty estimation to align model predictions with the solver's search space. Extensive experiments on standard benchmarks demonstrate that \mbox{Graph-RHO} establishes a new state of the art in solution quality and computational efficiency. Remarkably, it exhibits exceptional zero-shot generalization, reducing solve time by over 30\% on large-scale instances (2000 operations) while achieving superior solution quality. Our code is available at https://github.com/IntelliSensing/Graph-RHO. HistoVerify: A Token-Efficient Multi-Agent Framework for Historical Multimodal Misinformation Detection Yuanjian Zhao and Siyu Zheng (College of Computer Science,Sichuan University) Abstract Abstract Spanning centuries of human civilization, historical artifacts constitute irreplaceable records of cultural heritage. Despite the digitization of millions of such artifacts, verifying the authenticity of their associated metadata remains a laborious process, traditionally requiring domain experts to cross-reference extensive archival sources. The emergence of vision-language models presents a promising frontier for automated verification, yet these approaches face critical limitations when applied to historical content: excessive token consumption from information-dense imagery, and unconstrained reasoning that risks anachronistic hallucinations divorced from period-appropriate context. This paper introduces HistoVerify, a token-efficient multi-agent framework that addresses these challenges through attention-guided visual compression, claim-guided evidence extraction, and budget-adaptive reasoning with explicit constraints. To support this research direction, we construct HistoVerify-Bench, a benchmark comprising 6,000 samples across four historical domains with three verification categories. Extensive experiments demonstrate that HistoVerify achieves substantial token efficiency gains while maintaining competitive accuracy, charting a promising new course for AI-assisted historical artifact verification. Wednesday Virtual Room 7 IJCNN Paper LLM Evaluation and Benchmarking I Session Chair: Xiuwen Liu (Florida State University), Yufei Zeng (Beijing University of Posts and Telecommunications) TopoAtten: Probing the Internal Topology of Attention for Hallucination Detection in Large Language Models Yufei Zeng, Yukun Zhang, and Zhonghong Ou (Beijing University of Posts and Telecommunications) and Meina Song (Beijing University of Posts and Telecommunications, China University of Petroleum-Beijing at Karamay) Abstract Abstract Although Large Language Models (LLMs) show exceptional generation capabilities, their application is severely constrained by the challenge of ``hallucinations''. Existing white-box detection methods avoid the high costs of external retrieval by using internal states but are often confined to analyzing the static statistical properties of high-dimensional hidden states. Thus, they frequently fail when models generate high-confidence hallucinations. To address this limitation, we propose \textbf{TopoAtten}, a novel detection framework grounded in the topological structure of attention mechanisms, viewed through the lens of dynamical manifold evolution theory. Leveraging Topological Data Analysis (TDA), we first reconstruct the geometric skeleton of the attention space via a symmetrized exponential distance transform. Subsequently, we employ persistent homology to construct a topological persistence metric system comprising \textbf{semantic cohesion} ($\mathit{TP}_{\beta_0}$) and \textbf{logical self-consistency} ($\mathit{TP}_{\beta_1}$) to quantify the structural robustness of the attention mechanism. Finally, we design a lightweight, full-layer TopoAtten probe that precisely identifies hallucinations based on global topological features. Extensive experiments demonstrate that TopoAtten significantly outperforms existing methods across mainstream LLMs and multiple benchmarks, exhibiting superior generalization capabilities and unveiling the deep geometric mechanisms underlying hallucination generation. RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration Fabian Ridder, Laurin Lessel, and Malte Schilling (University of Münster) Abstract Abstract Retrieval-Augmented Generation (RAG) is widely used to augment the input to Large Language Models (LLMs) with external information, such as recent or domain-specific knowledge. Nonetheless, current models still produce closed-domain hallucinations and generate content that is unsupported by the retrieved context. Current detection approaches typically treat hallucination as a post-hoc problem, relying on black-box consistency checks or probes over frozen internal representations. In this work, we demonstrate that hallucination detection based on internal state representation can also serve as a direct training signal. We introduce RAGognize, a dataset of naturally occurring closed-domain hallucinations with token-level annotations, and RAGognizer, a hallucination-aware fine-tuning approach that integrates a lightweight detection head into an LLM, allowing for the joint optimization of language modeling and hallucination detection. This joint objective forces the model to improve the separability of its internal states regarding hallucinations while simultaneously learning to generate well-formed and meaningful responses. Across multiple benchmarks, RAGognizer achieves state-of-the-art token-level hallucination detection while substantially reducing hallucination rates during generation, without degrading language quality or relevance. Decoding Emotion in the Deep: A Systematic Study of How LLMs Represent, Retain, and Express Emotion Jingxiang Zhang and Lujia Zhong (University of Southern California) Abstract Abstract Large Language Models (LLMs) are increasingly expected to navigate the nuances of human emotion, yet their internal emotional mechanisms are still poorly understood. This paper investigates how, where, and for how long emotion is encoded in LLMs by introducing a novel, large-scale Reddit corpus of approximately 400,000 utterances, balanced across seven basic emotions through classification, rewriting, and synthetic generation. Using this dataset, we apply lightweight "probes" to analyze hidden states of Qwen3 and LLaMA models without modifying their parameters. Our findings show that LLMs develop a well-defined internal geometry of emotion that sharpens with model scale and significantly outperforms zero-shot prompting. Emotional signals emerge early in the network, peak in mid-layers, and remain detectable for hundreds of tokens, yet can be influenced through simple system prompts. We contribute our dataset, an open-source probing toolkit, and a detailed map of emotional representations, offering insights for developing more transparent and aligned AI systems. Numerical Instability and Chaos: Quantifying the Unpredictability of Large Language Models Chashi Mahiul Islam, Alan Villarreal, and Mao Nishino (Florida State University); Shaeke Salman (Florida State Unviersity); and Xiuwen Liu (Florida State University) Abstract Abstract As Large Language Models (LLMs) are increasingly integrated into agentic workflows, their unpredictability stemming from numerical instability has emerged as a critical reliability issue. While recent studies have demonstrated the significant downstream effects of these instabilities, the root causes and underlying mechanisms remain poorly understood. In this paper, we present a rigorous analysis of how unpredictability is rooted in the finite numerical precision of floating-point representations, tracking how rounding errors propagate, amplify, or dissipate through Transformer computation layers. Specifically, we identify a chaotic ``avalanche effect" in the early layers, where minor perturbations trigger binary outcomes: either rapid amplification or complete attenuation. Beyond specific error instances, we demonstrate that LLMs exhibit universal, scale-dependent chaotic behaviors characterized by three distinct regimes: 1) a stable regime, where perturbations fall below an input-dependent threshold and vanish, resulting in constant outputs; 2) a chaotic regime, where rounding errors dominate and drive output divergence; and 3) a signal-dominated regime, where true input variations override numerical noise. We validate these findings extensively across multiple datasets and model architectures. Wednesday Virtual Room 8 IJCNN Paper LLM Evaluation and Benchmarking II Session Chair: Boxun Li (North University of China), xueer wang (Xiangtan University) L-NCG: Narrative Consensus Graphs for Robust Stance Detection in Heterogeneous Social Networks Xueer Wang (Xiangtan University) Abstract Abstract Stance detection on social media is pivotal for understanding public opinion trends and mitigating misinformation. Existing approaches predominantly rely on GNNs to model social structures. However, these methods are often limited by the homophily assumption, indiscriminately aggregating features from heterogeneous interactions and introducing substantial structural noise. Conversely, while LLMs possess superior semantic reasoning capabilities, they struggle to capture global structural patterns due to computational constraints. To address these challenges, we propose the LLM-driven Narrative Consensus Graph (L-NCG) framework. This study posits that the essence of stance consistency lies in shared narrative frames rather than mere social connectivity. Specifically, we employ an LLM with a Probabilistic Chain-of-Thought mechanism to extract fine-grained narrative role distributions from unstructured text. These narrative features function as a semantic filter to reconstruct the graph topology: edges exhibiting high narrative consensus are reinforced, while those with conflicting logic or high uncertainty are suppressed via an entropy-based penalty mechanism. Finally, a GAT performs synergistic feature aggregation on this purified topology. Experimental results on multiple benchmark datasets demonstrate that L-NCG significantly outperforms state-of-the-art baselines, effectively mitigating structural noise and enhancing robustness in complex stance detection scenarios. Stance Scorer: Zero-shot Stance Detection on Social Media with Fine-grained Labels Qinlong Fan, Jicang Lu, Qiankun Pi, Yi Xia, and Yilin Liu (State Key Laboratory of Mathematical Engineering and Advanced Computing) Abstract Abstract Social media platforms serve as crucial arenas for the expression of public opinion, where user-generated content inherently conveys latent attitudes and stances toward specific entities and topics. Recent advancements in zero-shot stance detection have shown significant progress. However, most studies merely focus on the three-class stance, and the coarse-grained approach restricts both the depth of semantic comprehension and the optimization of inference strategies. To address this challenge, we introduce a novel Fine-grained Stance Scoring framework (Stance Scorer). Specifically, we first establish a quantifiable stance intensity metric through a meticulously designed fine-grained labeling system. Next, an LLM-powered data transformation pipeline is developed to systematically convert categorical labels into precise stance scores. Finally, we implement a model refinement strategy by fine-tuning the derived stance score metrics, with the optimized model serving as the stance scorer. Experimental results demonstrate that the proposed Stance Scorer outperforms GPT4, achieving average F1 score improvements of 6.03%, 3.91%, and 4.12% on the SemEval 2016 Task 6 A, P-Stance, and COVID-19 datasets, respectively. These results validate the effectiveness and robustness of our approach. FairGC: Fairness-aware Graph Condensation Yihan Gao, Chenxi Huang, and Wen Shi (Jilin University); Ke Sun (Dalian University); Ziqi Xu and Xikun Zhang (RMIT University); Mingliang Hou (Jinan University); and Renqiang Luo (Jilin University) Abstract Abstract Graph condensation (GC) has become a vital strategy for scaling Graph Neural Networks by compressing massive datasets into small, synthetic node sets. While current GC methods effectively maintain predictive accuracy, they are primarily designed for utility and often ignore fairness constraints. Because these techniques are bias-blind, they frequently capture and even amplify demographic disparities found in the original data. This leads to synthetic proxies that are unsuitable for sensitive applications like credit scoring or social recommendations. To solve this problem, we introduce FairGC, a unified framework that embeds fairness directly into the graph distillation process. Our approach consists of three key components. First, a Distribution-Preserving Condensation module synchronizes the joint distributions of labels and sensitive attributes to stop bias from spreading. Second, a Spectral Encoding module uses Laplacian eigen-decomposition to preserve essential global structural patterns. Finally, a Fairness-Enhanced Neural Architecture employs multi-domain fusion and a label-smoothing curriculum to produce equitable predictions. Rigorous evaluations on four real-world datasets, show that FairGC provides a superior balance between accuracy and fairness. Our results confirm that FairGC significantly reduces disparity in Statistical Parity and Equal Opportunity compared to existing state-of-the-art condensation models. CurParCL: A Curriculum-Guided Dual-Branch Graph–Text Contrastive Learning Framework for Parallelism Detection Boxun Li, Pinle Qin, Jianchao Zeng, and Yuanyuan Shen (North University of China) Abstract Abstract With the continued advancement of high-performance computing (HPC) platforms, automatic parallelization has emerged as a key enabler for high-performance software development. In the automatic-parallelization pipeline, parallelism detection is a key step in identifying candidate regions for parallel execution. Nevertheless, existing methods are limited in their ability to model structured program semantics, accurately identify complex dependencies, and generalize across datasets. To overcome these limitations, we propose CurParCL, a curriculum-guided dual-branch graph-text contrastive learning framework. Specifically, the framework introduces lightweight graph neural network to capture edge directionality and performs adaptive structural aggregation over data-flow graphs(DFGs). It further integrates difficulty-aware strategies into contrastive learning, improving label discriminability and training stability via adaptive scheduling. Moreover, we design a dual-branch graph–text contrastive mechanism that jointly encodes DFGs and source-code sequences. This design aligns sequential semantics with structural dependencies in a unified embedding space, enabling collaborative multimodal representation learning. Experiments show that CurParCL consistently outperforms baselines on multiple parallelism detection benchmarks and remains robust under cross-dataset evaluation, suggesting good practical applicability for automated parallelization. Wednesday Virtual Room 1 IJCNN Paper LLM Reasoning and Planning I Session Chair: Junhong Liang (MBZUAI), kaiyao Tan (sun yat-sen university) PERL: Pinyin Enhanced Rephrasing Language Model for Chinese ASR N-best Error Correction Junhong Liang (MBZUAI) and Bojun Zhang (Institute of Automation, CAS) Abstract Abstract Chinese ASR correction is challenging because errors are often \emph{phonetic} (many characters share similar Pinyin) while the correction model must also obey a \emph{length constraint} under noisy N-best hypotheses. Existing approaches either exploit Pinyin only at the prompt/feature level without integrating it into model representations, or rely on generative decoding that can drift in length. We propose \textbf{PERL}, a \textbf{constrained rephrasing pipeline} for Chinese N-best ASR correction that (i) predicts the target length and enforces it via mask budgeting, and (ii) fuses \emph{semantic} and \emph{phonetic} (Pinyin) representations through token-wise gates conditioned on sentence semantics. To improve robustness when more hypotheses introduce additional noise, we further introduce an \emph{adaptive hypothesis weighting} strategy that down-weights low-consensus hypotheses. Experiments on Aishell-1 and our new domain N-best benchmark \textbf{DoAD} show that PERL consistently reduces CER (29.11\% on Aishell-1 and up to $\sim$70\% on DoAD) while maintaining low latency. We also provide analyses of length generalization and phonetic--semantic interactions, showing when PERL relies on phonetic cues versus semantic constraints. Length-Aware Chinese Spelling Correction with Retrieval and Iterative Refinement Junhong Liang (MBZUAI) and Junnan Zhu, Yupu Liang, Bojun Zhang, Feifei Zhai, and Yu Zhou (Institute of Automation, CAS) Abstract Abstract Chinese Spelling Correction (CSC) traditionally assumes \emph{equal-length} character substitution and relies on pretrained language models. Recent Large Language Models (LLMs) improve correction fluency but often fail in \emph{domain-specific} settings (e.g., specialized terminology) and are unreliable under strict \emph{length constraints}. Moreover, real-world CSC frequently includes \emph{variable-length} phenomena such as ASR N-best correction and character splitting/merging, which are underexplored. We propose \textbf{RAIR} (\textbf{R}etrieval-\textbf{A}ugmented \textbf{I}terative \textbf{R}efinement), a model-agnostic framework that improves domain adaptation via (i) an adaptive multi-source retrieval corpus built from domain dictionaries and training sentences, (ii) a supervised fine-tuned retriever that is robust to misspellings and captures correction patterns, and (iii) \textbf{Multi-turn Length Reflection} (MLR) that iteratively refines LLM outputs to satisfy task-specific length constraints. Finally, an \textbf{Adaptive Selection} strategy switches between retrieval-augmented and direct generation. Experiments on domain CSC (ECSpell, LEMON) and variable-length settings (ChineseHP/Aishell-1, CSEC) show that RAIR consistently improves multiple LLM backbones and yields strong performance under both equal-length and variable-length scenarios. Reasoning to Find, Restraining to Fix: An Explainable Chinese Spell Checking Framework via Group Relative Policy Optimization Ruiqi Wang, Jindian Su, and Jiebin Huang (South China University of Technology) and Xiaobin Ye and Dandan Ma (Guangdong Unicomm) Abstract Abstract Chinese Spell Checking demands a delicate balance between detection sensitivity and correction faithfulness. Traditional discriminative models often lack the semantic reasoning to detect subtle errors, whereas generative Large Language Models (LLMs) suffer from severe over-correction. To bridge this gap, we propose $R^2$-CSC, a two-stage framework leveraging reasoning to find errors and restraining mechanisms to fix them. We first introduce a Recall-Oriented Supervised Fine-Tuning stage to enhance semantic sensitivity via Chain-of-Thought (CoT) instructions. Subsequently, we implement a Precision-Oriented Reinforcement Learning stage using Reference-Anchored Group Relative Policy Optimization. Unlike standard GRPO, RA-GRPO incorporates an identity baseline into advantage estimation. Coupled with Minimal Edit and Length Constraint rewards, this mechanism strictly penalizes unnecessary modifications by ensuring the model outperforms a do-nothing strategy. Extensive experiments demonstrate that $R^2$-CSC achieves state-of-the-art (SOTA) performance with a lightweight 4B parameter model. Notably, our method breaks the precision-recall trade-off, combining high semantic recall with BERT-level precision while offering interpretable reasoning paths. Latent Spectral Fourier Neural Operators kaiyao Tan (SUN YAT-SEN UNIVERSITY) and qingsong Zou (SUN YAT-SEN UNIVERSITY, Guangdong Province Key Laboratory of Computational Science) Abstract Abstract Fourier Neural Operators (FNOs) provide an effective backbone for PDE surrogate modeling, but their mode-wise spectral parameterization yields a parameter count that grows linearly with spectral resolution and typically requires frequency truncation, which degrades high-frequency accuracy. We propose the \textbf{Latent Spectral Fourier Neural Operator (LS-FNO)}, which preserves the global convolution structure of FNOs via the Fast Fourier Transform (FFT) while replacing discrete spectral weights with a \emph{continuous} frequency-conditioned generator. LS-FNO assumes spectral kernels lie on a compact \textbf{latent spectral manifold} and uses a \textbf{KAN}-based encoder--decoder to map frequency coordinates to a low-rank latent code and then to complex-valued kernels, promoting cross-frequency coherence and alleviating spectral bias. Across six benchmarks, LS-FNO achieves consistent improvements over strong baselines, particularly in high-frequency-dominated tasks (e.g., \textbf{42\%} lower relative error on 1D Compressible CFD) while reducing parameters by up to \textbf{77\%} compared to enlarged FNO variants. Wednesday Virtual Room 2 IJCNN Paper LLM Reasoning and Planning II Session Chair: Huang Shan (Tsinghua University), Jian Chen (Ningxia Research Institute of Transport Sciences) ARCE: Augmented RoBERTa with Contextualized Elucidations for NER in Automated Rule Checking Jian Chen (Ningxia Research Institute of Transport Sciences) and Jiabao Dou (Hong Kong Baptist University) Abstract Abstract Accurate information extraction from specialized texts is a critical challenge for automated rule checking (ARC) in the architecture, engineering, and construction (AEC) domain. While large language models (LLMs) possess strong reasoning capabilities, their deployment in resource-constrained AEC environments is often impractical. Conversely, standard efficient models struggle with the significant domain gap. Although this gap can be mitigated by pre-training on large, humancurated corpora, such approaches are labor-intensive and costly. To address this, we propose ARCE (Augmented RoBERTa with Contextualized Elucidations), a novel knowledge distillation framework that leverages LLMs to synthesize a task-oriented corpus, termed Cote, for incrementally pre-training smaller models. ARCE systematically explores the optimal strategy for knowledge transfer. Our extensive experiments demonstrate that ARCE establishes a new state-of-the-art on a benchmark AEC dataset, achieving a Macro-F1 score of 77.20% and outperforming both domain-specific baselines and fine-tuned LLMs. Crucially, our study reveals a less is more principle: simple, direct explanations prove significantly more effective for domain adaptation than complex, role-based rationales in the NER task, which tend to introduce semantic noise. ABEX-RAT: Synergizing Abstractive Augmentation and Adversarial Training for Classification of Occupational Accident Reports Jian Chen (Ningxia Research Institute of Transport Sciences) and Jiabao Dou (Hong Kong Baptist University) Abstract Abstract The automatic classification of occupational accident reports is pivotal for workplace safety analysis but is persistently hindered by severe class imbalance and data scarcity. In this paper, we propose ABEX-RAT, a resource-efficient framework that synergizes generative data augmentation with robust adversarial learning. Unlike computationally expensive large language models (LLMs) fine-tuning, our approach employs a two-stage abstractive-expansive (ABEX) pipeline: it first utilizes a prompt-guided LLM to distill label-critical semantics into concise abstracts, which are then expanded into diverse synthetic samples to balance the data distribution. Subsequently, we train a lightweight classifier using a random adversarial training (RAT) protocol, which stochastically injects perturbations to enhance generalization without significant computational overhead. Experimental results on the OSHA dataset demonstrate that ABEXRAT establishes a new state-of-the-art, achieving a Macro-F1 score of 90.32% and significantly outperforming both traditional baselines and fine-tuned large models. This confirms that targeted augmentation combined with robust training offers a superior, data-efficient alternative for specialized domain classification. Bridging the Domain Gap: Finding reliable Fine-Tuning Paths for GNNs Jiayi Que, Lian Shen, and Yue Hong (Xiamen University); Yuan Lin (Kristiania University College); and Juan Liu and Xiangrong Liu (Xiamen University) Abstract Abstract The pre-training and fine-tuning paradigms of Graph Neural Networks exhibit poor performance in out-of-distribution generalization. This limitation stems from negative transfer caused by distributional differences between source and target data. A promising solution is to sequentially fine-tune across auxiliary datasets, which act as “knowledge bridges” to mitigate the distribution shift. However, this strategy faces the challenge of finding the reliable auxiliary dataset sequence within a vast combinatorial space. To navigate this search problem, we develop Knowledge Bridge Tuning (KBT), which significantly reduces search complexity by using pre-computed gradients and a logistic regression approximation to efficiently estimate the effectiveness of different transfer paths, making the search for an reliable sequence computationally feasible. Extensive experiments demonstrate that our framework efficiently identifies transfer paths that achieve superior downstream performance, establishing new state-of-the-art results for both in-distribution and out-of-distribution generalization across multiple graph datasets. ViPFormer: Visual-Physics Reasoning for Safe Depalletizing via Segmentation-Guided Transformer Shan Huang (Tsinghua University); Yutian Zhang (Peking University); and Man Him Cheng, Guanqiu Guo, Lirong Che, Tian Gao, Junbo Tan, and Xueqian Wang (Tsinghua University) Abstract Abstract Automated depalletizing in unstructured warehousing environments faces a critical safety challenge: removing a key load-bearing box can trigger catastrophic stack collapse. Existing approaches primarily rely on geometric perception or static stability analysis, remaining "physics-blind" to the latent dynamic consequences of robotic interactions. To bridge this gap, we propose ViPFormer (Visual-Physics Reasoning Transformer), a novel end-to-end reasoning model that transitions depalletizing from passive perception to predictive physical reasoning. ViPFormer explicitly localizes potential instability risks and predicts the safety of the remaining stack after object removal, utilizing a segmentation-guided mechanism for precise spatial reasoning. To circumvent the high cost and risks of collecting real-world data with physical consequence labels, we introduce SafeDepal-200k, a large-scale synthetic benchmark featuring diverse stacking configurations and automatic counterfactual annotations. Extensive experiments demonstrate that ViPFormer learns generalized physical intuition, achieving an 89.58% F1-score on the simulation benchmark. Despite the domain gap, our method achieves a Sim-to-Real transfer with a 77.62% F1-score on a physical robot without fine-tuning, proving effective as a safety shield for autonomous logistics. Code and dataset are publicly available at https://github.com/YellowThree-HS/ViPFormer-Depalletizing. Wednesday Virtual Room 3 IJCNN Paper Learning Paradigms and Model Efficiency I Session Chair: Chin-Teng Lin (University of Technology Sydney), Huaijun Guang (Dalian Minzu University) Multiscale Spatiotemporal Ensemble Learning for Imbalanced Chiller Fault Diagnosis Using Generative Adversarial Networks Yuzhe Wang, Huaijun Guang, Fangbo Fu, Jinglong Liu, Ran Li, and Xiaodong Duan (Dalian Minzu University) Abstract Abstract Imbalanced class distributions and spatiotemporally coupled fault features in chiller operational data hinder reliable fault diagnosis. To address these challenges, this paper proposes a multi-scale spatiotemporal ensemble learning framework for imbalanced chiller fault diagnosis that integrates generative data augmentation with deep ensemble learning. Beyond Distribution Matching: Contrastive-Enhanced Diffusion for Imbalanced Tabular Data Augmentation Yuehang Ma, Siying Li, Linna Wang, and Li Lu (Sichuan University) Abstract Abstract Class imbalance in tabular data is a fundamental problem in real-world machine learning scenarios, often leading to suboptimal performance and generalization of downstream models on minority class. While existing data augmentation techniques partially alleviate sample scarcity, most fail to explicitly characterize intra-class consistency and inter-class discrepancy in data synthetic modeling. Consequently, ensuring that synthesized minority samples possess genuine discriminative power remains a formidable challenge. Concurrently, contrastive learning has been extensively validated in the vision and language domains for its ability to align positive pairs while segregating negative ones. Therefore, we propose CDTab, a contrastive-enhanced diffusion framework that models the feature distribution of minority samples while simultaneously capturing discriminative structures. Our method not only enhances the fidelity of generated data but also bolsters the semantic distinguishability of minority samples. Systematic experiments across six benchmark datasets using five mainstream classifiers demonstrate that our framework significantly outperforms competitive augmentation baselines in both generation quality and downstream predictive performance. Improving Time Series Generation and Self-Supervised Learning via Instance-Conditioned Contrasting Philipp Engler (Deutsches Forschungszentrum für Künstliche Intelligenz GmbH, RPTU Kaiserslautern-Landau); Nico Müller (RPTU Kaiserslautern-Landau); Ludger van Elst and Sheraz Ahmed (Deutsches Forschungszentrum für Künstliche Intelligenz GmbH); and Andreas Dengel (Deutsches Forschungszentrum für Künstliche Intelligenz GmbH, RPTU Kaiserslautern-Landau) Abstract Abstract Data sparsity is a major limiting factor for machine learning applications in industry. In order to make such applications feasible and also cost-effective, data synthesis strategies and self-supervised approaches can be employed, addressing the issue in different ways. We aim to advance both directions by combining generative approaches for time series synthesis and self-supervised learning techniques, yielding benefits for both. In this paper, we propose a self-supervised contrastive module for instance-conditioned generative adversarial networks (IC-GAN) and apply it to multiple time series datasets. This module improves the quality and diversity of generated time series without the need for data annotations. Simultaneously, the contrastive module can act as a self-supervised pre-training method of its own. Generated samples can also be used as augmentation in self-supervised techniques by injecting noise into the condition. We achieve significant improvements over the standalone IC-GAN, increasing the classification accuracy of a classifier trained on synthetic data by 5.5 percentage points on average across the 22 shortest time series datasets from the UCR classification archive. Even being trained without data annotations, our GAN is competitive with recent class-conditional time series generation models and generates samples orders of magnitudes faster than a diffusion model at comparable sample quality. Online Ensemble Learning Framework for Class Evolution under Incomplete Supervision Zhi Cao (Anhui Transport Consulting & Design Institute Co., Ltd; University of Science and Technology of China); Shuyi Zhang (Southern University of Science and Technology); Chin-Teng Lin (University of Technology Sydney); and Xin Yao (Lingnan University) Abstract Abstract The composition of classes in real-world data streams is constantly evolving. To maintain the classifier performance, labeling a limited number of instances for a novel class is more affordable and feasible than annotating the entire data steam. This poses a challenge for learning with class evolution under incomplete supervision. However, existing semi-supervised algorithms typically require storing past data to adapt classifiers to novel classes. Despite many online learning algorithms are proposed to tackle class evolution without retaining historical data, these methods are supervised.In this paper, we propose a novel Semi-supervised Online Ensemble Learning Framework (SOELF) to handle class evolution under incomplete supervision. This framework adapts in an online learning manner without the storage of past data. It integrates a data generation module based on an online ensemble clustering model to generate diverse instances to exploit the unlabeled instances for model adaptation. The decaying G-mean metric is employed to determine the pseudo-labels of these generated instances. Furthermore, an inverse-prior-ratio model adaptation strategy is designed to introduce more instances from existing classes for model adaptation when novel classes emerge with labeled instances, thereby balancing the exploitation between classes. Experimental studies on both synthetic and real-world data streams demonstrate that our method achieves higher accuracy compared to existing online learning algorithms and the state-of-the-art semi-supervised algorithm that requires data storage. Wednesday Virtual Room 4 IJCNN Paper Learning Paradigms and Model Efficiency II Session Chair: Shoji Toyota (Kyushu University), Zitai Kong (Zhejiang university) Self-Organizing Score-based Data Assimilation Yuma Yamaoka, Seiichi Uchida, and Shoji Toyota (Kyushu University) Abstract Abstract A state-space model is a statistical framework for inferring latent states from observed time-series data. However, inference with nonlinear and high-dimensional state-space models remains challenging. To this end, an approach based on diffusion models—a powerful class of deep generative models—has been developed, known as Score-based Data Assimilation (SDA). However, SDA cannot be directly applied when the latent-state transition depends on unknown parameters that must be inferred jointly with the latent states. To overcome this limitation, we propose a framework that enables SDA to handle latent states with unknown parameters. A key feature of the proposed method is the incorporation of the self-organization technique, which has been used in classical state-space modeling for the joint estimation of latent states and parameters. By integrating this classical technique into modern SDA, our method enables joint inference of latent states and unknown parameters while maintaining the high training efficiency of SDA. The effectiveness of the proposed approach is validated through numerical experiments on dynamical systems arising in neuroscience and atmospheric science. In addition, its scalability is demonstrated using a high-dimensional Kolmogorov flow, with the data dimension on the order of several hundred thousand. FDLC: Fast Decoupled Lossless Compression with State Space Models Pei Wu, Xueming Fu, and S.Kevin Zhou (University of Science and Technology of China) Abstract Abstract Abstract—Lossless image compression is essential for scientific and medical applications where preserving data accuracy is critical. Recent deep learning methods employ a decoupled framework, combining a lossy compressor with a residual compressor, to achieve high lossless compression rates while enabling efficient lossy reconstruction for practical applications. However, existing methods rely on computationally expensive Transformer components, hindering practical deployment. To address this limitation, we propose Fast Decoupled Lossless Compression (FDLC), the first architecture based on State Space Models that leverages their linear computational complexity. FDLC integrates Visual State Space blocks as the core component for effectively capturing global dependencies, generating compact latent representations, and accurately modeling the probability distribution of residuals. Experiments demonstrate that our model achieves a new state-of-the-art compression effectiveness on diverse datasets. Moreover, it achieves average speedups of 1.3× for compression and 1.6× for decompression compared to Transformer-based counterparts, demonstrating superior efficiency and effectiveness. Stochastic Adaptive Process State Space Model for Portfolio Management Based on Deep Reinforcement Learning Fengchen Gu, Xiaotian Ren, and Zhengyong Jiang (Xi’an Jiaotong-Liverpool University); Ángel F. García Fernández (Universidad Politécnica de Madrid); and Jionglong Su and Huakang Li (Xi’an Jiaotong-Liverpool University) Abstract Abstract Deep reinforcement learning (DRL) for financial portfolio management is challenged by the market's inherent volatility and non-stationary dynamics, which are often inadequately captured by models that treat financial data as generic sequences. To address this, we introduce a novel DRL framework, the Stochastic Adaptive Process State Space Model (SAP-SSM), that provides the crucial financial inductive bias lacking in prior work. The SAP-SSM explicitly models financial time series as a Hawkes jump-diffusion process, natively capturing volatility clustering and fat-tailed shocks. In contrast to prevailing models with static, deterministic mechanisms, its internal architecture features a dynamic state transition matrix that adapts to market regimes, a hierarchical state to differentiate between noise and trends, and stochastic transitions to represent market uncertainty. Evaluated on a portfolio of 30 Dow Jones Industrial Average stocks, our framework demonstrates competitive performance, achieving a cumulative return of 45.3% and a Sharpe ratio of 1.878. Generating Plausible Sequences of Protein Groups with Latent Optimal Transport Flow Matching Zitai Kong, Yinlong Xu, Xiaohong Jiang, Jian Wu, and Hongxia Xu (Zhejiang university) Abstract Abstract The de novo design of protein sequences with targeted functionalities is a significant task in bioengineering. Deep generative methods such as autoregressive models and diffusion models have promoted the rapid discovery of novel protein sequences. However, these algorithms mainly suffer from expensive computing costs, low inference efficiency, and underused global protein information on both sequences and space. To overcome these issues, we introduce PLFM, an optimal transport conditional flow matching-based protein sequence designer operating on latent embeddings of protein language models. Additionally, we develop a joint design pipeline for the design scene of multichain proteins. We apply PLFM to various protein design tasks, including representative protein types: general peptides, general proteins, antimicrobial peptides, and antibodies. Taking advantage of advanced flow matching and pLM embeddings, PLFM surpasses task-specific methods in these targeted applications, highlighting its significant potential and broad applicability in computational protein design and analysis. Wednesday Virtual Room 5 IJCNN Paper Learning Paradigms and Model Efficiency III Session Chair: Fang Wu (Kashgar University), Wei Wang (Yunnan University) Adaptive Weight Optimization for Ship Detection Fang Wu and Hui Peng (Kashgar University) Abstract Abstract —Ship detection is crucial for maintaining maritime sovereignty and monitoring ocean pollution. However, deploy ing this technology in complex ma rine environments presents significant challenges, especially when detecting small vessels. Their subtle features, multi-scale variations, and the interference of complex backgrounds often result in poor target localization and classification accuracy in existing models. To address these issues, this paper presents a ship detection model based on multi modal fusion. The model leverages pre-trained parameters from public datasets to extract features, enhances target identification through a cross-modal synergy mechanism, and introduces an uncertainty loss function to dynamically adjust loss weights, significantly improving detection accuracy across different ship sizes and complex backgrounds. Experimental re sults on the Levir-Ship dataset, which includes optical remote sensing images, demonstrate the model’s effectiveness with AP, AP50, AP75, and AR, scores of 33.7%, 84.8%, 16.1%, and 45.4%, respectively. These results validate the model’s superiority in ship detection, offering strong technical support for mari time surveillance and pollution monitoring, and paving the way for future ad vancements in marine monitoring technologies. An Improved Pest and Disease Detection Algorithm Based on YOLOv10 Binghui Liu (Xinjiang Teacher’s College) Abstract Abstract Existing cotton pest and disease detection algorithms face challenges in adapting to variations in target sizes and complex scenarios, with an inherent trade-off between detection accuracy and speed during the algorithm optimization process. To overcome these limitations, we propose MADN-YOLOv10, an enhanced algorithm based on YOLOv10 that introduces three key innovations: the Multi-Scale Dilated Fusion Attention (MDFA) module strengthens the feature representation of key pest and disease characteristics while improving adaptability to dynamic and complex environments, thereby enhancing detection precision; the DualConv module reduces the model’s complexity and shortens the detection time, making the model better suited for crop field pest and disease detection scenarios, and even allowing it to perform optimally on embedded systems with limited hardware resources; the MPDIoU function can mitigate bounding box distortion caused by the significant variability of pest and disease samples and improve the model’s robustness during detection. When evaluated on the CottonInsect dataset, our MADN model achieved remarkable results, with a mean Average Precision (mAP) of 95.9% (a 2.3% improvement) and an accuracy 2.0% higher than that of mainstream alternative al gorithms. Additionally, we compressed the parameters to 4.0MB and reduced the inference time to 0.9 milliseconds, representing a 30% speed increase compared to the original model. Intraday Price-Movement Prediction Using Technical Indicators with Feature Selection and Genetic Algorithm-Based Hyperparameter Optimization Vicenzo Copetti, Richard Pinto, and Bruno Dalmazo (Federal University of Rio Grande); Giancarlo Lucca (Federal University of Pelotas); Diego Bruno (São Paulo State University); Eduardo Borges (Federal University of Rio Grande); Fabian Cardoso (University of Rio Verde); and Viviane de Mattos and Rafael Berri (Federal University of Rio Grande) Abstract Abstract Intraday financial-market modeling remains a challenging task. In this context, technical indicators derived from price, volume, volatility, and trend have long been employed as proxies to capture different market behaviors and trading pressure. However, the large number of available technical indicators often leads to high-dimensional feature spaces, increasing redundancy and negatively impacting model generalization. This work proposes an end-to-end machine learning framework for binary next-interval price-movement prediction using intraday data, with a particular emphasis on the systematic selection of relevant technical indicators. Using the BovDBv2 dataset with 5-minute bars, a total of 61 indicators are computed and normalized, encompassing volume-based, volatility-based, and trend-based information. A Sequential Forward Selection (SFS) procedure, guided by a balanced Random Forest classifier and the binary F1-score, is applied to identify a compact and informative subset of features. And then, a Genetic Algorithm (GA) with binary encoding is employed to optimize the hyperparameters of two classification models: Random Forest (RF) and Multi-Layer Perceptron (MLP). Experimental results indicate that the GA-optimized MLP achieves an accuracy of 0.6269 and a F1-score of 0.6242, outperforming the GA-optimized RF, which attains an accuracy of 0.4786 and a F1-score of 0.4331. Overall, the findings suggest that combining feature selection over technical indicators with GA-based hyperparameter tuning improves generalization performance in intraday price movement prediction tasks. Enhancing Traffic Flow Prediction via Adaptive Spatial-Temporal Encoding and Attention Refinement Mingchao Zhang and Wei Wang (Yunnan University) Abstract Abstract Traffic flow prediction (TFP) is a critical component of modern traffic management systems, yet it remains challenging due to complex spatial-temporal dependencies. Recent studies have attempted to address this challenge by leveraging various advanced neural networks, such as graph neural networks and transformers. While these models may achieve high accuracy, they remain susceptible to multiple challenges, including the cumulative effects of individually minor attention weights. Furthermore, existing research has primarily focused on developing intricate TFP architectures, overlooking the importance of traffic observation encoding. To address these issues, we propose PARformer, a novel approach that introduces an adaptive spatial-temporal embedding module to capture evolving dependency patterns. Additionally, PARformer leverages controlled perturbations to encourage attention mechanism to actively calibrate the weight distribution, aiming to mitigate the adverse effects of minor attention weights. We validate the effectiveness of PARformer through comparative experiments on four real-world datasets collected from the California Department of Transportation. The results indicate that: (1) PARformer outperforms several recently proposed baselines in predictive accuracy. (2) The proposed perturbation mechanism enhances the calibration of attention weight distribution. Wednesday Virtual Room 6 IJCNN Paper Machine Learning Methods and Applications I Session Chair: Wei Guo (Shenyang Aerospace University, School of Computer Science), Guangchao Yang (Chongqing University) Dose-Aware Cold Diffusion with Physics Consistency for Generalizable Low-Dose CT Reconstruction Md Imam Ahasan (Chongqing University, Daffodil International University); Guangchao Yang (Chongqing University); A. F. M. Abdun Noor and S. M. Hasan Mahmud (Daffodil International University); and Md Mahfuzur Rahman (Chongqing University) Abstract Abstract Reducing radiation dose in computed tomography significantly degrades image quality and poses challenges for accurate and clinically reliable reconstruction. While recent approaches have shown promise for low-dose CT, they often struggle to generalize across continuous and previously unseen dose levels, leading to artifacts and loss of anatomical detail. To address these limitations, we propose Dose-Aware Cold Diffusion (DACD), a physics-consistent reconstruction framework that explicitly models radiation dose as a continuous latent factor within a cold diffusion process. The proposed DACD framework integrates image-based dose-aware perception, multi-scale structural prior extraction, and dose-calibrated step allocation to adaptively guide the denoising trajectory. In addition, an iterative forward-backprojection correction is incorporated into the reverse refinement process to enforce projection-domain data consistency. Extensive experiments on three public benchmarks, including Mayo-2020, Mayo-2016, and LoDoPaB-CT, demonstrate that DACD consistently outperforms state-of-the-art diffusion-based and physics-guided methods in both quantitative accuracy and visual fidelity, particularly under ultra-low-dose conditions. The results show that DACD achieves robust generalization across a continuous range of dose levels, including those unseen during training. Incorporating Dose Level Prior and Texture-Sensitive Contrastive Learning from Frequency Perspective for Multi-dose PET Reconstruction Yuchen Fei, Zhenghao Feng, Jiliu Zhou, and Yan Wang (Sichuan University) Abstract Abstract Due to the inherent radiation exposure of positron emission tomography (PET), reconstructing standard-dose PET (SPET) from low-dose PET (LPET) is becoming an attractive alternative to ensure both imaging quality and patient safety. In clinic, the noise levels of LPET images may vary significantly owing to the individual differences among patients, making it challenging to design a universal model for PET reconstruction across multiple low-dose levels. Nevertheless, most existing multi-dose-level PET image reconstruction methods incorporated limited dose-level information and overlooked the potential of frequency domain prior. In this paper, we demonstrate that the frequency domain possesses the capability to differentiate between SPET and LPET. Harnessing this crucial insight, we subtly design a high-frequency component classification network to embed dose prior into the network. Furthermore, considering the critical importance of texture details in clinical applications, we adopt a frequency contrastive learning paradigm that utilizes a hard negative sample strategy to enforce the reconstructed PET (RPET) to preserve more texture details. Meanwhile, we elaborately construct texture-sensitive embeddings that can perceive subtle differences for contrastive learning, enforcing more realistic and robust reconstruction. Extensive experiments conducted on the MICCAI 2022 UDPET dataset and the Phantom Brain Dataset have demonstrated the effectiveness of our proposed method. Joint Prediction with Dummy Node for Abdominopelvic Lymph Node Metastasis Xingyu Zou, Yiji Mao, Yuling Zheng, Xiao Yu, Yuxuan Ji, Mingxuan Tian, and Haixian Zhang (Sichuan University) Abstract Abstract Abdominopelvic lymph node (LN) metastasis prediction using preoperative computed tomography (CT) is a critical task for personalized disease progression assessment and treatment planning for colorectal cancer (CRC) patients. However, there are notable challenges in aligning CT images with pathology data. Moreover, manual annotation of LNs has inherent limitations, such as annotation omission, leading to low accuracy in prediction models. To address these issues, we propose a joint prediction framework guided by node-level and patient-level tasks, which leverages patient-level pathological labels to predict LN metastasis in CT scans. Key innovations in framework include: (1) Pretraining encoder using node discrimination from nine views, (2) Predicton with a dummy node, which simulates unlabeled LN to bridge the gap in annotation omission, (3) Intra-class contrastive learning based on neural memory ordinary differential equations (nmODE), which distinguishes subtle and continuous features from dynamic evolution, and (4) Prediction with inter-class contrastive learning based on gated attention for better nodes integration within patient. These techniques allow the model to make reliable joint prediction in such a difficult task and even in imperfectly labeled data. We validate the proposed framework on a private dataset of 517 CRC patients and a public chest CT dataset of 1,595 scans, showing superior performance. Additionally, we provide visualized explanations of the embedding expression. CACR-Net: Constraint-Aware Occlusion-Guided Dental Crown Reconstruction with Latent SDF Diffusion Wei Guo, Shilin Chen, and Chen Yu (Shenyang Aerospace University) and Ni An and Dan Meng (Capital Medical University) Abstract Abstract Computer-aided design (CAD) and computer-aided manufacturing (CAM) of dental crowns often rely on library templates and manual adjustments, while existing deep-learning methods still struggle to reconstruct anatomically accurate occlusal surfaces from intraoral scans. We propose a constraint-aware dental crown reconstruction network (CACR-Net), an occlusion-guided two-stage framework that integrates geometric constraints with diffusion modeling for anatomically accurate and functionally plausible reconstruction. In Stage-1, the Curvature-Guided Mamba Dental Crown Network (CMDen-Net) extracts multi-resolution geometric features via a Serialized-Mamba Network (SMN) to predict an initial crown point cloud, supervised by a curvature-guided occlusal constraint (CGOC) and an SDF-based non-penetration loss. In Stage-2, SDFDiff-Net performs latent-space diffusion conditioned on the stage-1 output and antagonist teeth. It then decodes the denoised representation into a continuous SDF, from which Marching Cubes (MC) extracts the zero-level set to form a closed crown mesh. Evaluation on the public Teeth3DS+ dataset shows average reconstruction errors of 0.74~mm (CD-L1) and 0.59~mm (CD-L2), with higher occlusal accuracy and geometric consistency than existing deep-learning methods. CACR-Net demonstrates potential for integration into CAD workflows and for reducing the need for manual post-editing. Wednesday Virtual Room 7 IJCNN Paper Machine Learning Methods and Applications II Session Chair: Handong Yao (University of Georgia), Ashitabh Misra (University of Illinois at Urbana Champaign) MotiMem: Motion-Aware Approximate Memory for Energy-Efficient Neural Perception in Autonomous Vehicles Haohua Que (University of Georgia; Infinity Exploration Robotics Technology Co., Ltd. (Infinity Robotics)); Mingkai Liu (Perking University); Jiayue Xie (Beijing Forestry University; Infinity Exploration Robotics Technology Co., Ltd. (Infinity Robotics)); Haojia Gao (Tsinghua University); Jiajun Sun (Shenzhen University); Hongyi Xu (Central Saint Martins, University of the Arts London; Infinity Exploration Robotics Technology Co., Ltd. (Infinity Robotics)); Handong Yao (University of Georgia); and Fei Qiao (Tsinghua University) Abstract Abstract High-resolution sensors are critical for robust au- tonomous perception but impose a severe ”memory wall” on battery-constrained electric vehicles. In these systems, data movement energy often outweighs computation. Traditional im- age compression is ill-suited as it is semantically blind and optimizes for storage rather than bus switching activity. We propose MotiMem, a hardware-software co-designed interface. Exploiting temporal coherence, MotiMem uses lightweight 2D Motion Propagation to dynamically identify Regions of Interest (RoI). Complementing this, a Hybrid Sparsity-Aware Coding scheme leverages adaptive inversion and truncation to induce bit- level sparsity. Extensive experiments across nuScenes, Waymo, and KITTI with 16 detection models demonstrate that MotiMem reduces memory-interface dynamic energy by ≈ 43% while retaining ≈ 93% of the object detection accuracy, establishing a new Pareto frontier significantly superior to standard codecs like JPEG and WebP. FeatureFence: A Regularization Approach for Energy-Efficient Secure Inference on Edge NPUs Sachintha Kavishan Jayarathne and Seetal Potluri (University at Albany, SUNY) Abstract Abstract Feature-snooping attacks (FSA) are very powerful for reverse engineering machine learning models running on neural processing units (NPUs). While memory encryption is an effective countermeasure for cloud devices, the increased data movement causes significant overheads, making it inefficient for edge devices. We make a crucial observation that features dominate the off chip memory accesses in edge NPUs and propose FeatureFence, which eliminates and compensates for feature encryption via a regularization approach to protect against FSA during inference. Our approach creates neuron pairs in the first layer called couples, and equates weights and biases of neurons within each couple, thereby making reverse engineering mathematically impossible beyond the first layer. During FeatureFence training, the nature of perturbations is gradually learnt across epochs, leading to graceful recovery of functional accuracy. When implemented across a wide range of neural network models mapped to the Eyeriss architecture, on average, FeatureFence is able to reduce energy overheads by ≈ 41% when compared to GuardNN. Spiking Neural Network-Based Radar Human Activity Recognition Haotian Zhang and Yu Zhou (Beijing Jiaotong University), Xuyang Zheng (Lancaster University), Fei Luo (Great Bay University), Tianwei Hou (Beijing Jiaotong University), and Anna Li (Lancaster.ac.uk) Abstract Abstract Radar-based human activity recognition (HAR) has become an important technique in a wide range of applications, from smart homes to healthcare monitoring. Compared with vision-based sensing, radar is privacy-preserving and robust to adverse lighting conditions, making it well suited for indoor environments. However, achieving both high accuracy and energy efficiency remains challenging due to the complexity of radar signals and the subtle variability of human motions. In this paper, we propose spiking-transformer-TopoLoss (STT), a topology-regularized spiking transformer model for micro-Doppler-based HAR. STT integrates block-wise spatiotemporal attention with a TopoLoss constraint to enhance both representation capacity and robustness. Specifically, the STAtten module captures long-range spatiotemporal dependencies, while Leaky Integrate-and-Fire neurons maintain event-driven sparsity for efficient computation. Meanwhile, TopoLoss encourages spatially organized intermediate representations, helping suppress spurious activations and preserve the continuity and periodicity of micro-Doppler signatures. Experimental results on a nine-class micro-Doppler kitchen dataset demonstrate that STT achieves 98.99% overall accuracy. Evaluations on public datasets further show that STT remains competitive in both accuracy and computational efficiency, indicating its potential for robust and scalable radar-based HAR in real-world applications. Adaptive Quantization of CNNs for Spectrogram-Based Foreground Activity Classification Ashitabh Misra, Madhav Agrawal, Tomoyoshi Kimura, Jinyang Li, Avaljot Singh, Arham Jain, and Tarek Abdelzaher (University of Illinois at Urbana-Champaign) Abstract Abstract Adaptive Quantization (AQ) enables mixed-precision computation in convolutional neural networks (CNNs) without retraining, reducing model size and inference cost for resource-constrained deployments. Existing AQ methods, however, typically assign a single bit-width to the entire input. While adequate for many vision tasks with approximate translation invariance, this assumption is ill-suited for spectrogram-based representations, where different frequency regions carry distinct semantic meaning. A single bit-width for the entire input can therefore misallocate precision, over-allocating bits for noise-dominant frequencies and under-allocating them for informative foreground signals. We propose FreqQuant, an adaptive quantization framework for CNNs operating on spectrogram inputs for foreground activity classification. FreqQuant incorporates input signal statistics into the quantization optimization to derive a preference metric that assigns frequency-band-specific bit-widths on a per-layer basis. This enables precision to be allocated where it most impacts performance without increasing computational cost. Experiments on vehicle and human activity classification demonstrate consistent improvements over state-of-the-art AQ methods across multiple CNN architectures under comparable compute budgets. Wednesday Virtual Room 8 IJCNN Paper Medical Image Analysis I Session Chair: Wu chaolin (JiNan University), Dmitrii Kaplun (Saint Petersburg Electrotechnical University "LETI") KGS-UNet: KAN-Gated Skip Connections for Small-Structure Medical Image Segmentation with Limited Data Wu Chao lin and Long Shun (JiNan University) Abstract Abstract Accurate segmentation of small anatomical structures remains challenging for U-shaped networks because skip connections often fuse encoder and decoder features without explicit selection. We propose KGS-UNet, which repurposes Kolmogorov–Arnold Networks (KANs) as learnable gates at skip interfaces. The KAN-Gated Skip (KGS) module performs pixelwise routing using spline-parameterized gates with noise-aware lower bounds to suppress inconsistent encoder activations while preserving decoding cues. In addition, the AttnDown module integrates KAN-driven channel attention and spatial gating to better retain thin structures during downsampling.Benchmarked against over 70 mainstream U-Net variants under a uniffed training framework, KGS-UNet secures the top rank on the challenging, small-scale DRIVE and CHASEDB1 datasets. Specifically, it surpasses competing mainstream models by margins of 1.5% in both IoU and F1 scores. Furthermore, compared with existing KAN-based U-Net variants that primarily utilize KANs as activation replacements, KGS-UNet achieves a substantial improvement of over 6 percentage points in IoU on DRIVE. These results suggest that spline-based gating at skip interfaces is an effective design choice for medical segmentation. Anonymous repository for reproducibility:https://anonymous.4open.science/r/KGS-UNet-C0EE SDG-UNet: Medical Image Segmentation Based on Structure-Guided Dynamic Convolution and Dual-Path Hybrid Attention Jiajian Hao, Jiaqing Mo, Gang Zhou, and Zifang Zhao (Xinjiang University) Abstract Abstract Medical image segmentation requires both global contextual reasoning and precise boundary preservation, while shallow skip features often introduce semantic noise. To address this, we propose SDG-UNet, which combines a dual-path context encoder for global-local representation learning with a structure-guided dynamic filtering module for skip calibration. We introduce a Dual-Path Context Encoder (DPCE), which integrates a Global-Cross Collaborative Module (GCCM) to capture long-range dependencies and an Edge-aware Local Activation (ELA) module to reinforce fine-grained boundaries, thereby jointly modeling multi-scale global semantics and fine-grained boundary information. Furthermore, we propose a Structure-Guided Hybrid Dynamic Filtering (SG-HDF) module in the decoder. This module utilizes high-level structural priors to generate adaptive kernels that calibrate shallow skip features, effectively suppressing noise and aligning semantics before fusion. We evaluated the proposed method on three different medical image segmentation benchmarks. Experiments on three medical segmentation benchmarks show that SDG-UNet achieves competitive or superior performance with a favorable accuracy-efficiency trade-off. CALDGSEG: A Medical Image Segmentation Framework via Interactive Feature Fusion and Dynamic Guidance Zihan Xiong and Siyan Xiao (Northeastern University, School of Computer Science and Engineering) Abstract Abstract Accurate segmentation remains challenging due to ambiguous boundaries and background interference, where foreground and background features are highly entangled. In medical image segmentation, this issue is further aggravated by soft boundaries and noise, while existing models typically address these problems in isolation and rely on fixed feature fusion strategies, resulting in limited robustness. We propose CALDGSeg, a neural network framework that explicitly restructures feature interaction and fusion through feature disentanglement, contrast-aware local attention, and dynamic guidance. A Foreground–Background Attention Disentangler first decouples foreground and background representations, reducing early-stage semantic interference. Subsequently, a Contrast-Aware Local Attention module performs foreground–background interactive weighting within windowed self-attention, enhancing discriminative boundary representation while suppressing background noise. Finally, a dynamically guided decoder learns per-pixel adaptive fusion weights with cascaded channel–spatial optimization to alleviate semantic conflicts during multi-scale decoding. Experiments on three public benchmarks demonstrate that CALDGSeg consistently outperforms state-of-the-art methods and exhibits strong generalization. This framework offers generalizable design insights for robust feature interaction and fusion in dense prediction networks. The code is available at https://github.com/zihan1215/CALDG SAGED-Net: Structural Adaptive Gated Encoder-Decoder Network for Nuclei Segmentation in Histopathology Images Arko Dasgupta and Arjeesh Palai (Jadavpur University), Alexander Voznesensky (Saint Petersburg Electrotechnical University "LETI"), Dmitrii Kaplun (Saint Petersburg Electrotechnical University), and Ram Sarkar (Jadavpur University) Abstract Abstract Standard deep learning models for nuclei segmentation in histopathology often struggle with overlapping boundaries while remaining computationally heavy and sensitive to noise. A key limitation in many existing architectures, such as U-Net, is their reliance on static, hard-wired skip connections that indiscriminately merge low-level noise with semantic features. To overcome this, we propose SAGED-Net, a lightweight and efficient architecture that replaces rigid connectivity with a learnable Adaptive Gating mechanism to actively suppress background artifacts. Our framework also introduces a Tri-Domain loss function to enforce anatomical accuracy in cell shapes and boundaries. Evaluated on various histopathological benchmarks, SAGED-Net consistently outperformed state-of-the-art methods, achieving notable segmentation accuracy and precise instance delineation on TNBC (Dice score: 0.9103), CPM-15 (Dice score: 0.9286), CPM-17 (Dice score: 0.9057), and PanNuke (Dice score: 0.9058), validating its utility as a robust method for computational pathology. Further details are available at: https://github.com/saged-net/SAGED-Net Wednesday Virtual Room 1 IJCNN Paper Model Compression and Quantization I Session Chair: DongChen Zhu (Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences), d Liu (Institute of Information Engineering, Chinese Academy of Sciences ; School of Cyber Security, UCAS) Error-Decoupled Learning for Knowledge Distillation Dongqin Liu, Hongchang Yang, Zhaoxing Li, Jiao Dai, and Jizhong Han (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, UCAS) Abstract Abstract Knowledge distillation aims to transfer the embedded knowledge of complex models (teacher models) to simpler models (student models). It is known that teacher models can produce both correct and incorrect predictions, making it essential to extract valuable knowledge from accurate predictions while minimizing the impact of errors. However, current research either ignores this problem or utilizes methods like temperature scaling to smooth the teacher's output, which may somewhat reduce the negative effects of errors but also results in a loss of detailed correct information from the teacher. In this work, we propose Error-Decoupled Learning for Knowledge Distillation (EDLKD), a novel framework that explicitly disentangles correct and incorrect teacher predictions during the distillation process. By reinforcing reliable signals while suppressing erroneous ones, EDLKD ensures reliable knowledge transfer while minimizing noise. Although incorrect predictions may contain limited information, our ablation studies confirm that their suppression consistently enhances student performance. Extensive experiments on image classification and object detection demonstrate that EDLKD achieves state-of-the-art results in logit-based distillation. Contextual Adversarial Consistency Distillation for Remote Sensing Yanze Gao, Changxin Rong, Lvzhou Chen, Xiangyu Wang, Qiuju Chen, and Huanhuan Chen (University of science and technology of China) Abstract Abstract Knowledge distillation has become a pivotal strategy for deploying efficient remote sensing classifiers on resource-constrained platforms. However, student models in existing methods primarily focus on mimicking the observational distribution of high-performance teacher models, inheriting the contextual biases embedded in them. These biases induce spurious correlations where models rely on statistical shortcuts from non-causal backgrounds instead of intrinsic object features, leading to poor discrimination and generalization. To overcome this limitation, a Contextual Adversarial Consistency Distillation (CACD) framework is proposed to actively intervene in the data generation process. Specifically, the framework first computes a soft inter-class affinity measure to uncover potential contextual biases of teacher models. Guided by this prior knowledge, we then employ a latent diffusion model to synthesize adversarial backgrounds, generating interventional samples where the foreground is preserved but the background is replaced with confusing textures. Finally, we transfer this causal invariance to student models through a dual-level consistency objective, which reinforces the stability of both probabilistic outputs and feature representations under different background interventions. Extensive experiments on multiple benchmark datasets demonstrate that our framework significantly outperforms state-of-the-art baselines, effectively severing spurious correlations and enhancing both classification accuracy and robustness against contextual variations. FedL2T: Personalized Federated Learning with Two-Teacher Distillation for Seizure Prediction Jionghao Lou, Jian Zhang, Zhongmei Li, Lanlan Chen, and Enbo Feng (East China University of Science and Technology) Abstract Abstract The training of deep learning models in seizure prediction requires large amounts of Electroencephalogram (EEG) data. However, acquiring sufficient labeled EEG data is difficult due to annotation costs and privacy constraints. Federated Learning (FL) enables privacy-preserving collaborative training by sharing model updates instead of raw data. However, due to the inherent inter-patient variability in real-world scenarios, existing FL-based seizure prediction methods struggle to achieve robust performance under heterogeneous client settings. To address this challenge, we propose FedL2T, a personalized federated learning framework that leverages a novel two-teacher knowledge distillation strategy to generate superior personalized models for each client. Specifically, each client simultaneously learns from a globally aggregated model and a dynamically assigned peer model, promoting more direct and enriched knowledge exchange. To ensure reliable knowledge transfer, FedL2T employs an adaptive multi-level distillation strategy that aligns both prediction outputs and intermediate feature representations based on task confidence. In addition, a proximal regularization term is introduced to constrain personalized model updates, thereby enhancing training stability. Extensive experiments on two EEG datasets demonstrate that FedL2T consistently outperforms state-of-the-art FL methods, particularly under low-label conditions. Moreover, FedL2T exhibits rapid and stable convergence toward optimal performance, thereby reducing the number of communication rounds and associated overhead. These results underscore the potential of FedL2T as a reliable and personalized solution for seizure prediction in privacy-sensitive healthcare scenarios. CMKD: Distilling Vision Language Models with a Student Model MobileVTL for Action Recognition Huiting Li, Yao Yao, and Lei Wang (Bio-vision System Laboratory, Science and Technology on Micro-system Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences); Xixia Xv (Bio-vision System Laboratory, Science and Technology on Micro-system Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences); and Dongchen Zhu and Jiamao Li (Bio-vision System Laboratory, Science and Technology on Micro-system Laboratory, Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Abstract Abstract Pre-trained vision-language large models (VLMs) have achieved strong performance in action recognition after fine-tuning. However, Transformer-based VLMs entail substantial computational costs, limiting their deployment on edge devices. Although recent efficient hybrid architectures combining CNNs and Transformers offer a compromise, they still underperform compared to fully fine-tuned VLMs. To narrow this gap, we propose a cross-modal knowledge distillation (CMKD) framework, which includes a lightweight hybrid-architecture student model called MobileVTL and three distillation strategies. CMKD distills knowledge from teacher into student through multi-level feature alignment—encompassing contrastive, temporal, and spatial features—thus reducing performance discrepancies between architectures. Additionally, we design a Hierarchical Video-Text Interaction Module (HVTI) to foster dynamic interaction between low-level and high-level semantic features of video and text, leading to more robust multimodal representations. Experiments on Kinetics-400, UCF-101, and HMDB-51 using VTR-B/16 as the teacher demonstrate that MobileVTL achieves highly efficient performance in both fully-supervised and few-shot video recognition, with minimal computational overhead. Wednesday Virtual Room 2 IJCNN Paper Model Compression and Quantization II Session Chair: Bin Ji (National University of Defense Technology), Saibal Mukhopadhyay (Georgia Institute of Technology) ORCA: Outlier-aware Rotational Cluster Factorization for Efficient Low-Rank LLM Compression Han Cho (Georgia Institute of Technology), Fernando Camacho (Laboratory for Physical Sciences), and Saibal Mukhopadhyay (Georgia Institute of Technology) Abstract Abstract The immense size and computational cost of Large Language Models (LLMs) present significant barriers to their widespread deployment. Low-Rank Approximation (LRA) offers a promising, hardware-friendly solution by factorizing large weight matrices into more compact forms. A key insight is that the accuracy of this factorization can be significantly enhanced by first applying a geometric transformation to the model's weights. In this work, we introduce ORCA (Outlier-aware Rotational Cluster Factorization), a novel LRA framework that uses structured clustering-based factorization of weight matrices. We first apply an orthogonal transform to restructure the weight geometry to be more suitable for clustering. We then apply a group-wise clustering algorithm to the transformed weights to achieve a precise approximation. Furthermore, we demonstrate that this factorized representation enables an optional clustered attention reformulation, which reduces the algorithmic complexity of attention computation by performing attention computations directly in the compressed domain. Through experiments on the LLaMA and OPT model families, we show that ORCA can compress models by 75% while retaining over 96% of the original zero-shot accuracy on LLaMA2-13B, achieving a competitive compression-accuracy trade-off. OBEQuant: Outlier-aware Batch-norm Equivalent Quantization for Large Language Models Ye Zhong, Bin Ji, Xiaodong Liu, Shasha Li, Jun Ma, and Jie Yu (National University of Defense Technology) Abstract Abstract Post-training quantization (PTQ) is pivotal for deploying large language models (LLMs) in resource-constrained environments, yet aggressive low-bit quantization (W4A4) leads to severe accuracy degradation due to channel-wise mean bias and dynamic residual errors. To address this, we propose OBEQuant, a collaborative error-compensation framework featuring a cascaded dual-module design: a Bias Mean Correction (BMC) module that first eliminates systematic per-channel mean bias after rotation, and a Quantization-Aware Efficient Fine-Tuning (EFT) module that then compensates for task-dependent dynamic residuals via an ultra-low-rank adapter. These modules synergize to form an error-suppression loop, enabling layer-adaptive error decoupling and compensation. Extensive experiments on LLaMA-2 and LLaMA-3 show that under the challenging W4A4KV4 configuration, OBEQuant achieves a zero-shot average accuracy of 64.46% on LLaMA-3-8B, surpassing OSTQuant by +0.23%, while maintaining competitive perplexity. Quantitative and qualitative analyses further validate the effectiveness of OBEQuant. Structured Sparse Deep Nonnegative Matrix Factorization with ℓ_{2,1}-norm and Graph Regularization for Clustering Weifeng Yang (Yunnan Earthquake Agency) Abstract Abstract Deep nonnegative matrix factorization (DNMF) is a powerful tool for feature extraction, and promoting the sparsity of the factor matrices in DNMF can improve both the quality of extracted features and model interpretability. However, most existing DNMF methods promote the sparsity of factor matrices by individually evaluating the importance of each element in factor matrices and pruning unimportant elements (e.g., via the ℓ1-norm). This element-wise sparsity fails to capture the interactions and dependencies between different features and neglects the inherent structure of features, impairing the quality of extracted features and interpretability. To overcome this drawback, we propose a novel structured-sparse and graph-regularized deep nonnegative matrix factorization with ℓ_{2,1}-norm regularization (ℓ_{2,1}-SSGDNMF) method. Our method employs the ℓ_{2,1}-norm to explicitly impose the group sparsity on factor matrices, which enables entire feature groups to be selected or discarded and effectively captures feature interactions and dependencies, thereby enhancing the feature extraction capability and interpretability. Additionally, ℓ_{2,1}-SSGDNMF employs graph regularization to preserve the geometric structure information of data. To solve ℓ_{2,1}-SSGDNMF, we propose an algorithm named the Alternating Proximal Linearized (APL) algorithm, which provides an efficient and practical convergent scheme for solving ℓ_{2,1}-SSGDNMF. Furthermore, we prove that the sequence generated by our algorithm is globally convergent to a critical point and analyze the per-iteration complexity of our algorithm. Experimental results on six real-world benchmark datasets demonstrate that our method outperforms several state-of-the-art methods for clustering. MESO: Memory-Efficient Speculative Offloading for Large Language Model Inference Qingxiao Zhang, Xiaopeng Li, Bin Ji, Xiaodong Liu, Jie Yu, Long Peng, and Hao Xu (National University of Defense Technology) Abstract Abstract In resource-constrained devices, Large Language Model (LLM) inference faces significant challenges due to limited GPU memory resources. Existing studies, such as SpecOffload, combine offloading with speculative decoding to utilize a draft model for generating candidate tokens, thereby improving LLM generation throughput. However, they require the draft model to reside permanently in GPU memory. In resource-constrained devices, it tends to introduce high fixed overheads, resulting in low memory efficiency, which can be measured by the throughput gain per unit of GPU memory usage. To address this limitation, we propose MESO, a framework that specializes in improving memory efficiency. Specifically, MESO achieves high memory efficiency by offloading the draft model, enabling fine-grained asynchronous component loading, and leveraging pinned memory optimization. Experimental results show that compared to the selected baselines, MESO achieves 1.35× memory efficiency on average across various scenarios, and delivers up to 2.25× memory efficiency. Additionally, MESO reduces peak GPU memory usage by more than 50% compared to SpecOffload. Quantitative and qualitative analyses further validate the effectiveness of MESO. Wednesday Virtual Room 3 IJCNN Paper Multimodal Representation Learning I Session Chair: Xing Yu (East China Normal University), qiaoming zhu (Soochow University) MDFDU: A Multi-Task Deceptive-Factual Dialogue Understanding Framework Yanqing Liu, Zhong Qian, Peifeng Li, and Qiaoming Zhu (Soochow University) Abstract Abstract Trustworthiness in dialogue systems fundamentally relies on Deception Detection in Dialogue (DDD), defined as the discernment of subjective deceptive intent, and Dialogue-based Fact-Checking (DialFC), which entails verifying objective factual accuracy. However, existing methodologies predominantly treat these tasks in isolation, neglecting their shared mechanism of detecting information incongruence. This separation impedes generalization, particularly within the DDD domain, where data scarcity frequently leads to overfitting. To mitigate these limitations, we propose Multi-Task Deceptive-Factual Dialogue Understanding (MDFDU), a unified framework designed to jointly model subjective and objective trustworthiness. Key to our approach is a Prompt-Modulated Shared Reasoning mechanism, which leverages prompts generated by Large Language Models (LLMs) to dynamically bias the attention of a shared encoder, effectively decoupling task-specific semantics from shared reasoning capabilities. Additionally, we incorporate an uncertainty-weighted consistency regularization to align the latent representations of semantically corresponding labels (i.e., mapping ``Lie'' to ``Refutes'' and ``Truth'' to ``Supported''), thereby enabling transfer from the resource-rich DialFC task to the data-constrained DDD domain. Extensive experiments on the Box of Lies and DialFact benchmarks validate the efficacy of our method, demonstrating that MDFDU achieves state-of-the-art performance on both tasks. DPMMA: Dual-Prompt MultiModal Alignment for Deception Detection in Dialogues Yanqing Liu, Zhong Qian, Peifeng Li, and Qiaoming Zhu (Soochow University) Abstract Abstract Deception Detection in Dialogues (DDD) aims to recognize deceptive communication by analyzing multimodal behaviors. However, existing methods often treat modalities independently before shallow fusion, failing to bridge the inherent semantic gap between high-level linguistic concepts and low-level visual signals. To address this, we propose Dual-Prompt for Multimodal Alignment (DPMMA). Unlike conventional approaches, DPMMA introduces a Dual-Phase Alignment strategy: it integrates LLaMA3-generated ``hard prompts'' for explicit semantic translation with learnable ``soft tokens'' for latent geometric adaptation. Furthermore, a Semantic Guidance Module explicitly encodes psychological priors (e.g., cognitive load indicators) to highlight subtle deceptive micro-cues. Finally, a Progressive Fusion Transformer equipped with Gated Residual Injection ensures deep integration while preserving modality-specific details. Experiments on the Box of Lies dataset demonstrate that DPMMA not only outperforms strong baselines but also significantly mitigates class bias, achieving robust detection for both deceptive and truthful statements. Multi-dimensional Hybrid Fact-Checking: Integrating Structured and Unstructured Information to Promote LLMs Yanqing Liu, Zhong Qian, Peifeng Li, and Qiaoming Zhu (Soochow University) Abstract Abstract Dialogue-based fact-checking (DialFC) presents distinct challenges due to contextual dependencies, informal phrasing, and the entanglement of claims within multi-turn exchanges, which conventional verification paradigms struggle to capture. Existing approaches face significant limitations: they often lack explicit reasoning modeling, underutilize structured knowledge, and remain vulnerable to hallucination and factual drift in Large Language Models (LLMs). To address these issues, we propose Multi-Dimensional Hybrid Fact-Checking (MDH-FC), a unified framework that fuses semantic reasoning with structured knowledge modeling. Specifically, our approach integrates (1) a multi-perspective Chain-of-Thought (CoT) prompting strategy to generate interpretable reasoning traces; (2) a graph-based Claim–Evidence Entity Relationship (ERE) module utilizing Graph Convolutional Networks (GCNs) to capture structural consistency; and (3) a Mixture-of-Experts (MoE) network to dynamically filter and fuse heterogeneous reasoning pathways. Comprehensive experiments on the DialFact and HEALTHVER benchmarks demonstrate that MDH-FC effectively mitigates hallucinations and substantially outperforms state-of-the-art baselines in complex conversational settings. UMA-AD: A Unified Multimodal Alignment Framework for Interpretable Autonomous Driving Guangqiang Li, Dehui Du, and Xing Yu (East China Normal University) Abstract Abstract End-to-end autonomous driving systems have demonstrated exceptional capability in complex environments but suffer from opacity, hindering user trust and safety assurance. While recent interpretability approaches attempt to align linguistic explanations with intermediate perception outputs, they often face two critical limitations: the loss of fine-grained semantics due to Bird's-Eye-View (BEV) compression and the lack of explicit supervision for structured reasoning. To address these challenges, we propose UMA-AD, a novel framework designed to unify multimodal alignment for interpretable autonomous driving. First, we design a Unified Token Aligner module to adapt heterogeneous intermediate outputs from the AD model for the language decoder. Within this module, we integrate a Dual-Branch Video-Enhanced BEV Former to explicitly inject multi-scale video features, effectively compensating for semantic loss in BEV representations. Furthermore, we construct a structured Chain-of-Thought supervision strategy to internalize the hierarchical logic of the ``Perception-Prediction-Planning'' pipeline, empowering the model to generate explicit reasoning steps, while synchronously predicting low-level control signals to ensure the physical grounding of generated explanations. Extensive experiments on the Nu-X, TOD3Cap, and NuScenes-QA benchmarks validate the effectiveness of our framework in generating logically rigorous linguistic explanations. Wednesday Virtual Room 4 IJCNN Paper Multimodal Representation Learning II Session Chair: Peiran Liang (Northwest A&F University), Zhihao Cai (Ocean University of China) ODF-PSN: An Outlier-Inlier Dual-Region Decoupling Feature-Guided Network for Photometric Stereo Zhihao Cai, Shiyu Qin, Yi Li, Lin Qi, and Junyu Dong (Ocean University of China) Abstract Abstract Photometric stereo aims to recover surface normals from images captured under varying lighting conditions. Although deep learning methods have achieved promising results, existing approaches often struggle to accurately estimate surface normals in scenes where outlier and inlier regions are interwoven, due to entangled optical effects due to complex local structures and surface reflections. This results in a degradation of the accuracy of the reconstruction. To address this issue, we propose ODF-PSN, a dual-branch frame work that separately extracts features from outlier and inlier regions. Equipped with our Partitioned Feature Fusion Module, the network effectively integrates features from both regions. Furthermore, we introduce an anomaly-guided feature fusion module to enhance the representation of anomalous regions, thereby constructing a more robust global representation. Experimental results demonstrate that ODF-PSN achieves state-of-the-art performance on the DiLiGenT dataset. DBFF: A Dual-Branch Feature Fusion Model for Image Manipulation Localization Chenqi Liu and Juanjuan Luo (Beijing University of Posts and Telecommunications) Abstract Abstract Image manipulation localization is a critical task in media forensics. It aims to ensure the authenticity of digital content by identifying tampered regions. However, traditional methods often struggle to balance global semantic consistency with local tampering traces. This imbalance often results in high false-positive rates in complex scenes or the failure to detect subtle, high-frequency artifacts, which limits the reliability of forensics in real-world applications. To address these limitations, we propose a novel Dual-Branch Feature Fusion (DBFF) network. First, an RGB branch extracts global semantic information, and a high-frequency branch captures subtle tampering traces via fixed high-pass filters. Unlike previous methods, we systematically fuse features from both branches at multiple scales to bridge the domain gap. Second, we introduce a Feature Pyramid Attention (FPA) module. This module employs Spatial Pyramid Pooling (SPP) to capture global context and Squeeze-and-Executation (SE) blocks within a top-down pathway to adaptively recalibrate manipulation boundaries. Finally, a Contrastive Learning module is incorporated to map pixel features into a discriminative embedding space. Experiments on four benchmark datasets—NIST16, Columbia, CASIA, and IMD2020—demonstrate that the proposed method outperforms state-of-the-art techniques and achieves improved stability. CURA-Stereo: Cross-Volume Consistency Uncertainty for Fusion-and-Rectification in Iterative Stereo Matching Qingqing Cheng (Nanchang Hangkong University); Dongyang Wang (Nanchang University, Nanchang Hangkong University); and Chao He, Zhen Chen, Xinping Mao, and Congxuan Zhang (Nanchang Hangkong University) Abstract Abstract In recent years, iterative stereo matching has become a widely adopted paradigm for dense disparity estimation by refining predictions with feature warping and recurrent updates. However, in ill-posed regions such as occlusions, weak textures, and depth discontinuities, inaccurate disparity estimates can cause unreliable warping, introducing noisy cues and accumulating errors across iterations. To address this issue, we propose CURA-Stereo, an uncertainty-driven iterative stereo framework that explicitly models alignment reliability inside the refinement loop. Specifically, Cross-Volume Consistency Uncertainty (CVCU) estimates a unified uncertainty map that guides Uncertainty-Conditioned Soft Fusion (UCSF) and Uncertainty-Driven Alignment Rectification (UDAR) to adaptively fuse features and rectify distorted alignments, thereby suppressing error propagation. Extensive experiments demonstrate that CURA-Stereo achieves 0.44\,px EPE on Scene Flow, improves KITTI 2012 2-noc to 1.56 and KITTI 2015 D1-fg to 2.36, respectively, and attains 0.32 Bad~2.0 on ETH3D. Without fine-tuning, it further reaches 3.1 on ETH3D and 5.0 on Middlebury at quarter resolution, indicating stable convergence and strong robustness under domain shifts with marginal computational overhead. A Registration-Aware Spatial-Spectral Diffusion Framework for Unregistered Hyperspectral-Multispectral Image Fusion Peiran Liang, Yifei Zhao, Jiexiao Peng, Qiman Li, Jiaxin Yu, and Jin Hu (Northwest A&F University) Abstract Abstract Hyperspectral and multispectral image fusion aims to combine a low-resolution hyperspectral image (LR-HSI) with a high-resolution multispectral image (HR-MSI) to generate a high-resolution hyperspectral image (HR-HSI). Existing methods predominantly rely on fixed degradation distributions during training; real-world data acquired from disparate sensors often exhibit pose and photometric misalignments that severely degrade fusion quality. To address this, we propose SSRAF-Diff, a diffusion-based registration-aware spatial-spectral fusion framework. Adopting a dual-branch collaborative architecture with a Pyramid Attention-Guided Denoising (PAGD) module as the core generative unit, the framework simultaneously learns spectral fidelity and spatial detail recovery during the iterative denoising process. Specifically, the spectral branch models consistency to provide stable references, while the spatial branch enhances structural textures via a selective detail injection mechanism to suppress noise. Furthermore, we introduce a Spatial Deformation Registration Fusion (SDRF) module that estimates deformation fields from cross-modal similarity maps, enabling robust alignment and fusion despite geometric and photometric inconsistencies. Extensive experiments demonstrate that SSRAF-Diff outperforms various state-of-the-art (SOTA) methods. Wednesday Virtual Room 5 IJCNN Paper Neural Architectures and Sequence Models Session Chair: Suhang Qian (Tianjin University), Qizhao Long (Zhejiang University) SHD-Mamba: A Spectral–Temporal Hybrid Mamba with Dual-Suppression Decoding for Satellite Image Time Series Qizhao Long (Zhejiang University) and Yunzhuo Dai and Xinming Chen (Territorial Consolidation Center in Zhejiang Province) Abstract Abstract Satellite Image Time Series (SITS) crop segmentation is crucial for precision agriculture and food security. However, existing methods still face three key limitations when handling irregularly sampled multispectral sequences: (i) spectral processing that lacks physical priors, (ii) difficulty in jointly modeling discrete phenological events and continuous growth dynamics, and (iii) limited receptive fields that hinder large-scale parcel segmentation. CascadeMambaSeg: Mamba-based Medical Image Segmentation for Prostate with Cascaded Decoding Hyeonwook Kim and Ming Ma (Yeshiva University) Abstract Abstract While Transformers have long been the standard, Mamba-based models are emerging as a powerful rival. They offer a major efficiency boost by processing data with linear time complexity, yet they face a specific hurdle: Mamba inherently views images as one-dimensional sequences. Hybrid Mamba-CNN framework have been shown to bridge this gap and avoid overly complex scanning techniques. These models blend Mamba’s computational speed with the Convolutional Neural Networks (CNNs) ability to local features in 2D structures. Meanwhile, loss aggregation from different stages of the decoder has been effectively used in medical segmentation for better stability and prediction power. Building on this progress, we propose a novel Mamba-CNN hybrid framework named CascadeMambaSeg with combined loss for prostate segmentation. Our architecture leverages a MambaMixer layer for efficient extraction of low-level features from prostate scans, while standard convolutional layers are utilized for capturing high-level features and forming the final decoder structure. Experimental results on four publicly available prostate datasets show that our method outperforms existing methods in various benchmarks. MSWMNet: Enhancing State-Space Models with Multi-Scale Sliding Windows for Polyp Segmentation Yang Meng (School of Computer Science, Shenyang Aerospace University); Guoxu Zhang (Department of Radiology, the People's Hospital of Liaoning Province); and Wei Guo, Zhaoxuan Gong, and Guodong Zhang (School of Computer Science, Shenyang Aerospace University) Abstract Abstract Accurate segmentation of diminutive polyps in colonoscopy images is crucial for early colorectal cancer screening. However, due to their small scale, low contrast, and blurred boundaries, such polyps are easily obscured by complex backgrounds, and the detail loss introduced by downsampling further exacerbates foreground–background confusion, leading to frequent missed detections. To address these challenges, we propose a novel Mamba-based network, termed MSWMNet (Enhancing State-Space Models with Multi-Scale Sliding Windows for Polyp Segmentation), which adopts a collaborative design that integrates uncertainty-aware local refinement, multi-scale context aggregation, and global structural modeling. Specifically, a Structure-aware Uncertainty Context Refinement (SUCR) module is introduced to enhance the representation of ambiguous regions; a Multi-channel Shifted Window Transformer (MSWinTransformer) module is designed to capture multi-scale local contexts and mitigate small-object feature dilution; and an Enhanced Dual-path Cross Attention (EDCA) module is employed to strengthen cross-level feature interaction and alleviate boundary over-smoothing. In addition, the Mamba module further reinforces global dependency modeling in a computationally efficient manner. Extensive experiments on five public datasets demonstrate the superiority of MSWMNet. On the challenging ETIS-LaribPolypDB dataset, it achieves a Dice score of 75.8%, surpassing previous state-of-the-art methods by up to 5.4%, highlighting its potential for reliable clinical polyp detection. Octree-MambaDiffuser: State-Space Diffusion with Octree Guidance for 3D Generation Suhang Qian and Chao Xu (Tianjin University); Yushi Li, Chengtao Ji, and Xiaobo Jin (Xi'an Jiaotong-Liverpool University); and Xuanmo Zhang (Tianjin Second High School) Abstract Abstract Diffusion model-based approaches for 3D shape generation have advanced significantly. However, existing methods often fail to preserve both fine-grained geometric details and long-range structural coherence in complex shapes. While state space models (SSMs) can improve geometric consistency, directly integrating them with diffusion architecture often introduces causal constraints and limited local perception. To address these issues, we propose Octree-MambaDiffuser, a hierarchical framework that integrates octree-guided spatial partitioning with state-space diffusion. At first, we bridge the structural gap between 1D SSMs and 3D point clouds through a z-order serialization of octree representations. This provides a spatially coherent 1D sequence that better aligns with SSMs' autoregressive generation assumptions and improves long-range dependency modeling compared to naive traversal orders. The sequential data is then effectively processed by our scale-flexible MambaBlock which is enhanced with a Feature Conditioning Module (FCM). Acting as a critical stabilizer, the FCM pre-processes the input feature to reconcile Mamba’s discrete-state dynamics with the continuous demands of the diffusion process. By applying a simple yet strategic sequence of pre-norm Layer Normalization and SiLU activation, the module ensures stable gradient propagation throughout the denoising trajectory. This not only primes the underlying state-space mechanism to model the evolving shape but also unlocks its ability to capture subtle geometric details. The subsequent dual-branch Mamba dynamically balances local feature with global structural context. Extensive experiments demonstrate that Octree-MambaDiffuser generates high-fidelity and diverse 3D shapes, outperforming existing methods in both visual quality and structural integrity. Wednesday Virtual Room 6 IJCNN Paper Neural Network Foundations and Optimization Session Chair: Jikun Wu (Stellaris AI Limited), yuhao zhang (BeiHang University) Geometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection Dongxin Guo (The University of Hong Kong), Jikun Wu (Brain Investing Limited), and Siu Ming Yiu (The University of Hong Kong) Abstract Abstract Expert specialization is fundamental to Mixture-of-Experts (MoE) model success, yet existing metrics (cosine similarity, routing entropy) lack theoretical grounding and yield inconsistent conclusions under reparameterization. We present an information-geometric framework providing the first rigorous characterization of MoE specialization dynamics. Our key insight is that expert routing distributions evolve on the probability simplex equipped with the Fisher information metric, enabling formal analysis via Riemannian geometry. We prove that standard heuristic metrics violate parameterization invariance (Theorem 1), establish that specialization corresponds to geodesic flow with quantified approximation bounds (Theorem 2), and derive a failure predictor with theoretical threshold justification (Theorem 3). The framework introduces two principled metrics: Fisher Specialization Index (FSI) achieving r=0.91±0.02 correlation with downstream performance, and Fisher Heterogeneity Score (FHS) predicting training failure at 10% completion with AUC=0.89±0.03—outperforming validation-loss-based early stopping by 23% while requiring 40× fewer compute cycles. We validate intervention protocols achieving 87% recovery rate when FHS>1 is detected. Comprehensive experiments across language modeling (WikiText-103, C4), vision MoE (ImageNet), and scaling studies (8–64 experts, 125M–2.7B parameters) validate our theoretical predictions. OT-Route: Provably Load-Balanced MoE Routing via Optimal Transport Dongxin Guo (The University of Hong Kong), Jikun Wu (Brain Investing Limited), and Siu Ming Yiu (The University of Hong Kong) Abstract Abstract Mixture-of-Experts (MoE) models achieve remarkable parameter efficiency by activating only a subset of experts per input, yet load balancing remains a critical challenge that limits scalability. Existing solutions rely on auxiliary losses requiring extensive task-specific hyperparameter tuning, without theoretical guarantees on load distribution or convergence. We formulate MoE routing as an entropy-regularized optimal transport (OT) problem and propose OT-Route, a principled auxiliary-loss-free framework with provable properties. Our contributions are fourfold: (1) we cast token-to-expert assignment as capacity-constrained optimal transport, unifying prior heuristics under a theoretical framework; (2) we develop a differentiable Sinkhorn-based algorithm with proven convergence rate O(e^{-tε/4M}), including rigorous analysis showing capacity clipping preserves contraction; (3) we establish the first formal load balancing guarantee for MoE routing, proving expert utilization concentrates around uniform distribution; (4) we demonstrate 2.5 perplexity improvement over Switch Transformer on WikiText-103 and 1.0% accuracy gain over Soft-MoE on ImageNet with only 3.8% computational overhead. Uniquely, OT-Route maintains stable performance as model depth increases from 2 to 16 layers, confirming effectiveness in mitigating expert collapse. Experiments across eleven baselines validate scalability to 128 experts. When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth Dongxin Guo (The University of Hong Kong), Jikun Wu (Brain Investing Limited), and Siu Ming Yiu (The University of Hong Kong) Abstract Abstract Early-exit neural networks enable adaptive computation by allowing confident predictions to exit at intermediate layers, achieving 2–8× inference speedup. Despite widespread deployment, their generalization properties lack theoretical understanding—a gap explicitly identified in recent surveys. This paper establishes a unified PAC-Bayesian framework for adaptive-depth networks. (1) Novel Entropy-Based Bounds: We prove the first generalization bounds depending on exit-depth entropy H(D) and expected depth E[D] rather than maximum depth K, with sample complexity O((E[D] · d + H(D))/ε²). (2) Explicit Constructive Constants: Our analysis yields the leading coefficient √(2ln2) ≈ 1.177 with complete derivation and formal justification. (3) Provable Early-Exit Advantages: We establish sufficient conditions under which adaptive-depth networks strictly outperform fixed-depth counterparts, with quantified improvement α(√K−√kE)√(d/n). (4) Extension to Approximate Label Independence: We relax the label-independence assumption to ε-approximate policies, broadening applicability to learned routing. (5) Comprehensive Validation: Experiments across 6 architectures on 7 benchmarks demonstrate tightness ratios of 1.52–3.87× (all p < 0.001) versus >100× for classical bounds. Importantly, unlike conformal methods that provide coverage guarantees, our bounds directly characterize the population-empirical loss gap. We demonstrate that bound-guided threshold selection matches validation-tuned performance within 0.1–0.3%, substantially reducing hyperparameter search when validation data is limited. SHAPE: Coalition-Aware Expert Pruning for Sparse Mixture-of-Experts LLMs yuhao zhang, HongXu Jiang, YiXiang Zhang, and Zhen Zhang (BeiHang University) Abstract Abstract Sparse Mixture-of-Experts (MoE) large language models achieve strong quality with low per-token compute, yet their deployment is often limited by the memory wall: the full expert pool must remain resident to support token-dependent routing. Expert pruning is a direct remedy, but prior criteria often score experts independently and ignore that MoE inference is inherently coalitional, where outputs arise from routed top-k expert combinations. We present SHAPE, a task-driven pruning framework that explicitly models intra-layer expert cooperation. SHAPE formulates routing traces from a small calibration set as an empirical cooperative game and assigns interaction-aware expert values via a Shapley-inspired attribution over observed top-k coalitions, enabling the identification of experts that are essential for high-utility collaborations rather than merely frequent. To preserve MoE topology under a global pruning budget, SHAPE further introduces a quality-coverage selection rule that retains, in each layer, the minimal expert subset covering an alpha fraction of non-negative Shapley mass, using bisection to match a target keep rate. Experiments on three modern MoE backbones (Qwen3-30B-A3B, GPT-OSS-20B, and DeepSeek-V2-Lite) across diverse benchmarks show that SHAPE improves robustness over global and layer-wise pruning variants, maintaining accuracy under 20% and 40% expert pruning without additional training and delivering clear reductions in peak GPU memory footprint. Wednesday Virtual Room 7 IJCNN Paper Object Detection and Recognition I Session Chair: xiyu pan (Central South University of Forestry and Technology), Wenzhu Yang (Hebei University; Machine Vision Engineering Research Center, Hebei University, Baoding, China) Cil-FPN: A New Paradigm for Efficient Feature Fusion for UAV Object Detection xiyu pan, Kai Xiong, and Jianjun Li (Central South University of Forestry and Technology) Abstract Abstract Object detection in UAV imagery faces persistent challenges from small objects, dense distributions, and complex backgrounds. Conventional detectors tend to lose fine-grained spatial cues due to repeated down-sampling, limiting small-object perception. To address this, we introduce the Cross-layer Lightweight Feature Pyramid Network (Cil-FPN), which enhances detection by combining shallow high-resolution features with deep semantic cues. Within Cil-FPN, the Multibranch Competitive Frequency-guided Module (MCFM) integrates global context pooling, frequency-based edge enhancement, and branch competition for efficient and discriminative feature fusion. Designed as a lightweight plug-and-play component, Cil-FPN can be seamlessly incorporated into mainstream detectors such as YOLO11, RT-DETR, and DEIM. Extensive experiments on four challenging benchmarks— VisDrone2019, CODrone, HIT-UAV, and CARPK—demonstrate improved accuracy with reduced computational cost. For instance, YOLO11 equipped with Cil-FPN and MCFM improves mAP@0.5 from 31.6 to 33.2 on VisDrone2019 with only 3.7M parameters. To facilitate reproducibility, the code is provided at https://github.com/ panxiyu2001/Cil-FPN. LGNet: A Lightweight Box-Guided Distillation Framework for Tiny UAV Object Detection Liguang Zhang, Renfang Wang, Xiaozhe Gu, Hong Qiu, and Yingying Huang (Zhejiang Wanli University) Abstract Abstract Detecting tiny objects in UAV imagery is challenging due to their extremely small size and dense spatial distribution, which make feature modeling and post-processing highly sensitive to localization noise. Although YOLO-style detectors adopt feature pyramids and naive multi-scale fusion, hard non-maximum suppression often lead to weak cross-scale interactions and excessive suppression of small objects. To address these limitations, we propose LGNet, a lightweight box-guided distillation framework built upon YOLOv11. Specifically, LGNet introduces (i) a high-resolution P2-oriented detection head for improved tiny-object representation, (ii) a gated IoU-aware confidence prior that reduces redundant suppression during NMS, and (iii) a localization-centric box-level knowledge distillation scheme to further improve the detection performance for small and dense objects. Experiments on the VisDrone benchmark show that LGNet achieves 44.6\% mAP\textsubscript{50} and 26.3\% mAP\textsubscript{50:95}, outperforming state-of-the-art competitors under comparable settings while maintaining a more compact model size, making it ideal for resource-constrained UAV platforms. Perception and Regression Network for Small Object Detection in UAV Aerial Photography Xiaofeng Wang and Wenzhu Yang (Hebei University) Abstract Abstract Unmanned Aerial Vehicles (UAVs) object detection has widespread applications in the real world. However, due to the small size, dense distribution and blurred features of UAV images, traditional object detection algorithms encounter significant challenges. To address these issues, this paper proposes a Perception and Regression Network (PRNet) built upon YOLOv11, which incorporates multi-scale feature fusion, spatial attention mechanisms, and regression optimization. First, we design the C3k2-MSF module, which enhances the model's ability to perceive objects at multiple scales. Second, we design the BG module that integrates the Bidirectional Feature Pyramid Network (BiFPN) module and the Global-to-Local Spatial Aggregation (GLSA) module, aiming to improve small object detection accuracy. Finally, we propose Coord-Dyhead, an advanced detection head that integrates our newly designed Coord-SE module into the Dynamic Head (Dyhead) architecture. Experimental results indicate that our PRNet achieves significant improvements of 10.8% and 7.4% in mAP over the baseline model on the VisDrone2019 and AI-TOD datasets respectively, while maintaining comparable model parameters. DH-DETR: A Dynamic Hypergraph Detection Transformer for Enhanced Tiny Object Detection in UAV Imagery Shengbo Wang, Wenzhu Yang, and Jinming Li (Hebei University) Abstract Abstract Tiny object detection in UAV imagery is inherently challenging due to extreme scale variations, limited pixel representations, and densely cluttered scenes. To address these challenges, we propose DH-DETR, a Dynamic Hypergraph Detection Transformer, specifically designed for tiny object detection in UAV imagery. DH-DETR incorporates three key components: a Hypergraph High-order Correlation Aggregator (HHCA) that captures cross-level and high-order dependencies to strengthen global structural reasoning; a lightweight Frequency–spatial Unified SElective (FUSE) module that enriches fine-grained details by injecting frequency-domain cues into shallow features while suppressing background noise; and a Dilated Depthwise Multi-Head Self-Attention (D2MHSA) block that efficiently enlarges the receptive field for long-range context modeling without quadratic complexity. Together, these components enable DH-DETR to preserve critical details, focus on discriminative regions, and achieve superior accuracy on challenging tiny object detection tasks. Extensive experiments on the VisDrone, UAVDT, and AI-TOD benchmarks demonstrate that DH-DETR achieves state-of-the-art performance, obtaining 28.8% AP on VisDrone, 19.1% AP on UAVDT, and 28.1% AP on AI-TOD, while maintaining competitive computational efficiency. These results highlight the effectiveness of integrating hypergraph-based high-order reasoning, frequency-domain enhancement, and efficient attention mechanisms for advancing high-resolution small-object detection in real-world UAV scenarios. Wednesday Virtual Room 8 IJCNN Paper Object Detection and Recognition II Session Chair: HaiDi Xu (Zhejiang Sci-Tech University), Tianzhu Xie (De Anza College) SF-YOLO: Scale-Aware and Frequency-Enhanced YOLO for Real-Time Aerial Object Detection Xiaoan Bao, Cheng Xu, Haidi Xu, Na Zhang, and Biao Wu (Zhejiang Sci-Tech University) and Qingqi Zhang (Hangzhou Institute of Medicine Chinese Academy of Sciences) Abstract Abstract The rapid development of Unmanned Aerial Vehicle (UAV) technology has established object detection in aerial imagery as a core supporting technology in fields such as urban surveillance and disaster relief. However, constrained by extreme scale variations, the irreversible loss of tiny object features, and the stringent latency constraints of onboard edge devices, real-time object detection in this domain still faces severe challenges. Although FBRT-YOLO has achieved excellent inference speeds through architectural pruning, its aggressive downsampling strategy trades efficiency at the cost of tiny object perception, leading to severe missed detections. To address this challenge, we propose a Scale-aware and Frequency-enhanced detector, SF-YOLO, aiming to break through the bottleneck of tiny object localization precision while maintaining real-time processing capabilities. First, we introduce a Selective Kernel Downsampling (SKD) module, which effectively circumvents feature aliasing and information loss caused by fixed convolutions by incorporating a dynamic receptive field mechanism in the shallow feature extraction stage. Second, addressing the resolution bottleneck, we design a lightweight High-Frequency Perception (HFPlite) module that utilizes high-pass filtering priors and a dual-path attention mechanism to efficiently recover fine-grained target textures while suppressing background noise. Finally, we construct a Weighted Inner-Wasserstein IoU (WIW-IoU) loss function, which fundamentally resolves the gradient vanishing problem caused by the sensitivity of tiny objects to positional deviations by synergizing distribution metrics with geometric constraints. Experimental results demonstrate that, compared with the baseline model FBRT-YOLO-S, SF-YOLO-S significantly improves mAP50 by 3.4% and 3.8% on the VisDrone2018 and AI-TOD datasets respectively, with a 6.9% reduction in parameters, successfully achieving the optimal balance between detection performance and inference speed. GS-YOLO: Lightweight and Highly Efficient Real-time Aerial Image Detector Rundong Gao and Yu Zhang (Shenyang University of Chemical Technology), Kailai zhuang (Tiangong University), Jiajie Fan (South China Normal University), Zeming Tian (wuhan University), and Runxiao Gao (Shandong University of Technology) Abstract Abstract Aerial image detection has extensive applications in real-world scenarios. Despite significant progress in object detection, extremely small targets remain difficult to detect due to sparse pixels and complex backgrounds. Existing high-performance models usually contain a large number of parameters, which makes lightweight deployment on edge devices difficult. To address these issues, we propose an innovative detection framework named GS-YOLO, which enhances the representation of small object features through two lightweight modules, the Edge-Aware Gaussian Downsampling Module (EAG-Stem) and the Gaussian Difference Calibration Module (GDCM). The EAG-Stem focuses on the edge information of small objects, while the GDCM models the continuity of local structures to improve the distinction between objects and background. Furthermore, a Scale-Adaptive Weighted Intersection Over Union(SA-WIoU) loss function is designed to handle multi-scale detection disparities by dynamically adjusting the loss weights of targets at different scales, thereby improving localization accuracy for small objects. Extensive experiments conducted on the Visdrone dataset demonstrate that GS-YOLO achieves an effective trade-off between accuracy and efficiency. Compared to the baseline model, GS-YOLO achieves an average improvement of 7.4% in Average Precision (AP) and an average reduction of 14.5% in FLOPs, with a particularly notable average reduction of 71.76% in Params. The proposed model also outperforms many other detectors on the SIMD dataset. MBCDE: Towards Plug-and-Play Spatial-Frequency Fusion for Aerial Object Detection Haodong Li and Haicheng Qu (the School of Software, Liaoning Technical University) Abstract Abstract The combination of computer vision and remote sensing imaging mechanisms drives the development of aerial object detection technology toward intelligence and precision. Existing methods are limited by the locality of spatial representation and insufficient use of frequency domain information, leading to the annihilation of small object features and sensitivity to noise interference. To address these limitations, we propose a multi-branch cross-domain enhancement (MBCDE) method, which incorporates three novel modules: an edge enhancement frequency-domain module to fuse multi-directional gradient and frequency features in shallow layers for refined texture detail, a Gaussian filtering frequency-domain module to suppress high-frequency noise in deep layers via a heat-diffusion Gaussian kernel, and a multi-branch large-Kernel frequency-domain attention module to mine and enhance shallow features during fusion for improved small object detection. Extensive experiments on the low-altitude drone dataset VisDrone2019 and the high-altitude satellite dataset DIOR demonstrate the effectiveness of MBCDE, showcasing its plug-and-play compatibility across diverse detection architectures. OVRD: Open-Vocabulary Rotated Tiny-Object Discovery via Selective Perception Lucas Bin Fan Wu and Tianzhu Xie (De Anza College) Abstract Abstract Open-vocabulary object detection (OVD) has achieved significant strides by leveraging vision-language models (VLMs) to recognize novel categories. However, discovering rotated tiny objects in aerial imagery remains a formidable challenge: the inherent misalignment between arbitrary-oriented targets and axis-aligned language priors, coupled with spectral noise in complex backgrounds, often leads to catastrophic semantic ambiguity for unseen classes. In this paper, we propose OVRD, a framework for Open-Vocabulary Rotated tiny-object Discovery via Selective Perception. To bridge the gap between geometric variability and semantic consistency, we introduce the Structure-Tensor Mixture of Experts (ST-MoE). By mimicking the mathematical properties of structure tensors, ST-MoE dynamically aligns visual features with their principal geometric components via an Anisotropic Tensor Encoder, ensuring that the visual representations of novel objects remain invariant to rotation and thus better match text embeddings. Simultaneously, an Isotropic Context Compensator preserves global texture coherence to maintain zero-shot generalization. To further purify the semantics of tiny objects, we develop the Spectrally-Decoupled Geometric Refinement (SDGR) module. SDGR employs a "coarse-to-fine" strategy using Haar wavelet-based spectral decoupling to isolate high-frequency boundary responses from background interference, followed by Orientation-adaptive Dynamic Convolution for sub-pixel rectification. Extensive experiments on the challenging CODrone benchmark demonstrate that OVRD achieves state-of-the-art performance with $42.0$ $AP_{50}$, surpassing existing methods by $0.9$ mAP. Our framework enables selective perception of geometric and spectral cues, significantly boosting novel target discovery within complex open-vocabulary aerial environments. Wednesday Virtual Room 1 IJCNN Paper Object Detection and Recognition III Session Chair: Tao Wu (Shanghai Institute of Technology, Faculty of Intelligent Technology), Jin Zhang (Changsha University of Science & Technology) PGA-YOLO: Toward Efficient and Robust Feature Interaction for Defect Detection in Complex Industrial Scenes Jin Zhang, Zhiwen Wu, Cheng Sun, Bin Hu, and Qi Cao (Changsha University of Science & Technology) and Tie Wang (Hunan Sanyue Shuwei Technology Co., Ltd.) Abstract Abstract Unlike conventional object detection tasks, industrial metal surface defect detection is challenged by severe background interference and extreme scale imbalance among defect instances. To address these challenges, we propose PGA-YOLO, a detection framework co-designed from the perspectives of feature interaction and detection architecture. Specifically, a refined feature interaction mechanism is introduced to enhance the discriminability of deep features while effectively suppressing background-induced confusion. Meanwhile, efficient convolutional operators together with an adaptive detection head are incorporated to improve small-object localization accuracy while maintaining a lightweight model profile. This synergistic design effectively alleviates the trade-off between detection precision and computational efficiency. Extensive experiments demonstrate that PGA-YOLO achieves a superior balance between accuracy and efficiency. With only 5.78M parameters, the proposed method attains 81.43\% mAP@50 and 50.67\% mAP@50:95 on the NEU-DET dataset. Furthermore, PGA-YOLO consistently outperforms state-of-the-art methods in terms of generalization capability and real-time performance on the GC10-DET and AL10-DET datasets. DSP-DETR: Towards Efficient Detail–Spatial Pyramid Modeling for Steel Surface Defect Detection Yihan Chai (Zhengzhou University); Qiming Yu (Zhengzhou Normal University, School of Information Science and Technology); and Chengming Liu (Zhengzhou University) Abstract Abstract Surface defect detection in steel is a fundamental task in industrial inspection systems. In complex inspection scenarios, defects tend to be weakened during cross-scale feature modeling. Moreover, spatial structure and defect morphology are not effectively captured, leading to degraded detection accuracy. These issues mainly result from diverse spatial defect distributions and the limited ability of existing frameworks to jointly model local details and spatial hierarchies. To address these challenges, this paper proposes an enhanced RT-DETR framework, termed Detail–Spatial Pyramid DETR (DSP-DETR). In this framework, the backbone integrates a Multi-Granularity Convolution Block (MGCBlock) to extract detailed features at multiple granularities, which enhances the representation of subtle defects while keeping the model lightweight. A Spatial-Aware Pyramid Fusion Network (SAPFPN) is further introduced to refine feature interaction in the pyramid and model spatial structures and defect morphology. Within this network, a spatial–channel attention (SCA) mechanism with directional operations and large receptive fields strengthens the perception of elongated and texture-sensitive defects, enabling more accurate detection of complex surface textures. Specifically, the proposed model achieves a 4.2% improvement in mAP₅₀ on the NEU-DET dataset and a 4.5% gain on the GC10-DET dataset compared to the baseline, validating its effectiveness for industrial surface defect detection. DS-YOLO-RC: Geometric-Aware Rotated Crack Detection via Over-Parameterized Dilated Convolution and Topology-Preserving Alignment Zhimin Yue and Li Li (Southwest University of Science and Technology) Abstract Abstract Standard attention-centric detectors, while powerful, often suffer from inductive bias mismatch when processing slender, high-frequency manifolds like hydraulic cracks. To address this, we propose DS-YOLO-RC, a lightweight rotated object detection framework tailored for geometric-aware feature modeling. By re-introducing morphological priors via Depth-wise Over-parameterized Dilated Convolution (DoDConv) and ensuring spatial fidelity via Dynamic Upsampling (DySample), we fundamentally reconstruct the feature fusion network. Specifically, DoDConv effectively elongates the receptive field along the crack trajectory to capture global topology, while DySample utilizes content-aware inverse reprojection to rectify feature misalignment caused by static interpolation. Extensive experiments on a self-constructed hydraulic concrete dataset and the CRACK500 benchmark demonstrate that our method consistently outperforms state-of-the-art detectors. Compared with the YOLOv12 baseline, DS-YOLO-RC achieves a 2.5% improvement in mAP@50-95 (reaching 45.1%) while reducing computational costs to 5.9 GFLOPs, striking an optimal balance between high-precision localization and real-time inference (96 FPS). ED-YOLO: An Algorithm for Detecting Questions and Visual Elements in Educational Documents Tao Lin and Tao Wu (Shanghai Institute of Technology) Abstract Abstract With the rapid development of digital education, the digital transformation of paper-based teaching resources has become a core task in the construction of educational informatization. Accurate identification and segmentation of questions, images, and tables in educational documents are essential for realizing the structured and intelligent utilization of resources. Existing methods for detecting questions in complex layouts, such as multi-column layouts, mixed text-image formats, and handwritten annotations, suffer from insufficient positioning accuracy and limited ability to suppress background interference. Additionally, the detection of diverse question types, such as multiple-choice, fill-in-the-blank, and solution questions, presents further challenges. To address these issues, this paper proposes an enhanced segmentation algorithm, ED-YOLO (Educational Document-YOLO), which improves upon YOLOv11 for detecting both questions and visual elements in educational documents. The algorithm introduces a Task-Aligned Dynamic Detection Head (TADDH) to enhance positioning and classification performance, optimizes feature fusion using a Context Guide Fusion Module (CGFM), and integrates DCNv4 with C3K2 blocks (DCNv4_C3k2) in both backbone and neck for adaptive feature extraction. Experimental results demonstrate that ED-YOLO achieves 98.32% mAP50 and 95.25% precision on a multi-subject dataset, showing significant robustness in complex document layouts. The proposed method provides valuable technical support for the efficient digital transformation of educational resources, enhancing both question detection and the identification of associated visual content such as images and tables. Wednesday Virtual Room 2 IJCNN Paper Object Detection and Recognition IV Session Chair: Congyu Liu (Changsha University of Science and Technology), Zhe Wang (East China University of Science and Technology) A Multi-Scale Semantic Feature based Relaxed Alignment Network for Visible-Infrared Person Re-identification Zelin Deng, Congyu Liu, Ke Nai, and Jiaxin Chen (Changsha University of Science and Technology) and Pei He (Guangzhou University) Abstract Abstract Visible-infrared person re-identification (VI-ReID) is a challenging image retrieval task that aims to match the same identity pedestrian images across visible and infrared modalities. Feature alignment is an effective method for identifying identical entities across modalities, yet mainstream models tend to be subject to rigid constraints during alignment, making them highly susceptible to alignment failure and degrading model performance. To address this issue, we propose a Multi-Scale Semantic Feature based Relaxed Alignment Network (MSRANet). Firstly, we propose a Multi-scale Semantic Correlation Enhancement module to fuse high similarity features in adjacent layers to obtain multi-scale semantic features, which can achieve better discriminative ability by utilizing low-level fine-grained features to enhance high-level semantic features. Secondly, by fusing semantic features of pedestrians with the same identity across different scenarios within a modality to enhance feature diversity, we develop a relaxed alignment strategy that achieves coarse-grained cross-modality feature alignment, thereby improving the model's adaptability. Finally, extensive experiments on two public datasets (SYSU-MM01 and RegDB) demonstrate that our MSRANet achieves state-of-the-art performance, significantly outperforming multiple popular methods. Reffusion: Self-Boosted Person Re-Identification via Pose-Guided Person Generation Yong Zhao, Yali Li, and Shengjin Wang (Department of Electronic Engineering, Tsinghua University; Beijing National Research Center for Information Science and Technology) Abstract Abstract Person re-identification (ReID) focuses on learning identity-discriminative and pose-invariant features. To enrich pose diversity of training data, person synthesis techniques are employed to generate images with novel poses. However, existing diffusion-based methods struggle to maintain identity consistency and decouple the person from background interference, often producing low-quality images which hinder ReID training. To address these limitations, we present Reffusion, a diffusion-based framework with an Online Filtering and Training Policy (OFTP). Reffusion concentrates on generating high-fidelity person images, while OFTP aims to train ReID model with generated persons. Specifically, Reffusion integrates Orthogonal Memory Module to capture decorrelated human patterns and Adaptive Masking Training to mitigate background interference. During inference of Reffusion, Instance Gradient Guidance is utilized to enhance identity consistency. When training ReID with these generated persons, OFTP adaptively filters out low-quality samples and retains informative ones for effective post-training. Experiments demonstrate that Reffusion yields an 8.008 FID on Market1501, while OFTP achieves 1.6% and 6.4% improvements in mAP over the strong baseline BoT on MSMT17 and CUHK03, respectively, significantly outperforming state-of-the-art methods. See Fine, Know Clear: Improving Video Object Segmentation via Fine-Grained Matching and MambaVision-Powered Semantics Dongpei Dong, Bo Li, Zhiheng Zhou, and DeLu Zeng (South China University of Technology) Abstract Abstract Memory-bank-based methods have attained promising performance in semi-supervised video object segmentation (SVOS). In most existing methods, the memory bank only stores single-scale coarse-grained features, thus leading to suboptimal performance in fine-grained segmentation. Meanwhile, memory matching is only conducted at the pixel level, which lacks high-level semantic modeling of objects as holistic entities, thus causing ambiguity in distinguishing similar objects. In this work, we propose an SVOS framework based on multi-scale memory representation and semantic enhancement. It stores multi-scale fine-grained features in the memory bank and fuses instance-level semantic information extracted from MambaVision, thereby boosting sensitivity to fine-grained object details while mitigating ambiguity in distinguishing similar objects. We thus name the proposed model See Fine, Know Clear (SFKC). This design faces a key challenge: elevated computational costs induced by fine-grained feature matching. To mitigate this issue, we design a Top-K-based filtering mechanism that optimizes the matching process and reduces computational overhead. Extensive experiments verify that our method enables accurate and robust foreground object segmentation, while retaining high efficiency when incorporating fine-grained memory features. RHNet: Residue Channel Prior-guided High-Low Feature Collaborative Network for Infrared Small Target Detection Ziheng Fang, Qian Zhang, Yunfei Tong, and Zhe Wang (East China University of Science and Technology) Abstract Abstract Following the detection-by-segmentation paradigm, U-Net and its variants have attained competitive performance in Infrared Small Target Detection (IRSTD) benchmarks. However, as typical CNNs relying on pure spatial-domain convolution stacking, they face limitations in utilizing differentiated high-low feature (semantic-spatial) collaboration and lack a global residue channel view. Such neglect of high-low feature synergy, coupled with the absence of residue channel-aware learning and redundant computations from excessive convolution stacking, induces representation bias that impairs detection reliability. This paper proposes a Residue Channel Prior-guided High-Low Feature Collaborative Network (RHNet), where Multi-Scale Feature Enhancement (MSFE) module and Prior-Guided Feature Boost (PGFB) module are advised to tackle the aforementioned issues. The low-level MSFE-D integrates bidirectional multi-scale and dilated convolutions to capture fine-grained contour details. The high-level MSFE-S employs saliency kernels to extract discriminative semantic information. The PGFB module leverages residue channel prior to guide feature optimization, compensating for the lack of global residue channel perspective and replacing redundant convolutions to suppress false responses. Extensive experiments on IRSTD-1K, NUAA-SIRST, and NUDT-SIRST demonstrate that RHNet outperforms 26 state-of-the-art methods with an IoU of 70.29\% on IRSTD-1K, while maintaining an ultra-lightweight architecture (3.79M parameters). Wednesday Virtual Room 3 IJCNN Paper Reinforcement Learning I Session Chair: Xiao Sun (Hefei University of Technology), Zilan Li (Guilin University of Electronic Technology) Hyper-DAG-Based Hierarchical Reinforcement Learning for Unifying Data and Resource Dependencies in Multi-UAV Rescue Zilan Li, Xuesong Wang, Zhongyi Zhai, and Lingzhong Zhao (Guilin University of Electronic Technology) Abstract Abstract Natural disaster rescue missions are characterized by extreme urgency, high complexity, and dynamic environmental uncertainties. In such scenarios,a fleet of Unmanned Aerial Vehicles (UAVs) act as critical assets for search and rescue, supply delivery, and communication recovery. However, traditional task scheduling approaches, often based on standard Directed Acyclic Graphs (DAGs), struggle to effectively model the implicit "resource coupling" (e.g., shared energy budgets and channel interference) alongside explicit data dependencies. This limitation frequently leads to task failures and resource imbalances under dynamic constraints. To address these challenges, this paper proposes a novel scheduling framework based on Hyper-DAG (Hypergraph Directed Acyclic Graph) and Hierarchical Reinforcement Learning (HRL). First, we introduce a Hyper-DAG modeling approach that unifies data, energy, and channel dependencies into a high-order graph structure, utilizing hyperedges to capture complex one-to-many resource constraints. Second, we develop an HRL-based multi-UAV scheduler integrated with the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm. This scheduler employs a graph attention mechanism to dynamically perceive critical dependency bottlenecks and generates optimal task execution sequences in real-time. Extensive experiments across five distinct disaster scenarios (earthquake, flood, wildfire, landslide, and typhoon) demonstrate that our method significantly outperforms baseline algorithms. The proposed framework effectively satisfies the "Golden Rescue Window" time constraints, ensures balanced energy consumption (Gini coefficient < 0.3), and maintains high system reliability (failure rate < 5%). Maskable Proximal Policy Optimization for Safe Control of a Silo Level Balancing Process Kerollan da Silva Ramos (Instituto Tecnológico Vale, Universidade Federal de Ouro Preto) and Thomás Vargas Barsante e Pinto and Gustavo Pessin (Instituto Tecnológico Vale) Abstract Abstract Maintaining balanced material levels across a set of silos is a safety-critical control task in industrial mineral processing plants, commonly addressed through tripper car positioning under strict operational constraints. This work investigates the application of Maskable Proximal Policy Optimization (MaskablePPO) as a Reinforcement Learning (RL) framework for safety-aware decision making, using the silo level balancing problem as an industrial inspired test environment. Invalid action masking is employed as a hard constraint mechanism to prevent invalid actions, while a curriculum learning strategy is introduced to improve training stability. The agent is trained and evaluated in a simulated environment characterized by complex physical constraints, stochastic dynamics, and interlock logics. Experimental results show that the MaskablePPO outperforms a standard Proximal Policy Optimization (PPO) baseline in terms of training convergence, computational efficiency, and constraint satisfaction, while eliminating invalid actions by design. These results highlight the effectiveness of MaskablePPO for RL in constrained industrial settings. Protocol Prior Guided Reinforcement Learning for Resilient Time-sensitive Networking Routing Chao Wang (SUN YAT-SEN UNIVERSITY), Chongwu Dong (Guangzhou University), Mu Yu (Tongji University), and Wushao Wen (SUN YAT-SEN UNIVERSITY) Abstract Abstract Time-Sensitive Networking (TSN) provides deterministic communication, yet the same determinism introduces failure modes that can mislead learning-based routing. When a persistent Gate Control List (GCL) misconfiguration occurs at a port, multiple flows that traverse the faulty port exhibit synchronized delay spikes, so a model-free reinforcement learning (RL) agent can confuse this fault with stochastic congestion if it relies only on end-to-end latency, repeatedly switching routes without resolving the root cause and destabilizing scheduled service. To address this correlation trap, we propose Protocol Prior-Guided Reinforcement Learning (PPRL), where we translate protocol defined determinism into a continuous prior signal. PPRL introduces a Spectral Protocol Prior Module (PPM) that quantifies cross flow synchronicity via the spectral entropy of a delay covariance matrix and outputs a Spectral Synchronicity Index. This spectral evidence constrains Proximal Policy Optimization (PPO) through an adaptive KL regularized objective that biases the policy toward a stability prior under strong fault evidence while naturally reverting to standard PPO updates when the evidence is diffuse. In ns-3 TSN simulations, PPRL reduces route switching by 84.1% under GCL faults and maintains scheduling efficiency comparable to strong baselines. ProgAttn-DAG: Progressive Attention-Based Modality Fusion in Directed Acyclic Graphs for Multimodal Conversational Emotion Recognition Yitong Yao and Xiao Sun (Hefei University of Technology) Abstract Abstract Affective computing constitutes a crucial branch of Natural Language Processing (NLP). Multimodal Emotion Recognition in Conversation (MERC) has emerged as a prominent task in multimodal emotion computing. It aims to process and fuse information from heterogeneous modalities (visual, acoustic, textual) to accurately recognize speakers' emotional states in conversations. Despite extensive research in emotion recognition, challenges persist. First, multimodal methods often employ simplistic single-step alignment and fixed-weight fusion strategies, inadequately addressing heterogeneity and failing to capture comprehensive cross-modal interactions. Second, conversational emotion recognition requires modeling complex contextual dependencies, as emotions evolve dynamically throughout dialogues. Third, benchmark datasets exhibit severe class imbalance with long-tail distributions, where minority emotion categories are underrepresented. To address these challenges, this paper proposes a novel Progressive Attention-Based Modality Fusion in Directed Acyclic Graphs (ProgAttn-DAG) framework for multimodal conversational emotion recognition. The framework integrates a progressive attention mechanism (ProgAttn) that hierarchically fuses text, audio, and visual modalities to address heterogeneity, multi-layer directed acyclic graph networks to explicitly model contextual dependencies, and curriculum learning strategies to optimize sample-difficulty-based training and mitigate data imbalance. Extensive experiments on two benchmark datasets demonstrate the effectiveness of ProgAttn-DAG, particularly in addressing data imbalance issues. Wednesday Virtual Room 4 IJCNN Paper Retrieval-Augmented Generation I Session Chair: Jundong Liu (Ohio University), Jingyu Wang (Beijing Institute of Technology) URA-NER: A Unified Retrieval-Augmented Framework with Retrieval Alignment and Uncertainty Reduction for Low-Resource NER Jingyu Wang, Shijie Wu, and Fusheng Jin (Beijing Institute of Technology) Abstract Abstract In-context learning (ICL) based on large language models (LLMs) has shown promising potential in alleviating performance bottlenecks caused by the limited availability of annotated data in Named Entity Recognition (NER). However, existing methods still face issues of retrieval misalignment and generation uncertainty, making their performance heavily dependent on the LLM's capabilities. As the parameter scale of LLMs decreases, their performance in few-shot settings deteriorates significantly. In this paper, we propose a novel unified retrieval-augmented framework, URA-NER, including three key components: Progressive Granularity Retrieval (PGR), Model-aware Representation Enhancement (MaRE), and Reason-aware Knowledge Verification. PGR is a two-stage retrieval mechanism that achieves stage alignment. It first retrieves demonstrations for span detection based on the query's global semantics, and then for type classification based on the specific entity context, providing fine-grained local information. Moreover, MaRE employs entity pre-recognition to guide the construction of representations, ensuring the query and demonstrations are aligned within the LLM’s semantic space and attention pattern. In addition, to mitigate generation uncertainty, we propose RaKV, a closed-loop "generation-retrieval-verification" process. It explicates the LLM's reasoning paths, leverages them for the retrieval of external knowledge, and reorganizes the knowledge into verification evidence aligned with the original reasoning paths. We conduct extensive experiments on multiple low-resource NER datasets. Results demonstrate that URA-NER significantly enhances the performance of LLMs under low-resource settings, with particularly pronounced gains for smaller LLMs, achieving new state-of-the-art results on several benchmarks. Heterogeneous Collaboration and Instance-Adaptive Constraints for Few-Shot Nested NER Zhiqiang Liu, Dan Wu, and Juncheng Jia (Soochow University) and Yang Yang (Suzhou City University) Abstract Abstract Few-Shot Nested Named Entity Recognition (NER) is dedicated to parsing complex hierarchical entity structures in extremely low-resource scenarios. However, the two existing mainstream approaches face distinct challenges: discriminative models are highly prone to overfitting in few-shot settings, while Generative Large Language Models (LLMs), constrained by their vast decoding spaces, frequently suffer from issues of entity boundary drift and hallucination. To this end, this paper proposes a heterogeneous collaborative framework, HeCoNER, designed to leverage discriminative models to extract entity structural features to compress the decoding space of LLMs. Specifically, we first guide the discriminative model to robustly extract entity boundary features via Dynamic Prompts. Next, we utilize a Heterogeneous Feature Alignment module to filter out noise from the entity boundary features and map them into the semantic space of the LLM. Serving as instance-adaptive constraints, these features are finally injected into the LLM, transforming open-ended generation into a constrained decoding process. Experimental results demonstrate that HeCoNER achieves significant performance improvements across multiple benchmark datasets, validating the effectiveness of the framework in complex nested scenarios. Vision-to-Text: Benchmarking Multimodal LLMs on Extremely Low-Resource Languages Shuoshuo Hou, Mieradilijiang Maimaiti, Zhexin Li, Jiaxin Wang, and Ahmad Hassan (Department of Computer Science, Xinjiang University); Kaishaer Jiapaer (School of Computer Science and Technology, Xinjiang University); Nilufar Abdurakhmonova and Urinbayeva Roza (Department of Computational Linguistics and Applied Linguistics, National University of Uzbekistan); Madina Mansurova (Faculty of Information Technology and Artificial Intelligence, AI-Farabi Kazakh National University); Shormakova Assem (Department of Information Systems, AI-Farabi Kazakh National University); Gulnar Murat (Kapshagay Bidai Onimderi); Wu Le (Integrated Laboratory for Space, Air, and Ground Systems); and Wushouer Silamu (School of Computer Science and Technology, Xinjiang University) Abstract Abstract Multimodal large language models (MLLMs) have emerged as a dominant paradigm for visually grounded language tasks. However, multimodal machine translation (MMT) remains constrained by the scarcity of multimodal data for low-resource languages (LRLs). Previously proposed approaches have also explored some strategies for machine translation with MLLMs under low-resource scenarios; however, they still face the critical bottleneck of a lack of high-quality multimodal datasets for extremely low-resource languages. To this end, we propose a novel and straightforward method for constructing a high-quality benchmark for MMT in extremely low-resource languages by harnessing LLMs and leveraging a human evaluation strategy. The presented approach for building the multimodal low-resource benchmark consists of three phases. First, we select an optimal visual captioner via metric-driven evaluation to ensure reliable grounding; second, we use a hybrid ensemble of diverse translation models to reduce bias and improve the fluency–faithfulness balance; finally, we apply a triangular quality estimation scheme—reference-free quality, semantic consistency, and visual alignment—to filter candidates. We conduct extensive experiments on multiple corpora using various evaluation metrics. The results show that strong baselines (e.g., Qwen3-VL-8B and LLaVA-v1.6-7B), when fine-tuned on our proposed benchmark, achieve substantial improvements not only on our benchmark but also on well-known multimodal datasets such as Visual Genome and Multi30K, with average gains of +7.67, +8.50, and +9.67 on COMET-Kiwi, CLIP, and BERTScore, respectively. The corresponding code and constructed high-quality multimodal low-resource languages benchmark are publicly available. Plug-and-Play KV Cache Reuse for Low-Latency Retrieval-Augmented Generation Ye Yue and Zhewei Wang (Ohio University), Fabio Agosto (Air Force Materiel Command), and Jundong Liu (Ohio University) Abstract Abstract Retrieval-augmented generation (RAG) systems often suffer from high latency during the prefilling stage, when large language models compute key–value (KV) caches for long prompts. Although caching shared documents can reduce redundant computation, directly reusing precomputed KV blocks typically causes substantial accuracy loss. Wednesday Virtual Room 5 IJCNN Paper Retrieval-Augmented Generation II Session Chair: Shiyu Zhu (Beijing Institute of Graphic Commmunication), Jiaxin Wu (Beijing Institute of Graphic Commmunication) Connection-Aware Graph Pruning in RAG for Multi-Hop Question Answering Jiaxin Wu and Miao Fan (Beijing Institute of Graphic Communication), Wenjie Zhou (Renmin University of China), Chen Ma (City University of Hong Kong), Xiangzeng Liu (Xidian University), and Haoyi Xiong (Independent Scholar) Abstract Abstract Retrieval-augmented generation (RAG) for multi-hop question answering is often compromised by the low signal-to-noise ratio in retrieved contexts, necessitating effective pruning. However, standard relevance-based pruning strategies frequently discard intermediate evidence that lacks high query similarity, inadvertently severing the reasoning chain, and leading to incorrect answers. To address this limitation, we introduce a connection-aware graph-pruning framework for post-retrieval processing in RAG that explicitly preserves the evidence-chain structure under a fixed token budget. We model the retrieved context by constructing a sentence-level evidence graph over the top-N retrieved documents, in which nodes represent candidate sentences and edges encode plausible multi-hop transitions. We then formulate pruning as budgeted subgraph selection, optimizing globally coherent pathways that link question-anchored nodes to potential answer candidates. Specifically, our method uses shortest-path augmentation to recover non-obvious bridge evidence required to repair fragmented reasoning routes, followed by connectivity-preserving pruning that eliminates redundancy to maximize information density. Experiments on HotpotQA and FanOutQA demonstrate that, under identical token budgets, our framework consistently surpasses relevance-based baselines in answer accuracy and supporting-fact retention, while exhibiting superior robustness against retrieval noise. StratGraph: A Cost-Efficient Graph-RAG Framework via Structural Stratified Sampling Chongchong Yang, Chaoqian Liu, Zhiqiang Yang, and Wenhao Zhu (Shanghai University) Abstract Abstract Graph-RAG enhances LLM reasoning but incurs prohibitive indexing costs. Existing cost-efficient methods, which rely on global centrality metrics like PageRank, introduce a systematic ”Hub Bias.” This bias favors information from dominant topics at the expense of discarding critical evidence from long-tail semantic clusters. To overcome this limitation, we propose StratGraph, a novel framework that prioritizes structural diversity over centrality through its core Structural Stratified Sampling (S3) strategy. S3 enforces proportional budget allocation across semantically partitioned communities, ensuring the retention of essential ”bridging evidence.” This evidence forms a robust knowledge skeleton within StratGraph’s dual-granularity index. Extensive experiments on the HotpotQA and MuSiQue benchmarks demonstrate that StratGraph consistently outperforms centrality-based baselines under equivalent budgets. Notably, StratGraph demonstrates superior cost-efficiency, outperforming the strongest baseline even with a 20% reduced indexing budget. On the challenging MuSiQue benchmark, it achieves a 16.1% relative improvement in Exact Match, highlighting the critical role of mitigating Hub Bias for reliable multi-hop reasoning. Retrieval-Augmented Knowledge Completion for Graph-Constrained Reasoning Wen Xu and Miao Fan (Beijing Institute of Graphic Communication), Wenjie Zhou (Renmin University of China), Chen Ma (City University of Hong Kong), Xiangzeng Liu (Xidian University), and Haoyi Xiong (Independent Scholar) Abstract Abstract Graph-constrained reasoning (GCR) ensures factual correctness by enforcing structural constraints. However, it often leads to reasoning failures when the underlying knowledge graph (KG) is incomplete. To address this fundamental challenge, this paper presents a retrieval-augmented strategy to complete the KG for graph-constrained reasoning (abbr. as RA-GCR). Unlike traditional methods on GCR that tend to be limited by missing connections in KGs, RA-GCR continuously repairs the incomplete KGs with an LLM during inference. Specifically, it proceeds in two steps as follows. 1) Given the topic entities extracted from the user query, we construct a retrieval-augmented subgraph. Specifically, we synthesize this graph by merging explicit KG neighbors with implicit candidates retrieved from the global entity index. 2) On this subgraph, we perform graph-constrained reasoning. At each hop, the model selects the next entity by ranking candidates using a fused score. This score combines graph-based evidence with the LLM’s semantic confidence derived from next-token predictions. We also employ vector retrieval to identify the target entity when explicit connections are missing, ensuring a faithful and continuous reasoning path. Extensive experiments on several KGQA benchmarks demonstrate that our approach not only achieves higher accuracy compared to SOTA methods, but also maintains stable performance in sparse graph settings, proving effective for reasoning over incomplete knowledge graphs. Agentic Memory Tagging: Scalable Yet Stable Indexing for Long-Term Memory Shiyu Zhu and Miao Fan (Beijing Institute of Graphic Communication), Wenjie Zhou (Renmin University of China), Chen Ma (City University of Hong Kong), Xiangzeng Liu (Xidian University), and Haoyi Xiong (Independent Scholar) Abstract Abstract Long-term memory is essential for multi-session interaction in large language model (LLM) agents. However, existing memory mechanisms lack a definitive indexing structure, forcing them to rely on inefficient global similarity searches during the retrieval phase. As memory accumulates throughout the agent-user iterations, this lack of structure results in severe storage redundancy and fragmented indexes, which compromise the coherence of the retrieved context. To address these structural bottlenecks, we propose agentic memory tagging, a training-free framework that constructs and maintains a stable memory index. Rather than deferring organization to retrieval, the agent assigns a primary semantic tag as an explicit storage key to each incoming memory, deterministically placing it into a corresponding tag box. To maintain index stability, the agent also aligns synonymous tags to shared canonical keys and merges redundant entries within each box when their count exceeds a set threshold. This structure allows queries to be quickly directed to their relevant tag boxes for efficient access. This design ensures a stable indexing structure even as memory grows continuously. Experiments on the LOCOMO benchmark show that our approach significantly reduces redundancy and fragmentation, enabling more coherent and efficient memory retrieval. Wednesday Virtual Room 6 IJCNN Paper Robustness and Adversarial Learning I Session Chair: Baolin Yan (Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences), Shuaiye Lu (Wuhan University) Transferable Targeted Attacks on Vision–Language Models via Multi-View Feature Optimization Baolin Yan (Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences); Xiaotian Ai (Shenyang Aircraft Design and Research Institute); and Yuxi Ma, Lingzhong Meng, Youdi Gong, and Guang Yang (Institute of Software Chinese Academy of Sciences) Abstract Abstract Vision–Language Models (VLMs) have achieved substantial progress across tasks such as image captioning, visual question answering, and multimodal reasoning, and are increasingly deployed in real-world applications. Despite these advances, recent studies show that VLMs remain highly susceptible to transferable adversarial examples even in black-box settings. Existing transfer-based attack methods typically optimize adversarial perturbations under a single or limited number of views, which tends to cause overfitting to surrogate models and consequently restricts cross-model transferability. To address these limitations, we propose MVM-Attack, a transferable targeted attack framework that performs multi-view feature optimization in the visual embedding space of VLMs. Our method leverages diverse geometric view transformations to enforce consistent adversarial representations, thereby enhancing robustness against feature variations. In addition, we introduce a memory-guided ensemble weighting strategy that adaptively emphasizes shared vulnerable regions across surrogate models while stabilizing optimization using historical gradients. Extensive experiments on both open-source and closed-source VLMs show that, under the same perturbation constraints, MVM-Attack achieves substantially higher targeted attack success rates compared to multiple strong baselines. Intriguing Equivalence Structures of the Embedding Space of Vision Transformers Shaeke Salman, Montasir Shams, Chashi Mahiul Islam, Mao Nishino, and Xiuwen Liu (Florida State University) Abstract Abstract Large foundation models play a central role in the recent surge of artificial intelligence, resulting in models with remarkable abilities when measured on benchmark datasets, standard exams, and applications. Due to their inherent complexity, these models are not well understood systematically; in particular, the structures of the representation space are not well characterized despite their fundamental importance. In this paper, using the vision transformers, we show via analyses and systematic experiments that the representation space largely consists of approximately piecewise linear subspaces where there exist very different inputs sharing the same representations, and at the same time, there are visually indistinguishable inputs having very different representations. Furthermore, we show how the vulnerabilities can affect the downstream applications significantly; using image retrieval, visual question answering, and image captioning applications, our results show that the outputs can be altered arbitrarily. Therefore, understanding the inherent risks of applications built on foundational models must consider this and similar vulnerabilities. Neighbor-Aware Token Reduction via Hilbert Curve for Vision Transformers Yunge Li and Lanyu Xu (Oakland University) Abstract Abstract Vision Transformers (ViTs) have achieved remarkable success in visual recognition tasks, but redundant token representations limit their computational efficiency. Existing token merging and pruning strategies often overlook spatial continuity and neighbor relationships, resulting in the loss of local context. This paper proposes novel neighbor-aware token reduction methods based on Hilbert curve reordering, which explicitly preserves the neighbor structure in a 2D space using 1D sequential representations. Our method introduces two key strategies: Neighbor-Aware Pruning (NAP) for selective token retention and Merging by Adjacent Token similarity (MAT) for local token aggregation. Experiments demonstrate that our approach achieves state-of-the-art accuracy-efficiency trade-offs compared to existing methods. This work highlights the importance of spatial continuity and neighbor structure, offering new insights for the architectural optimization of ViTs. BS-IG: Boundary-Aware Scale-Adaptive Integrated Gradients for Transferable Adversarial Attacks Shuaiye Lu, Linjiang Zhou, Chao Ma, Xiaochuan Shi, and Yuqing Li (Wuhan University) Abstract Abstract The transferability of adversarial examples remains a critical challenge, particularly in cross-architecture scenarios (e.g., from CNNs to ViTs). Existing Integrated Gradients (IG)-based attacks attempt to bypass high-curvature regions on the surrogate manifold via macroscopic path planning, yet they neglect the ubiquitous microscopic gradient oscillations caused by local ruggedness. We theoretically identify that such oscillations trap perturbations in surrogate-specific local optima, hindering effective transfer. Furthermore, we reveal that directly applying adaptive smoothing to mitigate this noise triggers a Boundary Collapse phenomenon, where the sampling region vanishes at data boundaries, leading to optimization failure during the crucial initialization phase. To address these issues, we propose Boundary-aware Scale-adaptive Integrated Gradients (BS-IG). Our method integrates a scale-adaptive sampling mechanism with a minimum sampling region constraint, which effectively smooths high-curvature fluctuations while preventing boundary collapse along the entire integration path. Extensive experiments on ImageNet demonstrate that BS-IG significantly enhances transferability across diverse black-box models, outperforming state-of-the-art IG-based methods by up to 5.9% and other state-of-the-art attacks by an average of 2.42%. Furthermore, when combined with other IG-based methods, the performance can be boosted by 12.94%. Wednesday Virtual Room 7 IJCNN Paper Robustness and Adversarial Learning II Session Chair: Jun Li (Capital Normal University), Lixiao Liang (Guangzhou University) Selection-Aware Poisoning: Boosting Clean-Label Backdoor Attacks via Distribution Deviation and Projection Residual Metrics Xin Wang, Liming Liu, Xu Han, Yuanbo Li, Jun Li, and Lei Zhang (Capital Normal University) Abstract Abstract Despite the unprecedented success of Deep Neural Networks (DNNs) in critical domains, they remain vulnerable to clean-label backdoor attacks—stealthy threats that endanger real-world deployments. Previous work confirms uneven contributions of training samples to backdoor injection, making sample selection pivotal to attack efficacy. Existing mining methods rely on local density metrics or auxiliary out-of-distribution (OOD) data, but overlook the global statistical distribution of the target class and incur high costs from surrogate model training. To address these issues, we propose two training-free sample scoring strategies focusing on global distributional analysis rather than local heuristics: the Distribution Deviation Metric, which uses robust global covariance to quantify sample deviation from the target class, and the Projection Residual Metric, which gauges structural divergence from the intra-class manifold via subspace reconstruction errors. Extensive experiments on two widely used datasets (CIFAR-10 and GTSRB) and three representative architectures (ResNet18, VGG19, and Vision Transformer (ViT)) demonstrate that our proposed strategies achieve superior attack performance. Furthermore, the strategies effectively bypass a comprehensive set of mainstream backdoor defense mechanisms, including STRIP, Neural Cleanse, Fine-Pruning, Activation Clustering, Spectral Signatures, and SPECTRE. This work provides a novel perspective for efficient backdoor attack sample selection and offers insights into safeguarding DNNs against stealthy clean-label backdoor threats. Our code is available at https://github.com/ByteTitan-star/ddm-prm-backdoor. Activation Differences Reveal Backdoors: A Comparison of SAE Architectures Sachin Kumar (LexisNexis) Abstract Abstract Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior when triggered by specific patterns. Detecting such backdoors through mechanistic interpretability remains an open challenge. We investigate two sparse autoencoder architectures---Crosscoders and Differential SAEs (Diff-SAE)---for isolating backdoor-related features in fine-tuned models. Using a controlled SQL injection backdoor triggered by year-based context (``2024'' triggers vulnerable code, ``2023'' triggers safe code), we evaluate both approaches across LoRA and full-rank fine-tuning regimes on SmolLM2-360M. We find that Diff-SAE consistently and substantially outperforms Crosscoders for backdoor isolation. Diff-SAE achieves a Backdoor Isolation Score (BIS) of 0.40 with perfect precision (1.0) and zero false positive rate across most experimental conditions, while Crosscoders fail almost entirely with BIS below 0.02 in most cases. This performance gap holds across multiple transformer layers (14, 18, 22, 26) and both fine-tuning regimes, with full-rank fine-tuning producing particularly clean backdoor signals. Our results suggest that backdoors manifest as directional activation shifts rather than sparse feature activations, making difference-based representations fundamentally more effective for detection. These findings have important implications for AI safety monitoring and the development of interpretability tools for detecting model manipulation. Lite-BD: A Lightweight Black-box Backdoor Defense via Reviving Multi-Stage Image Transformations Abdullah Arafat Miah and Yu Bi (University of Rhode Island) Abstract Abstract Deep Neural Networks (DNNs) are vulnerable to backdoor attacks. Due to the nature of Machine Learning as a Service (MLaaS) applications, black-box defenses are more practical than white-box methods, yet existing purification techniques suffer from key limitations: a lack of justification for specific transformations, dataset dependency, high computational overhead, and a neglect of frequency-domain transformations. This paper conducts a preliminary study on various image transformations, identifying down-upscaling as the most effective backdoor trigger disruption technique. We subsequently propose \texttt{Lite-BD}, a lightweight two-stage blackbox backdoor defense. \texttt{Lite-BD} first employs a super-resolution-based down-upscaling stage to neutralize spatial triggers. A secondary stage utilizes query-based band-by-band frequency filtering to remove triggers hidden in specific bands. Extensive experiments against state-of-the-art attacks demonstrate that \texttt{Lite-BD} provides robust and efficient protection. Codes can be found at \url{https://github.com/SiSL-URI/Lite-BD}. A Neural Model of Crossover Inhibition for Looming Motion Detection Lixiao Liang (Guangzhou University), Ziyan Qin (Guangdong Polytechnic Normal University), Shaobing Gao (Sichuan University), and Qinbing Fu (Guangzhou University) Abstract Abstract Looming motion detection is a prerequisite for collision avoidance of both biological and artificial entities. In primate retinas, crossover inhibition could play a key role in detecting looming motion by mediating inhibitory interactions between parallel ON and OFF pathways. This mechanism enhances sensitivity to transient luminance changes and supports reliable perception of approaching objects. However, most existing bio-inspired models either neglect or oversimplify crossover inhibition, which limits their ability to distinguish approaching motion from receding or translating motion and leads to unreliable responses in noisy environments. To address this limitation, we propose a neural model of crossover inhibition for looming motion detection. The model adopts a biologically plausible ON/OFF-pathway framework encoding brightness increment (ON)/decrement (OFF), separately. It firstly enhances motion edges using a spatial operator of Difference of Gaussians, then incorporates crossover inhibition to regulate interactions between ON-OFF/OFF-ON pathways, suppressing redundant luminance fluctuations and non-approaching motion responses. Experimental results show that the proposed model achieves strong selectivity for approaching motion while effectively inhibiting responses to receding and translating stimuli. Quantitative evaluations demonstrate performance improvements of 10--22% over comparative bio-inspired models in real-world scenarios, with stable behavior maintained under severe synthetic noise. These findings confirm the functional importance of crossover inhibition in looming motion detection and demonstrate its potential for improving robustness in machine vision systems. Wednesday Virtual Room 8 IJCNN Paper Semi-Supervised and Weakly-Supervised Learning I Session Chair: Donghai Zhai (Southwest Jiaotong University), Xinyu Li (Beijing Jiaotong University) Enhancing Unlabeled Data Efficiency for Semi-Supervised Domain Generalization via Reliability-Guided Energy Prior Panpan Ni, Kunlun Wu, and Donghai Zhai (Southwest Jiaotong University) Abstract Abstract This paper conducts an in-depth study of domain generalization (DG) to address distribution shifts when applying a model trained on a series of known target domains to unknown target domains. However, existing DG methods heavily rely on labels and high-confidence data, suffering from data efficiency issues and incurring expensive training costs. In this paper, we propose a novel multi-task semi-supervised domain generalization (RPCW) method that relaxes the strict requirement for plentiful labeled data in source domains while emphasizing superior generalization performance in unknown domains. This method comprises a contrastive task guided by reliability with energy prior and a semi-supervised task weighted by domain consistency and coherence, respectively, promoting the model to learn task-related features from unlabeled-confident and unlabeled-unconfident data and thereby jointly optimizing the covariate shift dilemma faced by the feature extractor. Experiments on two standard domain generalization datasets and a heterogeneous geological disaster dataset demonstrate that RPCW achieves better accuracy and generalization performance compared to state-of-the-art methods in scenarios with sparse labels. The code is available at https://github.com/2PNi/RPCW. Towards Adapting Vision-Language Models for Semi-Supervised Domain Generalization Xilin He (MBZUAI), XIaole Xian (Shenzhen University), Zhihua Liu (Edinburg University), Weicheng Xie (Shenzhen University), Xiangyu Yue (Chinese University of Hong Kong), and Muhammad Haris Khan (MBZUAI) Abstract Abstract Semi-supervised Domain Generalization (SSDG) offers a cost-effective solution for generalizing models to unseen domains with limited labels. While existing SSDG methods, mainly built upon small-scale backbones, struggle to match fully supervised DG performance, large-scale vision-language models like CLIP have shown remarkable generalization through downstream fine-tuning. However, adapting these models to SSDG remains underexplored. In this paper, we identify a critical issue: existing popular fine-tuning methods suffer from under-utilizing unlabeled data in the semi-supervised learning frameworks, thereby overfitting the limited labeled data, leading to training collapse and generalization ability degradation. To address these challenges, we propose two novel components: (1) the De-False-Correlation Adapter (DFC-Adapter), which reduces false correlations to refine visual features and (2) Learnable Multi-granularity Text-guided Embedding Augmentation (LMTEA), which synthesizes semantic-aligned but domain-perturbed augmented visual embedding for consistency regularization through multi-granularity text guidance and learnable style encoding. Moreover, we establish the first-ever benchmark for CLIP fine-tuning methods in SSDG, conducting extensive experiments across six DG datasets and two ImageNet variants. Our results demonstrate that our method significantly outperforms existing CLIP fine-tuning approaches and achieves performance comparable to even fully supervised DG methods in some cases. SABTM: Semi-Supervised Multimodal Emotion Recognition via Adaptive Barlow Twins and Mixture-of-Experts Tiantai zhai, Liang Luo, Yan Zhuang, and Fuji Ren (University of Electronic Science and Technology of China) Abstract Abstract Semi-supervised multimodal emotion recognition faces two critical challenges: inadequate disentanglement of shared versus modality-specific features, and modality dominance under limited labeled data. These issues lead to information loss and reduced performance. To address these problems, we propose Semi-Supervised Multimodal Emotion Recognition via Adaptive Barlow Twins and Mixture-of-Experts (SABTM). Specifically, SABTM comprises three key components: (1) Adaptive Barlow Twins (ABT) that dynamically disentangles shared and modality-specific representations, (2) Sample-wise Fusion Mixture-of-Experts (SF-MoE) module employs dynamic routing and load balancing to generate sample-specific weights, mitigating modality dominance while achieving dual-pathway fusion through feature-level and decision-level fusion to capture cross-modality interactions, and (3) Dual-Strategy Consistency Learning that enforces prediction coherence and improves pseudo-label reliability. Extensive experiments on CH-SIMS 2.0, CMU-MOSI, and CMU-MOSEI demonstrate state-of-the-art performance, with accuracy improvements of 4.62% on CH-SIMS 2.0, and 1.28% and 1.93% in 7-class accuracy on CMU-MOSI and CMU-MOSEI respectively, compared to supervised and semi-supervised learning methods. ICoRe-SSL: Information- and Confusion-Regularized Semi-Supervised Learning for Long-Tailed Recognition Xinyu Li (School of Software Engineering, Beijing Jiaotong University); Xiaoyu Guo (Beijing Institute of Astronautical Systems Engineering); and Xiang Wei (School of Software Engineering, Beijing Jiaotong University) Abstract Abstract Class-imbalanced semi-supervised learning (CISSL) aims to exploit abundant unlabeled data under long-tailed class distributions, but is often hindered by biased training dynamics and unreliable pseudo-labels, where head classes dominate optimization and errors on tail classes accumulate over training. Most existing methods address this issue through reweighting schemes or adaptive pseudo-label selection, while largely overlooking how redundant sample participation and opaque class-wise prediction dynamics jointly distort the optimization process under long-tailed distributions. We propose ICoRe-SSL, an Information- and Confusion-Regularized Semi-Supervised Learning framework that tackles these limitations by jointly reshaping optimization dynamics and enhancing class-wise interpretability. ICoRe-SSL introduces an information-aware sample participation mechanism that suppresses low-information samples with expectation-preserving rescaling, reducing redundant head-class updates, and incorporates a confusion-aware analysis module to explicitly track pseudo-label statistics and confusion evolution during training. As a result, ICoRe-SSL maintains more balanced pseudo-label distributions and enables stable tail-class learning. Extensive experiments on standard long-tailed benchmarks demonstrate that ICoRe-SSL consistently outperforms state-of-the-art semi-supervised methods, achieving significant gains in balanced accuracy and geometric mean with negligible computational overhead. Wednesday Virtual Room 1 IJCNN Paper Semi-Supervised and Weakly-Supervised Learning II Session Chair: Weiliang Meng (Chinese Academy of Sciences, Institute of Automation), Zhenming Yuan (Hangzhou Normal University, Hangzhou Hele Tech Co.LTD) Mamba-CAM: Mamba-Driven Class Activation Map Learning for 3D Medical Weakly-Supervised Segmentation Xiaofeng Han (Chinese Academy of Sciences, Institute of Automation); Ganggang Liu (Shanghai Institute of Aerospace Technical Foundation); Shunpeng Chen (Beijing University of Posts and Telecommunications); Mingming Yu (Beihang University); Duzhen Zhang (Institute of Automation, Chinese Academy of Sciences); Xinchen Li (Tongji University); Changwei Wang (Qilu University of Technology); and Weiliang Meng and Xiaopeng Zhang (Chinese Academy of Sciences, Institute of Automation) Abstract Abstract Weakly supervised medical image segmentation methods commonly rely on convolutional neural networks (CNNs) and class activation maps (CAMs). However, the inherent locality of CNNs often leads to suboptimal CAM quality and degraded segmentation performance. To overcome this limitation, we propose Mamba-CAM, a novel weakly supervised segmentation framework that integrates CAMs with the Mamba model. Leveraging State Space Models (SSMs) for efficient long-range contextual modeling, our Mamba-CAM generates more reliable pseudo-labels for training. To further enhance interpretability and detail preservation, we design a detail-aware visual state space (DVSS) block. Moreover, a correspondence learning mechanism is provided to refine pseudo-labels and improve feature representation. To the best of our knowledge, this is the first work to integrate Mamba with CAMs for medical image segmentation. Extensive experiments on the BraTS dataset demonstrate that our Mamba-CAM achieves state-of-the-art performance, highlighting its potential to advance medical image analysis while reducing annotation demands. A Novel Evaluation Metric for Class Activation Mapping Methods in Weakly Supervised Learning Qingdong Cai and Charith Abhayaratne (The University of Sheffield) Abstract Abstract Class activation mapping (CAM) methods are a fundamental component of weakly supervised learning tasks, and research on CAMs specifically tailored for these tasks has been steadily increasing. However, existing evaluation metrics primarily assess the faithfulness and interpretability of CAM algorithms. These metrics do not adequately assess whether an activation method can precisely activate the true target region and produce sufficiently large activation values, both of which are crucial for downstream weakly supervised learning tasks. Furthermore, CAM performance is also evaluated using metrics derived from downstream tasks; however, these metrics depend on manually selected thresholds, which can result in substantially different, or even contradictory, evaluation outcomes across thresholds. This threshold sensitivity introduces the risk of implicit cherry-picking, thereby undermining the fairness, robustness, and reproducibility of CAM evaluation. To address the inappropriateness and threshold sensitivity for existing metrics, we propose a new metric, termed Mask-Based Intersection Energy (MBIE), which is threshold-insensitive and leverages object masks to more accurately evaluate the activation outcomes of CAM methods. DCLIP:A Robust Dual-Modal Complementary Backbone for Weakly Supervised Semantic Segmentation Hongyang Chen, Lei Wang, and Hengyang Wang (Heilongjiang University) Abstract Abstract Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). In recent years, Contrastive Language-Image Pretraining (CLIP) has been introduced into WSSS to provide more accurate and robust semantic guidance for class localization. However, most existing approaches focus on image-text alignment at the CAM level. Although CLIP excels at global semantic matching, it struggles to extract sufficiently fine-grained local semantics from images. To address this issue, we propose the DCLIP framework, which improves CLIP’s image-text alignment process and compensates for its limited dense semantic modeling capability by integrating DINO’s stronger unsupervised spatial perception ability while optimizing CLIP’s own textual attribute descriptions. Specifically, we first introduce a Bimodal Complementary Learning module (BCL) that employs frozen DINO and frozen CLIP as parallel backbones, enabling more comprehensive yet low-cost semantic feature extraction through a shared decoder. Second, to further unlock CLIP’s dense knowledge and improve alignment quality, we design a Knowledge Graph-based text augmentation module (KG) that extends textual semantics with fine-grained attributes, thereby enhancing the mapping accuracy between visual and textual modalities. Extensive experiments demonstrate the effectiveness of DCLIP. While maintaining an end-to-end architecture and low training cost, DCLIP achieves $79.6\%$ mIoU on the PASCAL VOC 2012 test set and $51.4\%$ mIoU on the MS COCO val set, significantly outperforming existing methods. ResFTNet: Clinical Pregnancy Outcome Prediction via Multimodal Feature Fusion for Assisted Reproductive Technology Yang Wang (School of Information Science and Technology, Hangzhou Normal University); Xiaomei Tong (Assisted Reproduction Unit, Department of Obstetrics and Gynecology, Sir Run Run Shaw Hospital, Zhejiang University School of Medicine); and Zhenming Yuan (School of Information Science and Technology, Hangzhou Normal University; Hangzhou Hele Tech Co.LTD) Abstract Abstract Accurately predicting clinical pregnancy outcomes in assisted reproductive technology (ART) remains a critical and challenging problem, due to the complex interplay between embryo morphological characteristics and patient-specific clinical factors. This study proposes ResFTNet, an end-to-end multimodal fusion network for predicting clinical pregnancy outcomes (i.e., successful versus unsuccessful pregnancy after embryo transfer). In the image branch, a residual convolutional backbone enhanced with omni-dimensional dynamic convolution and squeeze-enhanced axial attention is employed to capture both fine-grained local morphology and global semantic representations from microscopic embryo images. In parallel, the tabular branch adopts an FT-Transformer to model nonlinear dependencies among structured clinical indicators. Based on these modality-specific encoders, ResFTNet jointly integrates visual and clinical representations through a Clinical Context Injection (CCI) strategy, which injects patient-specific clinical information into the visual feature space to alleviate cross-modal semantic misalignment. Experimental results on a private clinical pregnancy dataset denoted as the CP-ART dataset and a public blastocyst dataset demonstrate that ResFTNet consistently outperforms existing unimodal and multimodal approaches in terms of accuracy and robustness. Furthermore, interpretability analyses using Grad-CAM and SHAP indicate that the proposed model focuses on biologically meaningful embryonic regions and clinically relevant features, highlighting its potential as an effective and explainable decision-support tool for personalized treatment planning in reproductive medicine. Wednesday Virtual Room 2 IJCNN Paper Spatiotemporal Learning and Prediction I Session Chair: Ziyi Li (xinjiang university), Zhenli Qian (Shanghai Ocean University) MF-STGCN: Multimodal Fusion Sparse Spatio-Temporal Graph Networks with Physics-Aware Gating for Ocean Eddy Forecasting Yanling Du, Zhenli Qian, and Zhihong Huang (Shanghai Ocean University); Ke Chen (East China Sea Forecasting and Disaster Reduction Center, Ministry of Natural Resources); Dongmei Huang (College of Electrical Engineering, Shanghai University of Electric Power); and Qi He (Shanghai Ocean University) Abstract Abstract Accurate prediction of the trajectories of mesoscale ocean eddies is imperative for comprehending oceanic dynamic processes and ensuring offshore engineering safety. However, extant methodologies generally treat eddies as isolated systems, thereby constraining their capacity to capture the intricate and sparse interactions among multiple coexisting eddies. Moreover, unimodal trajectory data frequently prove inadequate in fully characterizing the nonlinear evolutionary patterns of eddies, which are subject to modulation by their intrinsic physical features as well as the surrounding environmental context. To address these challenges, this paper proposes the multimodal fusion sparse spatio-temporal graph convolutional network (MF-STGCN). Initially, the model designs a physical encoder and an adaptive gating fusion mechanism to dynamically adjust the contribution weights of physical features and environmental information. This approach effectively addresses the issues of multivariate heterogeneous data alignment and feature enhancement. Secondly, a sparse graph learning strategy is proposed based on asymmetric convolution and zero-softmax normalization. It enables the automatic mining of directed interactions with physical significance while effectively eliminating redundant noise. Finally, spatio-temporal features are aggregated in parallel through a dual-stream spatio-temporal graph convolutional network (STGCN) and a temporal convolutional network (TCN) to achieve trajectory prediction. Evaluated on the 1993–2022 South China Sea (SCS) mesoscale eddy dataset, MF-STGCN outperforms Transformer, ETPNet, and existing GNN baselines across multiple metrics. Visualization analysis further confirms the model's remarkable robustness and superiority in handling abrupt trajectory shifts and long-term predictions. End-to-End Flight Schedule Recovery via Decision-Focused Spatiotemporal Graph Networks Jianli Ding, Yifu Zhong, and Jing Li (Civil Aviation University of China) Abstract Abstract Flight disruptions cost airlines billions annually, yet current recovery methods struggle to balance computational tractability with operational feasibility. Traditional optimization approaches guarantee constraint satisfaction but timeout on large-scale networks, while machine learning methods achieve high prediction accuracy but generate infeasible schedules in over one-third of cases. A Generalized Multivariate Flight Data Prediction Framework with Dynamic Feature Decomposition Tao Hong, Qiaoli Zhou, Zhiqiang Chang, and Bin Guo (Shenyang Aerospace University) Abstract Abstract Accurate prediction of multivariate flight data is essential for aviation safety and operational optimization. However, flight sensor signals exhibit strong nonstationarity caused by environmental variations, abrupt maneuvers, and sensor noise, making it difficult for conventional models to capture both long-term trends and short-term fluctuations. To address this challenge, we propose a unified regression framework based on dynamic feature decomposition. The model adaptively separates multivariate time series into low-frequency trends and high-frequency fluctuations using a dual-branch architecture, where an auxiliary compensation branch enhances the representation of transient dynamics. A differencing-based preprocessing strategy improves sensitivity to rapid changes, while dynamic loss weighting balances learning across heterogeneous flight variables. In addition, a sigmoid-based normalization scheme stabilizes training under diverse feature distributions. Experiments on real-world flight datasets demonstrate consistent improvements or competitive performance against strong baselines across multiple flight-related variables, validating the robustness and generalization capability of the proposed method. GraFSTNet: Graph-based Frequency SpatioTemporal Network for Cellular Traffic Prediction Ziyi Li and Hui Ma (xinjiang university), Fei Xing (Xinjiang University), Chunjiong Zhang (Ajou University), and Ming Yan (Xinjiang University) Abstract Abstract With rapid expansion of cellular networks and the proliferation of mobile devices, cellular traffic data exhibits complex temporal dynamics and spatial correlations, posing challenges to accurate traffic prediction. Previous methods often focus predominantly on temporal modeling or depend on predefined spatial topologies, which limits their ability to jointly model spatio-temporal dependencies and effectively capture periodic patterns in cellular traffic. To address these issues, we propose a cellular traffic prediction framework that integrates spatio-temporal modeling with time–frequency analysis. First, we construct a spatial modeling branch to capture inter-cell dependencies through an attention mechanism, minimizing the reliance on predefined topological structures. Second, we build a time-frequency modeling branch to enhance the representation of periodic patterns. Furthermore, we introduce an adaptive-scale LogCosh loss function, which adjusts the error penalty based on traffic magnitude, preventing large errors from dominating the training process and helping the model maintain relatively stable prediction accuracy across different traffic intensities. Experiments on three open-sourced datasets demonstrate that the proposed method achieves prediction performance superior to state-of-the-art approaches. Wednesday Virtual Room 3 IJCNN Paper Time Series Forecasting I Session Chair: Mingshan Du (Chengdu University of Technology), Yiqi Yu (Zhejiang University) PMT-Former: Phase-Aware Multi-Scale Temporal Transformer for Robust Time Series Forecasting Mingshan Du and Ting Hou (Chengdu University of Technology) Abstract Abstract Time Series forecasting under complex and non-stationary dynamics remains challenging due to phase shifts, cross-scale coupling, and weak structure awareness in conventional Transformer models. To address these issues, we propose the Phase-Aware Multi-Scale Temporal Transformer (PMT-Former), a unified framework that integrates multi-scale rhythmic positional encoding, dual-attention fusion, and frequency-domain regularization for robust long-horizon forecasting. The proposed Multi-Scale Rhythmic Positional Encoding (MRPE) injects hierarchical periodic priors (e.g., daily and weekly rhythms) into the temporal embeddings, enabling the model to achieve phase-invariant and interpretable representations. A dual-attention fusion mechanism further coordinates local and global dependencies via dynamically gated cross-level attention, while a rhythmic alignment regularizer enforces temporal–spectral consistency between predicted and ground-truth signals. Extensive experiments on seven public benchmarks across energy, climate, and traffic domains demonstrate that PMT-Former achieves up to 12% lower MSE and 10% lower MAE than state-of-the-art Transformer, Mixer, and Linear baselines, while exhibiting superior robustness to temporal phase perturbations. This work establishes a rhythm-aware and phase-robust paradigm for long-term Time Series forecasting. DPTF: Dual-Path Time-Frequency Framework for Long-Term Time Series Forecasting Pengfei Ding, Wei Ding, Zehang Chen, Jingqi Gao, Lvqing Yang, and Fan Lin (xiamen university) Abstract Abstract Long-term time series forecasting plays a vital role in domains such as energy, transportation, and meteorology. However, real-world time series often exhibit fine-grained temporal variations alongside coarse-grained spectral patterns. Constrained by single-domain modeling paradigms, existing approaches fail to simultaneously capture these distinct yet complementary features, leading to sub-optimal performance. Motivated by the need to synergize these complementary perspectives, we propose the Dual-Path Time-Frequency (DPTF) framework, a novel architecture that effectively integrates local and global modeling. Specifically, the temporal branch adopts a hierarchical multi-scale multi-period design to capture the series' trend and seasonal components, thereby explicitly disentangling and fusing local features across different granularities and compositions. In parallel, the frequency branch employs a lightweight spectral transformation to extract global structural patterns. Finally, the two pathways are adaptively aggregated through a learnable weighting mechanism to produce robust forecasts. Extensive experiments on public benchmarks show that DPTF consistently outperforms state-of-the-art baselines, demonstrating superior accuracy and robustness in long-term forecasting tasks. RaPatch: Ratio-Based Patching with Spectral Fusion for Long-Term Time Series Forecasting Yiding Fu, Biao Luo, and Shengjie Gao (Hefei University of Technology) Abstract Abstract Scaling patch-based forecasting to long input sequences is sensitive to the choice of patch granularity because the appropriate temporal resolution changes with both context length and forecasting horizon, making fixed patch sizes or hand-picked scales suboptimal. We propose RaPatch, a ratio-based multi-scale patching framework that learns a set of bounded patch-length ratios relative to the input length, producing multi-scale patch tokens and a ratio-based scale descriptor to condition downstream spectral modeling. By parameterizing patch lengths as ratios of the input window, the model reduces horizon-specific patch retuning and maintains stable accuracy across different look-back lengths and forecasting horizons. We further introduce MS-SFB, a frequency-domain module that aligns multi-scale spectra onto a shared frequency grid and performs ratio-conditioned filtering together with cross-variable spectral mixing to better capture periodic structure. Results on multiple public long-term forecasting benchmarks indicate that RaPatch reduces forecasting errors relative to strong patch-based and frequency-aware baselines, and the advantage is more pronounced at longer prediction horizons. DDM-Mixer: Dual-Domain Multi-scale Mixer for Time Series Forecasting Yiqi Yu and Xuelin Cheng (Zhejiang University) Abstract Abstract Time series forecasting greatly benefits from synergistically fusing time-domain (local temporal dependencies) and frequency-domain (global periodic) representations. However, effectively integrating these spaces is often hindered by the ``Spectral Geometric Mismatch'' problem: when the historical input window and the desired output horizons differ, standard spectral features suffer from severe aliasing during mapping. To overcome this fundamental bottleneck, we propose the \textbf{DDM-Mixer (Dual-Domain Multi-scale Mixer)}, utilizing a \textbf{High-Density Spectral Interpolation} strategy via Extended DFT to rigorously align mismatched spectral dimensions into a unified representation. Furthermore, to effectively distinguish true, predictable periodicity from chaotic stochastic noise and prevent overfitting to non-periodic elements, we introduce an adaptive \textbf{Parameter-Free Information-Theoretic Gating Mechanism} based on Spectral Entropy. Extensive experiments conducted across seven diverse real-world benchmark datasets demonstrate that DDM-Mixer consistently achieves state-of-the-art performance, showcasing superior robustness in complex forecasting scenarios. Wednesday Virtual Room 4 IJCNN Paper Time Series and Temporal Modeling I Session Chair: Kasun Dewage (university of Central Florida), Jianbin gao (University of Electronic Science and Technology of China) Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models: A Comprehensive Ablation Study on High-Frequency Stock Prediction Kasun Dewage, Suranadi Dodampagamage, and Shankhadeep Mondal (university of Central Florida) Abstract Abstract Foundation models for time series forecasting demonstrate impressive zero-shot generalization but often underperform on specialized domains such as high-frequency finance. We present a comprehensive study of hybrid neural-classical correction for adapting frozen TimesFM (200M parameters) to stock return prediction during the volatile opening trading hour. We compare two neural correction architectures—AttnCorrect (multi-head self-attention, ~471K parameters) and GatedLinear (low-rank bilinear projection with gating, ~49K parameters)—each augmented with Random Forest residual learning. Through systematic ablation across 10 major technology stocks (NVDA, MSFT, AAPL, GOOG, GOOGL, AMZN, META, AVGO, TSLA, NFLX) spanning 2 million data points, we reveal critical insights: (1) The hybrid neural-classical approach achieves 0.597 pooled correlation and 6.4x mean per-day correlation improvement over frozen TimesFM; (2) Classical residual learning (Random Forest) provides the largest single-component contribution, matching or exceeding the neural correction component; (3) Simpler neural architectures surprisingly outperform complex ones when classical residual learning is removed; (4) Self-attention provides the largest neural-only contribution. GatedLinear+RF achieves best overall performance with 9x fewer neural parameters than AttnCorrect+RF. We report three complementary correlation metrics—mean per-day, cross-day cumulative, and pooled—to provide a complete picture of predictive quality. Our results provide practical guidance: effective foundation model adaptation requires careful integration of neural and classical components, with classical methods playing a crucial complementary role. Decomposing the Time Series Forecasting Pipeline: A Modular Approach for Time Series Representation, Information Extraction, and Projection Robert Leppich and Michael Stenger (University of Wuerzburg), André Bauer (Illinois Institute of Technology), and Daniel Grillmeyer and Samuel Kounev (University of Wuerzburg) Abstract Abstract Time series forecasting demands effective sequence representation, meaningful information extraction, and precise future projection. This work decomposes the forecasting pipeline into three core stages: input sequence representation, information extraction and memory construction, and target projection. Within each stage, we investigate architectural configurations to assess the effectiveness of various modules, such as convolutional layers and self-attention mechanisms, across diverse forecasting tasks on seven benchmark datasets. Our models achieve state-of-the-art forecasting accuracy on the majority of benchmarks while maintaining a competitive parameter count and inference speed relative to strong baselines. The source code is publicly available at https://github.com/RobertLeppich/REP-Net C2FAD: A Coarse-to-Fine Anomaly Detection for Time Series Wei Xia (Chongqing Normal University); Wang Chao (Chongqing Normal University; Yunnan Key Laboratory of Software Engineering, Yunnan University); and Songbo Su and Haoyue Zheng (Chongqing Normal University) Abstract Abstract Deep learning-based methods dominate time series anomaly detection. However, these methods typically rely on a single detection paradigm, which struggle to achieve an optimal balance between macro trends and micro fluctuations. To address this issue, this paper simulates a coarse-to-fine analysis logic and proposes a novel coarse-to-fine time series anomaly detection model (C2FAD). The first stage of this model employs a coarse-grained anomaly collector, CgAC, which efficiently filters out significant and coarse-grained macroscopic anomalies by capturing long-term dependencies and global context between sequences. In the second stage, we designed a fine-grained anomaly collector, FgAC, which extracts microscopic morphological features within local windows, deeply analyzes the local context, and identifies complex fine-grained anomalies that are not obvious in the global view but are incompatible with neighboring patterns. We conducted comparative experiments on seven datasets with ten mainstream anomaly detection methods and the proposed C2FAD. Experiments show that compared to the best baseline model, C2FAD improves the AUC score by 6.14%, F1 score by 9.02\%, and AFF-F1 score by 3.81%. Hybrid dynamic graph networks for vulnerability detection in blockchain transactions Hu Xia, Toussaint Elvis Compaore, Jianbin Gao, Qi Xia, Ansu Badjie, and Radwa El Bourari (University of Electronic Science and Technology of China) Abstract Abstract Blockchain technology enables decentralized digital transactions through peer-to-peer exchanges, yet this architecture introduces transaction-level security vulnerabilities exploited by adversaries for fraudulent activities. Existing graph neural network approaches for vulnerability detection exhibit fundamental limitations: static graph analysis neglects temporal transaction evolution, conventional architectures struggle with long-range dependencies, sender-receiver structural asymmetries degrade neighborhood aggregation, and comprehensive adversarial robustness evaluation remains absent. This paper proposes hybrid dynamic graph networks (HDGN), an intelligent system integrating graph attention networks, graph convolutional networks, and bidirectional encoder representations from transformers (BERT) transformers within a unified end-to-end pipeline for robust vulnerability detection in blockchain transaction networks. The framework introduces a missing information prediction module addressing sender-receiver asymmetry through role-aware differential estimation and node-specific affine transformations. A dual vulnerability scoring scheme classifies transactions into four behavioral patterns, enabling fine-grained, interpretable security assessments. Comprehensive adversarial evaluation examines robustness under perturbation edge attack, poisoning attack, and temporal evasion scenarios. Extensive experiments across four blockchain datasets demonstrate HDGN's effectiveness, achieving 93.5% average accuracy (97.7% on Bitcoin Alpha), 95.3% F1-score, and 94.7% AUC while consistently outperforming five state-of-the-art baselines. The framework exhibits strong resilience to structural perturbation attacks with accuracy degradation below 1.0 percentage points. Ablation studies confirm meaningful contributions from each architectural component, with graph attention networks providing the largest performance impact. HDGN advances blockchain security through a robust, interpretable intelligent system demonstrating potential for deployment pending scalability validation. Wednesday Virtual Room 5 IJCNN Paper Vision Domain Adaptation and Generalization I Session Chair: 劭卫 范 (河北工业大学), Zhikui Chen (Dalian University of Technology) Multi-View Structural Consistency Learning for Test-Time Adaptation in Medical Image Segmentation Shaowei Fan (Hebei University of Technology) Abstract Abstract Cross-domain distribution shifts substantially degrade the reliability of deep fundus image segmentation models in real clinical deployments. Test-time adaptation (TTA) offers an appealing solution by updating models online using unlabeled target data, yet existing segmentation TTA methods often rely on probability-only objectives that can over-smooth uncertain boundaries and may suffer from semantic drift under unconstrained updates. We propose \textbf{Multi-View Structural Consistency Learning (MVSCL)}, a structure-aware and parameter-constrained TTA framework for joint optic disc and optic cup segmentation. MVSCL constructs multiple geometric views of each test image via flips and multi-scale resizing, and enforces consistency under a teacher--student formulation across output probabilities, Sobel-based boundary responses, and decoder feature representations, thereby preserving anatomically meaningful morphology under domain shift. To stabilize online adaptation, we further restrict optimization to the Batch Normalization affine parameters in the decoder and segmentation head, while keeping the encoder frozen to retain generalizable representations. Extensive experiments on five public fundus datasets under a single-source training and multi-target testing protocol demonstrate that MVSCL consistently improves robustness and achieves state-of-the-art performance on the evaluated cross-domain benchmarks without requiring source data.The code is available at https://github.com/gelbartschiff-cloud/MVSCL.git Feature-Space Consistency for Semantics-Preserving Test-Time Adaptation in Medical Segmentation Jianyu Fang (Xi'an Jiaotong University) and Chengyuan Deng (Zhejiang University, Xi'an Jiaotong University) Abstract Abstract Test-time adaptation (TTA) offers a practical solution for deploying medical image segmentation models under cross-center domain shifts, where source data are often inaccessible and target annotations are unavailable. However, existing TTA approaches can be unstable in safety-critical settings due to semantic drift in latent representations induced by strong augmentations and confirmation bias from unreliable pseudo supervision, which together lead to noisy updates and error accumulation under continual exposure. In this paper, we propose a reliability-aware and semantics-preserving TTA framework for medical image segmentation. Our method introduces a feature-space consistency constraint that regularizes representation deviation between an input and its strongly augmented view, improving anatomical semantic preservation beyond output-level consistency alone. To further stabilize optimization, we incorporate an entropy-guided gradient-alignment regularization. Moreover, we develop a confidence-aware two-stage adaptation strategy that performs trusted updates first and conditionally refines on hard samples with an early-termination safeguard when reliable evidence is insufficient. We evaluate our approach on cross-domain optic disc/cup segmentation across five public retinal fundus datasets under an online, batch-of-one adaptation protocol that updates only BatchNorm affine parameters. Extensive experiments over 20 cross-domain transfer scenarios demonstrate consistent improvements over prior TTA baselines, achieving state-of-the-art performance under realistic cross-center shifts. Codes available https://github.com/JyFang999/Feature-Space-Consistency.git AIM: Augmentation-Invariant Memory for Robust Open-set Test-Time Adaptation Can Zhang and Yunfeng Liu (Beijing University of Chemical Technology), Ye Wang (Chongqing University of Posts and Telecommunications), and Ruirui Li (Beijing University of Chemical Technology) Abstract Abstract Vision-language models (VLMs) suffer severe performance degradation when the test distribution deviates from the pre-training data. Test-time adaptation (TTA) is an effective paradigm to mitigate such distribution shifts, among which cache-based methods have drawn increasing attention due to their training-free property and remarkable efficacy. However, these methods are vulnerable in open-set scenarios, where out-of-distribution (OOD) samples readily poison the cache and mislead subsequent predictions. We present AIM (Augmentation-Invariant Memory), a robust, open-set TTA framework that equips a VLM with an on-the-fly key-value cache. AIM admits a test sample only when its prediction is both high-confidence and consistent across augmentations, guaranteeing cache purity. Leveraging this clean memory, AIM refines logits via cache retrieval and simultaneously separates ID/OOD data through a lightweight linear head trained with progressively self-thresholded pseudo-labels. Evaluated on 20 ID-OOD pairs, AIM establishes a new state-of-the-art, raising mean harmonic accuracy from 66.1% to 70.5% on CUB-200-2011. Multi-Granularity Self-Learning Guided Deep Multi-View Clustering Zhikui Chen and Yuzhe Li (Dalian University of Technology) Abstract Abstract Although deep multi-view clustering achieves encouraging performance, there still remain two issues. (1) Most methods commonly utilize inter-view invariance to aggregate information across views, which results in the loss of useful complementary information in each view. (2) Most methods do not fully model inherent self-learning signals, resulting in suboptimal clustering results. To this end, this paper proposes a Multi-granularity Self-learning Guided Multi-View Clustering (MSG-MVC). Specifically, MSG-MVC devises a sample-guided information compression module to capture discriminant view-specific representations via self-learning compression with sample manifold preservation. Then, it designs representation-guided complementary integration to learn semantics-robust fusion representations via self-learning transformation with the conditional entropy minimization. Meanwhile, it introduces a cluster-guided consistent alignment module to capture a consensus cluster assignment from representations via self-learning partition with the multi-level semantic invariance. Experiments on eight datasets show that MSG-MVC achieves cutting-edge performance. The code is available at https://github.com/Yuzhe-Li123/MSG-MVC. Wednesday Virtual Room 6 IJCNN Paper Vision Domain Adaptation and Generalization II Session Chair: Yufan Yi (Huazhong University of Science and Technology), Huiwen Huang (Guilin University of Electronic Technology, School of Computer and Information Security) ID3A: A Multi-Source Domain Adaptation Framework with Identity Disentanglement for Cross-Subject Emotion Recognition Yufan Yi, Yan Tian, and Yiping Xu (Huazhong University of Science and Technology) and Bo Yang (North China Institute of Science and Technology) Abstract Abstract Cross-subject EEG-based emotion recognition remains a long-standing and challenging task due to the significant domain shifts in EEG signals across individuals, which severely hinders model generalization. To address this issue, we propose ID3A, a novel identity-disentangled multi-source domain adaptation framework. First, the ID3A model performs explicit disentanglement of identity-dependent and identity-independent features via an identity disentanglement (ID) module. Then, a subject domain adaptation (SDA) module maps identity-dependent features to subject-specific representations and aligns them across domains. Meanwhile, an emotion domain adaptation (EDA) module projects identity-independent features into the emotion space and performs cross-domain alignment of emotion-related features. Finally, the fused emotion features and subject-level features are combined to perform the final emotion recognition task. ID3A achieves 73.73% / 71.81% accuracy/F1 on SEED IV and 39.93% / 39.24% on FACED, outperforming state-of-the-art methods by 0.59% / 4.40% and 1.08% / 0.89%, respectively, and demonstrating strong generalization and efficiency under complex cross-subject scenarios. AOVANet: Adaptive One-vs-All Network for Universal Domain Adaptation yuntao du and Huining li (shandong university), mujie zhang (nanjing university), mingcai chen (Nanjing University of Posts and Telecommunications), and lizhen cui (shandong university) Abstract Abstract Universal Domain Adaptation (UniDA) aims at handling both domain and category shifts between different domains. Its primary goal is to transfer the learned knowledge of the source domain to the target domain and, moreover, identify ``unknown" classes in the target domain. Recently, methods leveraging multiple one-vs-all classifiers with open-set entropy minimization have emerged to detect unknown classes, eliminating the need to set thresholds or proportionally reject unknown classes manually. However, this method overlooks the fact that a target sample contributes differently to the one-vs-all classifiers, leading to sub-optimal adaptation performance. Besides, we experimentally found that the performance of this method is heavily influenced by the number of categories in the source domain, with larger class numbers resulting in degraded results. To address these limitations, we propose a simple yet effective solution: \textbf{A}daptive \textbf{O}ne-\textbf{V}s-\textbf{A}ll \textbf{Net}work (\textbf{AOVANet}). This method dynamically assigns weights to each one-vs-all classifier based on its importance to a given sample, enabling the model to focus more effectively on the more relevant classifiers adaptively. We conducted extensive experiments on both UniDA and open-set domain adaptation settings on three cross-domain image classification datasets. AOVANet achieves significant performance gains of $5.6\%$, $0.4\%$, and $7.1\%$ on Office-31, Office-Home, and VisDA than OVANet, validating its effectiveness. ComDet: Complex Domain Adaptation for AI-Generated Text Detection Lei Jiang and Desheng Wu (University of Chinese Academy of Sciences), Xiaolong Zheng (Chinese Academy of Sciences), and Cuicui Luo (University of Chinese Academy of Sciences) Abstract Abstract Labeled data scarcity, diverse feature distributions, and adversarial attacks represent three significant challenges in the AI-generated text detection (AGTD) task. These challenges demand detectors with strong generalization and robustness; however, prior works typically address only one of them in isolation. To bridge this gap and meet the demands of real-world applications, we introduce a novel AGTD scenario termed Complex Domain Adaptation (CDA), which integrates all three challenges into a unified task. Furthermore, we observe that most existing methods struggle to perform effectively in the CDA scenario. To address this limitation, we propose ComDet, a novel framework built on an unsupervised domain adaptation foundation. ComDet incorporates a reconstruction module to enhance robustness against adversarial attacks and a feature enhancement module to extract richer, domain-specific representations. Extensive experiments demonstrate that ComDet significantly outperforms all baseline methods in the CDA scenario while achieving comparable or superior performance in other real-world settings. Progressive Domain-Invariant Injection for Cross-Domain Few-Shot Object Detection Huiwen Huang (Guilin University of Electronic Technology, School of Computer and Information Security); Shenghao Fu, Kun-Yu Lin, Wei-Jin Huang, Shuxuan Li, and Nan Lei (Sun Yat-sen University, School of Computer Science and Engineering); Xiaonan Luo (Guilin University of Electronic Technology, School of Computer and Information Security); and Wei-Shi Zheng (Sun Yat-sen University, School of Computer Science and Engineering) Abstract Abstract Cross-domain few-shot object detection (CD-FSOD) aims to learn a detector that generalizes from a label-rich source domain to a novel target domain with disjoint categories and scarce annotations. A critical challenge is mitigating the domain gap between the source and target domains. To address this, we propose DP-CDFSOD, which progressively learns domain-invariant detection representations by leveraging cross-modal semantics as domain-invariant anchors. Our framework comprises three lightweight modules that inject cross-modal priors across the Encoder–Decoder–Head architecture. The Query-Side Semantic Calibration (QSC) and Support-Guided Query Refinement (SQR) modules inject text-aligned visual priors into the encoder and decoder, respectively. The Cross-Modal Alignment and Discrimination (CAD) module anchors the classification decision and realigns the visual priors with textual prototypes in the Head. This progressive cross-modal integration enables our framework to achieve state-of-the-art performance on the CD-FSOD benchmark, improving 3.7%mAP in the challenging 1-shot setting. Moreover, the framework supports efficient target-domain adaptation by fine-tuning a few lightweight projection layers with only 2.37M trainable parameters. It allows each domain to share the majority of the source-domain parameters while maintaining a small number of domain-specific ones. Code is available at https://github.com/HHuiwen/DP-CDFSOD. Wednesday Virtual Room 7 IJCNN Paper Vision-Language Models I Session Chair: YingChao Zeng (Beijing University of Posts and Telecommunications), hongye chen (Nanchang Hangkong University) LDO-SLAM: Hierarchical Semantic-Enhanced 3DGS-SLAM with Robust Loop Closure Jichao Jiao and YingChao Zeng (Beijing University of Posts and Telecommunications) Abstract Abstract Recent advances in 3D Gaussian Splatting (3DGS) have improved scene representation efficiency and fidelity, but existing 3DGS-SLAM systems suffer from tracking drift, limited global optimization, and limited online semantic reasoning. We present LDO-SLAM, a dense RGB-D SLAM framework that integrates 3DGS with hierarchical semantic inference and loop closure. Our system fuses closed-set object detection and open-vocabulary vision-language cues into 3D Gaussians, enabling semantic submap construction and a two-stage, semantic-guided loop closure with global graph optimization. This design improves robustness to perceptual aliasing while keeping the pipeline lightweight. Experiments on Replica, TUM-RGBD, and ScanNet show competitive localization and mapping quality, while maintaining an end-to-end throughput of $\approx$2 FPS on a single RTX 3090 GPU. HRTR: Hierarchical Reasoning Transformer for Scene Graph Generation Zengxun Jin and Yun Meng (Chang'an University) Abstract Abstract Scene Graph Generation aims to produce a fundamental structured representation for deep visual understanding and downstream reasoning tasks. However, current one-stage methods typically treat categories as flat, isolated labels, ignoring intrinsic semantic correlations. This semantic isolation restricts their generalization, causing failures in zero-shot and long-tailed scenarios. To address this, we propose the Hierarchical Reasoning Transformer (HRTR), an end-to-end framework that integrates structured knowledge priors into the one-stage detection pipeline. Specifically, we construct a Hierarchical Semantic Tree derived from WordNet to explicitly model parent-child dependencies, effectively overcoming the limitations of flat label spaces. Building on this structure, we propose a Semantic-Visual Unification module that dynamically aligns these static priors with visual features, bridging the semantic gap to enable robust structured reasoning. Furthermore, we design a Hierarchical Supervision Loss to enforce structural consistency during optimization, ensuring logical coherence across coarse-to-fine predictions. Extensive experiments on Visual Genome and Open Images V6 demonstrate that HRTR achieves a strong balance between efficiency and performance, with significant gains in zero-shot Recall and mean Recall. DK-VINS: Robust Deep Visual-Inertial Odometry for Embedded Devices Yibo Dou (Tongji University) and Qin Zhu (Hunan University) Abstract Abstract Deep learning-based local features offer great promise for improving the robustness of visual-inertial SLAM, yet their performance often degrades in textureless or motion-blurred scenes—particularly on resource-constrained platforms. To address this, we propose DK-VINS, a robust deep Visual-Inertial Simultaneous Localization and Mapping system that integrates the reinforcement learning-trained detector DISK and the graph neural network matcher LightGlue into the front-end of VINS-Fusion. On the most challenging sequences of the EuRoC dataset, DK-VINS achieves accuracy and robustness comparable to state-of-the-art VI-SLAM systems. By combining ONNX-based efficient inference with a DBoW3 loop closure dictionary built offline from DISK descriptors, DK-VINS runs efficiently on embedded Jetson platforms. Experiments demonstrate that our learning-driven front-end significantly enhances VI-SLAM robustness in real-world degraded environments while maintaining real-time performance. HG-DETR: Hierarchical-Guided Detection Transformer for Video Moment Retrieval and Highlight Detection Hongye Chen (Nanchang Hangkong University), Zheng Tang (Independent Researcher), and Tao Huang (Xidian Unviersity) Abstract Abstract Video Moment Retrieval (MR) and Highlight Detection (HD) are critical tasks in video understanding, yet existing methods predominantly rely on complex task interactions, multimodal fusion, or computationally heavy Large Language Models (LLMs), often overlooking the inherent hierarchical structure of video semantics. Videos naturally organize content into multi-level semantic hierarchies: frames form segments, segments compose scenes, and scenes interconnect to construct narratives. Conventional flat feature representations fail to capture this hierarchical organization, leading to suboptimal alignment between visual content and textual queries. To address this limitation, we propose HG-DETR, a Hierarchical-Guided Detection Transformer for Video Moment Retrieval and Highlight Detection. The core innovation lies in introducing a hierarchical video feature modeling method for the joint MR and HD task, constructing hierarchical semantic features spanning frames, segments, scenes, and stories. Furthermore, through hierarchical guidance loss and decoding mechanisms, hierarchical semantics explicitly guide and enhance the representation and decoding process of traditional global features, enabling fine-grained understanding of video semantics. Experiments on multiple benchmark datasets including QVHighlights demonstrate that HG-DETR effectively improves model performance, validating the effectiveness of our approach in MR and HD tasks. Wednesday Virtual Room 8 IJCNN Paper Vision-Language Models II Session Chair: Nan Mu (XinJiang university), Arunava Roy (University of Memphis) Counterfactual-Guided Diffusion Perception for Multimodal Hate Speech Detection Nan Mu (XinJiang university) and Wenzhong Yang, Yabo Yin, and Hongzhen Lv (XinJiang University) Abstract Abstract Detecting multimodal hate speech is impeded by two coupled challenges: macro-level spurious cross-modal correlations and micro-level noise within fused representations. Existing causal inference methods often overlook the latter, assuming high-quality feature representations despite the inevitable artifacts and background noise introduced during multimodal fusion. To address this, we introduce Counterfactual-Guided Diffusion (CF-DF), a unified framework that synergizes micro-level purification with macro-level debiasing. Specifically, we establish a ``Refine-then-Debias'' paradigm: at the micro level, a diffusion-based module leverages single-step Tweedie recovery to project noisy fused features back onto the high-density data manifold, thereby enhancing the signal-to-noise ratio. At the macro level, we construct text-absent and image-absent counterfactual paths upon these refined features to disentangle and suppress unimodal shortcuts via causal intervention. Optimized end-to-end, CF-DF demonstrates superior performance and stability compared to state-of-the-art baselines on the MAMI and HarMeme benchmarks, providing a more interpretable and trustworthy solution for multimodal content moderation. Causal Inference and Counterfactual Text-Debiasing for Multimodal Aspect-Based Sentiment Analysis Haolong Zheng (Shandong Normal University), Xinyue Wang (Shandong University of Finance and Economics), and Tenglong Zheng (Hebei University of Science and Technology) Abstract Abstract Multimodal Aspect-Based Sentiment Analysis (MABSA) aims to jointly extract aspect terms from image-text pairs and accurately predict their sentiment polarity. Existing research indicates that the text modality often dominates the sentiment prediction process. However, this modal imbalance tends to induce models to over-rely on statistical co-occurrences between inputs and labels during training, rather than learning true semantic reasoning. Consequently, models erroneously establish spurious correlations between textual features and sentiment labels, which severely impairs their generalization capabilities on out-of-distribution (OOD) data. To address this challenge, we propose a Causal Inference and Counterfactual Text-Debiasing (CCT) framework. The framework first constructs a causal model to dissect the complex causal mechanisms linking multimodal features to sentiment labels. Subsequently, we introduce an auxiliary textual branch combined with an attention mechanism to explicitly model and pinpoint specific textual features that induce spurious correlations. During the inference phase, CCT utilizes counterfactual reasoning to precisely disentangle biases stemming from textual shortcuts from the multimodal Total Effect (TE), while preserving and enhancing valid semantic cues essential for sentiment prediction. Extensive experiments on the Twitter2015 and Twitter2017 datasets demonstrate that the CCT framework can be integrated into existing mainstream models, significantly boosting their generalization performance and robustness. Sorting-Based Ordering Priors for Causal Structure Learning Changxin Rong, Yanze Gao, Taiyu Ban, Xiangyu Wang, and Huanhuan Chen (University of Science and Technology of China) Abstract Abstract Large language models offer a promising source of structural knowledge for causal structure learning. Among different priors, ordering priors are especially robust and actionable because they restrict feasible edge directions through a sequential order over variables. However, in iterative learning, eliciting globally consistent ordering priors under a limited query budget remains challenging. Querying only local relations in the current graph produces fragmentary constraints that can propagate structural bias, whereas extending queries to non-adjacent pairs incurs quadratic cost. We propose Sorting-Based Ordering Prior for Causal Structure Learning. The method reformulates structural verification as a path-based sorting problem. It decomposes the graph into maximal causal paths and exploits causal transitivity to rectify causal chains efficiently. In this way, it builds a globally consistent variable order with logarithmic query complexity. The elicited priors are injected through a constraint scheme that penalizes order-violating directions, while still allowing data-driven correction. The framework applies to both score-based and gradient-based methods. Experiments show improved structure recovery and faster convergence, yielding a better tradeoff between query cost and prior coverage. A Unified Diffusion–Autoregressive Framework for Permutation-Agnostic Graph Synthesis Arunava Roy (University of Memphis), Usman Ahmad Usmani (Xiamen University Malaysia), and Ari Happonen (Lappeenranta-Lahti University of Technology) Abstract Abstract Graph generative modeling faces a long-standing dichotomy: autoregressive (AR) methods offer high expressivity but suffer from node-order sensitivity, while diffusion-based models ensure permutation invariance yet lose directional coherence. We propose PARDiff, a novel Progressive Autoregressive Diffusion framework that unifies these paradigms through block-wise, order-agnostic synthesis driven by learned structural decomposition. Unlike prior heuristic-based or strictly sequential approaches, PARDiff simultaneously predicts block sizes, ranks node positions, and applies an equivariant diffusion process within each block, aligning the strengths of AR guidance with the robustness of score-based sampling. This formulation reframes graph generation as probabilistic inference over learned topological partitions, enabling scalable and semantically faithful synthesis across molecular and non-molecular domains, without relying on auxiliary node features or templates. Extensive experiments on molecular, and social graph benchmarks demonstrate that PARDiff achieves state-of-the-art performance in both fidelity and diversity metrics. Moreover, its modular and latency-aware design supports real-time deployment in high-impact settings such as drug–drug interaction prediction. We position PARDiff as a new foundation for permutation-robust, directionally controlled graph generation. In order to promote reproducibility, our codes are open-sourced at: https://github.com/llmresearch678/Pardiff_M_1. Thursday Virtual Room 1 IJCNN (Neural Networks) Industrial applications Session Chair: Zhen Zeng (Sun Yat-sen University), Chunling Xi (Google) Beyond Semantic Similarity: Multi-Stage Calibration for In-Context Learning Chunling Xi and Di Liang (Google) Abstract Abstract In-Context Learning (ICL) empowers Large Language Models (LLMs) to adapt to downstream tasks by conditioning on a few demonstrations. However, the efficacy of ICL is heavily contingent on the quality of selected examples. Existing methods predominantly rely on semantic similarity for retrieval, which suffers from two critical misalignments: (1) \textbf{Distributional Mismatch}: colloquial user queries often diverge from the canonical representation of demonstrations, impairing retrieval accuracy; and (2) \textbf{Utility Gap}: semantically similar examples do not necessarily contribute positively to the model's reasoning, and may even introduce noise. To address these challenges, we propose \textbf{Multi-Stage Calibration (MSC)}, a framework designed to align the retrieval process with the generation objective. MSC operates in three stages: first, it refines noisy queries into canonical forms via a perplexity-guided rewrite mechanism; second, it introduces a \textbf{contribution-aware re-ranker} that explicitly learns to predict the marginal utility of each demonstration rather than mere similarity; finally, it employs a generator-guided aggregation strategy to identify the optimal context. Extensive experiments across multiple benchmarks demonstrate that MSC significantly outperforms standard retrieval baselines, offering a robust solution for deploying ICL in open-domain scenarios. XSSplore: A Tree-Structured Prompting and Actor-critic Framework for Autonomous Red Teaming XSS Payload Generation Gang Yang, Bo Wu, and Tao Xia (Information Support Force Engineering University) Abstract Abstract Autonomous red teaming penetration testing has been of great significance in ensuring the security of Web systems, particularly in the face of cyber threats such as XSS attacks. However, in real-world scenarios, it often involves contending with complex Web Application Firewalls (WAFs), which necessitates substantial expert efforts for automated methods in practical applications. Consequently, autonomously generating bypassable XSS payloads capable of handling adversarial environments remains a challenge. To address this issue, this paper proposes XSSplore, a method that combines tree-structured prompting and an actor-critic architecture, based on plug-in knowledge and a corpus. By mapping knowledge onto tree-structured prompting paths, XSSplore can more effectively explore and exploit black-box environments while selecting appropriate bypass strategies. Additionally, due to the predefined corpus and the actor-critic mechanism, XSSplore can effectively ensure the executability of generated samples and mitigate the impact of hallucinations. We conducted experiments on 20 vulnerable Web applications within environments equipped with WAFs. The experimental results demonstrate that XSSplore can accomplish autonomous and effective XSS payload generation tasks, significantly outperforming existing methods. Multi-Granularity Reasoning for Natural Language Inference Chunling Xi and Di Liang (Google) Abstract Abstract Natural Language Inference (NLI) is a fundamental task in natural language understanding that requires determining the logical relationship between a premise and a hypothesis. Despite the remarkable success of transformer-based pre-trained models, most existing approaches primarily rely on the final-layer token representations, which are often insufficient for capturing the complex and hierarchical semantic interactions required for effective reasoning. In particular, fine-grained lexical cues, phrasal compositions, and higher-level contextual semantics are typically entangled or diluted in a single representation space. To address these limitations, we propose a novel \emph{Multi-Granularity Reasoning Network} (MGRN) that explicitly leverages hierarchical semantic features within an interactive reasoning space. The proposed framework mimics the human cognitive process of language understanding, which naturally progresses from shallow lexical matching to deeper semantic abstraction and logical reasoning. By integrating semantic information across multiple granularities in a progressive and structured manner, MGRN is able to uncover intricate semantic relationships underlying natural language expressions. Extensive experiments on multiple public benchmarks demonstrate that MGRN consistently outperforms strong baseline models, validating the effectiveness and robustness of the proposed approach. Semantic Feature-oriented Multi-Scale Dual Attention Network for Truck Re-Identification Zhen Zeng and Jiangtao Ren (Sun Yat-sen University) Abstract Abstract Truck re-identification (Re-ID), a crucial subtask of vehicle Re-ID in intelligent transportation systems, faces unique challenges like significant front-carriage appearance differences, diverse carriage types, and cargo variations. Existing methods often overlook these structural differences and struggle to suppress variable visual features. To address these issues, we propose a Semantic Feature-oriented Multi-Scale Dual Attention (SF-MSDA) network, which includes: 1) a Segmentation module that enables truck front and carriage feature extraction at two granularity levels;2) a Multi-Scale Dual Attention (MSDA) module that reconstructs front features by integrating spatial and channel dependencies, maximizing spatial distribution differences of each pair of features for fine-grained representation;and 3) a Part-Attention module that combines part features and global context to adaptively learn invariant features while suppressing cargo variations.Experiments on the Truck-ID dataset demonstrate that SF-MSDA outperforms both general vehicle and specialized truck Re-ID methods. Thursday Virtual Room 2 IJCNN Paper, IJCNN Late Breaking Paper IJCNN Various Tracks X Session Chair: Patrick Akinwumi (Clemson University) Neuro-Symbolic Reinforcement Learning for Cognitive Digital Twins: A Framework for Lifelong Intelligence Amplification Patrick O. Akinwumi and Meihua Qian (Clemson University) Whova Tag: Poster Presentation Abstract Abstract Current AI-based tutoring and personalization systems primarily optimize short-term task performance rather than long-term cognitive growth. Recent work on digital twins and cognitive digital twins (CDTs) suggests the possibility of high-fidelity virtual counterparts that continuously learn and adapt alongside humans. In parallel, neuro-symbolic AI integrates neural representation learning with symbolic reasoning for more interpretable and trustworthy decision-making, while reinforcement learning (RL) has been increasingly applied to intelligent tutoring and educational decision-making. This late-breaking paper proposes a novel computational framework: a Cognitive Digital Twin for Lifelong Intelligence Amplification (CDT-LIA). The CDT-LIA is a neuro-symbolic RL agent that maintains a latent cognitive state of an individual learner, reasons over a symbolic knowledge graph, and optimizes a new objective we call Cognitive Growth Rate (CGR), defined as the expected incremental gain in an intelligence functional over time. We formalize CGR, describe a three-layer neural-symbolic-RL architecture, and present a synthetic proof-of-concept experiment comparing CGR-optimized and performance-optimized policies. Preliminary results indicate that optimizing CGR can reduce plateau effects and improve long-term transfer in simulated learners, suggesting a promising direction for human-centred computational intelligence Dynamic Mixture-of-Head Attention via Feature Alignment and Conflict for Fake News Detection Rongxin Lin (National University of Defence Technology, Laboratory for Big Data and Decision); Meng Xing (Unit 32005 of the Chinese People's Liberation Army); Zihao Li (National University of Defence Technology, Laboratory for Big Data and Decision); Lincheng Jiang (National University of Defense Technology, College of Advanced Interdisciplinary Studies); and Shuohao Li and Jun Zhang (National University of Defence Technology, Laboratory for Big Data and Decision) Abstract Abstract Multimodal fake news detection, which fuses textual and visual information, has been proven to significantly boost the accuracy of fake news classification, thereby effectively curbing the proliferation of misinformation on social media platforms. Despite the fact that existing methods primarily focus on inter-modal and intra-modal fusion strategies, they still struggle to mitigate redundant attention patterns and lack feature learning mechanisms from the perspective of cross-modal conflicts. To address these limitations, we propose a Dynamic mixture-of-head attention via Feature Alignment and Conflict network (DFAC) for fake news detection. Specifically, we design a multi-scale feature attention to capture local fine-grained features and global correlations, enabling comprehensive extraction of discriminative information. Subsequently, we implement a dynamic expert-driven attention to adaptively select and aggregate the most discriminative sub-expert heads. Finally, we develop a feature alignment and conflict fusion module to adaptively modulate dynamic features and activate high-value semantic clues. Experiments demonstrate the superiority of our DFAC method on the Weibo, Twitter and Politifact datasets. Heat transfer enhancement model for the annulus region of heat pipe with combined convective and radiative losses using PINNs Himanshu Upreti, Ziya Uddin, and Mohd Vaseem (BML Munjal University) Whova Tag: Poster Presentation FDPTU: Federated Attack-aware Differentially Privacy Tree Unlearning framework for Multimodal Biometric Privacy Samyak Jain, Simran Gupta, and Pronaya Bhattacharya (Amity University, Kolkata, India); Sandip Roy and Sachin Shetty (Old Dominion University); Rutvij H. Jhaveri (Pandit Deendayal Energy University, India); and Saed Alrabaee (UAE University) Abstract Abstract Recently, multimodal biometric systems have begun to utilize Federated Learning (FL) to mitigate the risk of data exposure. However, FL remains vulnerable to privacy leakage, adversarial data poisoning, and the lack of data deletion upon user request. Recent schemes focus on the integration of unlearning mechanisms with the Differential Privacy (DP) mechanism. These are challenged by expensive retraining and coordinated attacks such as deepfakes and morph-based manipulation. To address this, the paper proposes a scheme, FDPTU, which is an attack-aware federated DP tree-based unlearning framework designed for multimodal systems. It integrates secure template generation, homomorphic encryption, and Gaussian DP to protect sensitive biometric features during distributed training. To address emerging threats, the framework incorporates a dedicated detection layer that identifies malicious contributions and constructs a dependency tree to trace their influence across federated rounds. Selective, shard-level unlearning is then applied to remove compromised or withdrawn data without global retraining, supported by cryptographic proof of removal for regulatory compliance. Extensive experiments on multimodal biometric and deepfake datasets demonstrate that FDPTU preserves verification accuracy while eliminating nearly 95% of residual influence, reducing unlearning cost by approximately 35-45% and achieves near random ROC value after unlearning. Thursday Virtual Room 3 IJCNN (Neural Networks) IJCNN Various Tracks XI Simple Averaging vs. Learned Stacking in K-Fold Ensembles: Evidence from Pore Pressure Prediction Pranav Patel (Northeastern University), Muhammad Raiees Amjad (Bahria Univeristy), and Rohan Benjamin Varghese and Tehmina Amjad (Northeastern University) Abstract Abstract Pore pressure prediction is critical for drilling safety, yet ensemble approaches lack principled methods for model selection and combination under cross-well distribution shift. This study presents a K-Fold stacking framework using well-based GroupKFold cross-validation, where models are ranked by out-of-fold (OOF) R², enabling leakage-free ensemble selection without ground truth from new wells. Five architectures (CNN, DFNN, RNN, Random Forest, XGBoost) are evaluated on 186,070 samples from 17 development wells in Pakistan's Potwar Basin. On 4 blind test wells (84,904 samples), the 5-model ensemble achieves R² = 0.8905, outperforming the best individual model by 7.03 percentage points. Simple averaging consistently outperforms Ridge meta-learner stacking across all ensemble sizes, with the gap widening from 0.41 to 2.36 percentage points as more models are included. Ridge concentrates weight on fewer models and assigns negative coefficients to others, discarding architectural diversity valuable under distribution shift. Top-K ablation confirms monotonic gains from K=1 through K=5, with the bulk of improvement from combining the first two architecturally distinct models. Physics-Guided Multi-Task Learning for Subgrid Scale Turbulence Parameterization: A Comparative Study of Physics Integration Strategies Sambit Kumar Panda, Todd Jones, and Muhammad Shahzad (University of Reading); Anna-Louise Ellis (Met Office); and Bryan Lawrence (University of Reading, National Center for Atmospheric Science) Abstract Abstract Neural network emulation of subgrid-scale (SGS) turbulence parameterizations in atmospheric models offers the promise of computational acceleration but suffers from brittle cross-regime generalization that limits operational deployment. We present a systematic evaluation of physics-guided multi-task learning strategies for SGS coefficient prediction, comparing six conditioning approaches across three neural architectures (MLP, ResMLP, TabTransformer) on contrasting atmospheric regimes: idealized tropical deep convection (Radiative Convective Equilibrium) and mid-latitude shallow convection (Atmospheric Radiation Measurement). Through evaluation of 222.6 million predictions (inference), we reveal dataset-dependent behaviour with important implications for neural network design. Simple MLPs with explicit Richardson number onditioning achieve improved cross-regime generalization on ARM data (R2 = 0.67 viscosity, 0.73 diffusivity), outperforming architecturally complex alternatives by 158% in viscosity R2 (0.432 vs 0.167) on dynamically important active turbulence regions. Residual networks that achieve strong performance on idealized RCE simulations (R2 > 0.78) produce predictions indicative of numerical instability on the ARM regime, producing negative R2 values with variance ratios exceeding 3× the ground truth. Kling-Gupta Efficiency decomposition reveals systematic positive bias across all model configurations, reflecting training data sparsity in active turbulence regions rather than fundamental physics incompatibility. These findings establish design principles for atmospheric ML parameterization development: architectural simplicity with targeted physics conditioning provides improved offline predictive robustness compared to complex architectures relying on implicit pattern learning. Online coupling experiments remain essential to validate operational stability under feedback dynamics. NeuroAPS-Net: Neuro-Anatomically Aware Point Cloud Representation for Efficient Alzheimer's Disease Classification Towhidul Islam and Mufti Mahmud (King Fahd University of Petroleum & Minerals) Abstract Abstract Alzheimer's disease (AD) is a progressive neurodegenerative disorder and a major cause of dementia. Structural MRI is widely used to analyse AD-related brain atrophy; however, most deep learning methods rely on computationally expensive 3D convolutional neural networks (CNNs), limiting deployment in resource-constrained settings. This work introduces two main contributions. First, we propose a pipeline that converts T1-weighted MRI into anatomically informed 2D point clouds using Anatomical Priority Sampling (APS), producing ADNI-2DPC, the first neuroanatomically labelled MRI-derived point cloud dataset. Second, we present NeuroAPS-Net, a lightweight geometric deep learning model that incorporates anatomical priors via region-aware feature encoding and ROI token aggregation. Experiments on ADNI-2DPC demonstrate that NeuroAPS-Net achieves competitive classification accuracy while significantly reducing inference latency and GPU memory compared to state-of-the-art point cloud methods. These results highlight the potential of anatomically guided point cloud learning as an efficient and interpretable alternative to voxel-based CNNs for AD classification. SwapGraph: Biologically Plausible Shortest Path Planning on Real-World Graphs Pallavi Das and Anup Das (Drexel University) Abstract Abstract We present SwapGraph, a biologically plausible shortest path algorithm that extends neuromorphic path planning to directed acyclic graphs with heterogeneous costs. Our approach replaces Euclidean distance-based predecessor selection with weight-augmented temporal scoring, implementing Bell- man’s optimality principle using using only spike arrival times and learned axonal delays. E-Prop based online learning adapts these axonal delays through local spike-timing information without global cost knowledge. We implement graph representation using CSR-CSC sparse format, maintaining separate encodings for outgoing neighbors during spike propagation and incoming neighbors during path extraction. Evaluation on nine benchmark datasets (168K–24M nodes) demonstrates 100% optimal path accuracy while consuming 5–8% of memory required by conventional representations, with 13–202× runtime improvements. These results establish the first biologically plausible shortest path algorithm for large-scale sparse DAG topologies. Thursday Virtual Room 4 IJCNN (Neural Networks) IJCNN Various Tracks XII FESFN: Frequency-Enhanced Supervised Fusion Networks for 3D object detection ShengFa Pei and YuanYao Lu (North China University of Technology) and Kamoliddin Shukurov (Tashkent University of Information Technology named after Muhammad al-Khwarizmi) Abstract Abstract In recent years, BEV-based 3D object detection has seen substantial progress in the context of perception for autonomous driving. To tackle the issues of inaccurate depth estimation, redundant fusion strategies, and imprecise classification supervision in existing LSS methods, this paper proposes a Frequency-Enhanced Supervised Fusion Network (FESFN) for LiDAR-camera fusion detection. Specifically, a Frequency Supervision Fusion Module (FSFM) is introduced to incorporate frequency information from LiDAR point clouds to guide the image-based LSS process, thereby improving depth estimation accuracy. A Progressive Fusion Module (PFM) is then designed to enhance cross-modal feature collaboration and reduce redundancy by combining dense and sparse feature fusion strategies. Additionally, an Instance-Aware Classification Head (IACH) is proposed to supervise classification only on predicted instances, enhancing the precision of classification guidance. Experimental results on the nuScenes dataset demonstrate that FESFN achieves 73.6% mAP and 75.5% NDS, outperforming current state-of-the-art methods. Thursday Virtual Room 1 IJCNN Paper Robustness and Adversarial Learning IV Session Chair: Yuanfang Guo (Beihang University), Sk. Subidh Ali (Indian Institute of Technology Bhilai) Generalized Adversarial Overwriting Attack on Deep Learning-Based Watermarking Under Restricted Black-Box Access Sudev Kumar Padhi and Sk. Subidh Ali (Indian Institute of Technology Bhilai) Abstract Abstract The increasing sophistication of AI-driven image manipulation and generation poses significant challenges to copyright protection and ownership verification. Deep learning-based watermarking techniques have emerged as a robust solution to embed imperceptible yet resilient watermarks within images. However, these techniques remain vulnerable to adversarial attacks, threatening their reliability in digital forensics and intellectual property protection. This work introduces a generalizable adversarial overwriting attack under restricted black-box access that systematically replaces an existing watermark with an attacker-specified one, effectively compromising ownership claims. Our technique is the first to successfully attack the decoders without any access (not even Oracle access). We leverage a surrogate model attack in a restricted black-box setting under no access to the decoder to generate imperceptible adversarial perturbations, forcing watermarking techniques to extract an attacker-specified watermark. Evaluated across several watermarking methods, our technique effectively overwrites watermarks while preserving image quality, making detection difficult. Our findings expose critical vulnerabilities in deep learning-based watermarking, highlighting the need for more resilient frameworks to safeguard copyright protection and prevent unauthorized ownership claims. FROM: A Framework for Recognizing the Origins of LLMs in Black-Box Conditions Hongru Wei, Qingyuan Hu, Linzhi Chen, and Yuqi Chen (ShanghaiTech University) Abstract Abstract Large Language Models (LLMs) exhibit remarkable capabilities across diverse domains, with their training requiring substantial computational resources and time. Fine-tuning has emerged as a standard practice for task-specific adaptation of LLMs, raising critical challenges in tracking model origins to ensure licensing compliance and prevent unauthorized use of proprietary base models. Additionally, fine-tuned models inherit characteristics from their original models, such as behaviors, vulnerabilities, and compatibility. Therefore, analyzing the original model can also help optimize performance and mitigate security risks in its fine-tuned versions. To address these challenges, we introduce FROM, an efficient and lightweight framework for tracing the origins of fine-tuned LLMs in black-box scenarios, which enhances accountability, reproducibility, and intellectual property protection. Experiments show that FROM achieves 94.4% accuracy in identifying original models, underscoring the importance of version awareness to safeguard model ownership and optimize deployment. Black-box Backdoor Attack on Continual Semi-Supervised Learning for Time-series Internet-of-Things Systems Thanh Cong Nguyen (VNU - University of Engineering and Technology); Hanrui Wang (National Institute of Informatics); and Isao Echizen (National Institute of Informatics, The University of Tokyo) Abstract Abstract Time-series Internet-of-Things (IoT) systems increasingly adopt continual semi-supervised learning (CSSL) to adapt to dynamic environments. However, the continual ingestion of unlabeled user data exposes such systems to backdoor injection threats. Existing backdoor attacks typically assume white-box access or manual control over labeled training data, assumptions that are often unrealistic in practical deployments. We present the first fully black-box backdoor attack against CSSL-based time-series classification that is gradient-free and annotation-free. Our approach constructs a surrogate model using only constrained query feedback from the target system. This surrogate is then used to optimize a trigger generator, enabling the injection of unlabeled poisoned samples that are automatically incorporated during CSSL updates. Under this challenging black-box setting, our attack achieves a high targeted attack success rate (average 78.4%) while inducing only a marginal drop in clean accuracy (average 2.5%). Moreover, the injected backdoor remains effective even when poisoning is performed only once every six system update rounds. These results demonstrate that CSSL pipelines are vulnerable under realistic black-box conditions, highlighting an urgent need for robust defenses. Distribution-Aware Black-Box Graph Injection Attack via Local Neighborhood Context Chengyu Zhu, Weina Xu, and Yufei Hou (Beihang University); Kahim Wong (University of Macau); Zeming Liu (Beihang University); Jiantao Zhou (University of Macau); and Yuanfang Guo (Beihang University) Abstract Abstract Graph injection attacks inject carefully crafted malicious nodes and edges into graph-structured data, posing a serious threat to Graph Neural Networks (GNNs) whose predictions heavily rely on topological aggregation. As GNNs are widely deployed in real-world applications, such as recommendation systems and security-sensitive domains, investigating such attacks is essential for assessing practical risks and improving the reliability of GNN-based systems. Existing methods have been proposed to analyze the vulnerability of GNNs. However, they still struggle to balance attack effectiveness and imperceptibility. To address this challenge, we propose DAGIA, a novel Distribution-Aware Graph Injection Attack framework that leverages local neighborhood context to achieve imperceptible and effective node injections. Specifically, DAGIA firstly identifies vulnerable nodes via a multi-metric assessment. Guided by Neighborhood Label Distribution (NLD), DAGIA then constructs injection topology by clustering vulnerable nodes, and initializes injected node features. Finally, a margin-based optimization refines injected features to enhance attack performance. Extensive experiments on benchmark datasets demonstrate that DAGIA achieves high attack effectiveness while maintaining strong imperceptibility, outperforming existing state-of-the-art methods. Thursday Virtual Room 2 IJCNN Paper Robustness and Adversarial Learning V Session Chair: Jianli Ding (Civil Aviation University of China), Sihan You (National University of Defense Technology) Cross-Trigger Generalization in Backdoor Defense: An Active Unlearning Approach for PLMs Jing Wang, Yuhao Zhou, and Jianli Ding (Civil Aviation University of China) Abstract Abstract To address the challenges of high trigger inversion costs and model performance degradation in backdoor defense for pretrained language models, this paper proposes the CTG-Unlearn method based on a minimax defense framework, breaking away from the traditional assumption of precisely recovering true triggers. The core innovation involves actively injecting simple proxy triggers to construct feasible backdoor solutions, reducing inner optimization complexity to O(1), and designing a pullback loss mechanism to achieve semantic feature alignment that blocks backdoor shortcuts while constraining poisoned samples to their original semantic neighborhoods. Experiments on SST-2 and AG News datasets show that CTG-Unlearn reduces average ASR to 3.24% against three attack types (72% lower than the D3 baseline) with only 0.61% ACC drop (8% less performance loss than MuScleLoRA). The method maintains robust performance on the complex Yelp dataset and RoBERTa architecture, validating its generalization capability. Permutation-Driven Mode Connectivity for Robust Backdoor Defense Jing Zhao, Jie Peng, Hongwei Yang, Hui He, Weizhe Zhang, and Hengji Dong (Harbin Institute of Technology) Abstract Abstract Backdoor attacks undermine the reliability of deep neural networks, while existing defenses often fail under low poisoning rates. To address this challenge, we propose Permutation-Driven Mode Connectivity (PDMC), a backdoor defense framework that exploits neuron permutation invariance and permutation-induced connectivity in the loss landscape to suppress backdoor behaviors. Our analysis reveals that backdoor behaviors are supported by fragile neuron co-adaptations confined to local low-loss basins. Building on this insight, PDMC applies neuron permutations to generate functionally equivalent yet structurally realigned models, and leverages permutation-guided mode connectivity to explore low-loss paths. PDMC selects model parameters that preserve clean accuracy while substantially increasing backdoor loss, thereby mitigating backdoor behaviors using only a small clean dataset. The proposed method is attack-agnostic and requires no prior knowledge of trigger patterns. We validate the effectiveness of PDMC on both object recognition and object detection tasks, demonstrating strong generalization across diverse backdoor attacks. BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models Jagadeesh Rachapudi, Ritali Vatsi, and Pranav Singh (Indian Institute of Technology Mandi) and Praful Hambarde and Amit Shukla (India Institute of Technology Mandi) Abstract Abstract In recent trends, one can observe Large Language Models (LLMs) are exposed to backdoor attacks where vicious triggers added during training or model editing to elicit harmful outputs on specific input patterns while maintaining clean performance on normal inputs. Legitimate watermarks used as ownership signatures share similar mechanisms to backdoors, creating a critical challenge: detecting and eliminating unknown backdoors without compromising watermark integrity. Existing defenses require prior knowledge of triggers or their payloads, depend on clean reference models, or sacrifice model utility without preserving the watermark. To address these limitations we introduce BackFlush and its variants, a unified framework for backdoor detection and elimination while preserving watermarks. We establish two novel observations: Backdoor Flushing Phenomenon, where injecting and unlearning auxiliary data eliminates pre-established backdoors, and Backdoor Susceptibility Amplification, enabling constant time detection independent of vocabulary size. BackFlush employs Rotation-based Parameter Editing (RoPE) Unlearning, a technique that preserves watermarks while eliminating backdoors by rotating the embeddings. Comprehensive evaluation across diverse trigger types over different architectures demonstrates BackFlush achieves ≈1% Attack Success Rate (ASR), ≈ 99% clean accuracy (CACC), and preserved watermarking capabilities in the realm where no existing method simultaneously provides these alongside maintaining model utility comparable to clean baselines. Codes are available at https://github.com/JagadeeshAI/BackFlush_IJCNN.git JanusGuard: A Proactive Video Defense by Disrupting Internal Spatio-Temporal Representations Sihan You, Yuchuan Luo, Ying Wang, Youhe Jiang, Fei Ding, Zhenyu Qiu, Lin Liu, and Shaojing Fu (National University of Defense Technology) Abstract Abstract Text-driven video editing models make convincing malicious edits and deepfakes increasingly accessible. Proactive defenses, which protect videos before release, are a promising countermeasure, but existing methods often rely on external target videos and show limited transferability. We propose JanusGuard, a target-agnostic framework that disrupts internal representations required for consistent editing. Specifically, JanusGuard combines (1) a Spatial Confusion Defense (SCD) in pixel space to weaken the VAE's foreground-background semantic separation, and (2) a Temporal Decoupling Defense (TDD) in latent space to disrupt motion consistency in the U-Net. A low-strength denoising refinement step preserves visual fidelity while retaining the protective signal. Experiments on two mainstream text-driven video editing models show that JanusGuard effectively degrades malicious edits while maintaining high visual quality and outperforming existing baselines. Thursday Virtual Room 3 IJCNN Paper Robustness and Adversarial Learning VI Session Chair: Mustafa Misir (Duke Kunshan University), Yuanbo Xie (Institute of Information Engineering, Chinese Academy of Sciences, China; School of Cyber Security, University of Chinese Academy of Sciences, China) Efficient Black-box Jailbreak for Diverse Generative AI Models with Neutral Padding Yuanbo Xie, Tianyun Liu, and Tingwen Liu (Institute of Information Engineering, Chinese Academy of Sciences, China; School of Cyber Security, University of Chinese Academy of Sciences, China) Abstract Abstract As generative models are widely deployed, their safety vulnerabilities have drawn increasing attention. Although prior work shows that prompt-based jailbreak attacks can undermine model alignment, most studies assume relaxed settings and ignore practical constraints such as external prompt guardrails, strict language-form restrictions, and prompt length limits. This paper investigates efficient black-box jailbreak attacks under Chinese-form prompts, and analyzes the vulnerability of generative models when multiple real-world constraints are jointly enforced. We propose VacuousAttack, a structured black-box framework that combines toxic-prompt rewriting and neutral padding to bypass both external filtering systems and internal safety mechanisms without access to model internals or gradients. Extensive experiments on commercial text-generation and text-to-image models show that VacuousAttack consistently outperforms state-of-the-art jailbreak methods in attack success rate, cross-model and cross-modality transferability under realistic deployment conditions. AS-SYNERGIC: Orchestrating Synergistic Safety and Availability Attacks on Large Language Models Yumeng Liu and Xin Wang (Qilu University of Technology), Zhenyong Zhang (Guizhou Univerisity), and Ming Yang and Xiaoming Wu (Qilu University of Technology) Abstract Abstract Large language models (LLMs) face dual threats to safety and availability, yet existing attacks often fail to balance jailbreaking success, resource consumption, and stealthiness. To address this, we propose AS-SYNERGIC, a unified framework that orchestrates safety-breaching and resource-exhaustion attacks. Our approach introduces a curriculum scheduler that implements a Breach-then-Flood strategy, dynamically transitioning from bypassing safety guardrails to maximizing inference resource consumption. We further integrate a style-constrained mutation mechanism to keep optimization within a natural language manifold. Extensive experiments on 12 black-box target models demonstrate that AS-SYNERGIC achieves a 69.0% average attack success rate and induces responses 2.3× longer than normal inputs. Furthermore, it maintains superior stealthiness, with a perplexity of 42.1, significantly lower than unconstrained baselines. Lezak: Explaining Complex Neural Networks through a Neuropsychological Lens Nima Mirnateghi and Sharjeel Tahir (Edith Cowan University); Syed Mohammed Shamsul Islam (Edith Cowan University, University of Western Australia); Joseph M. Barnby (Edith Cowan University, King's College London); and Syed Afaq Ali Shah (Edith Cowan University) Abstract Abstract Large language models (LLMs) are now widely used in interactive dialogues, where behaviour depends on context and conversational dynamics. Despite their impressive success, they are often treated as black boxes, with limited understanding of the internal mechanisms that maintain a consistent behavioural stance across turns. Inspired from neuropsychological approaches to understanding the human brain, we develop Lezak as a general-purpose ablation-based algorithm to probe black-box dialogue models. Lezak, named after the famous neuropsychologist, uses sparse autoencoders (SAEs) to decompose model behaviour before performing target ablations while preserving conversational context. We apply Lezak to therapeutic dialogues as a proof-of-principle for our approach, where context depends on more than semantic correctness. We find that different aspects of therapeutic behaviour evolve across deeper layers. Results suggest that Lezak can be useful as a causal tool to identify mechanistic regions of interest within LLMs, with future directions to expand our case-study into a generalised tool. Beyond Numerical Features: CNN-Driven Algorithm Selection via Contour Plots for Continuous Black-Box Optimization Yiliang Yuan (Mohamed bin Zayed University of Artificial Intelligence) and Xiang Shi and Mustafa Misir (Duke Kunshan University) Abstract Abstract The present paper introduces a new representation-driven approach to per-instance algorithm selection, applied to black-box optimization, for automatically choosing the most promising solver from a fixed portfolio. Prior work in continuous optimization largely relies on numerical descriptors, including Exploratory Landscape Analysis features and learned embeddings such as Deep-ELA. This work studies a complementary representation: contour-map visualizations of probed landscapes. A CNN regressor takes multiple instance-specific contour views (stacked or encoded per view and aggregated) and predicts per-solver performance, enabling selection by the predicted best value. On the standard BBOB 2009 single-objective protocol, the resulting selectors significantly outperform the single best solver (SBS) and are competitive with feature-based baselines. A subsequent bi-objective evaluation under the DeepELA setting further indicates that the same image-based principle can be competitive when using windowed contour views. Overall, the results suggest that simple vision models can exploit spatial structure in probed landscapes for algorithm selection without handcrafted ELA features. Thursday Virtual Room 4 IJCNN Paper Robustness and Adversarial Learning VII Session Chair: Siya Yao (Zhejiang Gongshang University), Erxiang Wang (ZZU) Buffer-Guided Differentiated Learning for Adversarial Fine-Tuning in Self-Supervised Pretrained Models Jiahui Kong, Siya Yao, and Xiaofeng Zhou (Zhejiang Gongshang University) and Xiaoyu Sean Lu (Nanjing University of Science and Technology) Abstract Abstract Self-supervised pretrained models are fundamental to modern visual recognition systems. However, their open-source nature makes them vulnerable to adversarial attacks, and their widespread reuse across diverse downstream tasks amplifies the impact of such attacks, particularly under universal adversarial perturbations (UAPs). Adversarial fine-tuning enhances robustness by integrating adversarial training into the downstream adaptation process and leverages the model’s redundant capacity to restore standard accuracy without compromising robustness. Existing methods typically freeze most network layers and finetune only a few isolated ones. However, this approach overlooks inter-layer dependencies and gradient propagation, resulting in a loss of coherence in the effective pretrained representations. To address this, we propose Buffer-guided Differentiated Learning for Adversarial Fine-tuning (BDAF), which enables more finegrained optimization to better balance robustness and accuracy. BDAF partitions network layers into robustness-sensitive layers, robustness-redundant layers, and buffer layers according to their contributions to robustness, and applies differentiated learning rates based on both functional roles and depth positions. We adopt a robustness–accuracy trade-off score to guide model selection. Extensive experiments on 7 self-supervised pretraining methods and 4 datasets, evaluated using 5 UAPs, demonstrate that BDAF consistently outperforms existing methods in both standard and robust accuracy, with strong cross-task generalization. Code is available at https://github.com/Xisi-7/BDAF. SAPatch: A Reinforcement Learning-based Framework for Spatial Adversarial Patch Attacks Longpeng Hao, Libing Wu, Zhuangzhuang Zhang, Jiaqi Feng, Li Yi, and Taiyu Bao (Wuhan University) Abstract Abstract With the extensive deployment of Deep Neural Networks in autonomous driving perception systems, the robustness of object detectors has become a crucial problem in public safety. Contemporary adversarial attacks targeting these detectors are predominantly categorized into global adversarial perturbations and adversarial patch attacks. Although adversarial patches have demonstrated significant potency against visual models, existing methodologies often rely on fixed strategies when handling discrete variables such as spatial positioning. Consequently, they struggle to reconcile attack success rates with physical realizability within complex, dynamic driving environments. To address this challenge, we propose a new attack framework named SAPatch, which adopts a bilevel optimization structure and combines reinforcement learning-based dynamic parameter tuning and adversarial patch texture generation. The adversarial patch within the effective target region is generated through a two-layer loop. In the outer loop, a reinforcement learning-based agent selects the optimal adversarial patch parameter settings according to a clean image to precisely attack the most sensitive semantic regions of object detectors. In the inner loop, a gradient-based optimization algorithm and a box-bound constraint mechanism are adopted to generate the adversarial texture with the best attack effect. The experiment results show that SAPatch has superior attack efficiency in various complex simulation environments highly simulated by CARLA and superior average attack accuracy on various mainstream object detectors, reflecting the severe vulnerability of existing autonomous driving perception systems. Unsupervised Frequency-Spatial Fusion for Detecting Adversarial Examples Canran Shen and Chenyi Huang (Sichuan University); Wenbo Fang (Southwest University for Nationalities); and Keying Li, Wengang Ma, Xiaolong Lan, and Junjiang He (Sichuan University) Abstract Abstract Deep neural networks (DNNs) are vulnerable to adversarial examples generated by adding imperceptible perturbations to the input. To address these threats, detecting adversarial examples has become a promising approach. It ensures that the model runs reliably, even in the presence of adversarial examples. However, existing methods struggle to detect adversarial examples generated by generative adversarial networks (GANs) and diffusion models. Because the perturbations generated by these attacks appear spatially natural, can circumvent detection based solely on the spatial domain. However, we found that these generative attacks consistently exhibit anomalous patterns in the frequency domain following a fast Fourier transform (FFT). Provides supplementary clues for detection. Based on this finding, we propose an unsupervised detector named SFDD. It operates in a Class- and Classifier-Free (CCF) setting. SFDD models frequency-domain distributions via normalizing flows and spatial-domain structures through an autoencoder. Then it combines these two anomalous signals for detection. Experiments on CIFAR-10 and Tiny ImageNet against 11 adversarial attacks show that SFDD achieves AUROC scores of 0.85 and 0.94, corresponding to relative improvements of 26.84% and 62.01% over existing methods. For five generative attacks on Tiny ImageNet, our method achieves relative gains of 34.45%, 56.97%, 71.77%, 102.5%, and 73.35%, respectively. TARSteg: Generative Image Steganography Based on Latent TARFLOW Erxiang Wang, Chunfang Yang, Long Yu, Yang Pei, Chengyu Mo, Xueyuan Fu, and Yu Shao (ZZU) Abstract Abstract Existing normalizing-flow-based generative image steganography methods face two key challenges: rapidly increasing training cost as the stego-image resolution grows, and limited robustness in payload extraction. To address these issues, we propose TARSteg, a generative image steganography framework built upon latent TARFLOW. TARSteg first encodes the payload with spherical codes to improve the noise robustness of the input vectors to the flow model. It then leverages TARFLOW to map the payload into the image latent space, which reduces training cost while increasing the effective payload capacity. Finally, through adversarial training, TARSteg learns a robust image decoder and payload extractor for high-quality stegoimage generation and accurate payload recovery. Extensive experiments on LSUN-Churches, LSUN-Bedrooms, and FFHQ show that TARSteg consistently outperforms representative generative image steganography methods, including IDEAS, S2IRT, and LDStega, in terms of stego-image quality, robustness, and hiding capacity. For example, under JPEG compression with a quality factor of 50, TARSteg improves the extraction accuracy from 60.20%, 70.46%, and 94.46% achieved by these methods to 96.44%. Thursday Virtual Room 5 IJCNN Paper Spatiotemporal Learning and Prediction II Session Chair: Yingchi Mao (Hohai University), Qingzheng Hu (INTI International University, Southwest Petroleum University) PAANet: Patch-Aligned Attention Network for Spatio-Temporal Prediction on Low-Quality Data Jiajun Ouyang, Yingchi Mao, Hongliang Zhou, Ling Chen, and Yi Rong (Hohai University) Abstract Abstract Spatio-temporal prediction enables reliability in climate monitoring and traffic management. Presently, the deployment of sensor devices generates massive data, which provides support for spatio-temporal prediction by extracting spatio-temporal dependencies in the data. However, making accurate spatio- temporal prediction is still challenging due to the presence of low-quality data, which generally includes two characteristics: 1) missing data, i.e., missing data disrupts the continuity of time series, and 2) noise data, i.e., deviations are introduced into data due to noise. Missing and noise data make it difficult for the model to accurately capture temporal dependencies between timesteps and the true spatial relationships between time-varying variables. To ensure prediction accuracy on missing and noise data, we propose a novel framework for spatio-temporal prediction, namely Patch-Aligned Attention Network (PAANet). In particular, PAANet is composed of two main parts, called Temporal Feature Mining (TFM) and Spatial Feature Mining (SFM) modules. To jointly address missing and noise data issues in the temporal dimension, the TFM module is designed on the basis of multi-scale temporal feature encoder (MSTFE) and wavelet temporal decomposition encoder (WTDE). More specifically, MSTFE is tailored to effectively to capture both local and global dependencies under missing data conditions. Then, WTDE is introduced to suppress high-frequency noise and stabilize temporal trends. Meanwhile, to handle both noise and missing observations in the spatial dimension, the SFM module leverages a missing-aware graph imputation (MAGI) layer and a gated graph attention network (GGAtt), which imputes missing node information and dynamically filters noise. Experimental results on two real-world datasets with missing and noise data demonstrate that PAANet is superior to advanced benchmarks, as evidenced by the higher prediction precision. Spatio-Temporal Attention Gated Transformer for Memory Uncorrectable Error Prediction Jianli Ding, Guohao Chao, and Jing Li (School of Computer Science and Artificial Intelligence, Civil Aviation University of China) Abstract Abstract Uncorrectable Errors (UEs) in memory are a critical trigger for hardware failures and service disruptions in modern data centers. Existing prediction methods typically rely on rule-based feature engineering or traditional machine learning models, which struggle to effectively capture the complex spatio-temporal evolution patterns of memory errors while ignoring the spatial correlation of memory errors in bit-level physical topology. To address this, this paper proposes an end-to-end deep learning model named Spatio-Temporal Attention Gated Transformer (STAGT). Firstly, the model adopts decoupled positional encoding to explicitly preserve physical topology information. Secondly, it constructs a spatio-temporal decoupled attention architecture, leveraging spatial attention and adjacency masks to accurately capture local correlations between error bits. On this basis, a spatial global gating mechanism is introduced to adaptively adjust feature weights according to the overall spatial state. Finally, temporal attention is integrated to model the progressive evolution patterns of failures. Experimental results on real production environment datasets demonstrate that STAGT outperforms existing mainstream methods in all performance metrics. VaT-BERT: Variational Tokenization and Correlation-Guided BERT for Spatio-Temporal Forecasting Haoyu Jin, Rongyan Chen, and Xinyu Gu (Beijing University of Posts and Telecommunications) Abstract Abstract Traffic flow forecasting requires modeling temporal dynamics and inter-sensor interactions. Recent work reuses pretrained backbones for spatio-temporal prediction, but serializing sensors as tokens and applying decoder-style causal attention can couple inherently order-free spatial interactions to an arbitrary token order, inducing order bias. We propose VaT-BERT, an adaptation framework that repurposes a pretrained bidirectional encoder, DistilBERT, for traffic forecasting via a node-as-token representation and non-causal self-attention to capture global inter-node dependencies with reduced sensitivity to node order. To improve token quality and stability, we perform variational tokenization by pretraining a Temporal variational autoencoder (VAE) on per-node historical segments and feeding continuous latent tokens to the encoder. We further introduce correlation-prior-guided contrastive regularization to inject long-range correlation structure and stabilize node representations under global attention. Experiments on PEMS03/04/07/08 achieve the best MAE on all four datasets and the best RMSE on three datasets, while remaining competitive in MAPE; ablation and robustness results on PEMS04 further support the contribution of each component. CFCST: A Cost-Efficient Spatio-Temporal Coupling Architecture for Multi-Task SHM Systems Qingzheng Hu (INTI International University, Southwest Petroleum University); Ningrong Lai (INTI International University); Yuansu Zou (Southwest Petroleum University); Deshinta Arrova Dewi and Siti Sarah Maidin (INTI International University); Luobing Pan (Southwest University); and Chong Zhang (Southwest Petroleum University) Abstract Abstract As a cornerstone of smart city ecosystems, Structural Health Monitoring (SHM) leverages extensive Internet of Things (IoT) networks to enable real-time health assessment and prediction-based proactive maintenance, ensuring the operational integrity of infrastructure and preventing catastrophic failures. However, achieving reliable SHM within AI-enabled IoT (AIoT) frameworks faces a dual challenge: (1) the prohibitive cost of dense physical sensor deployment required for comprehensive coverage, and (2) the insufficient accuracy of existing time-series forecasting algorithms under complex spatiotemporal dynamics, which leads to unreliable proactive maintenance decisions. Thursday Virtual Room 6 IJCNN Paper Spatiotemporal Learning and Prediction III Session Chair: wei Tang (Civil Aviation University of China), xiao yu (University of Jinan) STCMR-AVSS:Spatio-Temporal Context modeling and refining for Airport Video Semantic Segmentation caihua Liu, wei Tang, junhao Chen, and xia Feng (Civil Aviation University of China) Abstract Abstract Existing airport video semantic segmentation methods exhibit significant limitations in spatio-temporal context modeling. The extracted motion features often lack scale adaptability and inter-frame temporal consistency is not fully maintained. To effectively capture multi-scale motion features and fully leverage inter-frame information consistency, We propose a Spatio-Temporal Context Modeling and Refining method. Firstly, a Multi-scale Motion Dynamic Aggregation (MMDA) module is designed to flexibly learn the connections between temporal dynamics and multi-scale spatial information. It employs a joint Spatio-Temporal Context modeling strategy that not only integrates motion information across time dimensions but also incorporates multi-scale motion details. This effectively enhances the coherence and correlation between temporal dynamics and spatial features. Secondly, a Cross-frame Refinement (CFR) module is introduced to mine useful information from different frames, keeping inter-frame temporal consistency.By leveraging the synergistic effort of these two modules, we obtain comprehensive multi-scale spatial and temporal information to enhance the representation of the target frame. Finally, enhanced features are concatenated with original target frame features and fed together into a lightweight decoder for segmentation.Extensive experiments on the AVSS benchmark demonstrate significant performance gains across multiple metrics over state-of-the-art methods. Seeing Wider, Seeing Longer: OmniGait for Gait Recognition in the Wild with Re-parameterized Spatio-Temporal Large Kernel Convolution Kewen Dong, Zuhan Fan, Yitong Li, and Pan Dou (Hebei Normal University) Abstract Abstract Gait recognition, as a long-range, non-invasive human identification technology, has garnered increasing attention because of its convenience and effectiveness. Compared to methods that achieve significant results in laboratory settings, gait recognition methods in outdoor scenarios face a more urgent need for robust morphological feature extraction and effective modeling of long-range gait sequences. To capture more discriminative gait features from outdoor data, this paper proposes OmniGait, a method based on reparameterized large kernel convolution. Specifically, this method innovatively proposes a spatio-temporal large kernel convolutional module(ST-LK), which incorporate two decoupled branches: spatial and temporal. The spatial branch leverages the large effective receptive field(ERF) of 2D large kernel convolution to capture more discriminative pedestrian morphological features from sparse contour data. The temporal branch employs efficient 1D convolution to model long gait sequences, enabling the model to acquire near-global temporal receptive fields. This allows it to directly capture long-term motion patterns spanning dozens of frames. Additionally, this paper proposes a dynamic gating mechanism(DGM) that enables the spatio-temporal large kernel convolutional module(ST-LK) to adaptively adjust the weights of convolutional branches of different sizes based on the input data. This approach enhances the robustness of the gait features extracted by the model. Extensive experiments demonstrate that our proposed method achieves outstanding results on both the popular outdoor gait datasets Gait3D and GREW. The source code will be available. STHP-Net:A Spatio-Temporal Collaborative High-Frequency Purification Network for Intracranial Artery Segmentation in DSA Sequences Qianwen Zhao, Peng Zhang, Zeyu Lv, Fengbo Xie, Hongxin Dong, and Pinle Qin (North University of China) and Xiaoxia Zhao (Shanxi Provincial People's Hospital) Abstract Abstract Digital subtraction angiography (DSA) is regarded as the gold standard for the diagnosis of cerebrovascular diseases owing to its high spatial resolution and dynamic blood flow imaging capability. Accurate segmentation of intracranial arteries is critical for the diagnosis, treatment, and prognosis of ischemic stroke. Nevertheless, existing methods still suffer from several challenges: traditional convolutional attention mechanisms are prone to missing small vessels due to inherent limitations in modeling local features, and current noise suppression strategies can hardly maintain a balance between preserving fine vascular details and eliminating noise at the feature level.To address these issues, we propose a Spatio-Temporal Collaborative High-Frequency Purification Network (STHP-Net). The network first exploits cross-frame complementary information in DSA sequences to maintain the integrity of vascular structures. We then design a Spatial-Channel Combined Attention Module (SCAM) to enhance the weak feature representation of thin vessels. Based on this, a High-Frequency Adaptive Purification Module (HAPM) is introduced to achieve accurate decoupling of vascular details and high-frequency noise in the feature space. Finally, a full-resolution decoding network is employed to generate the segmentation result.Experiments on the public DIAS dataset show that the Dice and clDice coefficient of STHP-Net reach 0.7909 and 0.7386, respectively, significantly outperforming state-of-the-art methods in terms of small vessel detection rate and vascular topological connectivity. DynaMo-Diff: Spatio-Temporal Dual-Guided Latent Diffusion for Consistent Stochastic Human Motion Prediction Xiao Yu and Na Lyu (Shandong Key Laboratory of Ubiquitous Intelligent Computing, University of Jinan, China) Abstract Abstract Stochastic human motion prediction (SHMP) aims to generate multiple plausible future motions from a short observation sequence. Latent diffusion models (LDMs) have recently become a competitive paradigm by performing denoising in a compact motion manifold, offering a favorable fidelity efficiency trade-off. However, denoising purely in a compressed latent domain presents a persistent challenge in preserving articulated skeletal structures and stabilizing long-horizon trajectories, which can lead to spatio-temporal inconsistencies such as topology violations and cumulative drift. This paper proposes DynaMo-Diff, a spatio-temporal dual-guided latent diffusion framework that integrates task-aligned guidance into the denoising process. DynaMo-Diff incorporates two complementary interfaces: (i) Dynamic Graph Attention (DGA) for local spatial reasoning, which employs Top-k sparsification to adaptively model motion-dependent inter-joint dependencies; and (ii) a Trajectory Anchor Predictor (TAP) for global temporal guidance, which forecasts sparse pose-space waypoints to mitigate long horizon error accumulation. To ensure manifold stability and computational efficiency, we adopt a decoupled two-stage training strategy with a fixed motion encoder. Extensive experiments on Human3.6M and AMASS demonstrate competitive performance and improved long-horizon stability under an efficient sampling budget. Notably, DynaMo-Diff achieves the best FDE of 0.452 on Human3.6M in our comparison while maintaining robust performance on the challenging AMASS benchmark. Thursday Virtual Room 7 IJCNN Paper Spatiotemporal Learning and Prediction IV Session Chair: Yuping Wang (Beijing University of Posts and Telecommunications), Haowei Mei (Xi'an Jiaotong University) Beyond RNNs and CNNs: A Unified Transformer for Spatiotemporal Prediction Haowei Mei, Zhe Huang, Zhaoyu Qi, and Bo Li (Xi'an Jiaotong University) Abstract Abstract Spatiotemporal prediction has long been dominated by recurrent and convolutional neural networks. While Transformers have achieved remarkable success in other domains, their application as pure, standalone architectures for this task remains notably under explored. Most existing approaches merely augment RNN or CNN backbones with attention mechanisms, preserving their inherent inductive biases. In this work, we present MeldFormer, a simple yet powerful model built entirely upon Transformers. It introduces a token grouping and fusion mechanism that unifies spatiotemporal modeling within a single, coherent attention-based framework, eliminating the need for separate modules and enabling efficient global context modeling. Extensive experiments on four standard benchmarks, including Moving MNIST, WeatherBench, TaxiBJ, and KTH, demonstrate that MeldFormer achieves state-of-the-art prediction accuracy while maintaining significantly lower computational complexity. WaveSFNet: A Wavelet-Based Codec and Spatial–Frequency Dual-Domain Gating Network for Spatiotemporal Prediction Xinyong Cai and Runming Xie (College of Software Engineering, Sichuan University) and Hu Chen and Yuankai Wu (College of Computer Science, Sichuan University; Sichuan University National Key Laboratory of Fundamental Science on Synthetic Vision) Abstract Abstract Spatiotemporal predictive learning aims to forecast future frames from historical observations in an unsupervised manner, and is critical to a wide range of applications. The key challenge is to model long-range dynamics while preserving high-frequency details for sharp multi-step predictions. Existing efficient recurrent-free frameworks typically rely on strided convolutions or pooling for sampling, which tends to discard textures and boundaries, while purely spatial operators often struggle to balance local interactions with global propagation. To address these issues, we propose WaveSFNet, an efficient framework that unifies a wavelet-based codec with a spatial–frequency dual-domain gated spatiotemporal translator. The wavelet-based codec preserves high-frequency subband cues during downsampling and reconstruction. Meanwhile, the translator first injects adjacent-frame differences to explicitly enhance dynamic information, and then performs dual-domain gated fusion between large-kernel spatial local modeling and frequency-domain global modulation, together with gated channel interaction for cross-channel feature exchange. Extensive experiments demonstrate that WaveSFNet achieves competitive prediction accuracy on Moving MNIST, TaxiBJ, and WeatherBench, while maintaining low computational complexity. Our code is available at https://github.com/fhjdqaq/WaveSFNet. Self-Distilled Spatiotemporal Mamba Model for Efficient Video Prediction Jiaxin Li and Kai Liu (Chongqing University), Chunhui Liu (The Hong Kong Polytechnic University), and Yantao Li (Chongqing University) Abstract Abstract Recent advances in deep learning and large-scale models technologies have significantly improved the performance of video prediction. However, it is non-trivial to strike a balance between computational cost and predictive accuracy. In this regard, this work introduces MambaVP, a novel method that embodies the principle of lighter steps for better futures, aiming to achieve higher accuracy of video prediction with lower computational overhead. MambaVP models spatiotemporal dependencies by decoupling spatial and temporal modeling using dimension-specific scanning in STMamba modules. To further ensure that intermediate features retain maximal information useful for predicting the future, we adopt an information-theoretic perspective and introduce a self-distillation strategy, where deep-layer outputs guide shallow-layer representations. This approach effectively minimizes the conditional entropy between the prediction target and intermediate features, thereby reducing uncertainty in future frame prediction. Extensive experiments on MMNIST, Human3.6M, and TaxiBJ demonstrate that MambaVP consistently achieves state-of-the-art (SOTA) performance while significantly reducing computational cost, setting a new benchmark for efficient video prediction. DST: Differential Spatiotemporal Learning for Contactless Electrocardiogram Reconstruction Yuping Wang, Shuai Zhao, Qizheng Wang, Jianquan Zhang, and Junliang Chen (Beijing University of Posts and Telecommunications) Abstract Abstract Continuous cardiac monitoring has been a challenge in contactless vital sign detection using mmWave radar systems. However, existing approaches primarily focus on the heartbeat detection and often fail to achieve comprehensive electrocardiogram (ECG) reconstruction. Therefore, a synergistic approach is proposed that integrates signal differential algorithm and spatiotemporal learning model (DST) to enable high-precision, contactless ECG monitoring. Initially, a differential method is used to achieve cardiac acceleration signal from the radar signals. Subsequently, a cascaded spatiotemporal learning model is proposed to extract spatial and temporal features of the cardiac acceleration signal. Experimental results indicate that our DST approach achieves timing precision with 90-percentile mean absolute error of 5.0 ms and morphological accuracy with 90-percentile Pearson-Correlation of 0.981 and Root-Mean-SquareError of 0.17 compared to the ground-truth ECG. Thursday Virtual Room 8 IJCNN Paper Speech and Audio Processing Session Chair: Bao Thang Ta (Viettel AI), Xin Dong (Nanjing University of Aeronautics and Astronautics) MSE-TTS: Emotion-Controllable Text-to-Speech via Multi-Scale Vision Feature Xin Dong (Nanjing University of Aeronautics and Astronautics); Yong Yang (Nanjing University of Aeronautics and Astronautics; Beijing Kedong Electric Power Control System Co., Ltd.); and ZhiHao Li and Qun Yang (Nanjing University of Aeronautics and Astronautics) Abstract Abstract Vision-Guided Text-to-Speech aims to generate style-matched speech based on input reference images. Although existing methods can leverage facial images to control speech style, they struggle to fully capture the multi-scale characteristics of facial expressions as well as the fine-grained interactions between emotional features and textual content. To address this challenge, we propose a novel multimodal speech synthesis framework, MSE-TTS. First, we design a Multi-Scale Vision Emotion Encoder to capture richer emotional details. A dual-supervision strategy combining classification and contrastive losses is introduced to enhance feature discriminative power. Second, we introduce an Emotion-Text Fusion Module that employs a bidirectional attention mechanism to achieve deep interaction between emotion vectors and text sequences. Experimental results demonstrate that MSE-TTS outperforms baseline models in both naturalness and emotional expressiveness of synthesized speech. Samples are available at https://mse-tts.github.io/MSE-TTS/. RLAIF-SPA: Structured AI Feedback for Semantic-Prosodic Alignment in Speech Synthesis Qing Yang, Zhenghao Liu, Yangfan Du, Pengcheng Huang, and Tong Xiao (northeastern university) Abstract Abstract Recent advances in Text-To-Speech (TTS) synthesis have achieved near-human speech quality in neutral speaking styles. However, most existing approaches either depend on costly emotion annotations or optimize surrogate objectives that fail to adequately capture perceptual emotional quality. As a result, the generated speech, while semantically accurate, often lacks expressive and emotionally rich characteristics. To address these limitations, we propose RLAIF-SPA, a novel framework that integrates Reinforcement Learning from AI Feedback (RLAIF) to directly optimize both emotional expressiveness and intelligibility without human supervision. Specifically, RLAIF-SPA incorporates Automatic Speech Recognition (ASR) to provide semantic accuracy feedback, while leveraging structured reward modeling to evaluate prosodic–emotional consistency. RLAIF-SPA enables more precise and nuanced control over expressive speech generation along four structured evaluation dimensions: Structure, Emotion, Speed, and Tone. Extensive experiments on LibriSpeech, MELD, and Mandarin ESD datasets demonstrate consistent gains across clean read speech, conversational dialogue, and emotional speech. On LibriSpeech, RLAIF-SPA consistently outperforms Chat-TTS, achieving a 26.1% reduction in word error rate, a 9.1% improvement in SIM-O, and over 10% gains in human subjective evaluations. MSBFCodec: A Low-Bitrate Neural Speech Codec with Multi-Scale Back-Projection Feature Fusion Tianyu Cai, Ye Li, Peng Zhang, Shuxian Ren, and Jingxiang Wang (Qilu University of Technology) Abstract Abstract With the development of communication technology, narrowband speech coding plays an indispensable role in low-bandwidth or resource-constrained environments. Therefore, research on narrowband speech coding is of great significance. Recently, neural speech coding has rapidly developed, and the most advanced methods show superior compression performance compared to traditional methods. Despite significant progress, existing methods still struggle with reconstructing details, especially at low bitrates. In this study, we introduce MSBFCodec, a neural speech codec that achieves state-of-the-art performance in narrowband low bitrate speech coding. MSBFCodec employs a multi-scale back-projection feature fusion method, effectively extracting multi-scale information and addressing the inconsistency in semantic and hierarchical information caused by multi-scale feature fusion. Furthermore, we propose a convolutional feedforward network module designed to perform efficient local modeling and enhance local speech feature representations. Experimental results show that MSBFCodec outperforms mainstream neural speech codecs in terms of reconstructed speech quality at 0.6 kbps and 0.8 kbps, with even higher quality reconstructed speech at 0.8 kbps than LyraV2 at 3.2 kbps. Quality-Aware Contrastive Self-Supervised Learning for Speech Quality Assessment Minh Tu Le (Viettel AI); Bao Thang Ta (Hanoi University of Science and Technology, Viettel AI); Phi Le Nguyen (Hanoi University of Science and Technology); and Van Hai Do (Thuyloi University) Abstract Abstract Speech Quality Assessment (SQA) plays a critical role in evaluating the performance of communication systems, yet the scarcity of labeled data significantly limits the development of robust SQA models. In this paper, we propose Contrastive learning for Quality Awareness (CQA), a novel self-supervised learning (SSL) framework specifically designed for speech quality assessment. Unlike conventional SSL models that prioritize content modeling, CQA focuses on learning quality-discriminative features by constructing contrastive pairs based on shared distortion types. Notably, two segments with different content can still be considered a positive pair if they share the same type of distortion. This design allows the model to focus solely on distortion characteristics, rather than the content of the original speech samples, effectively disentangling quality-related features from semantic information. Additionally, we introduce PESQ score prediction as an auxiliary task to align the pretrained SSL model more directly with the downstream SQA objective. Extensive experiments on the NISQA benchmark demonstrate that CQA significantly outperforms existing state-of-the-art methods across five quality dimensions: Mean Opinion Score (MOS), Noisiness, Coloration, Discontinuity, and Loudness. Specifically, our approach achieves a 2.1–6.2% improvement in Pearson’s Correlation Coefficient (PCC) and a 7.5–14.4% reduction in Root Mean Squared Error (RMSE) compared to the best existing methods on the TEST_FOR and TEST_P501 datasets, setting a new state-of-the-art in speech quality assessment. Thursday Virtual Room 1 IJCNN Paper Text Generation and Multilingual NLP I Session Chair: Jiarui Zhang (Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China), Haixian Zhang (Sichuan University) Interpretable Semantic Denoising via Neuron-level Functional Separation for Low-resource Multilingual Translation Jiarui Zhang, Mingzhe Lu, and Yifan Deng (Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China) Abstract Abstract Low-resource multilingual neural machine translation is essential for extending language technologies to underserved communities, yet progress remains limited by scarce parallel data and fragile cross-lingual generalization. Existing solutions primarily rely on data augmentation or architectural modifications, which often fail to disentangle semantics from syntax, leaving noisy representations that degrade translation quality. We propose Neuron-Distinction-Based Semantic Denoising (NDS), a neuron-level framework that generates semantically equivalent data via entropy-guided controlled sampling, compares parallel model activations to identify syntactic-noise neurons, and prunes them before fine-tuning. Our findings reveal a functional dissociation between semantic and syntactic neurons, providing a novel perspective on how large-scale neural networks encode cross-lingual invariants. Experimental results demonstrate that NDS bridges the performance gap between low- and high-resource translation while remaining highly parameter-efficient. Enhancing Cross-Lingual Embedding Alignment with Additive Keywords for International Trade Product Classification Angga Wahyu Anggoro, Padraig Corcoran, and Yuhua Li (Cardiff University) Abstract Abstract Cross-lingual embedding alignment plays an important role in enabling effective multilingual classification tasks. Although multilingual pretrained language models and fine-tuning techniques are increasingly adopted, current approaches inadequately address specialised domains, where domain-specific terminology and mixed-language content present unique challenges that hinder classification accuracy. This work considers the problem of automatically classifying text-based descriptions of international trade transactions with respect to an international standard Harmonized System (HS) code taxonomy. We propose a novel method that incorporates mixed-language keyword embeddings to improve cross-lingual alignment, focusing on bilingual models, and subsequently leverages this alignment for downstream classification tasks, with particular applicability to low-resource domains. Using a supervised learning framework implemented through neural network architectures, the model is trained on pairs of product descriptions and their corresponding extracted keywords. Experimental results on benchmark bilingual datasets demonstrate significant and consistent improvements in classification performance over baseline models, including in low-resource target language scenarios. The findings demonstrate the effectiveness of incorporating additive keywords as a strategy for cross-lingual embedding alignment, thereby enhancing representation quality and improving classification accuracy. Retrieve, Refine, and Translate: LLM-Based Translation for Low-Resource Languages Shibo Zhang, Mieradilijiang Maimaiti, Zhengyi Guo, and Dezhi Wang (Xinjiang University, Department of Computer Science); Wu Le, Zhuofei Xie, and Jiawei Chen (Integrated Laboratory for Space, Air, and Ground Systems); and Wushouer Silamu (Xinjiang University, Department of Computer Science) Abstract Abstract Large language models (LLMs) have demonstrated promising performance in machine translation, and retrieval-augmented generation (RAG) can further improve translation quality by incorporating external examples and leveraging the in-context learning capabilities of LLMs. However, in low-resource scenarios, empirical studies indicate that generated translations still exhibit diverse and persistent errors, and effectively refining such errors remains a challenge. In this work, we present a two-stage refinement framework for low-resource machine translation. The first stage focuses on correcting major translation errors through implicit refinement that exploits retrieved parallel examples, together with auxiliary knowledge guidance. The second stage introduces quality-oriented feedback to identify and revise the remaining errors iteratively. By integrating implicit correction with explicit feedback, the proposed framework provides a structured refinement process to improve translation quality. Extensive experiments on three benchmarks (FLORES-200, NTREX-128, and TICO-19) covering 8 low-resource languages demonstrate the effectiveness of our framework, showing that our introduced approach achieves highly competitive performance against strong baseline systems with XCOMET and BLEURT. The code is publicly available. GRAF: a Geometry-Guided Adaptation Framework for Protein Representation in Isolated Low-Resource Tasks Cheng Bi, Yi Feng, and Mingyuan Yang (Sichuan University); Han Wang (State Key Laboratory of Natural and Biomimetic Drugs, School of Pharmaceutical Sciences, Peking University, Beijing, China); and Huawei Cai and Haixian Zhang (Sichuan University) Abstract Abstract Effectively learning protein representations is crucial in isolated low-resource tasks, which are characterized by data scarcity and high sequence heterogeneity. In these scenarios, the prevailing paradigm of adapting broad representations of protein language models (PLMs) to specific biochemical contexts through fine-tuning or meta-learning is often constrained by limited sequence homology. To address these constraints, we propose GRAF, a geometry-guided representation adaptation framework that leverages intrinsic protein geometry to steer the adaptation process. We first establish the scale-aware geometric grounding module to restrict the search space within noisy representations. Subsequently, we devise the surface-guided geometric mapping module to decouple functional surface from global structural stability, allowing the extraction of conserved patterns without relying on external supervision. Extensive experiments demonstrate that GRAF achieves state-of-the-art performance on bacterial effector identification and antimicrobial peptide functional classification, validating its superiority in resolving the inherent challenges of biological isolation and data scarcity. Thursday Virtual Room 2 IJCNN Paper Text Generation and Multilingual NLP II Session Chair: 武 庄 (Civil Aviation University of China), Yuyan Zhou (University of Science and Technology of China) Hyperbolic Hierarchical Topic-Based Keyphrase Generation 武 庄 (Civil Aviation University of Chi) Abstract Abstract Keyphrases facilitate rapid comprehension of the content derived from speech recognition. Keyphrases can also concisely describe high-level topics discussed in documents that usually possess hierarchical topic structures. Thus, it is crucial to understand the hierarchical topic structures and employ them to guide the keyphrase identification. However, existing works that integrate topic information into keyphrase generation model are still confined to Euclidean space. Their ability to capture hierarchical structures is limited by the nature of Euclidean space. To this end, we design a hyperbolic hierarchical topic-based keyphrase generation method (Hyper-HTKG) to effectively exploit hierarchical topics to improve keyphrase generation performance. To the best of our knowledge, this is the first attempt to explore a hyperbolic hierarchical topic-based network for keyphrase generation. Experiments on five benchmark datasets show that Hyper-HTKG outperforms seven baseline methods. Where are next facilities: Emergency planning in non-Euclidean space using graph-based reinforcement learning Xilong Zhao, Wenxuan Guo, Yaohui Jin, and Yanyan Xu (Shanghai Jiao Tong University) Abstract Abstract Facility siting is a common class of combinatorial optimization problems with widespread applications in real-world scenarios. In applications such as logistics and delivery, the platform needs to perform quasi-real-time planning to meet the dynamic needs of users. Moreover, the establishments of shelters and temporary medical centers in response to unexpected events also require quasi-real-time planning. However, when it comes to large-scale scenarios, traditional combinatorial optimization solvers are unable to provide a satisfactory solution within the required time frame. Unlike the typical p-median problem, real-world facility planning mostly builds one or more additional facilities upon existing infrastructure rather than originating from scratch. Meanwhile, the distances between locations are often computed within non-Euclidean spatial contexts due to the constraints of road networks. For this scenario, we propose an innovative reinforcement learning method to incrementally predict optimal sites for new facilities by leveraging graph neural networks in non-Euclidean space. Our approach exhibits adaptability to accommodate fluctuations in population dynamics. Experiments results on real-world cities show that our method achieves over 1000 times speedup compared to Gurobi on the Los Angeles dataset with 2656 nodes, with the gap of solution under 5\% on all datasets. DACTM: Enhancing Neural Topic Models through Denoising Autoencoders with Intelligent Semantic Masking and Contrastive Learning YiLong Kang (Southwest University of Science and Technology); Jie Yin (National Science Library (Chengdu), Chinese Academy of Sciences); LiJuan Peng and ShiDi Xie (Southwest University of Science and Technology); ChunJiang Liu (National Science Library (Chengdu), Chinese Academy of Sciences); and HaiYun Xu (School of Business, Shandong University of Technology) Abstract Abstract Topic modeling is a key technique for uncovering latent thematic structure from large-scale text corpora. However, existing neural topic models largely rely on sparse bag-of-words inputs, which makes it difficult to adequately capture intra-document semantic dependencies; moreover, they are prone to latent topic mode collapse, resulting in weak document--topic associations and limited interpretability. To address these issues, we propose DACTM, a neural topic model enhanced by semantic awareness and contrastive learning. On the encoder side, DACTM introduces a Semantic-Aware Module that constructs a document-specific sparse semantic subgraph to improve document--topic alignment. On the decoder side, DACTM incorporates topic-level contrastive regularization to impose geometric constraints in the topic embedding space, which (i) increases semantic separation among topics and (ii) aligns topic representations with genuine word clusters, thereby alleviating mode collapse. Experiments on six benchmark datasets covering both long and short texts demonstrate that DACTM consistently outperforms representative neural topic models in terms of topic coherence, clustering performance, and the quality of learned document representations. Large Language Model for Efficient Algorithm Design of Combinatorial Optimization Solver Yuyan Zhou and Jie Wang (University of Science and Technology of China) Abstract Abstract The algorithm design in exact combinatorial optimization (CO) solver plays a fundamental role in operations research. However, due to the extensive requirements on domain knowledge and the large search space for algorithm design, the refinement on these algorithms remains highly challenging for both manual and learning-based paradigms. To tackle this problem, we propose a novel machine learning framework—large language model for exact combinatorial optimization solver (LLM4Solver)—to efficiently design high-quality algorithms of the CO solvers. The core idea is that, instead of searching in the high-dimensional and discrete symbolic space from scratch, we can utilize the prior knowledge learned from large language models to directly search in the space of programming languages. Specifically, we first use an LLM as the generator for high-quality algorithms. Then, to efficiently explore the discrete and non-gradient algorithm space, we employ a derivative-free evolutionary framework as the algorithm optimizer. Experiments on extensive benchmarks show that the algorithms designed by LLM4Solver significantly outperform all the state-of-the-art (SOTA) human-designed and learning-based policies (on GPU) in terms of the solution quality, the solving efficiency, and the cross-benchmark generalization ability. The appealing features of LLM4Solver include 1) the high training efficiency to outperform SOTA methods within ten iterations, and 2) the high cross-benchmark generalization ability on heterogeneous benchmark (i.e. MIPLIB 2017). LLM4Solver shows the encouraging potential to efficiently design algorithms for the next generation of CO solvers. Thursday Virtual Room 3 IJCNN Paper Text Generation and Multilingual NLP III Session Chair: 健 陈 (Shenzhen University), Xianzheng Liu (Central China Normal University) Video-guided Machine Translation with Global Video Context Jian Chen and JinZe Lv (Shenzhen University) and Zi Long and XiangHua Fu (Shenzhen Technology University) Abstract Abstract Video-guided Multimodal Translation (VMT) has advanced significantly in recent years. However, most existing methods rely on locally aligned video segments paired one-to-one with subtitles, limiting their ability to capture global narrative context across multiple segments in long videos. To overcome this limitation, we propose a globally video-guided multimodal translation framework that leverages a pretrained semantic encoder and vector database-based subtitle retrieval to construct a context set of video segments closely related to the target subtitle semantics. An attention mechanism is employed to focus on highly relevant visual content, while preserving the remaining video features to retain broader contextual information. Furthermore, we design a region-aware cross-modal attention mechanism to enhance semantic alignment during translation. Experiments on a large-scale documentary translation dataset demonstrate that our method significantly outperforms baseline models, highlighting its effectiveness in long-video scenarios. MemCam: Memory-Augmented Camera Control for Consistent Video Generation Xinhang Gao, Junlin Guan, Shuhan Luo, Wenzhuo Li, Guanghuan Tan, and Jiacheng Wang (Guilin University of Electronic Technology) Abstract Abstract Interactive video generation has significant potential for scene simulation and video creation. However, existing methods often struggle with maintaining scene consistency during long video generation under dynamic camera control due to limited contextual information. To address this challenge, we propose MemCam, a memory-augmented interactive video generation approach that treats previously generated frames as external memory and leverages them as contextual conditioning to achieve controllable camera viewpoints with high scene consistency. To enable longer and more relevant context, we design a context compression module that encodes memory frames into compact representations and employs co-visibility-based selection to dynamically retrieve the most relevant historical frames, thereby reducing computational overhead while enriching contextual information. Experiments on interactive video generation tasks show that MemCam significantly outperforms existing baseline methods as well as open-source state-of-the-art approaches in terms of scene consistency, particularly in long video scenarios with large camera rotations. Generating Fine-Grained Scene Graphs via Relational Semantic Enhancement Networks Yantao Shang, Liyan Ma, Xiangfeng Luo, Shaorong Xie, Wendi Rao, and Chao Ma (Shanghai University) Abstract Abstract Scene graphs, with their highly structured visual content representation, are widely applied in downstream tasks (e.g., visual question answering and human-computer interaction) and have become indispensable for scene understanding. However, existing models often struggle with long-tailed distributions and coarse-grained labeling, leading to uninformative predicate predictions such as ``on'' or ``has''. To mitigate this limitation, we introduce the Relational Semantic Enhancement Network (RSEN), which generates refined scene graphs with enhanced semantic depth. RSEN integrates three complementary modules: specifically, the Hybrid Visual Alignment module, which supplements CLIP's semantic cues with fine-grained spatial information derived from convolutional layers via cross-attention; the Hierarchical Context Interaction module, which uses a gated mechanism paired with confidence feedback to aggregate global contextual relationships, effectively suppressing relational noise; and the Semantic Prototype Refinement module, which reconstructs the feature space through the CLIP semantic prior, achieving prototype-guided metric learning and thereby capturing fine-grained semantic details. Experiments conducted on the Visual Genome and GQA datasets have demonstrated that RSEN significantly outperforms existing mainstream methods. The semantically rich scene graphs generated by RSEN can provide strong support for downstream scene understanding tasks. Semantic Collaborative Re-Mining for Temporal Sentence Grounding with Relevance Feedback Xianzheng Liu, Mingyao Zhou, Hao Sun, and Wei Xie (Central China Normal University) Abstract Abstract Temporal sentence grounding with relevance feedback (TSG-RF) evaluates the relevance of a query to a video. If relevant, it predicts related segments, otherwise, it determines that the query is irrelevant. Existing methods usually simply use two separate modules for evaluation and localization after query and video feature extraction. However, these methods fail to suppress interference from irrelevant video content, which often causes erroneous relevance judgments for given queries due to neglected semantic synergy. To leverage the intrinsic connections, we propose a semantic collaborative re-mining (SCR) method. Firstly, a prediction module was built to generate an initial segment containing semantic information of the module. Secondly, an enhancement strategy is designed to reduce query-irrelevant information in video features through the initial segment, generating enhanced visual features for modality interaction. Finally, a multi-scale relevance evaluation module is built to facilitate module interaction through enhanced visual features, optimizing the final feedback results. Experimental results on Charades-RF and Activity-RF datasets demonstrate that SCR alleviates the limitations of existing methods on TSG-RF. Thursday Virtual Room 4 IJCNN Paper Text Generation and Multilingual NLP IV Session Chair: Zhe Yang (Computer Network Information Center), Shifali Agrahari (Indian Institute of Technology Guwahati, India) MoRE: A Mixture-of-Representation-Experts Model for LLM Text Detection Under PGD Attacks Shifali Agrahari, Subhashi Jayant, and Sanasam Ranbir Singh (Indian Institute of Technology Guwahati, India) Abstract Abstract The ability of modern LLMs to craft text that closely aligns with human writing styles has intensified discussions about academic integrity and the credibility of online material. Existing detectors perform well on directly LLM-generated text but degrade sharply under simple adversarial edits such as paraphrasing or grammar-based perturbations. To address these limitations, we propose MoRE, a Mixture of Representation Experts framework that integrates stylometric, semantic, and contextual experts combined with PGD-based adversarial augmentation and dual loss. We further curate a comprehensive essay dataset containing human-written essays and essays generated by six major LLM families (Gemini, LLaMA, Mistral, Phi, Gemma, and DeepSeek), along with four adversarial attack variants that emulate real student evasion strategies. Extensive experiments across eight datasets show that MoRE achieves state-of-the-art performance in both binary and multi-class settings, maintaining high F1 scores across attacks, LLMs, and cross-model evaluations. Overall, MoRE establishes a strong, practical foundation for adversarially LLM-generated text detection in education and real-world scenarios. ErrorBench: Fine-Grained Error Analysis of Multi-Family LLMs in Data-to-Text Generation Soumya Bharadwaj and Ashish Anand (Indian Institute of Technology Guwahati) Abstract Abstract Large language models (LLMs) increasingly power Data-to-Text (D2T) generation, yet their reliability varies widely across architectures, training regimes, and parameter scales. In this work, we present the first cross-family, multi-scale analysis of fine-grained generation errors in D2T outputs. We rigorously evaluate 27 variants from nine major LLM families (Meta, OpenAI, Google, Alibaba, DeepSeek, Mistral AI, Microsoft, Technology Innovation Institute (TII), and SmolLM) using triple-to-text setting with structured knowledge derived from DBPedia. We manually annotate model outputs with a 10-category span-level taxonomy, enabling rich error profiling beyond standard semantic faithfulness metrics, resulting in a reusable span-annotated error dataset. Our findings reveal that while model scale generally improves quality, this trend is highly non-uniform; certain intermediate-scale models exhibit catastrophic failures due to Prompt Echo. We identify Addition and Omission as the dominant global errors, and show that families like Mistral and Qwen-3 are the most reliable architectures for structured D2T. Complementing these findings, prompt ablation studies demonstrate that stricter instructional constraints can reduce specific errors in capable models but are ineffective against persistent failures in lower-capacity architectures, highlighting the primacy of model capacity over prompt design. Overall, this study quantifies distinctive error fingerprints across model lineages, offering actionable guidance for building more dependable structured-data generation systems. The dataset is publicly available at: \href{https://huggingface.co/datasets/soumyaBharadwaj/ErrorBench}{https://huggingface.co/datasets/soumyaBharadwaj/ErrorBench}. RelaClass: A Residual Tree with LLM-Augmented Prototypes for Hierarchical Text Classification Wenting He, Qiang Yan, Bing Li, Huaming Liao, Jiafeng Guo, and Xueqi Cheng (State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Abstract Abstract Hierarchical text classification (HTC) aims to assign documents to a predefined hierarchical taxonomy but often requires extensive manually labeled data for training. Although large language models (LLMs) show promise in zero-shot settings, they struggle to capture fine-grained hierarchical dependencies due to limitations in prompt expressiveness and structural inductive bias. To address these limitations, we propose RelaClass (A Residual Tree Framework with LLM-Augmented Prototypes), a novel weakly supervised framework that integrates the semantic richness of LLMs with an efficient, taxonomy-aware residual tree model, which relies solely on class names provided by the taxonomy as supervision. RelaClass addresses two core challenges: (1) structural alignment: by reformulating the taxonomy into a hierarchy of codebooks that recursively decompose document representations into taxonomic residuals, enabling coarse-to-fine classification; and (2) semantic anchoring: by leveraging LLMs to generate discriminative prototypes that provide high-quality supervisory signals without manual labeling. Experiments on two benchmark datasets show that RelaClass outperforms state-of-the-art weakly-supervised methods and zero-shot LLM prompting approaches. Code and data are available at https://github.com/hexiaoting/RelaClass. LLM-based Grant Proposal Structuring via Classification, Extraction, and Compression Zhe Yang, Wenjing Chang, Yue Wang, and Xiaojun Zhou (Computer Network Information Center) Abstract Abstract Research grant proposals contain rich scientific ideas and technical methodologies, serving as a critical resource for AI for Science. However, their multi-disciplinary nature, scattered distribution of core information, and heterogeneous writing styles pose significant challenges for neural document understanding and long-context reasoning. To address these chal- lenges, we propose GRIST (Grant pRoposal Intelligent STruc- turing), a structure- and semantic-aware LLM-based framework that transforms lengthy proposals into structured, high-density representations. GRIST integrates instruction-driven discipline classification, metadata extraction, and semantic compression, enabling automated analysis through in-context learning without task-specific training. To facilitate research in this domain, we construct and release CGP-Bench, a Chinese grant proposal benchmark comprising 300 anonymized proposals curated for privacy compliance while preserving scientific logic. We conduct a comprehensive empirical analysis of mainstream LLMs on CGP-Bench, providing insights into their generalization boundaries and reasoning capabilities in complex grant proposal processing. Thursday Virtual Room 5 IJCNN Paper Time Series Classification and Anomaly Detection Session Chair: zhang zejun (shenyang aerospace university), Hongbo Zhao (East China Normal University) HRD-Net: Hierarchical Relational Diffusion Network for Anomaly Detection in Multivariate Time-series Data Hongbo Zhao, Zhe Wang, Yinjie Teng, Yihong Huang, and Kai Zhang (East China Normal University) Abstract Abstract Multivariate time series anomaly detection is critical for ensuring the safety of complex systems with wide applications. Recently, Graph neural networks have emerged as a powerful tool for anomaly detection due to its ability to model intricate dependencies among different system components. However, the multi-layer information aggregation in GNNs often results in a receptive field that is insensitive to local topological contexts or ``edge densities''. This may introduce irrelevant nodes across community boundaries to contaminate message passing, degrade the prediction, and finally lead to false positive alerts. In this paper, we propose Hierarchical Relational Diffusion Network (HRD-Net) to build contextualized relational fields in GNNs for anomaly detection. HRD-Net uses random-walk based diffusion to recover the relational context for each node, with adaptive restart probabilities to balance the trade-off between local communities and global interactions under different scales. The resultant probability maps serve as a hierarchical ``topological firewall'' to suppress spurious correlations while preserving truly relevant neighbors to guarantee the quality of message passing. Furthermore, HRD-Net enables the construction of a more reliable multi-scale anomaly scoring system by integrating reconstruction errors from both individual scales and their aggregated representation. Experiments on six challenging benchmarks demonstrate the potential of HRD-Net over altogether 12 methods, with case studies to interpret and validate its capability to disentangle equipment faults from sensor drift artifacts in real-world maintenance logs. AMGNet: Adaptive Multi-Scale Gated GNN and Hierarchical Fusion for Time Series Forecasting Xin Gao, Tianyu Chang, Dongming Chen, and Dongqi Wang (Northeastern University) Abstract Abstract Real-world time series data contain multiple temporal patterns at different scales that are intricately intertwined, making forecasting extremely challenging. Recently, Graph Neural Networks (GNNs) have demonstrated significant potential in multivariate time series forecasting due to their ability to explicitly model inter-variable dependencies. However, existing methods typically rely on static and uniform aggregation rules, failing to adapt to the varying informational needs of different nodes across their multi-hop neighborhoods. This limitation hinders their effectiveness in modeling inter-series dependencies. Moreover, current models often suffer from error accumulation due to a lack of explicit cross-scale information interaction. To address these issues, we propose AMGNet. Our framework incorporates three key components: (1) A Node-Aware Gated Message Passing (NAGMP) module operating on scale-specific graph structures to adaptively capture node-specific correlations. This design effectively filters neighborhood noise and enhances the modeling of complex inter-series dependencies. (2) A coarse-to-fine hierarchical guidance structure that utilizes coarse-grained dominant trends to explicitly calibrate fine-grained feature learning, thereby facilitating cross-scale information interaction and mitigating error accumulation. (3) Finally, a Multi-Predictor Synthesis (MPS) module based on spectral energy is adopted to achieve a noise-resilient prediction ensemble. Extensive experiments on six datasets spanning diverse domains demonstrate the superiority of AMGNet. FAMNet: Frequency-Aware Adaptive Multi-Scale Network for Multivariate Time Series Classification Liyu Luo, Jiyao Gao, Jiaxuan Xu, Qingbao Guan, Jie Zuo, and Lei Duan (Sichuan University) Abstract Abstract In time series analysis, Multivariate Time Series Classification (MTSC) plays a crucial role in diverse real-world applications. Recently, integrating multi-scale analysis into deep learning models has been effective in improving performance for MTSC. However, existing multi-scale methods typically employ fixed scales, which prevents models from selecting the most discriminative scales for each input. To address this limitation, we propose the Frequency-aware Adaptive Multi-scale Network (FAMNet) for MTSC. Our method uses the Fourier transform to extract frequency characteristics of the input time series and scores candidate scales accordingly. To make scale selection trainable, we introduce Gumbel-Softmax, allowing the model to adaptively select multiple suitable scales per sample. The selected scales are then used for multi-scale patch embedding and integrated through cross-scale fusion for hierarchical representation. Experiments on 20 datasets demonstrate that the proposed FAMNet achieves the best accuracy compared to advanced baselines in MTSC tasks. TD-Hyper: Time-Density Guided Hypergraph Learning for Clinical Irregular Multivariate Time Series Guohui Ding and Zejun Zhang (Shenyang Aerospace University) and Jinwei Wang (Peking University First Hospital) Abstract Abstract Irregular multivariate time series (IMTS) are ubiquitous in clinical settings, where observations are asynchronously recorded at non-uniform intervals across variables.Most existing methods mitigate temporal irregularity through resampling, interpolation, or alignment, implicitly treating sampling density as a nuisance rather than an informative signal.However, in clinical practice, measurement frequency itself reflects patient state, giving rise to alternating dense and sparse sampling regimes along the timeline.In this paper, we propose TD-Hyper, a density-conditioned hypergraph framework that treats sampling density as a first-class structural variable governing both temporal neighborhood construction and information propagation.TD-Hyper models asynchronous observations with soft adaptive temporal hyperedges and introduces a density-gated dual-mode message passing mechanism to differentiate dense and sparse sampling regimes without resampling or explicit temporal alignment.Experiments on multiple clinical IMTS benchmarks demonstrate that TD-Hyper consistently outperforms strong baselines while maintaining a moderate computational footprint. Thursday Virtual Room 6 IJCNN Paper Time Series Forecasting II Session Chair: jiannan wan (Fuzhou University, College of Computer and Data Science), Wenhui Hu (Institute of Plasma Physics, Chinese Academy of Sciences) AMCRNet: Adaptive Multi-Cycle Residual Modeling for Long-Term Time Series Forecasting Jiannan Wan, Fei Chen, Xunxun Zeng, Hang Cheng, and Meiqing Wang (Fuzhou University) Abstract Abstract Time-series forecasting is fundamental to many real-world applications such as weather and transportation. In this paper, we propose AMCRNet, a long-term forecasting framework that integrates explicit multi-cycle modeling with correlation-aware residual learning in an end-to-end manner. Specifically, AMCRNet maintains multiple learnable cyclic prototypes associated with different cycle lengths. Periodic components are extracted via cyclic indexing, which enables direct and interpretable modeling of stable periodic patterns. A sample-wise gating fusion is applied to adaptively assign the explanatory contribution of each cycle to the input sequence. Meanwhile, the same cycle weights are reused to construct a consistent future-cycle compensation term for the forecasting horizon. This design prevents periodic information loss after decomposition. On the residual space, AMCRNet adopts a correlation-aware channel modeling to capture inter-variable relationships. A learnable fusion coefficient is introduced to adaptively balance the residual shortcut and correlation-enhanced representations. Extensive experiments on multiple benchmark datasets demonstrate that AMCRNet consistently achieves strong long-term forecasting performance with improved robustness under composite periodicities and non-stationary perturbations. TimeCF: A novel model with adaptive Convolution and Sharpness-Aware Minimization Frequency Domain Loss for long-term time series forecasting Bin Wang, Heming Yang, and Jinfang Sheng (Central South University) Abstract Abstract Recent studies have shown that by introducing prior knowledge, multi-scale analysis of complex and non-stationary time series in real environments can achieve good results in the field of long-term time series forecasting. However, this kind of models may produce suboptimal prediction results due to the correlation in time series data which in turn affects the generalization performance of the model. To address this challenge, we were inspired by the idea of Sharpness-Aware Minimization and the recently proposed FreDF method and designed a deep learning model TimeCF for long-term time series forecasting based on the decomposition-learning-mixing architecture, combined with adaptive convolution information aggregation module and Sharpness-Aware Minimization Frequency Domain Loss (SAMFre). Specifically, TimeCF first decomposes the original time series into sequences of different scales. Next, the same-sized convolution modules are used to adaptively aggregate information of different scales on sequences of different scales. Then, decomposing each sequence into season and trend parts and the two parts are mixed at different scales through bottom-up and top-down methods respectively. Finally, different scales are aggregated through a Feed-Forward Network. What's more, extensive experimental results on different real-world datasets show that TimeCF has excellent performance in the field of long-term forecasting. Our code is available at https://github.com/iHccob/TimeCF_IJCNN2026. PMRMamba: Progressive Multi-Resolution Mamba for Long-Term Time Series Forecasting Dejiao Niu, Yuxuan Yang, Tao Cai, Wenyi Xiao, Chengyu Zhang, and Liushan Zhang (Jiangsu University) Abstract Abstract Long-term time series forecasting remains a challenging task, as existing models struggle to capture long-range dependencies while maintaining computational efficiency. Recently, Mamba has demonstrated the ability to capture long-range dependencies with linear complexity, opening up new perspectives for time series modeling. Although Mamba excels at modeling long-range dependencies with linear complexity, the single-scale state propagation makes it difficult to simultaneously capture long-term trends and short-term variations within time series. In this paper, we propose a novel framework called Progressive Multi-Resolution Mamba (PMRMamba) which models time series through a coarse-to-fine progressive residual paradigm. PMRMamba consists of two tightly coupled modules that operate in conjunction with Selective SSM, progressively modeling time series at multiple resolutions to construct hierarchical temporal dependencies across different time scales. The Progressive Multi-Resolution Residual Decomposition progressively decomposes a time series into a set of complementary multi-resolution residual features through a stage-wise residual refinement mechanism, while the Temporal Resolution-wise Fusion adaptively integrates the multi-resolution residual features by explicitly estimating each resolution’s contribution to the overall temporal representation. Extensive experiments on eight real-world benchmark datasets demonstrate that PMRMamba consistently outperforms state-of-the-art models, achieving significant improvements in long-term forecasting accuracy. NJS-KAN: Neural Jacobi-Spectral KANs for Long-term Time Series Forecasting Yucheng Wang (University of Science and Technology of China; Institute of Plasma Physics, Chinese Academy of Sciences); Wenhui Hu (Institute of Plasma Physics, Chinese Academy of Sciences); and Qiping Yuan and Bingjia Xiao (Institute of Plasma Physics, Chinese Academy of Sciences; University of Science and Technology of China) Abstract Abstract Long-term Time Series Forecasting (LTSF) remains a challenging endeavor due to the intricate coupling of non- stationary dynamics and high-dimensional dependencies. While recent linear models prioritize efficiency, they often lack the capacity to model highly non-linear volatility, whereas standard Kolmogorov-Arnold Networks (KANs) face prohibitive parame- ter explosion and numerical instability when applied to multi- variate tasks. To bridge this gap, we propose the Neural Jacobi- Spectral Kolmogorov-Arnold Network (NJS-KAN), a novel archi- tecture designed to reconcile high-capacity non-linear modeling with computational efficiency. Our approach introduces a Low- Rank Jacobi-KAN layer that utilizes orthogonal Jacobi basis functions and imposes a CP-tensor decomposition on the weight space, effectively reducing parameter complexity from quadratic to linear while ensuring gradient stability. Furthermore, we design a Dual-Domain Gating mechanism that synergizes local temporal convolutions with global frequency-domain spectral features to overcome the spectral bias of local operators. Ex- tensive experiments on six real-world benchmarks demonstrate that NJS-KAN achieves state-of-the-art performance, consistently outperforming recent Transformer and linear baselines in both predictive accuracy and memory efficiency.Code is available at https://github.com/hyfqphy/NJS-KAN.git. Thursday Virtual Room 7 IJCNN Paper Time Series Forecasting III Session Chair: Guoqing Tang (University of Science and Technology of China), Yu Wang (Chengdu Institute of Computer Applications, University of Chinese Academy of Sciences) WPKAN-Mixer: Long-Term Time Series Forecasting with Adaptive Function Learning on Wavelet-Decomposed Components Guoqing Tang, Hongwei Zhao, Lei Liu, Tengyuan Liu, Jiahui Huang, and Bin Li (University of Science and Technology of China) Abstract Abstract Time series forecasting is widely applied in fields like weather, power, and traffic management, but the complex multi-scale dynamics of real-world sequences make accurate forecasting extremely challenging. Existing methods typically employ effective decomposition techniques for time series forecasting tasks, which simplify complex raw series into more tractable sub-series (e.g., trend, seasonality) for targeted modeling. However, these approaches remain subject to inherent limitations: their modeling units apply a unified representation paradigm to all components, which hinders the adaptive modeling of local nonlinear dynamics across components at multiple scales, thus limiting fitting accuracy. This paper introduces WPKAN-Mixer for long-term time series forecasting, comprising primarily a Wavelet Transform module, a Patching-Embedding module and a dual Kolmogorov-Arnold Networks (KAN) Mixers module. We leverage the time-frequency localization of the wavelet transform module to achieve fine-grained temporal-signal decomposition. Furthermore, we design a dual KAN Mixers module incorporating the spline-based learnable function space to ``tailor-make'' an optimal functional representation for each frequency component, enabling flexible capture of non-linear dynamics characteristics from both temporal and feature dimensions. Experiments show that WPKAN-Mixer significantly outperforms state-of-the-art models on public benchmarks, achieving an average MSE and MAE reduction of 3-5% on most datasets, while also demonstrating excellent robustness and superior computational efficiency. FreqMixer: Adaptive Cross-Frequency Learning for Time Series Forecasting Yi Zhou, Yaojun Liang, Chong Chen, Tao Wang, and Lianglun Cheng (Guangdong Provincial Key Lab of Cyber-Physical System Guangdong University of Technology) Abstract Abstract Real-world time series often exhibit complex and non-stationary temporal dynamics, posing substantial challenges for accurate forecasting. While existing models predominantly rely on simplified decomposition for trends and seasonality, they often struggle to disentangle the multifaceted frequency components inherent in real-world data, thereby limiting their representational power. To address this, we propose FreqMixer, a novel MLP-based framework designed to capture rich temporal patterns through a frequency-aware lens. The architecture begins by leveraging a Stationary Wavelet Transform (SWT) to decompose series into multi-frequency coefficients, enabling simultaneous modeling in both time and frequency domains. Building on this decomposition, a frequency-adaptive patching strategy is introduced to handle the heterogeneous characteristics of different bands, effectively differentiating long-term trends from short-term fluctuations. This hierarchical representation is further refined by a dual-level cross-frequency attention mechanism, which integrates local multi-scale information through MLP-based aggregation while capturing global dependencies across wavelet coefficients. Extensive experiments on diverse real-world datasets demonstrate that FreqMixer consistently achieves state-of-the-art performance, validating its effectiveness and robustness in complex time series analysis. WGMixer: Adaptive Time-Frequency Forecasting via Asymmetric Wavelet Mixture of Experts Chongfeng Liu, Jiwei Qin, and Dezhi Sun (Xinjiang University); Jungang Ma (Xinjiang Uyghur Autonomous Region Institute of Metrology and Testing); Xuesong Liu (Goldwind Science); and Duxiang Chen (Xinjiang University) Abstract Abstract Long-term time series forecasting faces the challenge of balancing global trends and transient variations in non-stationary signals. While wavelet-based methods offer time-frequency localization advantages, they suffer from the Static Basis Dilemma---relying on fixed basis functions that fail to adapt to dynamic signal evolution. To address this, we propose WGMixer, an adaptive framework based on an asymmetric wavelet mixture of experts. Instead of seeking a universal basis, WGMixer constructs two complementary paths: a Main Path using Symlet-4 to anchor long-term trends, and an Expert Path using Daubechies-4 to capture high-frequency abrupt changes. We introduce a MultiScaleSwitchGate to perform instance-level dynamic routing based on local signal energy. Experiments on 10 datasets (seven standard benchmarks and three complex environmental datasets) show that WGMixer matches the performance of state-of-the-art baselines on stationary tasks. On highly non-stationary environmental data, it reduces Mean Squared Error by an average of 5.6% against the strongest competitor, with a maximum reduction of 30.1% in transient-heavy scenarios. Visualizations confirm the model's ability to adaptively switch between trend preservation and anomaly detection, offering a robust and interpretable solution. Macro-Meso-Micro: A Hierarchical Scale Disentanglement Framework for Multiscale Dynamic Time Series Forecasting Yu Wang, Xiaolin Qin, and Jiacen Liu (Chengdu Institute of Computer Applications, University of Chinese Academy of Sciences) Abstract Abstract Real-world time series are often composed of complex, entangled nonstationary patterns and multiscale dynamics, posing significant challenges for forecasting. Existing homogeneous architectures are constrained by the Single-Mapping Bottleneck, failing to simultaneously capture differentiated dynamic characteristics. To address this, we propose Macro-Meso-Micro Hierarchical Scale Disentanglement (M3-HSD), a hierarchical scale disentanglement framework employing a Macro-Meso-Micro strategy. The core innovation lies in the Hierarchical Scale Disentangler (HSD), which orthogonally decomposes the series into macro-inertia, meso-manifolds, and micro-transients. To model these distinct dynamics, we design a Heterogeneous Expert Decoding Module (HEDM): (1) The Global Inertia Projector (GIP) leverages the global receptive field of MLPs to capture low-frequency inertial trends; (2) The Adaptive Manifold Fitter (AMF) innovatively incorporates Kolmogorov-Arnold Networks (KAN) convolutions, utilizing learnable spline functions to precisely fit complex nonlinear seasonal manifolds at the meso-scale; and (3) The Local Transient Encoder (LTE) employs multiscale CNNs to capture high-frequency micro-fluctuations. Finally, an Adaptive Gating Fusion mechanism dynamically integrates multistream outputs to enhance robustness. Extensive experiments on mainstream benchmarks demonstrate that M3-HSD significantly outperforms existing state-of-the-art (SOTA) methods, exhibiting superior performance, particularly in handling multiscale complex periodic patterns and non-stationary distributions. The source code is available at: https://github.com/WYyeshang/M3-HSD. Thursday Virtual Room 8 IJCNN Paper Time Series Forecasting IV Session Chair: Jiangrong Yang (Beijing Normal University), Dongwei Liu (South China Normal University) STENet: A Soft Translation-Equivariant Network for Time Series Forecasting Dongwei Liu, Weiping Zheng, and Huande Liu (South China Normal University) Abstract Abstract Time series forecasting plays a critical role in a wide range of real-world applications. However, translation equivariance remains an important yet relatively underexplored property in this field. Existing forecasting models typically neglect this property and are therefore forced to implicitly learn equivariant behavior from data, resulting in reduced model capacity. In this paper, we propose STENet, a novel forecasting architecture that explicitly encodes translation equivariance via a novel point-wise accumulation prediction paradigm, while adaptively relaxing the equivariance constraint when it is violated. Extensive experiments on 11 real-world multivariate time series datasets demonstrate that STENet achieves competitive performance with state-of-the-art methods. Moreover, quantitative analysis using an empirical equivariance metric provides empirical evidence for the effectiveness and usefulness of incorporating adaptive translation equivariance in time series forecasting models. Time-Prompt: Integrated Heterogeneous Prompts for Unlocking LLMs in Time Series Forecasting Zesen Wang, Yonggang Li, and Lijuan Lan (Central South University) Abstract Abstract Time series forecasting aims to model temporal dependencies among variables for future state inference, holding significant importance and widespread applications in real-world scenarios. Although deep learning-based methods have achieved remarkable progress, they often struggle with long-term forecasting and few-shot scenarios. Recent research demonstrates that large language models (LLMs) achieve promising performance in time series forecasting, but the full potential of LLMs in understanding time series remains largely untapped. To address this, we propose Time-Prompt, a framework for activating LLMs for time series forecasting. Specifically, we first construct a unified prompt paradigm with learnable soft prompts to guide the LLMs' behavior and textualized hard prompts to enhance the time series representations. Second, to enhance LLMs' comprehensive understanding of the forecasting task, we design a semantic space embedding and cross-modal alignment module to facilitate the fusion of temporal and textual data. Finally, we efficiently fine-tune the LLMs' parameters using time series data. Furthermore, we apply our method to carbon emission forecasting, contributing to the technical advancements aiding global carbon neutrality. Comprehensive evaluations on 6 public datasets and 3 carbon emission datasets demonstrate that Time-Prompt is a powerful framework for time series forecasting. Our code is publicly available at https://github.com/SanMuGuo/Time-Prompt. Forecasting with Guidance: Representation-Level Supervision for Time Series Forecasting Jiacheng Wang (Xijing University), Liang Fan and Baihua Li (Loughborough University), and Luyan Zhang (Northeastern University) Abstract Abstract Nowadays, time series forecasting is predominantly approached through the end-to-end training of deep learning architectures using error-based objectives. While this is effective at minimizing average loss, it encourages the encoder to discard informative yet extreme patterns. This results in smooth predictions and temporal representations that poorly capture salient dynamics. To address this issue, we propose ReGuider, a plug-in method that can be seamlessly integrated into any forecasting architecture. ReGuider leverages pretrained time series foundation models as semantic teachers. During training, the input sequence is processed together by the target forecasting model and the pretrained model. Rather than using the pretrained model's outputs directly, we extract its intermediate embeddings, which are rich in temporal and semantic information, and align them with the target model's encoder embeddings through representation-level supervision. This alignment process enables the encoder to learn more expressive temporal representations, thereby improving the accuracy of downstream forecasting. Extensive experimentation across diverse datasets and architectures demonstrates that our ReGuider consistently improves forecasting performance, confirming its effectiveness and versatility. AutoMixer: Automated Multi-Scale Mixing for Time Series Forecasting Based on Neural Architecture Search Jiangrong Yang (Beijing Normal University), Haodi Wang (City University of Hong Kong), Tangyu Jiang (Hong Kong Baptist University), and Maiyun Zhang and Rongfang Bie (Beijing Normal University) Abstract Abstract Time series forecasting has become an emerging paradigm for predicting future temporal variations. Real-world time series data exhibits heterogeneous fluctuations at different scales, thus demonstrating multi-scale characteristics. Previous works either rely heavily on the manual design of the featuremixing architecture or suffer from the high computational cost during the search procedure. Moreover, the performance of the existing methods can be further enhanced. To address the above issues, in this paper, we propose an automated multi-scale mixing approach for time series forecasting based on the Neural Architecture Search (NAS) called AutoMixer. Unlike the previous works, AutoMixer can automatically generate the adaptive fusion architecture for both multivariate and single-variate datasets. By proposing a customized search space and concrete training algorithm, AutoMixer realizes an efficient multi-scale featuremixing without invoking extra expert knowledge. The experimental results show that AutoMixer outperforms the state-of-the-art methods with 6 to 70 times faster search time compared to other NAS-based schemes with state-of-the-art accuracy. For instance, our method achieves 11.98 of MAPE on PEMS04, which is 4.3% better than the best baseline. Thursday Virtual Room 1 IJCNN Paper Time Series Forecasting V Session Chair: Jiayang Xu (Southern University of Science and Technology), Zenglin Xu (Fudan University) QuadFocusNet: Learning to Route Multi-Scale Patches for Multivariate Time Series Forecasting Jiayang Xu (Southern University of Science and Technology) Abstract Abstract Time series forecasting is essential yet challenging across various domains such as energy management, climate monitoring, and traffic flow prediction. Recent methods, such as PatchTST, have achieved strong performance. However, they rely on fixed-scale patching, which limits adaptability to non-stationary distributions and diverse temporal patterns. Other approaches, such as DLinear, are efficient but struggle to capture complex nonlinear dependencies. To address these limitations, we propose QuadFocusNet, a novel framework built upon multi-scale adaptive patching. Instead of using a static patch size, QuadFocusNet extracts multi-granular temporal representations at multiple scales in parallel. It then dynamically fuses them through a learnable router, enabling content-aware scale selection and efficient information flow. Moreover, we incorporate prompt learning to inject global context and improve representation robustness under distribution shifts. Extensive experiments on the ETT, UCI Household Power Consumption, and Jena Climate datasets demonstrate that QuadFocusNet achieves competitive state-of-the-art performance. The model consistently delivers superior or competitive results across diverse settings. These range from immediate prediction tasks to medium-term forecasting at 24- and 48-step horizons, up to long-term forecasting at a 96-step horizon. Code is available at https://github.com/David-SUSTech/QuadFocusNet.git. Multi-Scale Convolution with Optimal Transport Attention Effect on Multivariate Time Series HaoChong Fu (University of Macau) and Jian Xu (RIKEN AIP) Abstract Abstract The analysis of Multivariate Time Series (MTS) plays an important role in a lot of real-world practical applications, but it still remains some challenging problem about capturing multi-granularity structural patterns and suppressing noise appropriately. Multi-Scale Convolution with Optimal Transport Attention (MSC-OT) is proposed in this paper. MSC-OT is a useful architecture to optimize the attention mechanism. It combines multi-scale convolution with Sinkhorn optimal transport method based on inverted embedding. The inverted embedding approach embeds each variable as a token and allows the model to capture cross-variate relationships better. MSC-OT consists of two part: (1) Multi-Scale Convolution Enhancement, that applies multi-scale convolutions to attention score matrices based on inverted embedding, capturing local structural patterns in the variate-interaction space induced by compressed temporal representations; (2) Sinkhorn Optimal Transport Regularization, that formulates attention computation as an optimal transport problem and employs iterative matrix scaling to ensure balanced information flow across variates. Adaptive Fusion Strategy utilizes softmax-normalized learnable weights to dynamically combine base attention, convolution-enhanced, and OT-regularized scores. Experiments on widely-used datasets, including ETT, Electricity, Traffic, Solar-Energy, and Exchange-Rate, show that MSC-OT achieves well performance in both short-term and long-term forecasting tasks. Ablation experiments further validate the effectiveness of each proposed component and their synergistic contributions to improving prediction accuracy for multivariate time series forecasting. Enhancing Time Series Forecasting via Distribution Calibration and Prototypical Pattern Modeling Tielin Yin and Ping Li (Changsha University of Science and Technology) Abstract Abstract Multivariate Time Series Forecasting (MTSF) is crucial in domains such as energy dispatch and traffic management. Recently, the inverted embedding paradigm that maps variable sequences as independent tokens has significantly enhanced models' representational capacity. However, practical applications still face two major bottlenecks: (1) feature distribution distortion caused by noise interference leads to normalization failure, preventing models from capturing stable data distributions; (2) traditional value embedding ignores behavioral pattern disparities among heterogeneous variables, resulting in blind coupling between different variables. To address these challenges, this paper proposes the AW-PVA model. First, we design the Adaptive Weighted Reversible Instance Normalization module (AW-RevIN), which dynamically identifies reliable signals through adaptive weighting to suppress outlier interference and enhance distribution fidelity; meanwhile, sparse regularization is introduced to guide the model in accurately anchoring key distributional characteristics. Second, we design the Prototypical Vector Augmentation (PVA) module, which constructs a library of typical behavior patterns through K globally shared prototypical vectors, enhancing the discrimination of heterogeneous variables while significantly improving the modeling precision of complex variable relationships. Experiments demonstrate that AW-PVA outperforms existing state-of-the-art models in long-term forecasting tasks. FreDC: A Frequency-Enhanced Dynamic Clustering Network for Multivariate Time Series Forecasting Mengna Hu (College of Computer Science and Technology, Harbin Institute of Technology); Jinghua Wang (College of Computer Science and Technology, Harbin Institute of Technology; State Key Laboratory of Smart Farm Technologies and Systems); and Zenglin Xu (Artificial Intelligence Innovation and Incubation (AI³) Institute, Fudan University) Abstract Abstract Multivariate time series forecasting plays a critical role in various real-world applications. However, most existing approaches for modeling inter-variable dependencies overlook the temporal dynamics and structural sparsity. Meanwhile, existing frequency-domain methods primarily focus on amplitude information and neglect phase information, which is crucial for preserving temporal alignment. To address these limitations, we propose FreDC, a lightweight yet effective model that integrates Frequency-domain enhancement with variable Dynamic Clustering. Our FreDC exploits complementary amplitude and phase information in the frequency domain to construct informative prior representations. To capture time-varying and sparse inter-variable correlations, we introduce a Mixture-of-Experts (MoE)–based variable dynamic clustering module. The MoE consists of multiple private experts for modeling cluster-specific variable dynamics and a shared expert for capturing global dependency patterns. Furthermore, we design an inter-variable association consistency loss to explicitly align the dependency structure of predictions with ground-truth correlations, providing effective structural supervision. Extensive experiments on multiple real-world datasets demonstrate that FreDC consistently achieves state-of-the-art performance, while maintaining competitive computational efficiency. Thursday Virtual Room 2 IJCNN Paper Time Series Forecasting VI Session Chair: Jiameng Chen (Beijing University of Posts and Telecommunications), Hongkai Jiang (Tsinghua University) Stock Price Forecasting Using a Transformer with Time-Interval-Based Attention and Trainable Stock Selection Hongkai Jiang, Baiting Wu, and Xiaolin Hu (Tsinghua University) Abstract Abstract Recently, there has been growing interest in predicting stock price movements by exploring the relationships between stocks. However, existing methods focus on the correlation between time steps and aggregate features from all stocks, regardless of the number of stocks, which leads to several limitations. First, in time-series data, especially stock data, the correlation between time intervals is more indicative of future trends than correlations at single time steps. Second, not all features contribute equally to stock price movements. Finally, not all stocks are correlated, and aggregating information from unrelated stocks introduces noise into the prediction. To address these challenges, we introduce TITAN, a Transformer with Time-Interval-Based Attention and Trainable Stock Selection, enabling the model to capture cross-time correlations, focus on the most informative features, and dynamically select the most relevant stocks. TITAN simulates complex stock correlations by aggregating information across time, feature, and stock dimensions. While validated on stock forecasting, the proposed network advances neural network architectures for multivariate time-series modeling more broadly, particularly in domains with noisy, high-dimensional, or relational data. Experiments on stock market indices demonstrate the effectiveness of TITAN compared to current state-of-the-art methods and highlight the importance of each module through ablation studies. StockLTG: Multi-Scale Stock Trend Forecasting via Lagged Co-Trending Dynamic Graph Networks Shiwei Pu and Chuanchang Liu (Beijing University of Posts and Telecommunications, State Key Laboratory of Networking and Switching Technology) Abstract Abstract In the realm of stock forecasting, traditional relational modeling approaches are constrained by two significant limitations. First, relying solely on static attributes to describe stock relationships makes it challenging to capture the dynamic changes in these relationships over time. Second, these methods often overlook the interaction between the return rate of an individual stock and its lagged relation. To address these issues, this paper proposes a novel modeling framework based on return rate trend decomposition, which captures the dynamic changes in stock relationships and lagged interactions through multi-scale trend decomposition and lagged cross-correlation algorithms. Empirical analysis on two major datasets from Chinese and American stock markets demonstrates that the proposed model's information coefficient on Chinese stock market data is 80\% higher than that of the baseline model, with significantly higher returns compared to the benchmark model. Additionally, the performance of the proposed framework on certain datasets surpasses that of current mainstream stock prediction models. Beyond Regression: Binary Encoding Classification with Confidence for Stock Index Prediction Junzhe Jiang and Chang Yang (The Hong Kong Polytechnic University), Xinrun Wang (Singapore Management University), and Bo Li (The Hong Kong Polytechnic University) Abstract Abstract Stock market indices serve as fundamental market measurements that quantify systematic market dynamics. However, accurate index price prediction remains challenging, primarily because existing approaches treat indices as isolated time series and frame the prediction as a simple regression task. These methods fail to capture indices' inherent nature as aggregations of constituent stocks with complex, time-varying interdependencies. To address these limitations, we propose Cubic, a novel end-to-end framework that explicitly models the adaptive fusion of constituent stocks for index price prediction. Our main contributions are threefold. i) Fusion in the latent space: we introduce the fusion mechanism over the latent embedding of the stocks to extract information from a large number of constituents. ii) Binary encoding classification: since regression tasks are challenging due to continuous value estimation, we reformulate the regression into a classification task, where the target value is converted to binary and we optimize the prediction of the value of each digit with cross-entropy loss. iii) Confidence-guided prediction and trading: we introduce the regularization loss to address market prediction uncertainty for the index prediction and design the rule-based trading policies based on the confidence. Extensive experiments across multiple stock markets demonstrate that Cubic achieves superior performance on both forecasting accuracy and trading profitability compared to baseline approaches, with performance gains varying across different market conditions and model architectures. MHG-FAN: Friction-Aware Heterogeneous Graphs for Stock Forecasting Jiameng Chen, Zhongliang Yang, and Linna Zhou (Beijing University of Posts and Telecommunications) Abstract Abstract Financial markets are driven by multimodal information streams where asset movements are shaped by complex intra- and inter-firm dependencies. However, traditional forecasting models struggle to characterize the complex mechanisms of risk transmission, often overlooking latent dependencies and the time-lagged nature of shock propagation. To address this challenge, we propose MHG-FAN, a risk-oriented framework that integrates multimodal market knowledge to construct a dynamic heterogeneous graph of financial entities. Specifically, we construct an adaptive dependency network that infers time-varying inter-firm relations from multimodal market signals, allowing the relational structure to evolve with changing market conditions. Building on this structure, we introduce a friction-aware relational module to model heterogeneous and asynchronous diffusion of market shocks, explicitly decoupling structural connectivity from transmission lags. Extensive experiments demonstrate that our approach consistently outperforms state-of-the-art baselines, offering an interpretable and effective tool for investment decision-making. Thursday Virtual Room 3 IJCNN Paper Time Series Forecasting VII Session Chair: Hongnian Wang (North Sichuan Medical College), Hongbo Zhao (East China Normal University) STAM-Net: Spectral-Temporal Adaptive Memory Network for Financial Time Series Kai-Ye Hu, Wenyun Xiao, and Ziyuan Liu (Jinan University) and Hongnian Wang (North Sichuan Medical College) Abstract Abstract Stock return prediction from financial time series remains challenging due to the dynamic and non-stationary nature of financial markets. Different market conditions require different temporal horizons and frequency components, yet exist- ing methods typically process temporal and spectral information separately. To address this, we propose STAM-Net, a spectral- temporal adaptive memory network where spectral characteris- tics guide temporal memory aggregation, and the resulting tem- poral representations in turn refine frequency filtering. Specifi- cally, global spectral statistics adapt the memory retention of the temporal encoder to balance short-term fluctuations and long- term dependencies, while the learned temporal representations guide spectral gating to suppress noise and preserve informative components. Experiments on the CSI300 and CSI500 benchmarks show that STAM-Net achieves the best overall performance among the compared methods in both accuracy and stability, suggesting the value of adaptive spectral-temporal coupling for modeling non-stationary financial time series. DSFormer: Dimension-Segment Transformer with Cross-Stock Correlation Modeling for Stock Prediction Wenyun Xiao, Kai-Ye Hu, and Ziyuan Liu (Jinan University) and Hongnian Wang (North Sichuan Medical College) Abstract Abstract Stock return prediction requires extracting predic- tive signals from high-dimensional and noisy financial time series. Existing methods typically project heterogeneous factors into a shared representation space, which can entangle semantically distinct features and hinder dynamic cross-stock correlation mod- eling. To address this issue, we propose the Dimension-Segment Transformer (DSFormer), a framework that combines dimension- wise temporal encoding with cross-stock correlation modeling. DSFormer first applies dimension-segment embedding to learn multi-scale intra-stock representations while preserving factor- specific semantics. It then uses data-driven inter-stock attention to capture market-wide co-movements from these disentangled features. Experiments on the CSI300 and CSI500 benchmarks show that DSFormer achieves the best IC and Rank IC among the compared methods, showing the benefit of combining dimension- wise temporal modeling with cross-stock dependency learning for stock return prediction. A Dynamic Factor Gating Architecture with Market Regime Awareness for Stock Return Forecasting Jiacheng Wang (Xijing university), Liang Fan and Baihua Li (Loughborough University), and Luyan Zhang (Northeastern University) Abstract Abstract Accurate stock return forecasting remains a central challenge in quantitative finance, as it directly informs the construction of portfolios and the management of risk. Although traditional static factor models are widely used, they are limited by manual factor selection and fixed weight assignments, which makes them vulnerable to evolving market conditions and regime shifts. To overcome these limitations, we introduce Market Regime Aware-Augmented Attention GRU (MRA-AGRU), an automated dynamic factor gating framework that adaptively reweights factors in response to market regime signals. By integrating an attention-enhanced GRU network, MRA-AGRU effectively suppresses obsolete or noisy factors while amplifying those most relevant to the prevailing environment, thereby capturing nuanced temporal and cross-factor dependencies. Extensive experiments on the CSI 300 and NASDAQ 100 demonstrate the superior performance of MRA-AGRU, highlighting machine-driven factor modulation's role in improving robustness to structural breaks and reducing bias in factor engineering. Benchmarking Multimodal Financial Time-Series Forecasting with Self-Evolving Captions Yinjie Teng and Hongbo Zhao (East China Normal University) Abstract Abstract Textual information is increasingly incorporated into financial time-series forecasting, yet the lack of reliable and predictive textual annotations has hindered the development of robust multimodal benchmarking. Existing benchmarks often rely on static or weakly grounded descriptions that fail to reflect the non-stationary and regime-dependent nature of financial markets, limiting the credibility of their empirical conclusions.In this work, we introduce FinEvolve, a framework that enables the scalable construction of high-quality multimodal financial time series by reformulating financial captioning as a dynamic,state-adaptive semantic alignment problem. FinEvolve employs structured multi-agent reasoning together with a self-evolving optimization mechanism to generate captions that are faithful to market dynamics, logically consistent, and predictive of future behavior. By treating textual generation as an adaptive and forward-looking process, FinEvolve serves as a principled foundation for multimodal financial benchmarking. Building on FinEvolve, we construct Fin-MM900K, a large-scale dataset comprising 900K multimodal time-series–caption pairs spanning diverse asset classes, temporal resolutions, and market regimes. We further establish a comprehensive benchmarking suite for multimodal financial forecasting, covering point prediction, regimespecific evaluation, and cross-modal retrieval. Experiments reveal the benefits of textual supervision are highly regime-dependent, with substantial gains under crisis and structural break scenarios but limited effects in noise-dominated markets. Thursday Virtual Room 4 IJCNN Paper Time Series Forecasting VIII Session Chair: yiren zhou (sichuan university), Kaixuan Chen (Zhejiang University) Progressive Dynamic Graph Sparsification for Spatiotemporal Traffic Forecasting yiren zhou, shiyong lan, yao ren, and zicheng sun (Sichuan University) Abstract Abstract Accurate traffic forecasting is the cornerstone of Intelligent Transportation Systems (ITS), yet it remains a formidable challenge due to the complex non-linear spatiotemporal dependencies. Existing methods often struggle to balance physical topological constraints with dynamic semantic correlations. Furthermore, temporal modeling frequently employs fully connected structures or standard attention mechanisms, which lead to noise propagation from irrelevant historical time steps and lack multi-scale hierarchy. To address these issues, we propose a novel framework named PDGCN (Progressive Dynamic Graph Sparsification Network). Spatially, we introduce a Dual-View Fusion mechanism that combines static physical diffusion with dynamic semantic interaction to capture complementary spatial dependencies. Temporally, we design a Progressive Temporal Graph Sparsification strategy. By utilizing a hierarchical Top-K masking mechanism, the model adaptively filters temporal noise, achieving a curriculum learning process from High-Salience Core Features to Global Comprehensive Semantics. Experiments on real-world datasets (PeMS03, PeMS04, PeMS07, and PeMS08) demonstrate that PDGCN achieves state-of-the-art performance, and ablation studies validate the effectiveness of the proposed sparsification strategy. The source code will be made publicly available on GitHub upon the acceptance of this paper. Dynamic Spatial-Temporal Pattern-Enhanced Graph Network for Traffic Forecasting Zilong Wu, Junwei Yang, and Hongyan Mao (East China Normal University) Abstract Abstract Traffic prediction remains a critical challenge in urban transportation systems. While existing approaches leveraging Graph Neural Networks (GNNs) and attention mechanisms have demonstrated promising results, two key limitations persist: 1) The dynamic evolution of spatial dependencies with respect to both time and traffic context is often overlooked; 2) There is a lack of explicit modeling of complex traffic conditions and their interactions in forming diverse traffic patterns. To address these challenges, we propose a novel Dynamic Spatial-Temporal Pattern-Enhanced Graph Network (DSTPEGN). To tackle the first limitation, we propose a Time-Aware Graph Convolution Network (TAGCN) that captures dynamic spatial dependencies conditioned on traffic states. For the second limitation, we propose a novel Pattern-Enhanced Graph Convolution Network (PEGCN) to model fundamental spatiotemporal conditions and dynamically retrieve them to represent complex traffic patterns. We also simulate realistic traffic pattern propagation to capture real-world traffic nature. Additionally, we design a self-supervised task to regularize the learned pattern representations. A Multi-Graph Interactive Learning (MGIL) module is proposed to effectively fuse the output of multiple graph convolutions, enabling effective interaction between time-aware spatial correlations and traffic pattern representations. Extensive experiments on three real-world traffic datasets demonstrate that DSTPEGN achieves state-of-the-art performance. STAR2: Spatio-Temporal Adaptive Retrieval and Refinement Transformer for Traffic Prediction Xinyi Bao and Yu Fang (Tongji University) Abstract Abstract Accurate traffic flow forecasting is a cornerstone of Intelligent Transportation Systems (ITS). Although advanced parametric models have achieved significant progress, they encounter two critical bottlenecks: (i) optimization bias, where models prioritize routine patterns to minimize global error, rendering them inadequate in capturing sharp fluctuations; and (ii) a persistent long-tail error distribution, where a minority of "hard samples" contribute the bulk of the overall prediction error. To address these issues, we propose STAR^2, a novel Spatio-Temporal Adaptive Retrieval and Refinement framework. Our architecture integrates a coarse-to-fine two-stage retrieval strategy that reduces computational complexity from O(N * M) to O(M + N * K_g) to ensure real-time efficiency, an error estimator that quantifies backbone uncertainty via high-level spatio-temporal embeddings, and an adaptive refiner that dynamically fuses historical patterns with a frozen backbone to prevent blind refinement. Extensive evaluations on six real-world datasets indicate that STAR^2 yields outstanding predictive capabilities, notably for hard samples, while promoting higher computational efficiency to facilitate robust forecasting in complex road networks. Thursday Virtual Room 5 IJCNN Paper Time Series and Temporal Modeling II Session Chair: Mingyue Qin (Shanghai Jiao Tong University), Rui Yan (Zhejiang University of Technology) Inter‑Layer Recurrent Hebbian Feedback Enhances Neural Network Robustness Yan Li, Mingyue Qin, Shuyu Yin, Peilin Liu, and Fei Wen (Shanghai Jiao Tong University) Abstract Abstract The human brain exhibits extraordinary robustness and adaptability in noisy and out‑of‑distribution (OOD) scenarios. Prior studies suggest that this capability is largely attributable to rich feedback pathways and dynamic synaptic plasticity. Inspired by these neuro-scientific insights, in this paper we introduce ReHNet, which is a recurrent neural model featuring inter-layer feedback connections and partially plastic fast weights. Unlike traditional feedforward-only networks, ReHNet introduces inter-layer feedback in certain layers. Additionally, ReHNet is a partially plastic model, in which the synaptic weights are decomposed into a fixed component and a data‑dependent short-term plastic component that is dynamically updated online via the Hebbian rule. We show that such local Hebbian plasticity, combined with feedback connections, can enhance the network’s sensitivity to correlated activations and enhance resilience to distributional shifts. This mechanism is especially beneficial under noisy and OOD conditions, where feature representations are prone to degradation. On corrupted datasets like CIFAR‑10C and CIFAR‑100C, ReHNet significantly improves the performance and even surpasses domain adaptation methods that rely on backpropagation-based optimization. In OOD evaluations on the PACS dataset, ReHNet also demonstrates excellent robustness. Overall, our results show that incorporating biologically inspired feedback loops and online Hebbian learning can markedly enhance the robustness to noise and OOD data, which may provide insights for both neuroscience research and the development of more robust and adaptive neural networks. Fully distributed adaptive pinning control for fractional-order multiplex networks Yujuan Han and Wenjun Wang (Shanghai Maritime University) and Lili Wang (Shanghai University of Finance and Economics) Abstract Abstract This work studies distributed adaptive strategies for achieving intra-layer synchronization in fractional-order multiplex networks via pinning control.Two distributed adaptive strategies are introduced: a node-based strategy, where the coupling strength of each node is dynamically updated using local state information from both the node and its neighbors; and an edge-based strategy, in which the weight of each intra-layer edge is updated according to the relative states between the two linked nodes. Additionally, the pinning gains applied to a subset of nodes are adaptively adjusted based on the discrepancy between each pinned node and its reference state. Rigorous theoretical analysis and numerical simulations are provided to demonstrate the effectiveness of the proposed algorithms. Fractional-Order Differential Equation-Driven Transformer for Time Series Forecasting via Gauss-Jacobi Quadrature Shengxiang Zhu (College of Software Engineering, Sichuan University) and Jiuhong Luan, Zhonglian Wei, Chong He, Zhicheng Zhang, and Junjie Hu (College of Computer Science, Sichuan University) Abstract Abstract Long-term time series forecasting faces the challenge of efficiently capturing long-range dependencies and adaptively forgetting historical information in fields such as finance, meteorology, energy, and healthcare. Traditional integer-order models, which rely on exponential decay mechanisms, struggle to retain long-range correlations and exhibit poor robustness to short-term noise. To address these issues, fractional-order differential equations (FDEs) naturally model non-Markovian memory effects through power-law decay mechanisms, enabling gradual forgetting of irrelevant short-term fluctuations while enhancing the memory of key historical patterns. This significantly improves the modeling capacity for complex dynamics. In this paper, we propose FODEformer, an architecture that integrates Gauss-Jacobi quadrature accelerated fractional-order differential equations into the self-attention framework. By utilizing the Caputo fractional-order derivative in its integral form and employing a variable transformation to map historical dependencies to a reasonable interval, FODEformer ensures robustness through fixed-node spectral approximation. Experimental results on 8 datasets show that FODEformer achieves significant performance improvements on multiple benchmark datasets, especially in terms of long-range dependency capturing and noise robustness, outperforming existing models. A Hippocampus-Inspired Associative Memory Model Based on Spiking Neural Networks Xinhua Bao, Jiaqiang Jiang, Junwei Cheng, and Rui Yan (Zhejiang University of Technology) Abstract Abstract Associative memory is a fundamental cognitive function of the hippocampus and has been widely modeled using spiking neural networks (SNNs) combined with spike-timing-dependent plasticity (STDP). Existing memory models achieve effective associative recall through recurrent connectivity and synaptic plasticity. In this work, we propose a hippocampus-inspired associative memory model based on SNNs, incorporating three complementary mechanisms to regulate neuronal interaction and competition. We introduce a Top-K excitation circuit to selectively enhance excitatory coupling among recently active neurons. Second, the variable threshold mechanism dynamically adjusts neuronal excitability during learning. Third, an adaptive lateral inhibition mechanism is introduced, where neurons with higher accumulated activity receive stronger inhibition, while less active neurons are suppressed more weakly. Collectively, these mechanisms regulate neuronal interaction and competition within the memory module while remaining compatible with STDP. Experimental results demonstrate that the proposed model achieves robust associative memory retrieval under noisy and incomplete cues. Further analysis indicates that the introduced mechanisms promote more balanced neuronal activity and sparser synaptic utilization. Ablation studies demonstrate that the Top-K excitation circuit, the variable threshold mechanism, and the adaptive inhibition mechanism are complementary. Thursday Virtual Room 6 IJCNN Paper Time Series and Temporal Modeling III Session Chair: huanlan yan (tongji university), Shaoqi Tan (University of Electronic Science and Technology of China) NIMO: Module-Level Interpretability for Time Series X-formers via Mask Joint Optimization Shaoqi Tan, Yong Wang, and Hongwei Zhu (University of Electronic Science and Technology of China); Zhicheng Zhang (National University of Defense Technology); and Wen Yin and Ruizheng Huang (University of Electronic Science and Technology of China) Abstract Abstract Transformer-based models, denoted as X-formers, employ complex architectures to capture intricate temporal dependencies. However, their internal decision-making processes lack transparency, hindering the understanding of how specific modules interact with distinct time-series patterns. To address this challenge, we introduce Neural Network Input and Module Mask Optimization (NIMO), a unified and fully differentiable framework designed to provide fine-grained module-level interpretability. Unlike traditional methods that rely on unstable discrete searches or input-only attribution, NIMO leverages a Temporal-Module Fusion Network (TMFN) and a joint optimization strategy to establish a direct mapping between dynamic data features and static model modules. This methodology explicitly unveils the functional roles within X-formers, revealing that Position Embeddings and Attention are critical for capturing seasonality, while Temporal Embeddings and Feed-Forward Networks primarily model trends. Extensive experiments on 12 representative X-formers validate the fidelity of our explanations. Furthermore, we demonstrate that this interpretability is actionable: by pruning identified non-salient modules, NIMO achieves a 46.40% improvement in computational efficiency and a 42.35% reduction in model parameters without compromising forecasting performance. Multi-Atlas Static–Dynamic Collaborative Diagnostic Network for Autism Jingxia Chen and Yuhan Shi (School of Electronic Information and Artificial Intelligence, Shaanxi University of Science and Technology); Huiru Zheng (School of Computing,Ulster University); and Pengwei Zhang and Haifeng Chen (School of Electronic Information and Artificial Intelligence, Shaanxi University of Science and Technology) Abstract Abstract Functional brain networks derived from functional magnetic resonance imaging (fMRI) exhibit complex spatiotemporal patterns that are critical for accurate identification of autism spectrum disorder (ASD). Existing diagnostic methods primarily rely on single-atlas modelling and static functional connectivity, which limits their ability to capture complementary information across different brain parcellation schemes as well as the temporal variability of brain activity. To address these limitations, we propose a static–dynamic multi-atlas feature fusion network for automatic ASD diagnosis. The proposed framework incorporates static and dynamic functional connectivity features to comprehensively characterise the temporal diversity in brain functional activity. In addition, a collaborative multi-atlas feature fusion strategy is introduced to adaptively integrate feature representations from different brain atlases, thereby effectively exploiting complementary information under diverse spatial partitioning schemes and enhancing the discriminative representation of functional brain networks. Extensive experiments on the public ABIDE I dataset demonstrate that the proposed method achieves a classification accuracy of 81.64% and an AUC of 81.53%, outperforming existing baseline approaches. Furthermore, analysis of discriminative brain regions highlights the interpretability of the proposed multi-atlas fusion framework and provides new neurobiologically meaningful insights into ASD-related functional patterns. Engram Memory Network: Brain-Inspired Prototype Explanations Hyunjun Kim (Seoul National University) and Myoung Hoon Ha (Korea Advanced Institute of Science and Technology) Abstract Abstract Artificial intelligence systems are increasingly being deployed in high-stakes decision-making scenarios. In response, prototype-based models in explainable AI have attracted growing attention for their ability to retrieve visually similar examples as faithful explanations. Although these models achieve competitive classification accuracy, they often exhibit degraded explanation quality across datasets with varying characteristics, thereby limiting their practical applicability. To address this limitation, we draw on insights from cognitive neuroscience, particularly the human memory system, which encodes input-relevant information into representations that support recognition. In this paper, we propose an Engram Memory Network (EMN), a model that emulates the ventral visual stream and inferior temporal cortex in the human brain. Central to this model is a Hopfield-inspired memory module that iteratively retrieves representative prototypes. These memory-refined latents are used for both reconstruction and classification, encouraging the selected prototypes to capture input-relevant information across diverse datasets. To address the lack of standardized evaluation for prototype explanations, we assess three complementary aspects: perceptual similarity (LPIPS and DISTS) for perceived resemblance, subclass alignment, measuring whether prototypes match the input's fine-grained category beyond supervised labels, and information preservation (NMI), measuring how much input information is captured by selected prototypes. Experiments show that EMN improves visual similarity by 17.06%, subclass alignment by 104.57%, and NMI by 42.42% on average across CIFAR-10/100, MNIST, HAM10000, and CUB-200, while achieving the highest average classification accuracy (87.62%) among all methods. These findings underscore the effectiveness of neuroscience-inspired explanation mechanisms in advancing the interpretability and reliability of prototype-based models. HFAN-ITF: Hierarchical Fusion Wasserstein Generative Adversarial Network for Interbank Network Temporal Forecasting Huanlan Yan, Yijun Chen, Xinjia Shi, and Zhijun Ding (Tongji university) Abstract Abstract Temporal forecasting of interbank networks is crucial for financial system risk assessment, but is challenged by the feature capture of complex dynamic graphs as well as the training instability and convergence difficulties of generative models like Generative Adversarial Networks (GANs). To address these limitations, we introduce HFAN-ITF, a novel generative adversarial framework for interbank network forecasting. Within the generator, a Hierarchical Fusion Network (HFN) captures multi-level structural features, where HFN extracts global and local structural features and captures inter-graph temporal dependencies for spatio-temporal modeling. Furthermore, a novel Self-adaptive Edge Density Control (SAEDC) mechanism, an adaptive Top-K strategy, guides the network toward accelerated learning. The entire process is steered by a Wasserstein distance objective to ensure stable convergence. Through extensive evaluations on our proposed interbank network dataset, we demonstrate that HFAN-ITF consistently outperforms current Widely-adopted techniques across the comprehensive tasks of node attribute, link, and weight forecasting, exhibiting a particularly significant advantage in predicting interbank lending weights. Thursday Virtual Room 7 IJCNN Paper Trustworthy and Safe Language Models Session Chair: Jianyuan Ni (Juniata College), Jixuan Guo (Beijing University of Technology, School of Information Science and Technology) Fed-MultiFND: Federated Multimodal Learning for Imbalanced Short Video Fake News Detection Jixuan Guo, Boyue Wang, Yihan Gao, Dabao Zhang, and Yikun Liu (Beijing University of Technology, School of Information Science and Technology) Abstract Abstract Short video platforms have become a dominant medium for news dissemination, significantly amplifying the spread and impact of misinformation. Compared to traditional text-based news, short video fake news detection is more challenging due to heterogeneous multimodal data and stringent requirements on both accuracy and efficiency. Moreover, real-world deployment is constrained by data privacy regulations that prohibit centralized cross-platform data collection, while platform-specific data often exhibits uneven distributions of fake and real news, further hindering model learning. Motivated by these challenges, we propose a unified framework,Fed-MultiFND, for multimodal short video fake news detection that jointly addresses modeling complexity and cross-platform learning constraints. The framework integrates a lightweight yet effective multimodal base model with a federated learning architecture, enabling collaborative training without raw data sharing. Specifically, hierarchical attention is employed to progressively fuse textual, auditory, and visual information, while federated learning facilitates knowledge aggregation across platforms under skewed data distributions. Through this integrated design, the proposed approach achieves a practical improvement between detection effectiveness, robustness to distributional imbalance, and real-world deployment feasibility. MEAFS: Event-level Aggregation with MLLMs for Fake News Detection in Short Videos Haisong Gong (Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences); Zhibo Liu (University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences); and Qiang Liu, Shu Wu, and Liang Wang (Institute of Automation, Chinese Academy of Sciences) Abstract Abstract Short video platforms have become a popular medium for information sharing, yet they also facilitate the rapid spread of misinformation. Detecting fake news in short videos is challenging, as malicious creators can closely mimic legitimate content, making single-video analysis unreliable. Existing methods mainly rely on intra-video signals, limiting robustness against sophisticated forgeries. To address this, we propose MEAFS (MLLM-powered Event-level Aggregated Fake news detection in Short videos), a framework that incorporates event-level context. MEAFS employs multimodal large language models (MLLMs) to generate objective summaries of videos from the same event, which are aggregated into an event-level representation. A discrepancy feature is then derived by comparing the target video with the aggregated context, highlighting potential inconsistencies. Combined with multimodal features, these representations enable more reliable classification of real versus fake content. Experiments on public datasets show that MEAFS substantially improves detection accuracy over baseline approaches. MultiPress: A Multi-Agent Framework for Interpretable Multimodal News Classification Tailong Luo (New York Institute of Technology); Hao Li (University of Arizona); Rong Fu (University of Macau); Xinyue Jiang, Huaxuan Ding, Yiduo Zhang, and Zilin Zhao (Peking University); Simon Fong (University of Macau); Guangyin Jin (Chang'an University); and Jianyuan Ni (Juniata College) Abstract Abstract With the growing prevalence of multimodal news content, effective news topic classification demands models capable of jointly understanding and reasoning over heterogeneous data such as text and images. Existing methods often process modalities independently or employ simplistic fusion strategies, limiting their ability to capture complex cross-modal interactions and leverage external knowledge. To overcome these limitations, we propose MultiPress, a novel three-stage multi-agent framework for multimodal news classification. MultiPress integrates specialized agents for multimodal perception, retrieval-augmented reasoning, and gated fusion scoring, followed by a reward-driven iterative optimization mechanism. We validate MultiPress on a newly constructed large-scale multimodal news dataset, demonstrating significant improvements over strong baselines and highlighting the effectiveness of modular multi-agent collaboration and retrieval-augmented reasoning in enhancing classification accuracy and interpretability. Thursday Virtual Room 8 IJCNN Paper Vision Domain Adaptation and Generalization III Session Chair: Junsong Leng (Huazhong University of Science and Technology, School of Artificial Intelligence and Automation), Feifei Zhang (Qilu University of Technology; Shandong Provincial Key Laboratory of Industrial Network and Information System Security, Shandong Fundamental Research Center for Computer Science, Jinan, China) A Source-Free Universal Domain Adaptation Method for Fault Diagnosis Based on Dual-View Disentanglement and Orthogonal Decomposition Huijuan Hao, Feifei Zhang, Jinqiang Bai, Qingyan Ding, and Huanqing Xu (Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center (National Supercomputing Center in Jinan), Qilu University of Technology (Shandong Academy of Sciences), Jinan, China; Shandong Provincial Key Laboratory of Industrial Network and Information System Security, Shandong Fundamental Research Center for Computer Science, Jinan, China) Abstract Abstract Cross-operating-condition fault diagnosis in real industrial scenarios typically encounters two major challenges. First, source-domain data are often inaccessible due to privacy and compliance constraints, which hinders the deployment of conventional domain adaptation methods that rely on joint training with both source and target data. Second, the relationship between the source and target label spaces is unknown a priori, which often leads to misalignment between target samples and source categories. To address these issues, this paper focuses on Source-Free Universal Domain Adaptation (SF-UniDA) for fault diagnosis and proposes a two-stage target adaptation framework. In the source training stage, we leverage RandMix-based data augmentation to construct a second view and impose a supervised contrastive learning objective to enforce cross-view consistency among samples of the same class, thereby encouraging domain-invariant representation learning and feature disentanglement.In the target adaptation stage, we first construct three criteria to determine whether unknown classes exist in the target domain; we then propose a pseudo-label construction method that integrates DBSCAN clustering and orthogonal feature decomposition by fusing cluster-level and instance-level information. Experiments on the Case Western Reserve University (CWRU) and Paderborn University (PU) datasets under universal-domain settings demonstrate the effectiveness of the proposed fault diagnosis model. Uncertainty and Temporal Consistency for Domain-Adaptive Bearing Fault Diagnosis Chenxuan Huan and Rui Xi (University of Electronic Science and Technology of China), Shilei Zhou (Xi'an Rainbow Intelligent Engineering CO. Ltd.), and Mengshu Hou (University of Electronic Science and Technology of China) Abstract Abstract Bearing fault diagnosis is essential for industrial machinery reliability, but domain shifts from varying operational conditions often degrade model performance. While domain adaptation techniques mitigate this issue, challenges like unreliable pseudo-labels and lack of sample evaluation mechanisms remain critical bottlenecks. This paper proposes UTC-DA, a novel approach leveraging uncertainty and temporal consistency for domain-adaptive bearing fault diagnosis. UTC-DA integrates two key modules: (1) a data augmentation-based pseudo-label voting strategy, enhancing robustness through majority voting on perturbed samples; (2) an uncertainty-aware sample selection mechanism, balancing sample discriminability and diversity by dynamically weighting predictions based on entropy and neighborhood consistency; and filtering high-confidence samples by tracking temporal stability. Experiments on PU and JNU bearing datasets show that UTC-DA outperforms existing state-of-the-art methods in diverse cross-domain tasks, validating UTC-DA’s generalization capabilities and practical utility in real-world industrial applications. Cross-domain Bearing Fault Diagnosis Based on a Triple Attention Multi-scale Spatiotemporal Convolutional Network Haojie Deng, Yanrong Hao, Xin Wen, Jing Bian, Jin Li, and Rui Cao (Taiyuan University of Technology) Abstract Abstract Critical to mechanical systems, rolling bearings face severely degraded fault diagnosis performance under complex industrial conditions due to domain distribution shifts. To address the above issues, we propose a domain-adaptive deep transfer learning model for fault diagnosis, which incorporates a Triple-Attention Mechanism-Based Multi-Scale Spatiotemporal Network (TAM-MST) to enhance feature representation and facilitate effective knowledge transfer. Concurrently, we design a multi-layer maximum mean difference to improve distribution similarity between source and target domains, thereby strengthening the model’s transfer learning capabilities. Our approach overcomes the low diagnostic accuracy of traditional models under varying operating conditions, endowing the model with adaptability to unknown operating scenarios. Finally, to validate the model’s transfer learning capability, we compared it with several classical algorithms on the CWRU and PU datasets. Experimental results show diagnostic accuracies of 99.86% and 98.88% on the two datasets, respectively, providing a reliable research direction for bearing fault diagnosis in industrial equipment. Uncertain Estimation Based on Prediction Probability and Feature Prototypes for Universal Domain Adaptation Junsong Leng, Zhong Chen, and Guoyou Wang (Huazhong University of Science and Technology, School of Artificial Intelligence and Automation) Abstract Abstract Universal domain adaptation (UniDA) aims to transfer knowledge from the source to the target domain without restrictions on label sets. The key challenge is to maintain robust classification for known classes while identifying unknown classes in the target domain. We propose a model based on uncertainty estimation to address this issue. It integrates a teacher–student network, domain adversarial training, and uncertainty estimation module. The student network provides pseudo-labels for the teacher network, with parameters updated via exponential moving average. Source features initialize class prototypes, which are updated during training. Finally, uncertainty estimation determines whether target samples belong to known classes. Experiments on multiple datasets show that our model outperforms existing baselines. Thursday Virtual Room 1 IJCNN Paper Vision Domain Adaptation and Generalization IV Session Chair: zhiyong zheng (University of Science and Technology of China; School of AI and Data Science, University of Science and Technology of China), Shayok Chakraborty (Florida State University) Multi-target Deep Domain Adaptation for Deepfake Detection Md Shamim Seraj and Shayok Chakraborty (Florida State University) Abstract Abstract While the recent progress of generative AI technology (primarily through diffusion models) has revolutionized AI research, they have also posed significant challenges for real-world deepfake detection. Existing detection techniques demonstrate promise in detecting deepfakes on which the model has been trained; however, their performance drops significantly when applied to detect forgeries created using other manipulation techniques, on which the model has not been sufficiently trained. Thus, there is a pressing need for a technology that can detect deepfakes created using unknown faking techniques, without losing prior knowledge about already learned faking techniques. In this paper, we propose a novel multi-target deep domain adaptation framework to address this practically relevant and timely problem. Our framework can leverage a large amount of annotated data (fake/genuine) generated using a particular faking technique (source domain) and a small amount of labeled data generated using different unknown faking techniques (target domains) to induce a deep neural network with good generalization capability on the source domain, as well as all the target domains of interest. Further, our framework can efficiently utilize unlabeled data in the target domains, which are more readily available than labeled data. We design a novel loss function specific to the multi-target domain adaptation task and use the SGD method to optimize the loss and train the deep neural network. Our extensive empirical studies on benchmark datasets, using multiple types of deepfakes, corroborate the promise and potential of our framework for real-world applications. To the best of our knowledge, this is the first research effort to develop a multi-target deep domain adaptation framework for deepfake detection. Our code is open-sourced at https://github.com/shuvornb/MultiTargetDomainAdaptation DAMamba State-Space Modeling for Robust Unsupervised Domain Adaptation yuqi he (College of Computer Science and Technology,Jilin University) Abstract Abstract Unsupervised domain adaptation (UDA) remains challenging due to severe distribution shifts between the labeled source and unlabeled target domains. Existing UDA approaches are based on Convolution Neural Networks (CNNs) or Vision Transformers (ViTs), which suffer from limited receptive fields or quadratic complexity issues. While recent State Space Models (SSMs) like Mamba offer linear complexity and strong long-range modeling, their recursive state updates make them vulnerable to amplified style drift and misaligned local semantics under domain shift. To address these issues, we propose DAMamba, a novel unsupervised domain adaptation framework tailored for Mamba-based visual backbones. DAMamba introduces two key modules: a State-Aware Style Memory (SASM) that models global and class-wise target statistics and performs dynamic style transfer to suppress state drift, and a Bidirectional Patch Alignment (BPA) that enforces fine-grained semantic correspondence through saliency-guided patch matching. Extensive experiments demonstrate that DAMamba significantly enhances cross-domain robustness while preserving computational efficiency. Class Bias-Aware Source-Free Domain Adaptation with Multi-Center Contrastive Learning Hua Tong, Siya Yao, and Kaibo Zhou (Zhejiang Gongshang University) and Xiaoyu Sean Lu (Nanjing University of Science and Technology) Abstract Abstract Source-free Domain Adaptation (SFDA) transfers knowledge from a pre-trained source model to an unlabeled target domain with a different data distribution, without accessing source data. A key challenge in SFDA methods is the transfer difficulty across classes, leading to class bias and pseudolabel noise. This issue is pronounced for hard-to-transfer classes, where feature discrepancies result in unreliable pseudo-labels. Consequently, the pseudo-label distribution becomes biased: easily transferable classes dominate, while harder classes suffer from scarce or noisy labels, degrading model performance. These challenges highlight the need for robust methods addressing classwise adaptation complexity and pseudo-label instability. Therefore, we propose a Class-Difficulty-aware and Minority-class Enhancement framework (CDME). First, we model intra-class compactness based on target domain feature pairs to quantify per-class transfer difficulty, and adjust pseudo-label confidence to suppress overconfidence in hard-to-transfer classes. Second, we introduce a dynamic pseudo-label strategy with multi-center prototype contrastive learning. This enhances intra-class stability and boosts pseudo-label reliability. Extensive results show our method achieves stable and competitive performance across multiple SFDA tasks. Code is available at https://github.com/Huatong9074/CDME. FDSense: Prior-Guided Cross-Room Domain Generalization in WiFi Sensing Zhiyong Zheng and Ke Xu (Suzhou Institute for Advanced Research of USTC; School of AI and Data Science, University of Science and Technology of China); Shijia Liu (Southwest Jiaotong University, The Institute of Smart City and Intelligent Transportation); and Jiangtao Wang (Suzhou Institute for Advanced Research of USTC; School of AI and Data Science, University of Science and Technology of China) Abstract Abstract WiFi sensing has become increasingly valuable in daily life by enabling the perception of human activities through the analysis of Channel State Information (CSI). However, its effectiveness is often hindered by domain shift. Existing approaches to mitigate domain shift either rely on target domain data for domain adaptation or overlook prior knowledge inherent to WiFi sensing. To address this gap, we propose FDSense, a novel framework that integrates signal processing with domain generalization guided by CSI priors. Specifically, we first employ signal processing techniques to denoise and standardize CSI signals. Next, we introduce a Selective Fourier Domain Mixing (SFDM) module, which encourages the model to capture domain-invariant features. Finally, deep neural networks are used for feature extraction and classification. Extensive cross-room experiments on a public dataset demonstrate that FDSense consistently outperforms state-of-the-art methods, highlighting its effectiveness and robustness. Thursday Virtual Room 2 IJCNN Paper Vision-Language Models III Session Chair: Zicheng Wang (Shandong Jianzhu University), Xiangfeng Luo (Shanghai University, School of Computer Engineering and Science) Bridging Naive Semantics through Proxy Injection and Matrix Fusion for Open-Vocabulary Semantic Segmentation Zicheng Wang, Xiushan Nie, Yang Ning, Jingyuan Fang, and Runhu Zhao (Shandong Jianzhu University) Abstract Abstract Open-Vocabulary Semantic Segmentation (OVSS) aims to achieve pixel-level classification under arbitrary categories. Although recent methods leverage foundation models like SAM to enhance segmentation capabilities, weakly supervised OVSS still faces three critical challenges: (1) intermediate-layer semantic compression that loses fine-grained noun distinctions, (2) rigid feature fusion strategies that fail to adaptively balance semantic and boundary cues across diverse scenes, and (3) single-point semantic attraction where models attend only to the most salient instance when multiple objects share the same noun. To address these issues, we propose BNS-Seg, bridging naive semantics through proxy injection and matrix fusion for open-vocabulary segmentation. First, Semantic Proxy Injection (SPI) selects proxy tokens from attention key/value pairs and injects noun-level semantics into intermediate visual layers via differential attention, enhancing fine-grained semantic discrimination. Second, Doubly Stochastic Matrix Fusion (DSF) predicts sample-adaptive fusion weights under doubly stochastic constraints, enabling flexible yet stable integration of CLIP semantic features and SAM boundary features. Third, Peak-Anchored Query Diversification (PQD) identifies multiple local peaks on text-conditioned similarity maps as spatial anchors, generating diversified queries to enable effective multi-instance segmentation under the same noun. Extensive experiments on six benchmark datasets demonstrate that BNS-Seg achieves state-of-the-art performance, with improvements of up to 3.6\% mIoU over existing methods. LoCoFuse: Locality-Constrained Dual-Branch Inference for Training-Free Underwater Open-Vocabulary Semantic Segmentation Yuhang Zhang, Weidong Tang, and Yinxue Shi (China Agricultural University) Abstract Abstract Underwater open-vocabulary semantic segmentation (UOVS) is a key technology for underwater robotic perception and marine ecological monitoring. However, multimodal foundation models are trained mainly on terrestrial imagery, and the resulting distribution gap severely limits their generalization in underwater scenes; absorption and scattering further degrade images and exacerbate this challenge. We propose LoCoFuse, a training-free inference framework for UOVS that keeps both the vision--language model and the geometric encoder frozen, and improves robustness for UOVS under underwater image degradations through three complementary modules: LoCoMix introduces locality-constrained geometry-guided token mixing to curb long-range semantic leakage; Text-CPC mitigates domain bias via prompt contrast to calibrate category prototypes and fuses image-conditioned cues from a multimodal large language model (MLLM), using these MLLM cues to strengthen text representations; DualMix leverages token-level uncertainty and boundary cues to fuse global and local branches in probability space, balancing region consistency and boundary accuracy. Extensive experiments on UOVSBench demonstrate notable performance gains with efficient inference. Decoupling Geometry from Semantics: Open-Vocabulary Grasping via SE(3) Diffusion Yichen Xiao, Haoning Wu, Youyuan Tu, Xiaoping Wu, and Xiaoguang Niu (Wuhan University) Abstract Abstract Open-vocabulary grasping requires learning physically valid actions on the non-Euclidean SE(3) manifold while aligning them with ambiguous, long-tailed semantic instructions. Existing approaches either bias grasp generation through end-to-end semantic conditioning, compromising geometric stability, or rely on task-specific semantic supervision that limits generalization. We propose a geometry-semantics decoupled framework that explicitly separates generative modeling from semantic decision-making. An SE(3)-aware diffusion model first learns a geometry-only conditional distribution, generating diverse, physically valid grasp candidates from point clouds. A training-free vision-language pipeline (CLIP+SAM) then projects open-vocabulary 2D segmentations into 3D semantic target constraints. Finally, a post-hoc joint ranking strategy combines geometric quality and semantic relevance through a controllable utility function, enabling inference-time adjustment of risk preferences without re-sampling or re-training. Extensive simulations show that our method achieves a 90.7% success rate on language-conditioned manipulation tasks, outperforms GraspNet in geometric metrics, and surpasses planning-based baselines such as Code as Policies and VoxPoser in semantic alignment. Semantic-Guided Weakly Supervised Open-Vocabulary Object Detection Chao Ma, Liyan Ma, Xiangfeng Luo, Shaorong Xie, Yantao Shang, and Wendi Rao (Shanghai University) Abstract Abstract Weakly Supervised Open-Vocabulary Object Detection (WOVOD) focuses on detecting novel classes with only image-level labels of base classes. However, existing methods suffer from three main limitations: insufficient semantic guidance for localization tasks, underutilized features, and noise in pseudo-labels. To address these issues, we propose a Semantic-Guided Weakly Supervised Open-Vocabulary Object Detection (SGWOVD) framework. Our method includes three key innovations: (1) a Semantic-Guided SAM Proposal Generator uses CLIP's semantic understanding to guide SAM in generating semantic proposals; (2) a Semantic Gating Proposal Feature Enhancement module combines image, text, proposal and dataset-aware features to enhance the semantic representation of targets; and (3) a Proposal Noise-Aware Multi-Instance Learning module refines pseudo-label quality by assessing noise across multiple detection heads, alongside spatial and semantic dimensions. Extensive evaluations on OV-COCO, OV-LVIS, and Pascal VOC 2007 demonstrate that SGWOVD consistently outperforms existing state-of-the-art methods. These results validate SGWOVD's strong semantic understanding capability and robustness to noise. Thursday Virtual Room 3 IJCNN Paper Vision-Language Models IV Session Chair: An Zhao (Institute of Artificial Intelligence (TeleAI), China Telecom), Ye Shen (Shanghai Jiao Tong University) MaskRAG: Mask Retrieval Augmented Generation for MLLM-based Referring Expression Segmentation Zhongjiang He (Beijing University of Posts and Telecommunications; China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd); An Zhao (China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd); Canhui Tang (Xi’an Jiaotong University; China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd); Hao Sun, Hongbo Sun, and Ye Yuan (China Telecom Artificial Intelligence Technology (Beijing) Co., Ltd); and Kongming Liang and Zhanyu Ma (Beijing University of Posts and Telecommunications, Beijing Key Laboratory of Multimodal Data Intelligent Perception and Governance) Abstract Abstract Multimodal Large Language Models (MLLMs) have significantly advanced referring expression segmentation (RES) by understanding complex language instructions. Current approaches typically employ a [SEG] token as a textual prompt to bridge the MLLM and downstream segmentation models (e.g., SAM). However, due to the weak prompt nature of the [SEG] token, existing frameworks often suffer from the “segmentation hallucination” issue, where the model confidently predicts structures or objects that are absent from ground truth. In this paper, we propose MaskRAG, a multimodal RAG-like framework that advances RES through adaptive mask retrieval and augmentation mechanisms. Specifically, we propose a mask retrieval module that efficiently encodes region features and embeds those features with a customized language template, enabling finer-grained perception of scenes and discrimination among candidate regions. Furthermore, we propose a mask augmentation module with multi-granularity semantic fusion and adaptive routing mechanisms, which effectively leverages the strengths of both [SEG]-based segmentation and our retrieval approaches. Experiments demonstrate that MaskRAG achieves state-of-the-art performance across multiple benchmarks. GLA-MLLM: A Flowchart Understanding Method Integrating Global-Local Awareness With Multimodal Large Language Model Generation Xinyue Liang and Mengyu Han (Shanghai University), Hang Xiao (Shanghai Minhang Polytechnic), and Wenhao Zhu (Shanghai University) Abstract Abstract Existing Multimodal Large Language Models (MLLMs) demonstrate excellent performance in handling visual tasks in everyday scenarios, but they struggle to understand structured images such as flowcharts and organizational charts. The complex non-linear structure and rich textual content of these images often cause structural hallucinations, resulting in logically inconsistent interpretations. In this paper, we propose GLA-MLLM, a novel flowchart understanding framework integrating global-local awareness with MLLM generation. GLA-MLLM operates in three stages: local structural perception, global semantic understanding, and MLLM-based generation. First, visual primitive detection, text recognition, and a geometry-aware Graph Neural Network (GNN) are integrated to transform raw flowchart images into high-fidelity semantic triplets. Subsequently, an open-source Vision-Language Model (VLM) is leveraged to extract a global semantic summary, providing high-level guidance and steering the model toward a coherent logical interpretation. Finally, these multi-scale cues are serialized into a structured prompt to fine-tune the multimodal model via Low-Rank Adaptation (LoRA), enabling grounded and topologically faithful logical reasoning. We evaluated GLA-MLLM on the Computerized Block Diagrams (CBD) dataset and the FC\_B handwritten flowchart dataset, where it achieved state-of-the-art (SOTA) performance, demonstrating strong robustness to diagram complexity and improved cross-domain generalization. Visual-Prior Guided and MLLM-Enhanced Dual Retrieval for Knowledge-based VQA ji yan, jingzi gu, ruicheng xiong, Gengqi yang, guimiao yang, dayan wu, peng fu, and zheng lin (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Abstract Abstract Knowledge-based Visual Question Answering (KB-VQA) necessitates retrieving external knowledge to answer questions about images. Despite the progress of Retrieval-Augmented Generation (RAG), existing paradigms face significant limitations. In particular, visual-only retrieval often neglects valuable semantic information in questions, whereas employing Multimodal Large Language Models (MLLMs) to generate semantically relevant queries from images is prone to severe entity hallucinations due to the lack of grounding. To address these challenges, we propose a novel dual-stream retrieval framework, where the core module of the text-to-text stream is Visual-Prior Guided Retrieval-Augmented Captioning (ViP-RAC). Specifically, ViP-RAC leverages retrieval results from the image stream as visual priors to anchor MLLMs, effectively mitigating hallucinations and generating factually aligned search queries. Furthermore, we employ a Weighted Reciprocal Rank Fusion (RRF) algorithm to synergistically integrate the visual-based and text-based retrieval streams, ensuring robust knowledge acquisition. Extensive experiments on the Encyclopedic-VQA and InfoSeek benchmarks demonstrate that our approach achieves new state-of-the-art performance. A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation Ye Shen (Shanghai Jiao Tong University, Shanghai Artificial Intelligence Laboratory); Junying Wang (Shanghai Artificial Intelligence Laboratory, Fudan University); Farong Wen and Yijin Guo (Shanghai Jiao Tong University, Shanghai Artificial Intelligence Laboratory); Qi Jia and Zicheng Zhang (Shanghai Artificial Intelligence Laboratory); and Guangtao Zhai (Shanghai Jiao Tong University, Shanghai Artificial Intelligence Laboratory) Abstract Abstract The rapid progress of Multi-Modal Large Language Models (MLLMs) has ushered in an explosion of benchmarks to evaluate and rank their capabilities. However, traditional full-coverage Question and Answering (Q\&A) evaluation strategies suffer from high redundancy and low efficiency. Drawing inspiration from real-world human interview protocols, we propose the \textbf{multi-to-one interview paradigm} to break the limitations of conventional MLLM assessment. Our paradigm comprises three synergistic components, which are (i) a two-stage interview strategy with a pre-interview for initial proficiency calibration followed by a formal inquiry phase, (ii) dynamic adjustment of interviewer weights, which adaptively calibrates the influence of multiple evaluator models to mitigate individual biases and ensure a fair assessment, and (iii) an adaptive difficulty mechanism that iteratively selects questions based on real-time performance to probe the model’s capability boundaries. Ablation experiments demonstrate that each component plays a role in the overall paradigm. Extensive experiments across diverse benchmarks demonstrate that the proposed paradigm achieves significantly higher correlation than random sampling, with improvements of up to 17.6\% in PLCC and 16.7\% in SRCC, while reducing the number of required questions. These findings demonstrate that our paradigm provides an adaptive, reliable and efficient alternative for large-scale MLLM benchmarking. Thursday Virtual Room 4 IJCNN Paper Vision-Language Models V Session Chair: Haowen Zheng (Central University of Finance and Economics), Jiajie Fan (South China Normal University) KDKD: Cross-Architecture Knowledge Distillation for Few-Shot Unsupervised Vision-Language Models Kailai zhuang (Tiangong University), Jiajie Fan (South China Normal University), Rundong Gao (Shenyang University of Chemical Technology), Zeming Tian (Wuhan University), and Qingzeng Song and Yongjiang Xue (Tiangong University) Abstract Abstract Vision-Language Models (VLMs) like CLIP have utilized large-scale pre-training to achieve impressive zero-shot capabilities. However, adapting these heavy models to downstream tasks in data-scarce scenarios remains challenging. While existing unsupervised self-training methods have shown promise on Vision Transformers (ViTs), they suffer from catastrophic performance degradation when applied to Convolutional Neural Networks (CNNs), even falling below zero-shot baselines. To bridge this gap, we propose KDKD (Kendall-Driven Knowledge Distillation), a novel framework designed for robust cross-architecture adaptation. We introduce a bidirectional adaptive temperature mechanism guided by the Kendall Correlation Coefficient, which dynamically aligns the prediction rankings between teacher and student models. Furthermore, we propose a soft-weighting strategy, as opposed to hard truncation, to adjust the loss based on sample quality. Extensive experiments on 9 downstream datasets demonstrate that KDKD not only boosts ViT students but effectively rescues ResNet students from degradation, achieving state-of-the-art performance in few-shot unsupervised settings with less computational overhead. Our codes are available at https://github.com/zhuangkailai/KDKD HAC: Hierarchical Attention Consensus for Visual Token Pruning in Large Vision-Language Models Jingru Li (China University of Geosciences) and Haowen Zheng (Central University of Finance and Economics) Abstract Abstract Visual token pruning has emerged as an effective strategy for accelerating large vision-language models (LVLMs) by reducing the number of tokens processed during inference. Existing methods typically estimate token importance based on attention patterns from a single layer of the language model backbone, often selecting shallow layers for computational efficiency. However, we observe that attention patterns vary substantially across network depth: shallow layers tend to focus on low-level visual features such as edges and textures, while deeper layers attend more to semantically meaningful regions. This divergence suggests that single-layer pruning may discard tokens that are important for high-level reasoning despite being less salient in early processing stages. We propose Hierarchical Attention Consensus (HAC), a simple framework that aggregates attention information from multiple layers—shallow, intermediate, and deep, to identify tokens that are consistently important across the processing hierarchy. By preserving tokens that maintain attention across layers, HAC better captures semantic relevance while filtering out those with only superficial saliency. The method requires no additional training and introduces minimal computational overhead. Experiments on vision-language benchmarks show that HAC consistently improves performance over single-layer baselines while using the same computational budget, with consistent gains on tasks requiring semantic understanding. Thursday Virtual Room 5 IJCNN Paper Vision-Language Models VI Session Chair: Hao Kong (shanghai university), Ju Wang (Macao Polytechnic University, Faculty of Applied Sciences) Preference-Guided Diffusion for Embodied Robot Morphology Evolution Liming Xin, Yaowen Zhang, Bin Sheng, and Hao Kong (Shanghai University) Abstract Abstract Morphology optimization is critical for embodied robots but remains challenging due to the vast design space and the high cost of evaluating candidates with policy training. Recent diffusion-guided evolutionary methods improve sample efficiency by learning morphology priors. However, they typically rely on thresholded binary preference supervision and a fixed mutation-diffusion ratio. These choices yield coarse guidance and discard relative performance information, thereby failing to adapt to changing search dynamics. To address these limitations, we propose Preference-Guided Diffusion Evolution (PGDE), a framework that improves guidance reliability and search efficiency within a unified evolutionary loop. PGDE introduces rank-normalized soft preference learning to provide smoother targets and finer-grained guidance signals. This is complemented by a confidence-aware guidance mechanism that scales guidance strength based on the uncertainty of preference predictions to prevent misdirected sampling. To better match exploration and refinement across search stages, PGDE adopts an adaptive mutation-diffusion ratio that dynamically balances mutation-based local search and diffusion-based global search through online fitness feedback. Experiments on diverse EvoGym tasks show that the proposed framework consistently achieves higher reward compared to baseline approaches. LC-DPO: Length-Controlled Direct Preference Optimization via Dynamic Token Selection Jinlei Dong, Bin Zhang, and Jialong Zhang (Xi’an Jiaotong University) Abstract Abstract Direct Preference Optimization often exhibits a pronounced verbosity bias, favoring longer responses even when quality does not improve. We propose LC-DPO, a simple extension of DPO that mitigates this bias by reweighting token level contributions on the longer response based on a policy reference loss gap criterion. Intuitively, LC-DPO emphasizes tokens that are more discriminative for preference learning under this criterion, while softly downweighting tokens that contribute weaker preference signal and are more correlated with verbosity. We evaluate LC-DPO on UltraFeedback with multiple backbones, including Llama-3-8B, Qwen2.5-7B, Mistral-8B, and Qwen2.5-1.5B. Results on MT-Bench and AlpacaEval2 show that LC-DPO consistently outperforms standard DPO and representative length control baselines, while producing shorter responses on average. These findings suggest LC-DPO is an effective and practical approach for reducing verbosity bias in preference alignment. SKIP: a Self-knowledge-guided Step-wise Preference Learning Framework for Concise Reasoning Qinhong Lin, Yuhao Zhang, Yinglun Feng, Zhongliang Yang, and Linna Zhou (Beijing University of Posts and Telecommunications) Abstract Abstract While Chain-of-Thought (CoT) reasoning has been proven to be effective, it often leads to overthinking, resulting in computational overhead, inference latency, and even degraded performance in large language models (LLMs). Existing concise reasoning frameworks significantly compromise accuracy while compressing the length of output. In this paper, we propose SKIP, a self-knowledge-guided step-wise preference learning framework. Starting with lightweight fine-tuning to adjust the model’s output style, SKIP introduces a carefully designed knowledge probing mechanism to guide model to output an answer at each reasoning step. Based on the correctness of intermediate steps, we construct preference data that guide the model toward more efficient and correct reasoning by leveraging DPO. Experimental results demonstrate that our method effectively improves reasoning compression while mitigating performance degradation after fine-tuning. Besides, SKIP shows strong generalization ability on out-of-distribution datasets. We further conducted ablation studies on the component parameters of our framework. Uncertainty-Guided Diffusion Model for Biomedical Image Segmentation Ju Wang, Ying Wang, and Xiaochen Yuan (Macao Polytechnic University, Faculty of Applied Sciences) and Guoheng Huang (Guangdong University of Technology, School of Computer Science and Technology) Abstract Abstract Biomedical image segmentation is crucial to precisely diagnose lesions on medical images. Recently, diffusion probabilistic models (DPMs) have proven to be effective for medical segmentation tasks. However, most diffusion-based methods apply uniform guidance across the entire image, which treats all regions equally during optimization and neglects regional differences in prediction difficulty. To address this issue, we propose UGD-Seg, a two-stage uncertainty-guided diffusion segmentation framework. In the first stage, an evidential network is employed to estimate a pixel-level uncertainty map while extracting multi-scale features from the input image. In the second stage, the multi-scale features are integrated and decoupled into certain and uncertain components under uncertainty guidance, which are then used to guide a latent diffusion model (LDM) for progressive segmentation. Extensive experiments on the ISIC2016 and CVC-ClinicDB datasets demonstrate that UGD-Seg consistently outperforms state-of-the-art methods, achieving more accurate boundary delineation and more stable segmentation predictions. Thursday Virtual Room 6 IJCNN Paper Vision-Language Models VII Session Chair: Shuxiang Song (Guangxi Normal University), Zhenjun Tang (Guangxi Normal University) GazePyFormer: Capturing Multi-Scale Spatial Contexts for Gaze Following Haiying Xia, Ruihan Yang, Yumei Tan, and Shuxiang Song (Guangxi Normal University) Abstract Abstract Gaze following aims to predict the target location of a person’s gaze by integrating scene context with head cues. While multi-scale representations are essential for handling targets of varying sizes, existing Transformer-based methods often rely on fixed-resolution feature maps or aggregate features in a post-hoc manner. Such designs fail to explicitly capture spatial dependencies across different scales, leading to sub-optimal localization in complex scenes where targets may be distant or small. To address this, we propose GazePyFormer, a Multi-scale Pyramid Pooling Transformer that performs cross-scale reasoning within the encoding process. Specifically, we develop a GazePyFormer Encoderthat embeds hierarchical pyramid pooling into the Transformer blocks to extract features across diverse receptive fields. This is followed by a GazeModulator, which uses Spatial-FiLM to adaptively align gaze priors with multi-level scene features. Finally, a Multi-scale Fusionmodule integrates these coarse-to-fine representations to produce precise gaze heatmaps. Experiments on the GazeFollow and VideoAttentionTarget benchmarks demonstrate that GazePyFormer consistently outperforms state-of-the-art methods, showing superior robustness and accuracy in challenging environments. Dual-Interactive Coordinate Aggregation Network for Fine-Grained Visual Classification Xuanyi Wu (School of Computer and Computing Science, Hangzhou City University); Zixuan Yan (Zhejiang University); and Jinling Wei (School of Computer and Computing Science, Hangzhou City University) Abstract Abstract Fine-grained visual categorization (FGVC) remains challenging due to subtle inter-class differences and large intra-class variations. Although Vision Transformer (ViT) has demonstrated strong global modeling capabilities, its patch-based representations often weaken local spatial awareness and restrict effective interaction across feature channels. To address these issues, we propose the Dual-Interactive Coordinate Aggregation Network to enhance feature representations. Specifically, we introduce a Coordinate-Aware Global Aggregation (CAGA) module that performs horizontal and vertical coordinate pooling to capture long-range spatial dependencies while preserving precise positional cues for accurate localization. Furthermore, to alleviate channel-wise information isolation, a Global Channel-Shuffle Attention (GCSA) module is designed by integrating grouped channel shuffling with depth-wise separable convolutions, facilitating dense global cross-channel interactions. Additionally, a multi-view feature fusion strategy is employed to jointly leverage enhanced local features and global semantic representations. Extensive experiments on CUB-200-2011, Stanford Cars, Stanford Dogs, and NABirds demonstrate that the proposed method consistently outperforms strong ViT baselines and achieves competitive performance against SOTA approaches. Cross-Scale Collaborative Dual-Resolution Network for Shadow Removal Wanzhi Hong, Qingxuan Shi, Fang Yang, Tong Wang, Jing Zhang, and Ming Cheng (Hebei University) Abstract Abstract Image shadow removal is inherently challenging due to its spatially non-uniform and context-dependent nature: degradation is local, while accurate correction requires reliable guidance from distant non-shadowed regions. Existing approaches struggle to balance this duality. Global or attentionbased models often introduce irrelevant information that contaminates shadow regions, whereas purely local methods fail to fully exploit long-range contextual cues. To address this issue, we propose a Dual-Resolution Collaborative Encoding (DRCE) framework that explicitly models fine-grained structural details and global context through two interacting resolutions. DRCE employs a high-resolution pathway to preserve detailed spatial structures and a low-resolution pathway to capture stable global illumination context. Their interaction is enabled by a CrossScale Collaborative Fusion (CSCF) module, which performs content-adaptive cross-scale token interaction to selectively integrate complementary information. Extensive experiments on multiple public benchmarks demonstrate that the proposed method achieves strong quantitative performance and high visual quality. Triplet-Based Feature Refinement and Quality Difference Perception Network for No-Reference Image Quality Assessment Chunyu Wu, Yihua Chen, Kejing Wu, and Zhenjun Tang (Guangxi Normal University) Abstract Abstract No-Reference Image Quality Assessment (NR-IQA) is an important technology in the field of computer vision. Most existing NR-IQA methods are constrained by the absence of reference images. These issues lead to the limited IQA performance. To address this, we propose a novel method called Triplet-Based Feature Refinement and Quality Difference Perception Network (TRQDNet) for No-Reference Image Quality Assessment. It introduces the restored image and the degraded image as dual reference images to form a triplet. The method consists of a multi-scale spatial-channel refinement module, a difference-guided distortion perception module and a quality prediction module. Firstly, the multi-scale spatial-channel refinement module refines the features extracted by the pre-trained Swin Transformer from the triplet. Secondly, the difference-guided distortion perception module constructs restoration and degradation difference features and performs cross-attention across the triplet features. It captures quality differences and learns difference-aware representations effectively. Finally, the quality prediction module predicts the score through weighted feature aggregation. Experimental results demonstrate that the TRQDNet outperforms some excellent NR-IQA methods in both prediction accuracy and generalization ability. Thursday Virtual Room 7 IJCNN Paper Vision-Language-Action and Robotic Learning I Session Chair: Qian Zhang (National University of Defense Technology), Yanglan Dong (University of Science and Technology of China; SKL of Processors, Institute of Computing Technology, CAS) DWIP: Efficient Depth–Width Integrated Pruning for Vision-Language-Action Models in Robotic Manipulation Qian Zhang, Rongchun Li, and Peng Qiao (National University of Defense Technology) Abstract Abstract Vision-Language-Action (VLA) models demonstrate excellent performance in robotic manipulation tasks, but their massive parameter scale leads to high computational and storage costs, posing deployment challenges. Existing acceleration approaches either rely on lightweight architectures that require costly retraining and exhibit limited generalization, or focus solely on inference acceleration without effectively reducing model size. To address these limitations, we propose DWIP, a depth–width integrated pruning strategy for VLA models. This method removes redundant layers from VLA models through dynamic layer pruning or task-driven adaptive layer pruning, enhancing inference efficiency while maintaining model performance. By further integrating width pruning, it constructs lightweight models that balance parameter size, accuracy, and speed under high pruning rates. Experimental results on the LIBERO benchmark demonstrate that DWIP enables lossless acceleration on simple tasks, and on complex tasks, achieves a 42.7% reduction in model parameters and a 1.82× inference speedup, while preserving 98.1% of the original model performance. Compressor-VLA: Instruction-Guided Visual Token Compression for Efficient Robotic Manipulation Juntao Gao (Beijing University of Technology), Feiyang Ye (Li Auto Inc.), and Jing Zhang (Beijing University of Technology) Abstract Abstract Vision-Language-Action (VLA) models have emerged as a powerful paradigm in Embodied AI. However, the significant computational overhead of processing redundant visual tokens remains a critical bottleneck for real-time robotic deployment. While standard token pruning techniques can alleviate this, these task-agnostic methods struggle to preserve task-critical visual information. To address this challenge, simultaneously preserving both the holistic context and fine-grained details for precise action, we propose Compressor-VLA, a novel hybrid instruction-conditioned token compression framework designed for efficient, task-oriented compression of visual information in VLA models. The proposed Compressor-VLA framework consists of two token compression modules: a Semantic Task Compressor (STC) that distills holistic, task-relevant context, and a Spatial Refinement Compressor (SRC) that preserves fine-grained spatial details. This compression is dynamically modulated by the natural language instruction, allowing for the adaptive condensation of task-relevant visual information. Experimentally, extensive evaluations demonstrate that Compressor-VLA achieves a competitive success rate on the LIBERO benchmark while reducing FLOPs by 59% and the visual token count by over 3x compared to its baseline. The real-robot deployments on a dual-arm robot platform validate the model's sim-to-real transferability and practical applicability. Moreover, qualitative analyses reveal that our instruction guidance dynamically steers the model's perceptual focus toward task-relevant objects, thereby validating the effectiveness of our approach. TAR: Time-Adaptive Reweighting for Data-Efficient Vision-Language-Action Model Post-Training Chang-Tao Zhao, Tian-Yu Xiang, and Xiao-Hu Zhou (Institute of Automation, Chinese Academy of Sciences) and Mei-Jiang Gui, Xiao-Liang Xie, Shi-Qi Liu, Ao-Qun Jin, and Zeng-Guang Hou (Institute of Automation, Chinese Academy of SciencesInstitute of Automation, Chinese Academy of Sciences) Abstract Abstract Supervised Vision-Language-Action (VLA) post-training typically aggregates per-action imitation losses with uniform temporal averaging, treating all actions equally despite success hinging on brief critical moments. To address this limitation, a method called Time-Adaptive Reweighting (TAR) is proposed. It changes only how action losses are aggregated over time during training, without altering the policy architecture or inference procedure. TAR derives a weight from each action’s loss relative to other actions within the same segment, where segments are sampled from trajectories. It then replaces uniform temporal averaging with a weighted mean of the original per-action losses to form the segment loss. TAR improves the overall success rate by \textbf{3.04\%} on MetaWorld and \textbf{2.16\%} on LIBERO under the same demonstration data and evaluation protocol, averaged over five representative VLA models. It is also data-efficient by reducing horizon-normalized interaction steps under the same demonstration budget. The learned weights place more training emphasis on a small set of success-critical moments, which typically occur during early approach and around interaction onset (first contact and early interaction). For long-horizon multi-stage tasks, an additional peak can appear near completion. This shift in training emphasis enables a more effective policy, improving success rate and evaluation efficiency while making better use of limited demonstrations when data collection is costly. PatchAlign: Fine-Grained Semantic–Geometric Alignment for Robotic Manipulation Yanglan Dong (University of Science and Technology of China; SKL of Processors, Institute of Computing Technology, CAS); Jiaming Guo (SKL of Processors, Institute of Computing Technology, CAS); Yunkai Gao and Siming Lan (Institute of Al for Industries); Shaohui Peng (Intelligent Software Research Center, Institute of Software, CAS); Zihao Zhang (Institute of Al for Industries); and Rui Zhang and Xing Hu (SKL of Processors, Institute of Computing Technology, CAS) Abstract Abstract Robot manipulation policies that generate actions from visual observations and language instructions have achieved promising progress, yet robust scene understanding remains a major challenge, especially for precise manipulation tasks. Methods relying solely on 2D visual representations capture high-level semantics but lack spatial awareness, while existing 3D-enhanced approaches fail to preserve fine-grained geometric details and local semantic–geometric correspondences. This limitation suggests that effectively combining semantic information from pretrained visual encoders with detailed local 3D geometry is crucial for enabling precise and reliable manipulation. To address these limitations, we propose PatchAlign, a patch-level cross-modal fusion framework that explicitly aligns RGB image patches with local 3D geometric regions. Semantically rich image patches are extracted using a frozen pretrained visual encoder, while 3D geometry is captured via a lightweight clustering strategy. By establishing patch-level semantic–geometric correspondences, our method enables precise local reasoning without compromising the generalization of pretrained vision-language models. We evaluate our method on multiple manipulation benchmarks across diverse tasks. Experimental results demonstrate that the proposed framework consistently improves manipulation performance and robustness, particularly in scenarios requiring fine-grained spatial reasoning and precise object interactions. We believe this work presents a effective framework for fine-grained cross-modal fusion, offering new insights into leveraging 3D geometry in vision-language-action learning. Thursday Virtual Room 8 IJCNN Paper Vision-Language-Action and Robotic Learning II Session Chair: Jian Xue (University of Chinese Academy of Sciences), Weixuan Liu (Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China) ACT-US: End-to-End Visuomotor Policy Learning for Respiratory-Adaptive Robotic Ultrasound Xingyu Zhang, Jian Xue, Hongjuan Pei, and Qingyuan Liu (University of Chinese Academy of Sciences); Yihua Shao (Institute of Automation, Chinese Academy of Sciences); Xiang Luo (University of Chinese Academy of Sciences); and Ke Lu (University of Chinese Academy of Sciences, Peng Cheng Laboratory) Abstract Abstract Autonomous robotic ultrasound (RUS) focuses on the standardized acquisition of diagnostic images. Traditional control-based methods excel in maintaining stable contact but operate under a restrictive ''Initialization Assumption,'' often neglecting the complex visual-guided approach phase. In contrast, pure end-to-end learning methods offer global trajectory planning capabilities but often struggle to guarantee patient safety and comfort amidst physiological motion such as respiration. Therefore, we propose ACT-US, a hierarchical autonomy framework that combines generative imitation learning and respiratory-adaptive control to perform full-cycle autonomous scanning. Specifically, ACT-US first employs a Conditional Variational Autoencoder (CVAE) with Action Chunking to model multi-modal expert behaviors, generating nominal trajectories that seamlessly bridge free-space navigation and soft-tissue manipulation. Then, it utilizes a weakly-supervised quality estimator to learn a ''comfort-aware'' force policy by correlating contact pressure with image clarity without manual labeling. Finally, it integrates a respiratory-adaptive admittance controller that uses feedforward visual tracking to actively compensate for thoracic motion. We conducted extensive experiments on human subjects. The results demonstrate that ACT-US achieves expert-level dexterity in unstructured environments, significantly outperforming baselines that lack physiological adaptation. REACT: Responsive Execution-aware Action Coordination And Tuning For Moving Object Manipulation Xiaofei Qin (University of Shanghai for Science and Technology) and Jianyu Zhang and Anluo Yi (University of Shanghai for Science and Techology) Abstract Abstract Learning to reliably and stably manipulate moving objects with bimanual agents remains a challenging task. Many real-world manipulation tasks involve dynamic scenes, where objects of interest may be in motion. While existing visual manipulation methods primarily focus on manipulating static objects, leading to limited applicability when dealing with moving objects. These methods struggle to meet the strict demands on real-time responsiveness and robustness required in moving object manipulation. Furthermore, large-scale datasets for dynamic manipulation settings remain scarce. To alleviate these challenges, we introduce REACT (Responsive Execution-aware Action Coordination and Tuning for moving object manipulation), a novel sim-to-real method for manipulation of moving objects. Within REACT, we adopt a simulation-based dynamic scene dataset generation strategy for bimanual agents and employ a Slow–Fast policy to achieve closed-loop control and real-time refinement of action trajectories. This design enables zero-shot transfer from simulation to real-world dynamic settings, substantially improving stability and responsiveness in manipulation. Extensive experiments on real-world moving object manipulation tasks demonstrate that REACT outperforms existing visual imitation learning methods. LUNA: Low-Light Robust Panoptic Lifting for Adverse Robotic 3D Scene Perception Ahalya Ravendran (CSIRO); Léo Lebrat and Rodrigo Santa Cruz (Queensland University of Technology); and Hu Zhang, Lars Petterson, Dadong Wang, and Xun Li (CSIRO) Abstract Abstract Robotic perception in low-light environments remains challenging due to severe noise and motion blur that degrade RGB imagery and hinder reliable scene understanding. Although recent 3D panoptic reconstruction methods unify semantic and instance segmentation, they largely assume high quality inputs - an assumption that fails under real-world conditions. We present LUNA, a geometry-aware panoptic lifting framework that integrates depth-guided consistency and adaptive multitask optimization for robust reconstruction under degraded visual inputs. We further introduce a systematically degraded Replica dataset with progressive noise and blur to evaluate robustness. Experiments show that LUNA consistently outperforms restoration-based baselines in panoptic segmentation and volumetric reconstruction across all degradation levels. Quantization-Aware Super-Resolution for Real-Time Embodied Visual Perception Weixuan Liu, Jintao Li, and Yun Li (Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China) Abstract Abstract In an embodied intelligence system such as a mobile robot, relatively low-quality images captured by onboard cameras and their subsequent visual perception must run in real time with limited computing resources and power, while processing low-quality images captured by onboard cameras. Using super-resolution (SR) as an on-device perception front-end can improve visual quality for downstream perception tasks, but existing methods often fail to balance image quality, processing speed, and efficient deployment. In this work, we design a Dual-Cross Gate Network (DCGNet), as a hardware-aware SR architecture for real-time embodied visual perception. The DCGNet adopts a dual-stream structure that separately processes spatial and channel information, thus enabling efficient feature extraction. The network is trained with quantization-aware strategies and a lightweight attention mechanism to ensure robustness under a robotic low-precision inference. Then it is deployed on 8-bit integer arithmetic for real-world humanoid applications to test accuracy losses. Experimental results show that the DCGNet achieves real-time performance of 106 frames per second at 2K resolution, while maintaining image quality comparable to its full-precision version. These results verify that DCGNet can serve as a practical and deployable visual enhancement module for embodied perception under real-world system constraints. Thursday Virtual Room 1 IJCNN Paper Visual Anomaly and Defect Detection I Session Chair: Chaoli Wang (上海理工大学), Chengming Liu (Zhengzhou University, School of Cyber Science and Engineering) FastFlow-DETR: An Efficient Steel Surface Defect Detection Algorithm Based on RT-DETR Yihan Chai (Zhengzhou University); Qiming Yu (Zhengzhou Normal University, School of Information Science and Technology); and Chengming Liu (Zhengzhou University) Abstract Abstract Surface defect detection in steel is critical for industrial quality control. However, existing detection frameworks tend to dilute fine-grained defect information during multi-level feature processing and introduce unnecessary computational overhead. This leads to missed detection of weak defects, unstable performance, and limited real-time applicability in industrial scenarios. In this paper, we introduce FastFlow-DETR, a lightweight and efficient framework built upon RT-DETR, designed to address these challenges. The backbone incorporates a Selective Spatial Block (SSB) to preserve critical local features, while a Branch Interweave Module (BIM) is used during feature fusion to align shallow detail features with deep semantic representations. In addition, a Dual Dynamic Multi-Scale Relay (D²MR) encoder is proposed to enhance feature interaction across different representation levels, within which a Self-Scaled Tanh Normalization (STN) stabilizes feature modulation and prevents excessive suppression of weak defect responses. Experimental results on GC10-DET and NEU-DET show that FastFlow-DETR achieves higher detection accuracy than RT-DETR, with mAP@0.5 gains of 2.9 and 2.4 percentage points, respectively. Meanwhile, the proposed method reduces computational complexity by 32.1% in FLOPs and model parameters by 29.4%. These results show that the proposed method effectively balances detection accuracy and computational efficiency for industrial steel surface defect detection. PGS-DETR:A Lightweight Detector for Multi-Class Steel Surface Defects in Complex Backgrounds Yongfeng Qiu (Hunan University of Technology, Hunan Zhuzhou 412007, China.; Hangzhou Huaxin Electrical & Mechanical Engineering Co., Ltd, Zhejiang Hangzhou 310030 , China.) and Shiyao Liu, Zhibing Wang, Bing Tang, Kaixi Luo, Lanlin Liu, and Ning Liu (Hunan University of Technology, Hunan Zhuzhou 412007, China.) Abstract Abstract Steel surface defect detection is critical for online quality control, yet small-scale and low-contrast defects remain challenging under complex textures and reflections. To address the challenges that steel-surface defects often appear at small scales, exhibit low contrast, and are easily confused with the background in complex scenarios, we propose PGS-DETR, a lightweight end-to-end defect detection algorithm. In PGS-DETR, a C2f-PFD module is introduced into the backbone to enhance the extraction of fine-grained characteristics for tiny defects. Moreover, we design a polarity convolution gated attention module(AIFI-PCGA), to suppress disturbances caused by specular reflections and texture noise. Furthermore, we develop a lightweight cross-scale fusion structure(Slim-Neck), which improves cross-channel interaction efficiency and strengthens multi-scale information aggregation. Experimental results in the NEU-DET dataset show that PGS-DETR achieves an mAP50 of 76.5% and an mAP95 of 43.2%, outperforming the baseline by 5.1% and 2.6%, respectively, while reducing the number of parameters and computational cost by 40.3% and 35.2%. In addition, cross-dataset generalization on DeepPCB yields an mAP50 of 97.6%, further demonstrating the effectiveness and deployability of the proposed method in industrial visual inspection. PAA-YOLO: A Steel Surface Defect Detection Method Based on Progressive Axial Attention Yunfei Xiong (Jiangnan University); Kaining Liu (Huzhou University); Jingyue Sun (University of Science and Technology of China); and Chenglong Fu, Jian Yao, Chuang Wang, and Pengjiang Qian (Jiangnan University, The PRC Ministry of Education Engineering Research Center of Intelligent Technology for Healthcare) Abstract Abstract To address the industrial demand for higher accuracy and efficiency in steel surface defect detection, this paper proposes PAA-YOLO, an improved model based on YOLOv11. To enhance the model’s extraction of linear defects (e.g., scratches), we design the Progressive Axial Attention (PAA) module. Using the PAA module, we restructure the backbone’s C3k2 and C2PSA modules into C3k2-PAA and C2PSPAA, boosting the backbone’s feature extraction capability. Finally, we optimize the neck network by integrating the PAA module and hypergraph-based semantic collection and scattering framework (HGC-SCS), inserting a hypergraph computation intermediate layer between the backbone and neck to strengthen high-order semantic feature correlations and improve the neck’s feature fusion performance. Evaluations on the public NEU-DET dataset show PAA-YOLO achieves an mAP50 of 81.0% (3.1% higher than the baseline), with precision and recall improved by 4.6% and 4.0% respectively. On GC10-DET, it gains a 3.3% mAP50 improvement over the baseline. Experimental results confirm PAA-YOLO effectively enhances steel surface defect detection accuracy. CAM-DETR: An Efficient End-to-End Transformer for Real-Time Metal Surface Defect Detection Zheyuan Lu, Chaoli Wang, and Zhanquan Sun (University of Shanghai for Science and Technology) Abstract Abstract Real-time and precise metal surface defect detection is crucial for industrial quality control. Existing deep learning methods face challenges in practical applications: redundant backbones limit edge inference efficiency; attention mechanisms suffer from high computational overhead and low robustness; and insufficient cross-layer interaction hinders multi-scale defect handling. To address these, we propose CAM-DETR, a lightweight end-to-end detector based on RTDETR. It features three core modules: (1) CE-RGELAN, a lightweight backbone that reduces complexity while preserving fine-grained features; (2) Additive Token Mixer (ATM), an encoder attention mechanism replacing quadratic self-attention with linear-complexity operations to cut costs and boost noise robustness; (3) MANet, a multi-scale fusion network that enhances localization via parallel branches and depthwise convolutions. Experiments show CAM-DETR reduces parameters by 10.97% and computation by 15.26%. It achieves 79.17% mAP@0.5 on NEU-DET (+2.29%, +57.0% speed) and 74.52% on GC10-DET (+5.59%, +89.3% speed), demonstrating an excellent accuracy-latency balance for industrial deployment. Index Terms—Object detection, real-time detection, end Thursday Virtual Room 2 IJCNN Paper Visual Anomaly and Defect Detection II Session Chair: Mingle Zhou (Qilu University of Technology; Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, China.), Lei Duan (Sichuan University) PPRL:Prototype-driven Periodic Representation Learning for Multivariate Time Series Anomaly Detection Mingle Zhou, Shijie Miao, Yuan Gao, Min Li, and Delong Han (Qilu University of Technology) Abstract Abstract Multivariate time series anomaly detection is critical in many real-world applications, where periodic structure is pervasive and anomalies often appear as subtle violations of temporal regularities. Existing reconstruction-based methods, even when they incorporate frequency-domain modeling, typically capture periodic patterns only implicitly. As a result, they may over-reconstruct anomalous segments and wash out the residual cues needed to detect periodic anomalies. To tackle this, we propose PPRL, a self-supervised framework that makes periodicity an explicit structure through learnable periodic prototypes. PPRL builds a multi-scale periodic prototype representation and introduces two training-only objectives: prototype-driven periodic contrastive learning and prototype-driven periodic anomaly regularization. These objectives better align representation learning with periodic structure violations. The learned periodic representations are injected into a time--frequency reconstruction backbone, while inference-time scoring still relies solely on reconstruction residuals. Extensive experiments on four public benchmarks show that PPRL delivers consistent gains, especially under affiliation-based metrics, and mitigates the over-reconstruction of periodic anomalies without increasing inference-time complexity. MSDDS: Multivariate Time Series Anomaly Detection via Multi-Scale Decomposition with Fusion and Dual-Stream Channel Transformer Yuying Ma (Xinjiang University, School of Computer Science and Technology); Jiong Yu and Shu Li (Xinjiang University); and Yong Hu (Beijing University of Posts and Telecommunications) Abstract Abstract Multivariate time series anomaly detection is widely used in applications such as industrial monitoring, financial risk control, and cyber-physical system security. However, real-world multivariate time series usually exhibit non-stationary temporal dynamics and complex channel correlations, making anomaly detection particularly challenging. Surpassing existing time-frequency integrated approaches that typically process dual-domain features in parallel, we revisit anomaly detection from a structure-enhanced perspective, where temporal dynamics are first explicitly refined through multi-scale decomposition to serve as a robust basis for subsequent frequency-domain modeling. Accordingly, we propose MSDDS, which integrates a Multi-Scale Decomposition with Fusion (MSDF) module that leverages a dynamic adaptive strategy to integrate diverse temporal patterns, and a Dual-Stream Channel Transformer (DSCT) for modeling complementary global and local channel dependencies in the frequency domain. Extensive experiments on seven widely used benchmark datasets demonstrate that MSDDS achieves state-of-the-art performance on most benchmarks, validating its effectiveness and robustness in unsupervised multivariate time series anomaly detection. TF-ChebKAN: Time-Frequency Coupled Chebyshev Kolmogorov-Arnold Networks for Multivariate Time Series Anomaly Detection Feng Wang (Southern University of Science and Technology, Shenzhen University of Advanced Technology); Xiaoheng Wang (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences); Shuaipeng Wu (Southern University of Science and Technology; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences); and Kejiang Ye (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; Shenzhen University of Advanced Technology) Abstract Abstract Multivariate time series anomaly detection requires a balance between model capability and computational efficiency. Many existing methods struggle to simultaneously capture complex nonlinear patterns and multivariate dependencies. In this paper, we propose a lightweight two-stream architecture, TF-ChebKAN, which integrates a Chebyshev II-enhanced Kolmogorov-Arnold network to address the boundary instability problem commonly encountered in multivariate time series anomaly detection. Furthermore, we use a multivariate frequency mixer to capture global cross-channel correlations by processing spectral components with linear complexity. Finally, an adaptive fusion mechanism dynamically adjusts the balance between the two modules. Experiments on MSL, PSM, SMD and SWaT show that TF-ChebKAN achieves strong performance, while using fewer parameters than baselines. Furthermore, TF-ChebKAN significantly reduces training and inference times, demonstrating efficient anomaly detection capabilities in resource-limited scenarios. DRIFT: A Dual-Branch Time-Frequency Transformer for Time Series Anomaly Detection Jiaxin Wu, Jiaxuan Xu, and Lei Duan (Sichuan University) Abstract Abstract Although existing reconstruction-based methods have significantly improved state-of-the-art results for time series anomaly detection, they still struggle with distribution shifts in time series exhibiting complex patterns. To address the issue, this paper proposes a novel transformer-based model that is robust to distribution shifts, called Dual-bRanch tIme-Frequency Transformer (DRIFT). DRIFT proposes a dual-branch attention layer that focuses on the inherent periodic patterns in time series. The core principle is that even when distribution shifts occur, the periodic structure remains stable. Specifically, one branch employs Fourier transforms to extract global periodic patterns, while the other branch uses wavelet transforms to capture local periodic patterns. Finally, the contrastive structure learns pattern consistency across both branches, reinforcing the consistency of normal data while highlighting the distinctiveness of anomalies. Extensive experiments on real-world time series datasets demonstrate the effectiveness of DRIFT. Thursday Virtual Room 3 IJCNN Paper Visual Anomaly and Defect Detection III Session Chair: Guitao Cao (Shanghai Key Laboratory of Trustworthy Computing, East China Normal University), Jiong Yu (Xinjiang University, Department of Computer Science and Technology) ReMAP-AD: Unifying Visual Memory Reconstruction and Prompt Learning for Few-Shot Anomaly Detection Anshuo Yin and Jiong Yu (Xinjiang University, Department of Computer Science and Technology) and Jiangqi Shi (Kashi University, School of foreign language) Abstract Abstract Few-shot industrial anomaly detection is typically trained with only a small $K$-shot set of normal images, where reliable localization is easily confounded by repetitive textures and background patterns.Recent CLIP-based prompting introduces semantic priors through text, yet its semantic responses can vary substantially with prompt templates, leading to over-activation and spatial drift. This paper presents ReMAP-AD, an evidence-calibrated vision--language framework that constrains prompt-driven semantics with explicit residual visual evidence derived from the normal support set.ReMAP-AD first stabilizes the abnormal semantic anchor by aggregating complementary abnormal prompts into prototypes.It then constructs a reconstructive visual memory over multi-scale patch features to softly reconstruct query patches from the $K$-shot normal gallery, and converts reconstruction discrepancies into a residual evidence map.Guided by this evidence, a residual evidence-guided semantic alignment module calibrates the raw semantic anomaly map via an evidence-adaptive gate, and the final prediction is obtained by harmonic-mean fusion to emphasize agreement between semantic and evidence cues.A lightweight routing head further outputs defect-type likelihoods for explanation without modifying the detection pathway. Experiments on MVTec AD and VisA under few-shot protocols demonstrate consistent gains in image-level detection and pixel-level localization, and robustness evaluations under template perturbations show reduced performance dispersion and smaller worst-case degradation. CoL-SAM : Fine-Grained Zero-Shot Anomaly Segmentation via Collaborative CLIP–SAM Yuhang Xie and Xiaofei Ma (University of Jinan); Lin Wang (University of Jinan, Quan Cheng Laboratory); and Bo Yang (Quan Cheng Laboratory, University of Jinan) Abstract Abstract Zero-shot anomaly segmentation (ZSAS) aims to localize defect regions on unseen industrial products without target-category training data, reference samples, or dense pixel-level annotations. Recent foundation models such as CLIP and SAM enable training-free transfer, yet their direct use remains challenging: CLIP often produces spatially coarse and noisy anomaly cues due to its global contrastive alignment, while SAM may generate redundant or over-segmented masks when driven by unreliable prompts, leading to non-trivial post-processing. In this work, we propose CoL-SAM, a collaborative CLIP-SAM framework for fine-grained ZSAS. First, we follow a WinCLIP-style multi-scale window scoring scheme to extract anomaly evidence and introduce an Adaptive Score Fusion (ASF) module to replace the fixed, manually designed scale fusion used in WinCLIP. ASF performs fusion with spatially varying weights over multi-scale anomaly maps, producing sharper anomaly guidance with reduced noise propagation.Second, we transform the fused anomaly guidance into compact prompt constraints that steer SAM toward boundary-accurate defect masks while suppressing redundant proposals. Experiments on MVTec-AD and VisA demonstrate that CoL-SAM consistently outperforms CLIP-only, SAM-only, and WinCLIP-based baselines, validating the effectiveness of adaptive multi-scale fusion and prompt-constrained refinement for zero-shot industrial anomaly segmentation. Zero-Shot Industrial Anomaly Detection Based on Category-Aware Adaptive Vision-Language Agent Qianlin Lei, Li Li, and Zhanpeng Wang (Southwest University of Science and Technology,School of Computer Science and Technology) Abstract Abstract While Large Vision-Language Models (VLMs) demonstrate remarkable cognitive capabilities in general scenarios, they face fundamental perceptual barriers when handling physics-constrained industrial extremes, such as micron-level deformations and high-frequency illumination noise. This paper identifies a critical “Semantic-Resolution Gap” in current generalist models. To address this, we propose a Category-Aware Hybrid Perception Framework (HPF). This framework constructs an adaptive agent capable of dynamically restructuring inference paths based on the physical attributes of the object: (1) utilizing a single-shot reference mechanism to counteract reflective noise on metal surfaces; (2) employing a positional prior-guided physical zooming mechanism to overcome perceptual bottlenecks in microstructures; and (3) combining signal enhancement with semantic parsing to handle low-contrast defects. Furthermore, we introduce a Neuro-Symbolic Logic Calibrator to map probabilistic textual outputs into deterministic industrial-grade verdicts, effectively mitigating the model’s “Response Inertia.” Experiments on the MVTec AD dataset show that our HPF method, based on the lightweight Phi-3.5-Vision (3.8B), achieves an average accuracy of 74.67%, with a peak accuracy of 80.87% on challenging categories like Metal Nut, significantly outperforming the 7Bparameter SOTA baseline LLaVA-1.5. The proposed method enables cold starts with only a single golden sample and achieves efficient quasi-real-time inference on resource-constrained edge computing devices, reducing memory usage to ∼3.2 GB. Uncertainty-Guided Dual-View Distillation for Industrial Anomaly Detection Conghua Wei, Xidong Xi, Guitao Cao, and Mengshang Nie (Shanghai Key Laboratory of Trustworthy Computing, East China Normal University) Abstract Abstract Teacher-Student (T-S) architectures are promising for industrial anomaly detection but often suffer from ``over-generalization'', where students learn trivial local identity mappings that bypass global spectral constraints, allowing them to undesirably replicate high-frequency anomalies. To address this issue, we propose a novel Dual-View Distillation Framework (DVDF). First, we propose an Omni-Frequency Student (OFS) that enforces consistency across spatial and frequency domains, effectively inhibiting the replication of anomalous patterns. Second, we design a Cross-scale Efficiency Block (CEB) to integrate multi-scale features, preserving salient fine-grained details often lost in standard bottlenecks. Third, we employ an Uncertainty-Guided Distillation (UGD) strategy via Monte Carlo Dropout to suppress unreliable supervision signals, preventing the student from mimicking noisy teacher predictions. Results on HSS-IAD and MVTec AD benchmarks indicate that our method achieves highly competitive performance compared to current state-of-the-art methods. Thursday Virtual Room 4 IJCNN Paper Visual Anomaly and Defect Detection IV Session Chair: Chen Ling (Center for Information Research, Academy of Military Sciences, Beijing 100142, China), Xiaowu Liang (Jinan University) LogRD: A Robustness-Enhanced Framework for Log Anomaly Detection with Dynamic Masking Xiaowu Liang, Xianxia Zou, Jie Ren, Zijiong Su, Cenyu Zheng, and Weiwu Xu (Jinan University) Abstract Abstract Log anomaly detection is vital for system reliability, yet existing methods struggle with high labeling costs or limited semantic adaptability. This paper proposes LogRD, a novel semi-supervised framework that leverages only normal logs for training. To address the “mask position dependency” prevalent in current masked language models, LogRD introduces a dual-masking strategy: dynamic masking during training and a unique odd-even masking scheme at inference. By strengthening semantic understanding and ensuring comprehensive sequence evaluation, LogRD demonstrates superior robustness. Experimental results on three benchmark datasets show that LogRD consistently outperforms state-of-the-art baselines in detection performance. A Target-Attenuation and Mask-Purification Network for Hyperspectral Anomaly Detection Ke Wu, Song Liu, Congxuan Zhang, and Zhen Chen (Nanchang Hangkong University) Abstract Abstract Hyperspectral anomaly detection (HAD) aims to identify targets with spectral characteristics distinct from the background. Reconstruction-based methods are commonly used and effective in HAD, but they are susceptible to the influence of anomalous targets during background generation, which can compromise the purity of the background and reduce detection performance. To address this issue, we propose the Target-Attenuation and Mask-Purification Network (TAMP-Net). Our framework first employs a pre-detection module for initial target localization, followed by a decaying background replacement strategy to attenuate the influence of anomalies. Subsequently, a high-ratio masked autoencoder is employed, applying random masking to the remaining regions to further purify the background. Finally, the anomaly map is generated by computing the residual between the original image and the reconstructed background. Experiments conducted on multiple hyperspectral datasets have demonstrated the effectiveness and superiority of the proposed method. Towards Cross-Scene Hyperspectral Anomaly Detection via a One-Step Generative Framework Cuiwei Liu, Yujing Zhao, Zhaokui Li, and Huaijun Qiu (Shenyang Aerospace University) Abstract Abstract Hyperspectral Anomaly Detection (HAD) aims to identify small and sparse targets that significantly deviate from the abundant background spectra, without requiring prior knowledge. Most state-of-the-art methods follow a reconstruct-then-detect paradigm, primarily learning scene-specific background distributions. Consequently, the well-trained models suffer from severe performance degradation when applied to unseen scenes. To this end, this study introduces a novel one-step detection paradigm that formulates HAD as a generative conditional diffusion process. We develop a Conditional Diffusion Model-based Hyperspectral Anomaly Detection framework (CDM-HAD), which leverages multi-scale hyperspectral features as conditional guidance to progressively reconstruct anomaly maps from noisy inputs. The iterative denoising process implicitly learns generalizable anomaly-background deviation patterns and alleviates overfitting to training-specific distributions. To overcome data scarcity, the proposed CDM-HAD incorporates an Anomaly Simulation Strategy (ASS) for generating high-quality hyperspectral anomaly data. This approach perceives inherent anomalies in the training scene and adaptively injects cross-scene anomalous spectra, thereby synthesizing diverse and realistic anomaly samples. Experimental results highlight the superior cross-scene generalization of the proposed CDM-HAD, achieving robust anomaly detection across diverse unseen scenes when trained on a single scene. K-LAD: Knowledge-Enhanced Log Anomaly Detection via Syntax-Aware Verification Lizhu Mi (North China University of Technology; Center for Information Research, Academy of Military Sciences, Beijing 100142, China) and Hongbin Zhang, Lu Li, Wei Gu, Feng Tian, and Chen Ling (Center for Information Research, Academy of Military Sciences, Beijing 100142, China) Abstract Abstract Large Language Models (LLMs) in anomaly detection often exhibit a Defensive Over-reaction on imbalanced data, while standard retrieval mechanisms relying on generic semantic models (e.g., BERT) suffer from a Semantic-Syntactic Gap by treating logs purely as natural language. To address these challenges, we propose K-LAD, a data-efficient framework grounded in the insight that logs are weakly-structured code derivatives. K-LAD employs a coarse-to-fine cascade architecture: (1) A High-Recall Anomaly Candidate Generator (LoRA-tuned LLaMA-3 with Focal Loss) designed to maximize anomaly capture; and (2) A Syntax-Aware Knowledge Verification stage utilizing UniXcoder to rigorously correct false positives via machine syntax alignment. Extensive experiments on HDFS, BGL, and Thunderbird demonstrate that K-LAD achieves SOTA performance using only 2,000 training samples, empirically validating the superiority of modeling logs via code-derived syntax over traditional NLP approaches. Thursday Virtual Room 5 IJCNN Paper Visual Anomaly and Defect Detection V Session Chair: Hongtao Wang (Department of Computer, North China Electric Power University; Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education), Xiaolong Zheng (University of Chinese Academy of Sciences) CDAD: A Multi-Scale Feature Fusion Model for Cross-Domain Anomaly Detection Hui Zhang and Hongyu Wang (Department of Computer, North China Electric Power University); Jianyong Zhu (Department of Computer, North China Electric Power University; Hebei Key Laboratory of Knowledge Computing for Energy & Power); and Hongtao Wang (Department of Computer, North China Electric Power University; Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education) Abstract Abstract Anomaly detection aims to identify deviations by learning normal data distributions, mainly covering industrial, semantic, and medical domains. Most existing methods, however, are domain-specific and degrade when transferred. To address this limitation, we proposed CDAD, a unified Cross-Domain Anomaly Detection model. CDAD leverages a pre-trained feature extractor with a feature adaptor to align cross-domain features and tighten the decision boundary of normal samples. A hybrid loss emphasizes hard-to-classify anomalies, while a Multi-Scale Feature Fusion Discriminator integrates global and local features. Extensive experiments on five datasets across three domains demonstrate that CDAD consistently outperforms recent state-of-the-art approaches, highlighting its robustness and effectiveness in cross-domain anomaly detection. The code is available at https://github.com/Stardust457/CDAD. Elastic One-Class Domain-Adversarial Neural Network for Trustworthy Anomaly Detection Yang Liu, Jianliang He, Long Chao, and Wentao Mao (School of Computer and Information Engineering Henan Normal University) Abstract Abstract Deep transfer learning has shown substantial potential in developing the decision capability of one-class anomaly detection. However, existing SVDD-like methods adopting hypersphere-formed decision mechanism with hard boundary generally neglect the transitional characteristics from weak anomalies (e.g., incipient faults of mechanical supporting element) to genuine anomalies. Also, these methods lack uncertainty quantification for weak anomalies near decision boundary, thus leading to inadequate decision confidence. To address these concerns, this paper proposes an elastic one-class domain-adversarial neural network for trustworthy anomaly detection. By adopting the classical Deep SVDD as the baseline architecture, a deep elastic SVDD model with hypersphere annulus is built to capture weak anomalies’ transitional characteristics, while the anomaly probability of samples within the annulus is calculated using the Gumbel Copula function. On this basis, an elastic one-class domain-adversarial neural network with hypersphere annulus is developed to seek domain-invariant representation of one-class discriminative information. By forcing geometric overlap of the hyperspheres in source and target domains, not only the detection accuracy for genuine anomalies but also the identification capability for weak ones are both enhanced. Experimental evaluation is conducted on two typical anomaly detection problems, i.e., image recognition detection on the MNIST~USPS dataset and early fault detection on the IEEE PHM Challenge 2012 bearing dataset. The results show that the proposed model is capable of reaching higher detection accuracy on data of various forms with lower false alarm rate and better detection confidence for weak anomalies. Towards Open World Anomaly Detection in Tabular Data: A Method for Adapting to Normality Shift Shu Li, Yi Lu, and Jiong Yu (Xinjiang University) Abstract Abstract Current tabular anomaly detection methods can effectively identify unknown and diverse anomalies by learning general patterns of normal data and detecting deviations from these patterns. However, their strong performance relies on a closed-world assumption that the distribution of normal data remains stable over time. In open-world scenarios, external factors such as environmental changes often cause shift in the normal data distribution, known as normality shift, which can lead to significant performance degradation. Thus, we propose a tabular Anomaly Detection method for Adapting to Normality Shift (ADANS). During training, ADANS learns perturbation-invariant representations of normal data to improve robustness and adaptability to normality shift. During testing, it employs an adaptive representation revision process to align the representations of normal data between the training and testing sets. This process is guided by a dynamic weight adjustment mechanism, which prioritizes likely normal samples while suppressing the influence of anomalies. Extensive experiments on four real-world benchmarks demonstrate that the proposed ADANS outperforms ten state-of-the-art methods in adapting to normality shift. Breaking the Evidence Trap: Trustworthy Graph Anomaly Detection via Multi-View Evidential Alignment Songran Bai, Xiaolong Zheng, and Daniel Zeng (Institute of Automation, Chinese Academy of Sciences) Abstract Abstract Unsupervised graph anomaly detection is fundamentally challenged by the difficulty of distinguishing inherent data noise from genuine out-of-distribution anomalies. Although evidential deep learning offers a computationally efficient framework for uncertainty quantification, its direct adaptation to graph anomaly detection exposes three unresolved pathologies. First, continuous attribute regression suffers from evidence contraction, where suppressed evidence counts induce vanishing gradients that degrade shared encoder representations. Second, topology reconstruction exhibits structural confidence bias, as extreme edge sparsity drives the model toward globally overconfident predictions for edge absence. Third, existing anomaly scoring relies on linear aggregation of reconstruction error and epistemic uncertainty, which neglects their nonlinear interaction and renders the score unreliable under moderate uncertainty. To address these limitations, we propose \textbf{MVEA}, a Multi-View Evidential Alignment framework that integrates multi-view evidential learning with robust anomaly evaluation. We introduce a gradient compensation mechanism via latent evidential clustering to mitigate the risk of regression collapse. For structure learning, we design a cost-sensitive evidential strategy with dual-branch decoding and contrastive ranking to rectify confidence bias. We further propose a neighborhood distribution alignment module to enforce local semantic consistency beyond individual node reconstruction. Finally, we devise an uncertainty-rectified scoring function with a U-shaped gating mechanism to suppress ambiguous noise while amplifying reliable anomaly signals. Extensive experiments on benchmark datasets demonstrate that MVEA achieves superior detection accuracy and robustness compared to state-of-the-art methods. Thursday Virtual Room 6 IJCNN Paper Visual Anomaly and Defect Detection VI Session Chair: 书一 尚 (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences), Xuanning Liu (Beijing University of Posts and Telecommunications) RESTAD: Residual-Enhanced Spatio-Temporal Modeling with Multi-Scale Attention for Multivariate KPI Anomaly Detection Shuyi Shang (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Abstract Abstract Multivariate Key Performance Indicators (KPIs) are essential time-series data for characterizing the real-time operational status of Internet-based services, where precise anomaly detection is fundamental to maintaining system stability. Although current unsupervised reconstruction models learn normal patterns by mining spatio-temporal dependencies, they struggle to capture fine-grained features amidst noise interference. Addressing the limitation where current models bias toward fitting low-frequency trends while neglecting high-frequency fluctuations, we observe that the high-frequency components, which are easily relegated to the residual space by base learners, actually carry critical spatio-temporal correlation patterns that reflect system state transitions. Specifically, a base learner leveraging spatial attention mechanism is first employed to filter out residuals. Subsequently, a dual-scale temporal attention learner models both low-frequency trends and short-term fine-grained features within the residuals. Furthermore, we introduce a memory-enhanced cross-attention mechanism to refine residual representations. Simultaneously, a spatial consistency constraint based on KL divergence is incorporated to strengthen the modeling of topological associations among multivariate KPIs. Extensive experimental results demonstrate that RESTAD outperforms state-of-the-art methods across four real-world datasets, exhibiting a superior capability to identify anomalies from intricate high-frequency backgrounds. FSA-AD: Frequency-Semantic Anomaly Detection with Adaptive Patching and Orthogonal Prototypes Limin Liu, Yiran Qian, Qianqian QI, Jun Long, and Yueyi Luo (Central South University) Abstract Abstract Abnormal patterns in time series often manifest heterogeneously across frequency components: low-frequency signals encode long-term trends and regime shifts, whereas high-frequency signals capture short-term fluctuations and local structural variations. However, prevailing unsupervised anomaly detection methods typically employ uniform modeling granularities across frequency components, leading to suboptimal feature representation and semantic entanglement. To address these limitations, we propose FSA-AD, a Frequency-Semantic Adaptive framework. Central to our approach is a frequency-adaptive patching mechanism that explicitly aligns modeling resolution with spectral densities, enabling the simultaneous capture of global regimes and fine-grained morphological transients. Furthermore, we introduce a dual-stream orthogonal prototype memory to enforce a disjoint information bottleneck. This design effectively mitigates the "identity mapping" paradox and prevents semantic leakage between frequency components. Extensive experiments on real-world and synthetic benchmarks demonstrate that FSA-AD achieves highly competitive performance by effectively disentangling frequency-specific anomaly patterns. Differential Contrastive Representation Learning with Frequency-Domain Attention for Anomaly Detection in Aerospace Cyber-Physical Systems Jiandun Li, Sha Zhou, and Shuo Zhang (School of Electronic Information Engineering, Shanghai DianJi University) Abstract Abstract Anomaly detection in aerospace cyber-physical systems (CPS) is crucial for ensuring spacecraft safety and mission reliability. However, multivariate telemetry data exhibit strong cross-variable coupling, pronounced non-stationarity, and severe environmental noise, which significantly challenge existing unsupervised detection methods. To address these issues, we propose DCRFA, a differential contrastive representation learning anomaly detection framework with frequency-domain attention. Our differential enhancement module explicitly separates high-frequency dynamics from low-frequency drifts, mitigating the influence of non-stationary trends. Meanwhile, the frequency-domain attention mechanism operates in the spectral domain, providing a global receptive field to capture periodic patterns and transient energy variations while suppressing local noise and improving computational efficiency. Furthermore, we introduce a KL-divergence-based consistency contrastive objective with a stop-gradient strategy to stabilize the learning of nominal physical representations in latent space. Extensive experiments on multiple aerospace telemetry datasets demonstrate that DCRFA consistently outperforms nine representative baseline methods, validating its effectiveness and robustness in complex aerospace anomaly detection scenarios. Modeling Diachronic Semantics in Classical Chinese via Multi-Scale Fourier Analysis Xuanning Liu and Bin Wu (Beijing University of Posts and Telecommunications) Abstract Abstract Modeling diachronic semantics in Classical Chinese is challenging due to the millennial span and the intricate balance between cultural continuity and linguistic shift. Existing methods often rely on discrete temporal labels, failing to capture the inherently continuous nature of semantic evolution. To bridge this gap, we propose the Multi-Scale Fourier Transformer (MSFT), which reformulates diachronic representation learning from a frequency-domain perspective. By decoupling textual features into modulus and phase components, MSFT utilizes the Semantic Inheritance Intensifier (SII) and Evolutionary Trajectory Gating (ETG) modules to independently model stable semantic inheritance and dynamic linguistic drift. Specifically, low-frequency signals capture long-term thematic stability, while high-frequency signals encode localized semantic innovations. A Cross-Scale Resolution Attention (CRA) mechanism further enables the model to adaptively integrate evolutionary patterns across different historical resolutions. Experiments on the Temporal CCL-WSD dataset demonstrate that MSFT consistently outperforms state-of-the-art baselines. Furthermore, qualitative analysis confirms the structural interpretability of our framework, showing that internal phase dynamics strongly correlate with external semantic distance, providing a principled methodology for quantifying the millennial metamorphosis of the Chinese language. Thursday Virtual Room 7 IJCNN Paper Visual Anomaly and Defect Detection VII Session Chair: Zewen Wang (Xinjiang University), Na Liu (Inner Mongolia University of Technology) MDFE-Net: Multi-level Degradation-aware Feature Enhancement Network Yi Liu, Na Liu, Long Yang, Xufei Zhuang, Wenhong Wu, and Guiping Liu (Inner Mongolia University of Technology) Abstract Abstract Detecting small-scale defects in complex industrial environments remains a challenging problem in computer vision, primarily due to progressive feature degradation during down-sampling, background clutter interference, and insufficient multi-scale feature fusion. To overcome these obstacles, we propose MDFE-Net, a Multi-level Degradation-aware Feature Enhancement Network designed to enhance the robustness of small-defect representations by explicitly addressing feature degradation and enabling efficient multi-scale feature interactions. MDFE-Net consists of three key components: a degradation-aware feature enhancement module to strengthen weak small-object representations, a background-aware attention mechanism to suppress background noise, and a structure-preserving multi-scale fusion strategy to alleviate scale inconsistency. Notably, MDFE-Net is detector-agnostic, allowing for seamless integration into existing one-stage detection frameworks. We conduct experiments on several challenging datasets, including two wind turbine blade defect datasets with complex industrial backgrounds, as well as additional datasets from other visual domains characterized by small targets and cluttered scenes. These experiments assess both in-domain and cross-domain generalization. The experimental results demonstrate that MDFE-Net consistently improves detection accuracy for small and densely packed targets while maintaining high computational efficiency across diverse scenarios. A Weak-Signal-Aware Framework for Subsurface Defect Detection: Mechanisms for Enhancing Low-SCR Hyperbolic Signatures Wenbo Zhang (Zhejiang Normal University, College of Engineering); Zekun Long (Griffith Univeristy); Zican Liu (Zhejiang Normal University, College of Engineering); Yangchen Zeng (southeast university, School of Cyber Science and Engineering); and Keyi Hu (Hainan Normal University, School of Artificial Intelligence) Abstract Abstract Subsurface defect detection via Ground Penetrating Radar (GPR) is challenged by "weak signals"—faint diffraction hyperbolas with low signal-to-clutter ratios (SCR), high wavefield similarity, and geometric degradation. Existing lightweight detectors prioritize efficiency over sensitivity, failing to preserve low-frequency structures or decouple heterogeneous clutter. We propose WSA-Net, a framework designed to enhance faint signatures through physical-feature reconstruction. Moving beyond simple parameter reduction, WSA-Net integrates four mechanisms: Signal preservation using partial convolutions (PConv); Clutter suppression via heterogeneous grouping attention (LWGA); Geometric reconstruction (SCConv) to sharpen hyperbolic arcs; Context anchoring (CAA) to resolve semantic ambiguities.Evaluations on the RTSTdataset show WSA-Net achieves 0.6958 mAP@0.5 and 164 FPS with only 2.412 M parameters. Results prove that signal-centric awareness in lightweight architectures effectively reduces false negatives in infrastructure inspection. Robust Tiny Insulator Defect Detection under Complex UAV Backgrounds with State Space Modeling ChuanXing Geng, Xiaojun Xue, Heng Li, and hui liu (Faculty of Information Engineering and Automation, Kunming University of Science and Technology) Abstract Abstract Insulator defect detection in UAV inspection is challenging due to tiny defect scales, complex backgrounds, multi-view variations, and long-chain structural dependencies. To address these issues, we propose PRD-DFINE, a lightweight and robust end-to-end detection framework based on DFINE. The framework introduces three key components: (1) a Progressive Injection Feature Pyramid (PI-FPN) to enhance fine-grained representation for tiny defects, (2) a Context-Driven Robustness Enhancement module (CRFE) to suppress background interference during multi-scale fusion, and (3) a Decoding-Enhanced Non-Causal Visual State Space Model (DSSM) for efficient global context modeling. Experiments on IFDD and InsPLAD show that PRD-DFINE outperforms mainstream detectors while remaining lightweight, achieving strong robustness and generalization under complex UAV scenarios. With only 9.1M parameters and 29.67 GFLOPs, it is suitable for real-world deployment. Furthermore, qualitative heatmap visualizations confirm that our method produces more compact and discriminative responses on defect regions, indicating improved resistance to background distractions.Code is available at https://github.com/GcxMaxx/PRD-DFINE. U2TSG-Net: Infrared Small Target Detection via Nested U-Structure with Target-Sensitive Attention and Gated Interaction Zewen Wang, Boyuan Li, Shuai Li, Shengbin Hao, Kurban Ubul, and Alimjan Aysa (Xinjiang University) Abstract Abstract Infrared small target detection (IRSTD) plays a pivotal role in precision guidance and surveillance systems. The submersion of weak infrared signals by complex background clutter results in extremely low signal-to-noise ratios (SNR), thereby making the accurate detection of dim and small targets a significant challenge. Existing U-shaped networks often suffer from deep feature dilution during continuous downsampling, causing faint targets to be submerged. Furthermore, direct feature fusion via skip connections often introduces background clutter from shallow layers into the decoder, leading to false alarms. To address these dual challenges, we propose U2TSG-Net, a novel nested U-structure network. Specifically, in the encoding stage, we propose a target sensitive selective attention (TSSA) module. By incorporating intensity priors with sparse attention, TSSA enforces target-sensitive feature preservation, effectively preventing small targets from being lost in deep layers. In the decoding stage, we propose a gated cross-level feature interaction (GCFI) module. It generates a spatial gate under the guidance of deep semantics, thereby selectively injecting target context into shallow features and suppressing activations caused by clutter. Extensive experiments on public benchmarks, including NUAA-SIRST, NUDT-SIRST, and IRSTD-1k, demonstrate that U2TSG-Net surpasses state-of-the-art methods, fully validating the superior robustness and practical value of the proposed architecture in complex backgrounds. Thursday Virtual Room 8 IJCNN Paper Visual Representation Learning Session Chair: Feiyu Chen (Chongqing Normal University), Xiaolin Xiao (South China Normal University) Relational Contrastive Learning For Multi-View Clustering Zhonghua Li, Zhehan Zhou, Na Tang, and Xiaolin Xiao (South China Normal University) Abstract Abstract Multi-View Clustering (MVC) has recently benefited from contrastive learning, which enhances representation discriminability by pulling positive pairs together and pushing negative pairs apart. However, existing contrastive MVC methods fail to preserve the consistency of pairwise similarities across samples. They either treat different views of the same instance as positives, overlooking intra-cluster relations, or rely on pseudo-labels that uniformly assign positives and negatives, neglecting the pairwise similarities. To address these issues, we propose a relational contrastive learning for multi-view clustering (RCMVC) framework that models the similarity distribution between an anchor and all other samples, explicitly preserving pairwise consistency. This design enables soft positive and negative assignments, capturing fine-grained similarities while maintaining coherent cluster structures. Moreover, RCMVC incorporates both instance-level and cluster-level modules to jointly exploit detailed and semantic relations. Extensive experiments demonstrate the importance of preserving the consistency of pairwise similarities for multi-view clustering. Bayesian Debiasing for Robust Contrastive Multi-view Clustering Zhehan Zhou, Zhonghua Li, Na Tang, and Xiaolin Xiao (South China Normal University) Abstract Abstract Contrastive Multi-view Clustering (MvC) learns discriminative representations by distinguishing positive and negative pairs, but its performance often suffers from false positives and negatives. Existing approaches mitigate this issue with k-means pseudo-labels, yet these deterministic labels are prone to misclassifying hard samples near cluster boundaries. To address this issue, we propose a Bayesian Debiasing framework for robust contrastive multi-view clustering (BDMvC). First, BDMvC employs a Bayesian Gaussian mixture model to generate pseudo-labels, enabling soft assignments of hard samples. Second, two mixing strategies are designed to synthesize diverse hard samples, ensuring smoother and more robust cluster boundaries. Finally, unlike prior methods, BDMvC provides the first unified treatment of both hard positives and negatives within a principled framework, substantially enhancing the robustness of contrastive learning. Experiments on benchmark datasets demonstrate that BDMvC consistently outperforms state-of-the-art methods. ISCST: Improved Supporting Clustering Based on Sentence-Transformers with Contrastive Learning Kongqiang Wang, Xuejie Zhang, Jin Wang, and Xiaobing Zhou (Yunnan University) Abstract Abstract Unsupervised clustering faces challenges when semantic categories overlap in the representation space, limiting the effectiveness of distance-based methods. We propose an Improved Supporting Clustering based on Sentence-Transformers with Contrastive Learning (ISCST), a novel framework combining bottom-up instance discrimination with top-down clustering objectives. ISCST leverages contrastive learning to simultaneously enhance categorical separation and optimize intra-cluster compactness and inter-cluster discrimination. Evaluated on multiple short text clustering benchmarks, ISCST achieves substantial improvements over state-of-the-art methods: 5\%-8\% Accuracy enhancement and 3\%-5\% Normalized Mutual Information improvement. A comprehensive quantitative analysis validates ISCST's effectiveness in addressing category overlap, producing semantically meaningful, well-separated clusters in the learned representation space. Our joint optimization approach synergizes contrastive learning with clustering objectives to achieve superior performance. Uncertainty-Aware Graph Contrastive Clustering with Coupled Dual-Path Augmentation Yanpei Xiao, Shangshang Zhao, Feng Liu, Zhongyang Zhou, and Feiyu Chen (Chongqing Normal University) Abstract Abstract Graph contrastive clustering has emerged as a promising paradigm for unsupervised graph representation learning, which leverages contrastive objectives to learn discriminative embeddings for high-quality clustering. Despite its effectiveness, most existing methods rely on predefined graph augmentation strategies and do not explicitly take the downstream clustering task into account. In practical scenarios, this inevitably introduces noise or distorts the underlying semantics. Such inherent limitations often lead to semantic drift between augmented graph views, ultimately degrading clustering performance. To address these issues, we propose a Uncertainty-Aware Graph Contrastive Clustering with Coupled Dual-Path Augmentation method, termed UGCC. The proposed method constructs multi-view contrastive learning objectives through learnable attribute augmentation and adaptive structure augmentation, enabling the generation of high-quality augmented samples and the learning of more discriminative representations for graph contrastive clustering. Furthermore, we introduce an uncertainty-aware clustering alignment objective to adaptively regulate the optimization process of unsupervised clustering. Extensive experiments conducted on five benchmark datasets demonstrate the effectiveness of UGCC. Thursday 0.01 London IJCNN Paper Smart Energy, Grid, and Infrastructure including IJCNN SS35 Computational Intelligence Techniques for Observable Smart Grid and Sustainable Energy Systems Session Chair: Giulia Tanoni (Università Politecnica delle Marche, Ancona, Italy) Surrogate-Assisted Fictitious Play for EV Charging Station Pricing Game Bo-Lin Zheng, Feng-Feng Wei, and Wei-Neng Chen (South China University of Technology) Abstract Abstract Pricing competition among Electric Vehicle Charg- ing Stations (EVCSs) under Dynamic Traffic Assignment (DTA) poses a challenging game-theoretic problem with continuous action spaces and black-box objectives. Existing methods rely on static traffic models or discretized pricing, failing to capture the coupling between pricing strategies and congestion dynamics. This paper proposes a Surrogate-Assisted Fictitious Play (SAFP) framework that integrates belief-based Fictitious Play with Multi- Agent Deep Reinforcement Learning. The learned critic network serves as a differentiable surrogate to approximate best responses, enabling gradient-based NashConv estimation for convergence monitoring. An asynchronous master-worker architecture further decouples training from simulation to mitigate DTA latency. We construct DTA-integrated testbeds by extending benchmark traffic networks with EV charging scenarios, and validate the framework’s convergence and economic interpretability with three MADRL instantiations (MADDPG, MFDDPG, IDDPG). Dynamic Aggregate Signal Reduction for Augmented Non-Intrusive Load Monitoring Giulia Tanoni, Enrik Xhani, Muhammad Affan Khan, Stefano Squartini, and Emanuele Principi (Università Politecnica delle Marche) Abstract Abstract Non-Intrusive Load Monitoring (NILM) typically relies on one-to-many or one-to-one algorithms to disaggregate energy usage of multiple appliances. Considering the wide variety of devices installed in residential and industrial buildings, we asked why not leverage the disaggregated power signals of one or more appliances to facilitate the disaggregation of devices with more complex activation profiles. In this way, both the complexity of the aggregate signal and the number of simultaneous operating appliances can be reduced. Thus, we propose a dynamic aggregate signal reduction framework for augmented NILM that takes advantage of predictions from multiple one-to-one networks during their training process to simplify the aggregate signal for other networks. This strategy enhances the reconstruction of activation profiles and reduces false positives and negatives by systematically removing appliance activations at each training step. The eligibility criterion for subtraction is based on the Signal-to-Noise Ratio estimated on each training instance. Real-world NILM applications often lack sufficient data to handle disaggregation effectively and our proposed subtraction mechanism addresses this issue by providing an augmented view of the same power signal that dynamically evolves at each epoch, depending on the activations of subtracted devices. To validate the methodology, two public residential datasets are used, with experiments designed to evaluate the reliability of the networks on unseen power signals and among several training scenarios. The framework demonstrates significant improvements, achieving up to 67.3% better performance in Mean Absolute Error and 57.4% in Normalized Error in assigned Power compared to the state-of-the-art. Two-Stage Prediction Intervals via Spiking Neural Network-Gated State-Specific Residual Models for Solar Power Generation Thomas Aguilera, Oscar Cartagena, and Alex Navas-Fonseca (University of Chile); Alfredo Núñez (Delft University of Technology); and Doris Sáez (University of Chile) Abstract Abstract A two-stage framework has been developed for reliability-aware day-ahead forecasting. This framework includes a forecasting backbone that generates multi-step point predictions. From these outputs, an aligned residual is defined at a selected horizon and used to identify significant forecast error events. This identification is accomplished through a validation-calibrated threshold that ensures the desired prevalence of these events. In the proposal, an event detector based on a surrogate-gradient spiking neural network (SNN), implemented in snnTorch, estimates the probability of entering a significant error state. To achieve that, it uses features available at the time of the forecast, as the predicted trajectory shape and various calendar descriptors, and combines them with a validation-calibrated operating threshold that converts the detector's probability into a decision regarding the state. The resulting method effectively combines SNN-based state classification with state-dependent to construct prediction intervals for asymmetric residual modeling. The experimental results demonstrate strong performance in event detection and provide practical trade-offs between coverage and width for prediction intervals, supporting the framework's suitability for reliability-aware energy forecasting in microgrids. Forward--Forward Learning for Imbalanced Tabular Predictive Maintenance on a Real-World Smart-Grid Fault Dataset Enrico De Santis, Danial Zendehdel, Gialuca Ferro, and Antonello Rizzi ("Sapienza" University of Rome) Abstract Abstract Forward-Forward learning replaces global error back-propagation with a layer-local objective based on the goodness of forward activations. This paper evaluates a stabilized, Trifecta-inspired Forward-Forward formulation for imbalanced tabular predictive maintenance on a real-world smart-grid fault dataset. A supervised label-embedding Forward-Forward network is assessed under stratified i.i.d. train/test splits, with cross-validation on the training portion and final reporting on held-out test sets. Comparative experiments against standard tabular baselines show that Forward-Forward is competitive under class imbalance and remains close to back-propagation multi-layer perceptrons, while random forests provide the strongest overall performance. Probability calibration is also examined to support risk-aware decision making, and goodness-margin dynamics together with single-factor ablations highlight the role of LayerNorm as a key stabilizer for reliable training. Timing measurements quantify training and inference cost. Simulation of Microgrid Energy Management under Battery Degradation Costs: a PPO-Based Reinforcement Learning Approach Gianluca Ferro and Alessio Orlandi ("Sapienza" University of Rome), Francesco Giuseppe Quilici and Giovanni Lutzemberger (University of Pisa), and Enrico De Santis and Antonello Rizzi ("Sapienza" University of Rome) Abstract Abstract Residential microgrids with photovoltaic generation and battery storage require energy management strategies that reduce grid costs while limiting long-term battery degradation. We present a degradation-aware simulator in which the storage system is managed by a Battery Management System module based on an equivalent circuit model (ECM) with SoH-dependent parameters. Aging is updated from experimentally identified SoH-Ah throughput curves for a commercial NMC 18650 cell, and a DoD-based wear cost is included in the objective. A PPO controller is trained to select charge/discharge setpoints and is tested on real household data against a rule-based controller and an oracle MPC benchmark. The best policies consistently outperform the rule-based baseline and approach oracle MPC performance at medium battery utilization. Twofold Cross-Validation Is Better than Five- or Tenfold Shigeo Abe (Kobe University) Abstract Abstract Model selection by k-fold cross-validation is a very useful way of realizing a support vector machine with a high generalization ability. The number of folds is usually five or ten. And thus, cross-validation is time consuming especially for large-size problems. In this paper first we show that under some assumption, for any classifier the cross-validation accuracies increase as the number of folds increases. (This does not imply that the larger number of folds realizes the better generalization ability.) Then, we propose using twofold cross-validation instead of five- or tenfold. We discuss the advantage of twofold cross-validation from the speed-up, and adjustment of parameter values when underfitting occurs. By computer experiments using two-class and multiclass problems, we demonstrate that the generalization abilities by twofold cross-validation are statistically comparable to five- or tenfold cross-validation with a considerable speed-up. Thursday 0.02 Berlin IJCNN Paper, FUZZ-IEEE Position Paper, CEC Late Breaking Paper, CEC Paper, FUZZ J2C Presentation, CEC J2C Presentation, FUZZ-IEEE Paper, CEC Position Paper, IJCNN J2C Presentation, IJCNN Position Paper, IJCNN Late Breaking Paper, FUZZ-IEEE Late Breaking Paper Student Best Paper Award On the Structural (Dis)Agreement of Landscape Representations in Black-Box Optimization Sara Gjorgjieva, Eva Tuba, and Barbara Koroušić Seljak (Jožef Stefan Institute); Carola Doerr (Computer Science department LIP6, Sorbonne Université, CNRS); and Tome Eftimov (Jožef Stefan Institute) Abstract Abstract Landscape feature representations play a central role in automated algorithm selection and meta-learning for black-box optimization, yet little is known about how different representations agree (or disagree) in the structures they impose on problem spaces. This paper presents a systematic unsupervised evaluation of four state-of-the-art representations (ELA, DeepELA, TransOptAS, and DoE2Vec) using a diverse set of affine combinations of BBOB functions (MA-BBOB). By applying extensive clustering analyses, coverage-based stability measures, and cross-representation similarity assessments, we show that each representation organizes the same problems in markedly different ways: ELA and TransOptAS form compact geometric structures, DeepELA provides a balanced intermediate view, and DoE2Vec achieves strong semantic alignment but with substantial fragmentation. Our results reveal that no single representation dominates; rather, they capture complementary aspects of the underlying landscapes. These findings highlight the importance of multi-view analyses for understanding representation behavior and offer guidance on selecting or combining representations in downstream meta-learning and algorithm selection tasks. In addition, across two different algorithm families (Differential Evolution and Particle Swarm Optimization), we show that landscape representations face an inherent trade-off in how well they align structural landscape descriptions with observed performance, indicating that no single representation can fully capture algorithm performance. LLM-to-Phy3D: Physically Conform Online 3D Object Generation with LLMs Melvin Wong (NTU); Yueming Lyu (Centre for Frontier AI Research, Agency for Science, Technology and Research); Thiago Rios and Stefan Menzel (Honda Research Institute Europe); and Yew Soon Ong (NTU) Abstract Abstract The emergence of generative artificial intelligence (GenAI) and large language models (LLMs) has revolutionized the landscape of digital content creation in different modalities. However, its potential use in Physical AI for engineering design, where producing physically viable artifacts is paramount, remains largely underexplored. The lack of physical knowledge in existing LLM-to-3D models often results in outputs that are detached from real-world physical constraints. To address this gap, we introduce LLM-to-Phy3D, an online 3D object generation method that enables existing LLM-to-3D models to produce physically conforming 3D objects on the fly. LLM-to-Phy3D introduces a novel online black-box refinement loop that empowers large language models (LLMs) through synergistic visual and physics-based evaluations. By delivering directional feedback in an iterative refinement process, LLM-to-Phy3D actively drives the discovery of prompts that yield 3D artifacts with enhanced physical performance and geometric novelty bias relative to user reference objects, marking a substantial contribution to AI-driven generative design. Systematic evaluations of LLM-to-Phy3D, supported by ablation studies in vehicle design optimization, reveal various LLM improvements gained over 78% in producing physically conform target domain 3D designs over conventional LLM-to-3D models. The encouraging results suggest the potential general use of LLM-to-Phy3D in Physical AI for scientific and engineering applications. Compare Similarities Between DNA Sequences Using Permutation-Invariant Quantum Kernel Chenyu Shi (Leiden University), Gabriele Leoni (European Commission Joint Research Centre), Mauro Petrillo (Seidor), Antonio Puertas Gallardo (European Commission Joint Research Centre), and Hao Wang (Leiden University) Abstract Abstract Computing the similarity between two DNA sequences is of vital importance in bioscience, yet it can be computationally expensive on classical hardware. For example, the edit distance with move operations (EDM), a DNA similarity measure of interest in biology, is proven to be NP-Complete to compute exactly on classical hardware. Recently, applied quantum algorithms have been anticipated to offer potential advantages over classical approaches. In this paper, we propose a novel variational quantum kernel model served as a surrogate model for estimating similarity between DNA sequences defined by EDM. Since the EDM metric exhibits a pairwise permutation-insensitive property, we incorporate a permutation-invariant structure into the variational quantum kernel to approximate this symmetry. Furthermore, to encode the four nucleotide bases as quantum states, we introduce a theoretically motivated encoding scheme based on symmetric informationally complete positive operator-valued measure (SIC-POVM) states. This encoding ensures mutual equivalence among bases, as each pair of symbols is mapped to quantum states that are equidistant on the Bloch sphere. We experimentally show that, equipped with the permutation-invariant circuit design and mutual-equivalence encoding, the proposed quantum kernel model achieves strong performance in approximating the similarity defined by EDM. Compared with classical kernel learning methods, our quantum approach achieves significantly higher accuracy while using substantially fewer trainable parameters. A Geodesic‑Aware Quaternion Neural Network for Orientation Estimation Oumar Butt and Alin Tisan (Royal Holloway, University of London); Danilo Mandic (Imperial College London); and Clive Cheong Took (Royal Holloway, University of London) Abstract Abstract Quaternion-valued neural networks are the natural candidates for orientation estimation due to their numerous advantages in modelling rotations. However, existing quaternion algorithms ignore the curved geometry of the hypersphere S³. Perhaps, this is because most quaternion neural networks have been designed for general multi-dimensional processing, making the implicit assumption of Euclidean space. To address this shortcoming, this paper proposes a quaternion neural network with a geodesic loss that accounts for the curved structure in S³. More importantly, our loss formulation explicitly accounts for the ambiguous equivalence q ≡ −q in quaternion rotation. Simulation studies on both synthetic and the real-world aerial vehicle dataset demonstrate an order-of-magnitude reduction in mean angular tracking error compared to Euclidean methods. Explainable Object Detection in 360° Images Through Fuzzy Logic Systems and Vision-Language Models Integration Amer Harfoush and Hani Hagras (University of Essex) and Hugo Leon-Garza and Anasol Pena-Rios (British Telecom) Abstract Abstract Panoramic images (360-degree) contain significant amount of information since they manage to take a snapshot of the complete surroundings at once which makes them a great choice in multiple fields. However, they pose special challenges for object detection due to significant geometric distortions coming from equirectangular projection. Current methods are inadequate for critical applications since they do not have the needed accuracy and provide no explanations for their decisions. This paper presents a novel framework that integrates Fuzzy Logic Systems (FLSs), an object detection model, and Vision-Language Models (VLMs) to achieve both good detection performance and explainability. Our method uses a fine-tuned YOLOv9 object detector enhanced by a VLM-based validation engine. The fuzzy logic system module evaluates image quality and adapts processing parameters, providing linguistic explanations that bridge numerical processing with human understanding. Testing on industrial datasets from British Telecom demonstrates robust performance on previously unseen object categories, validating real-world applicability. The proposed framework achieves up to 124.67% improvement over previous state-of-the-art methods while providing explanations for each detection decision Interval-Valued Epistemic Reasoning for Hallucination-Aware Language Models Taniya Seth and Pranab K. Muhuri (South Asian University) Abstract Abstract LLMs tend to generate fluent but epistemically dubious responses, which can be observed as hallucinations, because of the lack of a principled representation of truth and uncertainty. This paper presents a novel type-2 fuzzy epistemic control framework for modelling truth in language generation as an interval-valued variable, representing uncertainty about uncertainty, rather than confidence values. The internal model signals of token entropy and logit margin are transformed using an interpretable fuzzy inference system to obtain the bounds of epistemic truth without any retraining or parameter modification of the model. Experiments on the TruthfulQA dataset with various transformer-based models show that the proposed method conservatively controls epistemic trust, distinguishes model reliability, and resists false confidence with weak evidence, outperforming entropy-based and scalar uncertainty methods. The proposed method thus validates type-2 fuzzy logic as a sound and interpretable basis for hallucination-robust and truth-sensitive language modelling. Thursday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 19 - Related Topics IV Session Chair: Juan J. (University of Granada) Evolutionary Refinement of Generative Graph Topologies: A Hybrid WGAN-GA Approach James Sargant, Seyedeh Ava Razi Razavi, Renata Dividino, and Sheridan Houghten (Brock University) Abstract Abstract Generating realistic graph-structured data is challenging due to discrete connectivity, varying graph sizes, and class-specific structural patterns. Recent Generative Adversarial Networks (GAN)-based graph generation methods improve edge modeling by learning connectivity and matching class-specific density distributions. However these models still exhibit noticeable deviations such as in degree and spectral distribution when compared to real graphs, indicating that important structural properties are not fully preserved. This work aims to reduce these deviations by refining the graphs produced by an existing GAN-based graph generator framework with a Genetic Algorithm (GA). In the GAN framework, the generator produces both node features and connectivity patterns, while a GNN-based critic evaluates graph realism and class consistency to ensure global structural and class alignment. Building on this foundation, we apply a GA to refine the edges of generated graphs. The refinement process guides synthetic graphs toward closer agreement with real data, while preserving diversity and novelty. Experimental results show that the GA refinement consistently lowers combined Maximum Mean Discrepancy (MMD) compared to the base model, leading to graphs that more closely match real structural patterns. This demonstrates that evolutionary refinement is an effective and flexible way to correct residual structural deviations in GAN-based graph generators, improving their suitability for realistic graph synthesis and data augmentation. From Configuration to Evolution: Corporate Social Responsibility Practices and Supply Chain Effects on Environmental Innovation Huzhi Xue (Beihang University, Beijing Institute of Mathematical Sciences and Applications) and Haihua Xie (Beijing Institute of Mathematical Sciences and Applications) Abstract Abstract This paper employs the stakeholder theory and practice-based view to examine how Corporate Social Responsibility (CSR) practices combine to drive environmental innovation in manufacturing firms under different supply chain network contexts. Leveraging data from Refinitiv and FactSet, this study uses fuzzy-set qualitative comparative analysis (fsQCA) to identify three equifinal CSR configurations associated with high environmental innovation. Firms positioned with high centrality and few structural holes in their supply chains display distinctive CSR patterns. Extending this analysis, an N-K fitness landscape model simulates the dynamic evolutionary paths of CSR implementation, revealing that the order in which CSR practices are adopted critically affects environmental innovation outcomes. Shareholder-oriented practices lay the foundation. They are followed by environmental and workforce-oriented practices in sequence. Customer-oriented practices build on these prior implementations, whereas supplier-oriented practices exert influence only after the preceding practices are established. By combining configurational analysis with an evolutionary perspective, this study clarifies both the joint effects of CSR practices and the role of implementation sequence in shaping environmental innovation outcomes. Multi-objective Optimisation of Traffic Light Control for Fast and Safe Traffic Incident Recovery Qihong Yang, Zhuowei Zhao, Dong Zhao, Kai Qin, and Yu Sun (Swinburne University of Technology) Abstract Abstract In post-incident traffic management, traffic light control is one of the most widely adopted approaches. Although many existing studies can optimise traffic signal control patterns using multi-objective evolutionary optimisation techniques, the commonly used objective functions are not specifically designed for post-incident management and thus cannot effectively facilitate efficient traffic recovery. In this work, we first formulate traffic recovery as traffic conditions returning to normal within a predefined road area, and then propose two recovery-oriented objectives: Relaxed Traffic Recovery Time, which aims to minimise the time required for traffic to return to normal, and Overall Recovery Count, which seeks to improve recovery stability at the network level. The proposed recovery-oriented objectives, together with a safety-related objective (speed variance), are incorporated into NSGA-II and evaluated via traffic simulation. This results in a bi-objective optimisation approach for deriving traffic light control patterns that enable fast and safe post-incident traffic recovery. Experiments on both synthetic and real-world road networks demonstrate that, compared to traditional congestion-based objectives, the proposed recovery-oriented objectives lead to more effective traffic light control for post-incident traffic recovery. Can we measure energy consumption in population-based metaheuristics? Juan J. (University of Granada) and Cecilia MereloMolina (Zenzorrito) Abstract Abstract As we proceed with incorporating energy efficiency into the design of our algorithms, establishing a robust methodology for energy profiling and the eventual comparison of algorithms or implementations is essential for a solid foundation for any further work. The main issue to overcome is the lack of specific per-process measurement of energy consumption, and above that, the fact that we are going to be measuring a system that is actively trying to optimize energy consumption itself, and doing so for the whole system, above and beyond the workload we are trying to measure. In this paper, we propose a two-stage methodology for energy profiling in population-based metaheuristics. The first stage will consist of an experimental design, i.e., how to run a series of experiments to measure the energy consumption of a specific configuration and implementation of the algorithm. The second phase will consist of the statistical processing of these results to obtain a measure of energy consumption with a certain degree of certainty. Finally, we will apply these measurements to a type of evolutionary algorithm, called the Brave New Algorithm, written in the Julia language, to determine the impact of changes at different levels on energy consumption. Bioinspired strategy for coordinating multiple unmanned aerial vehicles for surveillance tasks Sara Saori Satake and Guilherme Henrique de souza nakahata (University of Tsukuba) and Rodrigo Calvo (State University of Maringa) Abstract Abstract Among the various unmanned autonomous systems (UAS), unmanned ground vehicles (UGVs) and unmanned aerial vehicles (UAVs) stand out due to the great possibility of applications. The latter, despite being used in the military environment, currently, research involving their use in everyday situations is growing and gaining popularity in various purposes such as surveillance, agriculture, traffic, among others. Bioinspired algorithms exert a strong influence on the development of navigation strategies, in particular the ant colony optimization (ACO), which has variations with satisfactory results for the coordination of multiple robots, as is the case of the Inverse Ant System-Based Surveillance System (IAS-SS). This approach has already been applied to UGVs, presenting good results when compared to traditional ACO. The present work aims to develop an autonomous navigation strategy for multiple UAVs capable of performing surveillance tasks. The experiments were divided into two strands: the basic ones, where there are variations in the number of different heights for each agent; and the adaptive ones, which cover the addition and removal of an agent during the simulation. Both are performed by virtual simulators, and statistical metrics were used for performance analysis. The computational results show that the repulsive characteristic of the pheromone manages to readapt the agents to the environment, as observed when an agent is added and the number of cycles generated is notorious when associated with other experiments. Decision-Making Policies under Stigmergic Communication: A Pheromone-Based Multi-Agent Study Sara Saori Satake, Guilherme Henrique de Souza Nakahata, and Claus Aranha (University of Tsukuba) Abstract Abstract We address the challenge of investigating how different decision-making policies operate under stigmergic communication mediated by pheromones in multi-agent systems, using early hominin behavior as a motivating case study. We propose an agent-based model that incorporates a computational abstraction of pheromones, inspired by Ant Colony Optimization (ACO), enabling agents to coordinate indirectly through attraction to resources and avoidance of predators. Our model extends the HOMINIDS framework, originally focused on tool use, by integrating sensory-driven reactive behaviors and contrasting deterministic, weighted-random, and heuristic decision policies. Experimental results reveal trade-offs between risk exposure and food acquisition that depend on both the decision policy and pheromone persistence: while pheromone-guided strategies reduce predator encounters, they may also limit caloric intake compared to unguided exploration. These findings highlight how the interaction between decision-making policies and stigmergic communication shapes emergent collective behavior, offering insights into coordination, robustness, and adaptation in multi-agent systems operating under different time-related environmental signals. Thursday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 20 - Algorithms V Session Chair: Diego Oliva (Universidad de Guadalajara) Automated Algorithm Design of Tailored Metaheuristics for Photovoltaic Parameter Estimation in Single- and Double-Diode Models Daniel Fernando Zambrano Gutierrez (Tecnologico de monterrey), Grecia C. Duque-Gimenez (Universidad Autónoma de Nuevo León), José Carlos Ortiz-Bayliss (Tecnologico de monterrey), Itzel Aranguren and Diego Oliva (Universidad de Guadalajara), and Jorge M. Cruz-Duarte (Centre Inria de l'Université de Lille) Abstract Abstract Automated Algorithm Design and Configuration (AADC) reduces the manual effort required to select and tune population-based optimizers for engineering tasks. We apply AADC to photovoltaic parameter estimation under the Single-Diode (SD) and Double-Diode (DD) models via a two-stage pipeline that comprises a tailored metaheuristic and a subsequent refinement of its hyperparameters. In the design stage, a sampling-based hyper-heuristic with local search builds MHs by selecting and ordering search operators from a predefined heuristic space, and then tunes the selected structure using Optuna. Experiments on the R.T.C. France cell dataset show that DD renders a more stable design outcome than SD, while both models share the same best structure, a particle-swarm-based operator followed by a differential mutation. After fixing this structure, further tuning improves performance by lowering the median objective value and reducing dispersion across runs. The final SD and DD fits closely match the measured I–V and P–V characteristics, yielding physically plausible parameter estimates. In-Loop Mechanical Screening for Multiphysics Multi-Objective Electric Motor Optimization Federico Valpiani, Alessandro Niccolai, and Sonia Leva (Politecnico di Milano) Abstract Abstract Multi-objective optimization of electric motors based solely on electromagnetic criteria neglects structural constraints, yielding Pareto-optimal solutions that may be mechanically infeasible. This paper presents a multiphysics, multi-objective framework for IPMSM rotor design, where electromagnetic and structural analyses are coupled within the same evolutionary loop to enforce mechanical admissibility. The rotor geometry, defined by a 52-parameter asymmetric double-layer V-shaped configuration, is optimized to maximize torque and minimize torque ripple at two operating speeds. Electromagnetic performance is evaluated via finite elements, while structural feasibility is assessed in-loop under overspeed through a stress-based criterion. Results show that neglecting mechanical constraints distorts the Pareto front: 32\% of solutions from electromagnetic-only optimization are infeasible a posteriori, indicating that post-processing is insufficient. The admissible front instead comprises physically realizable designs with improved performance near structural limits. Bi-Scale Evaluation Particle Swarm Optimization for Neuron-Level Enhanced Backdoor Attacks Huan-Yu Chen, Guanxiong Ha, Chunfu Jia, Jun Zhang, and Zhi-Hui Zhan (Nankai University) Abstract Abstract Existing backdoor attacks usually suffer from poor generalization across different neural network architectures and datasets. Enhancing the generalization capability of existing attack models is termed as enhanced backdoor attacks problem (EBAP). The essence of EBAP is to precisely select a subset of key neurons from the attack model and perturb their parameters, so as to enhance the performance of the attack model. To efficiently solve the EBAP, this paper models the EBAP as a 0-1 integer programming problem for fine-tuning neural-level parameters, and proposes a bi-scale evaluation particle swarm optimization (BSEPSO) algorithm based on a bi-scale evaluation (BSE) mechanism. As evaluating the fitness value of an EBAP solution is expensive, the BSE mechanism adopts a fast evaluation and a full evaluation to work cooperatively to reduce the computational cost. This way, the BSEPSO algorithm achieves efficient optimization in high-dimensional neural spaces while effectively balancing evaluation accuracy and computational overhead. Experimental results demonstrate that across extensive testing on 2 datasets and 15 mainstream attack models, BSEPSO significantly enhances attack success rate (ASR) while strictly controlling clean accuracy (CA) fluctuations within safety thresholds, validating its effectiveness as a universal attack enhancement framework. A Multiobjective Evolutionary Feature Selection Framework for End-Point Molten Steel Temperature Prediction in Ladle Furnace Refining Yi Yang, Chang Liu, Lixin Tang, and Te Xu (Northeastern University) Abstract Abstract Accurate prediction of the end-point temperature of molten steel in the ladle furnace (LF) is crucial for ensuring product quality and stable production, yet the complex high-temperature physical and chemical reactions lead to high-dimensional, strongly nonlinear, and highly coupled variables that make data-driven modeling challenging. In this paper, a multiobjective evolutionary feature selection (MOFS) framework is proposed for end-point molten steel temperature prediction in LF refining, which uses a unified multiobjective evolutionary optimization framework to jointly optimize feature selection and model hyperparameters, balancing prediction accuracy, feature subset size, and model complexity. Experiments on an industrial dataset demonstrate superior performance with fewer features and lower complexity, confirming its effectiveness for practical applications. Towards Standardized Evaluation of Feasible Region Identification in Constrained Engineering Design Ioana Nikova (Siemens Digital Industries Software, Ghent University); Yashesh Dhebar (Siemens Digital Industries Software); and Sebastian Rojas Gonzalez, Tom Dhaene, and Ivo Couckuyt (Ghent University) Abstract Abstract Many benchmarks exist for constrained optimization, which is crucial when it comes to comparing constrained optimization algorithms. However, in the field of Feasible Region Identification (FRI), where we are solely interested in discovering and sampling in feasible regions, no standardized way of benchmarking exists. Therefore, we propose a set of constrained problems that covers feasibility properties like number, location, shape and size of feasible regions, including presence of convex or non-convex regions. Another challenge in FRI is the lack of a standard way to measure FRI performance. In case of model-based FRI methods, the model's performance, e.g., accuracy, is often used to gauge the algorithm's performance. However, we show that this overlooks the fundamental aspect in real world use cases where engineers are interested in both: a good model and a diverse set of feasible designs that have been returned by the algorithm. Therefore, we introduce a new performance metric based on Inverted Generational Distance (IGD) that measures the diversity of the acquired feasible points. By providing a collection of benchmark functions and metrics, we aim to contribute towards a more standardized approach for FRI benchmarking. To demonstrate the benchmark collection, we compare several state-of-the-art data-efficient and model-based FRI approaches to show their performance on the proposed benchmarks. FedEvo: Evolutionary Population-Based Federated Learning Minji Park, Seunghyun Yoon, and Hyuk Lim (Korea Institute of Energy Technology) Abstract Abstract Federated Learning (FL) enables collaborative model training across distributed clients without sharing raw data. In conventional FL, client updates are aggregated to construct a single global model for all participants. However, because data distributions often differ across clients, a single global model may not be optimal for every client. In this paper, we propose FedEvo, a population-based FL framework that maintains multiple candidate models on the server while preserving the standard uplink communication pattern from clients. At each round, participating clients evaluate the broadcast candidates on their local validation split, select one candidate, train it locally, and upload the resulting model update. The server then updates the population through candidate-wise aggregation, anchor-based population management, usage-based recombination, and diversity-enhancing mutation. On CIFAR-10 with Dirichlet-partitioned clients, FedEvo achieves higher accuracy than conventional FL under matched communication rounds and local optimization steps. Thursday 0.10 Sydney IJCNN Position Paper Reliable, Robust, and Adaptive AI Systems (Position Track) Session Chair: Annabel Latham (Manchester Metropolitan University), Akira Hirose (The University of Tokyo) Position Paper: Minimizing Risks in Artificial Intelligence Projects Fernando Pereira dos Santos, Bruno Miguel de Souza, and Maurício Schiezaro (Venturus) Abstract Abstract To successfully develop an Artificial Intelligence (AI) project, a deep knowledge of learning-based algorithms is essential. However, understanding the concepts of innovation, scientific methodology, and project management can be key to a smooth design. The risks in AI projects are numerous, ranging from data availability, tight budgets, and deadlines to a lack of awareness of how AI actually works among stakeholders. These setbacks are likely to jeopardize progress and may even lead to project dissolution. Therefore, accurate and early management of these elements can increase the likelihood of success. In accordance, this position paper lists the main liabilities in AI development and discusses ways to anticipate and minimize them. Our goal is to enhance the discussion on how we can effectively manage AI projects, highlighting issues that might go unnoticed during the entire project. By planning for these risks, AI projects are more likely to be successfully deployed. Our paper presents vital concepts of learning-based algorithms, as well as project management, innovation, and scientific methodology. Finally, we discuss the major risks found in the literature and the possible solutions for fluid development. Position Paper: Post-Solve Robustness in Decision Engines: Feasible Regions and Smoothness Under Perturbations Yi-Xiang Hu (University of Science and Technology of China) Abstract Abstract Mixed-Integer Linear Programming (MILP) decision engines routinely output nominally optimal plans for high-stakes industrial systems. Yet deployment rarely matches solve-time assumptions: small perturbations in costs, demands, or resource availability can invalidate feasibility or trigger discontinuous shifts to qualitatively different solutions. We argue that this post-solve robustness gap is a missing layer in today's optimization pipelines and a missing evaluation dimension for learning-enabled decision systems. Rather than replacing robust optimization or stochastic programming, the proposed layer audits a solved incumbent and returns solver-backed evidence about how far that solution can be trusted. We formalize two central objects: (i) an $\epsilon$-near-optimal feasible neighborhood in parameter space, capturing when an incumbent remains feasible and near-optimal under perturbations, and (ii) solution smoothness in decision space, capturing whether nearby alternatives with small combinatorial edits remain competitive. We then synthesize the most relevant partial answers from sensitivity and stability analysis, robust optimization, neighborhood search, adversarial testing, and learning-based enhancements, and articulate an agenda for a unified post-solve robustness layer. Concretely, we call for certified inner approximations around the incumbent, probabilistic robustness estimation with calibrated uncertainty, adversarial robustness margins, and learning-based prediction and explanation aligned with solver-backed verification. We conclude with a compact reporting template and evaluation protocol that would make robustness a first-class output of decision engines. Position Paper: Uncertainty Quantification in Deep Learning Is Unsatisfactory for Clinical Applications and Complex Decision Making Ciaran Bench (National Physical Laboratory) Abstract Abstract Deep learning has significant potential to enhance medical services, but low tolerance for error and the risk of poor performance on unseen data has slowed the widespread adoption of models in practical settings. In principle, uncertainty quantification (UQ) may be used to evaluate the trustworthiness of predictions, facilitating the effective use of models in medical applications. UQ techniques in deep learning aim to reliably express the doubt in a measurement/prediction. However, common UQ techniques and evaluation metrics/measures in deep learning only consider uncertainty reliability as viewed from relatively simple measurement frameworks where contextual factors relevant to complex medical decision making can not be easily integrated. But even for cases where these factors may be quantified and considered, common methods do not achieve/assess all aspects of reliability that are relevant to clinical applications. We describe these shortcomings, and propose research priorities to help improve the effectiveness of UQ for medical applications, and realise the positive impact deep learning could have on patient outcomes. Position Paper: Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods Shaina Raza (Vector Institute); Rizwan Qureshi (University of Central Florida); Azib Farooq (University of Cincinnati, Vector Institute); Marcelo Lotif (Vector Institute); Aman Chadha (Independent Researcher); Deval Pandya (Vector Institute); and Christos Emmanouilidis (University of Groningen) Abstract Abstract Large language models (LLMs) reproduce misinformation by learning the linguistic patterns that make falsehoods persuasive, such as hedging, false presuppositions, and citation fabrication, rather than merely memorizing false facts. We propose model immunization: supervised fine-tuning on curated (false claim, correction) pairs injected as small "vaccine doses" (5–10\% of tokens) alongside truthful data. Unlike post-hoc filtering or preference-based alignment, immunization provides direct negative supervision on labeled falsehoods. Across four open-weight model families, immunization improves TruthfulQA accuracy by 12 points and misinformation rejection by 30 points with negligible capability loss. We outline design requirements, which includes, dosage, labeling, quarantine, diversity and call for standardized vaccine corpora and benchmarks that test generalization, making immunization a routine component of responsible LLM development.The project webpage is available at https://github.com/shainarazavi/ai-vaccine. Position Paper: From Edge AI to Adaptive Edge AI Fabrizio Pittorino and Manuel Roveri (Politecnico di Milano) Abstract Abstract Edge AI is often framed as model compression and deployment under tight constraints. We argue a stronger operational thesis: \emph{Edge AI in realistic deployments is necessarily adaptive}. In long-horizon operation, a fixed (non-adaptive) configuration faces a fundamental failure mode: as data and operating conditions evolve and change in time, it must either (i)~violate time-varying budgets (latency/energy/thermal/connectivity/privacy) or (ii)~lose predictive reliability (accuracy and, critically, calibration), with risk concentrating in transient regimes and rare time intervals rather than in average performance. If a deployed system cannot \emph{reconfigure} its computation - and, when required, its model state - under evolving conditions and constraints, it reduces to static embedded inference and cannot provide sustained utility. This position paper introduces a minimal Agent-System-Environment (ASE) lens that makes adaptivity precise at the edge by specifying (i)~what changes, (ii)~what is observed, (iii)~what can be reconfigured, and (iv) which constraints must remain satisfied over time. Building on this framing, we formulate ten research challenges for the next decade, spanning theoretical guarantees for evolving systems, dynamic architectures and hybrid transitions between data-driven and model-based components, fault/anomaly-driven targeted updates, System-1/System-2 decompositions (anytime intelligence), modularity, validation under scarce labels, and evaluation protocols that quantify lifecycle efficiency and recovery/stability under drift and interventions. Position Paper: The Significance of Class Imbalance in Online Semi-Supervised Data Stream Learning Hadi Talal Jaafar Al-Kadhimi (University of Birmingham, Middle Technical University) and Leandro L. Minku (University of Birmingham) Abstract Abstract An important challenge in machine learning is learning from data streams in changing environments, especially when the data is imbalanced and only a few labels are available. A useful approach is semi-supervised online learning, which can take advantage of many unlabelled examples to help the model adapt over time. However, standard methods often have difficulty when imbalance occurs, since the majority class tends to dominate and the minority class is often ignored. In this paper, we show that the challenge of class imbalance is exacerbated by lack of labels in semi-supervised data stream learning, such that tackling class imbalance is an even more significant issue in this context than in supervised data stream learning. However, very little has been done in this area so far. Our study alerts the community of the need for novel semi-supervised strategies that can benefit from unlabelled examples to cope with class imbalance in streaming scenarios, rather than propagating the bias towards the majority class through these examples. Thursday 0.11 Cape Town IJCNN Paper IJCNN SS07 Quantum Machine Learning Algorithms and Applications I Session Chair: Samuel Yen-Chi Chen (Wells Fargo) Compact Feedback-Enhanced Quantum Fast Weight Programmer for Time Series Forecasting Francesco Sanetti and Andrea Ceschini (Sapienza University of Rome); Andrew Hornback (Georgia Institute of Technology,; Wallace H. Coulter School of Biomedical Engineering); Samuel Yen-Chi Chen (Wells Fargo); and Antonello Rosato and Massimo Panella (Sapienza University of Rome) Abstract Abstract Quantum recurrent models are principled for temporal learning, yet their reliance on backpropagation through time and deep unrolling is poorly matched to Noisy Intermediate-Scale Quantum constraints. Quantum Fast Weight Programming alleviates this depth bottleneck via shallow, online parameter adaptation, but remains sensitive to noise-induced variance and gradient degradation, especially when the programmer has limited feedback about the current fast-weight state. This work tries to overcome these challenges by introducing the Compact Feedback-Enhanced Quantum Fast Weight Programmer, a hybrid architecture that strengthens quantum fast-weight programming with an explicit state-aware controller while preserving shallow, hardware-efficient quantum execution. To maintain parameter efficiency, the controller generates rank-1 parameter increments via the outer product of layer-wise and qubit-wise modulation vectors, avoiding full-matrix generation while supporting expressive adaptation. The proposed approach is assessed on three real-world forecasting benchmarks spanning disparate fields and dynamics, consistently improving predictive accuracy over both quantum fast weight programming approaches and quantum deep recurrent neural networks, while requiring also substantially less training time than quantum baselines. Multi-Split Quantum–Classical Normalizing Flows via Analytically Invertible Quantum Transforms Achal jain, Abhijith Nair, and Dinesh Singh (Indian Institute of Technology Mandi) Abstract Abstract Normalizing flows are a powerful class of generative models that offer exact likelihood estimation and tractable sampling via invertible transformations. Simultaneously, quantum machine learning (QML) promises to enhance expressivity through high-dimensional Hilbert space mappings. However, integrating quantum neural networks (QNNs) into normalizing flows presents a fundamental theoretical conflict, the Measurement Paradox. Standard quantum measurement is a non-invertible operation, breaking the strict diffeomorphic constraints required by flows. Existing literature relies on approximate reconstruction (e.g., quantum variational autoencoders), which sacrifices the exact log-likelihood guarantee that defines normalizing flows. In this work, we propose the multi-split quantum-classical normalizing flow (MSQCNF), a hybrid architecture designed to mitigate the effects of this paradox. We introduce two novel contributions: 1) The Cosine Chain, a mathematical formulation that directly addresses the measurement paradox by constructing quantum ansatzes whose post-measurement mappings remain analytically invertible with a tractable Jacobian, thereby recovering analytic invertibility under ideal expectation-value assumptions. 2) A multi-split coupling architecture that introduces auxiliary assistant pathways to stabilize gradient flow and mitigate the barren plateaus common in hybrid training. Theoretical analysis reveals that deep quantum circuits in this context converge to truncated fourier series, offering a pathway for quantum-inspired classical transforms. Experimental validation on tabular benchmarks and image generation tasks demonstrates that optimized multi-split configurations achieve performance competitive with classical baselines on image density modeling, while remaining less effective on tabular benchmarks—highlighting both the promise and current limitations of the proposed transform. Federated Learning with Quantum Enhanced LSTM for Applications in High Energy Physics Abhishek Sawaika (The University of Melbourne), Durga Pritam Suggisetti (BITS Pilani Dubai), and Rajkumar Buyya and Udaya Parampalli (The University of Melbourne) Abstract Abstract Learning with large-scale datasets and information-critical applications, such as in High Energy Physics (HEP), demands highly complex, large-scale models that are both robust and accurate. Even though the computing capacity of current supercomputers is increasing day by day, there is a huge cost associated with the energy consumption of such systems. To tackle this issue and cater to the learning requirements, we envision using a federated learning framework with a quantum-enhanced model. Specifically, we design a hybrid quantum-classical long-shot-term-memory model (QLSTM) for local training at distributed nodes. It combines the representative power of quantum models in understanding complex relationships within the feature space, and an LSTM-based model to learn necessary correlations across data points. Given the computing limitations and unprecedented cost of current stand-alone noisy-intermediate quantum (NISQ) devices, we propose to use a federated learning setup, where the learning load can be distributed to local servers as per design and data availability. We demonstrate the benefits of such a design on a classification task for the Supersymmetry(SUSY) dataset, having 5M rows. Our experiments indicate that the performance of this design is not only better that some of the existing work using variational quantum circuit (VQC) based quantum machine learning (QML) techniques, but is also comparable (Delta ~ +- 1%) to that of classical deep-learning benchmarks. An important observation from this study is that the designed framework has <300 parameters and only needs 20K data points to give a comparable performance. Which also turns out to be a 100X improvement than the compared baseline models. This shows an improved learning capability of the proposed framework with minimal data and resource requirements, due to the joint model with an LSTM based architecture and a quantum enhanced VQC. Comparing Simulated versus Hardware Training Data for Learning-Based Quantum Error Mitigation Seyed Mohamad Ali Tousi and G. N. DeSouza (University of Missouri, Vision-Guided and Intelligent Robotics Lab (ViGIR)) Abstract Abstract Learning-based Quantum Error Mitigation (QEM) has emerged as a promising approach for reducing noise-induced bias on near-term Noisy Intermediate-Scale Quantum (NISQ) hardware. While simulated noise models enable scalable training of learning-based mitigators, real hardware data more accurately reflect device-specific error mechanisms, but are expensive and sparse. Understanding how these two data regimes interact is therefore critical for the practical deployment of learning-based QEM. In this paper, we present a systematic and empirical analysis on the nature of the data (simulated versus hardware noise) for learning-based QEM. Using a recently developed attention-based graph transformer for small- and large-scale quantum error mitigation, we compare various training regimes, starting from exclusively simulated noise and progressively integrating more measurements from real quantum hardware. Our study spans two distinct circuit families: 1) randomly generated circuits with relatively homogeneous structure; and 2) targeted quantum circuits derived from a specific problem, in this case the one-dimensional Burgers’ equation. Across five-fold cross validation, we analyze how hardware data affects generalization performance as measured by mean absolute error (MAE). Our results demonstrate that the impact of hardware data strongly depends on circuit structure and diversity. For random circuits, simulated data already guarantees lower absolute error, while hardware data yields consistent, but moderate improvements. In contrast, for targeted circuits, like the structured Burgers’ circuits, simulated noise can be less representative of the actual (testing) quantum error, and incorporating hardware data produces substantially larger relative reductions in MAE -- despite the intrinsically higher MAE. These findings indicate the value of the nature of the training data in learning-based QEM, which is governed not only by the data ability to approximate reality -- as expected -- but also by how difficult it may be to capture circuit-induced error mechanisms in quantum simulations. In general, this work highlights the urgency of quantum circuit coverage in designing scalable and cost-effective learning-based QEM pipelines for devices in this new NISQ-era. Scaling Laws for Hybrid Quantum Neural Networks: Depth, Width, and Quantum-Centric Diagnostics Danil Vyskubov and Kirill Vyskubov (University of Sharjah) and Nouhaila Innan and Muhammad Shafique (New York University Abu Dhabi) Abstract Abstract Hybrid quantum neural networks are increasingly explored for classification, yet it remains unclear how their performance and quantum behavior scale with circuit depth and qubit count. We present a controlled scaling study of hybrid quantum-classical classifiers along two axes: (1) increasing the number of quantum layers L at fixed qubits Q, and (2) increasing the number of qubits Q at fixed depth L. Across multiple datasets, we evaluate predictive performance using Accuracy, PR-AUC, Precision, Recall, and F1, and track quantum-specific metrics (QCE, EEE, QGN) to characterize how quantum properties evolve under scaling. Our results summarize scaling trends, saturation regimes, and dataset-dependent sensitivity, and further analyze how quantum metrics relate to predictive performance. This study provides practical guidance for selecting (Q,L) in hybrid QNN classifiers and establishes a consistent evaluation protocol for scaling analysis. Design Space Exploration of Hybrid Quantum Neural Networks for Chronic Kidney Disease Muhammad Kashif, Hanzalah Mohamed Siraj, Nouhaila Innan, Alberto Marchisio, and Muhammad Shafique (New York University (NYU) Abu Dhabi, UAE) Abstract Abstract Hybrid Quantum Neural Networks (HQNNs) have recently emerged as a promising paradigm for near-term quantum machine learning. However, their practical performance strongly depends on design choices such as classical-to-quantum data encoding, quantum circuit architecture, measurement strategy and shots. Hybrid Quantum Neural Networks (HQNNs) have recently emerged as a promising paradigm for near-term quantum machine learning. However, their practical performance strongly depends on design choices such as classical-to-quantum data encoding, quantum circuit architecture, measurement strategy and shots. In this paper, we present a comprehensive design space exploration of HQNNs for Chronic Kidney Disease (CKD) diagnosis. Using a carefully curated and preprocessed clinical dataset, we benchmark 625 different HQNN models obtained by combining five encoding schemes, five entanglement architectures, five measurement strategies, and five different shot settings. To ensure fair and robust evaluation, all models are trained using 10-fold stratified cross-validation and assessed on a test set using a comprehensive set of metrics, including accuracy, area under the curve (AUC), F1-score, and a composite performance score. Our results reveal strong and non-trivial interactions between encoding choices and circuit architectures, showing that high performance does not necessarily require large parameter counts or complex circuits. In particular, we find that compact architectures combined with appropriate encodings (e.g., IQP with Ring entanglement) can achieve the best trade-off between accuracy, robustness, and efficiency. Beyond absolute performance analysis, we also provide actionable insights into how different design dimensions influence learning behavior in HQNNs. Thursday 0.14 Singapore IJCNN Paper Reinforcement Learning, Control, and Autonomous Systems Session Chair: Jen-Tzung Chien (National Yang Ming Chiao Tung University), Jiajie Zhang (Technical University of Munich) A Performance Model for Deadline-Aware Off-Policy Reinforcement Learning with Vectorized Environments Klavdiya Bochenina, Volodymyr Beimuk, and Laura Ruotsalainen (University of Helsinki) Abstract Abstract Vectorized environments in modern reinforcement learning (RL) frameworks are widely used to accelerate training for computationally expensive simulators. In many real-world applications, RL policies require periodic retraining under strict operational constraints. To ensure timely deployment, it becomes essential to accurately predict RL training time under limited computational resources. Reducing Experience-Level Non-Stationarity in Multi-Agent Reinforcement Learning via Policy Constrained Replay. Manas Shil, G. N. Pillai, and Manu Kumar Gupta (Indian Institute of Technology Roorkee) Abstract Abstract Replay buffer sampling in multi-agent off-policy learning suffers from experience-level non-stationarity, where outdated transitions no longer align with the current joint policy, leading to biased and unstable updates. Existing approaches address this issue through importance sampling or recency-based replay; however, importance sampling corrects gradient bias without restricting replay selection, whereas recency-driven methods prioritize freshness over relevance and may discard older transitions that remain informative. To address this gap, we propose policy-consistent prioritized experience replay (PC-PER), a trust-region experience replay framework that constrains replay based on policy consistency. PC-PER assigns each replay transition a policy-drift score measuring the mismatch between the stored behavior policy and the current policy. A PPO-style clipping mechanism then adapts the trust-region radius online, resulting in a policy-consistency window of transitions eligible for replay. Within this window, transitions are prioritized based on their temporal-difference error to emphasize samples with higher learning value. Across a broad and diverse suite of multi-agent continuous-control benchmarks from MuJoCo and PettingZoo, PC-PER consistently outperforms all existing baselines, achieving an average performance improvement of 14.6%. PC-Track: Prior-Conditioned 3D Multi-Object Tracking with Cross-Level Fusion for Autonomous Driving Jiajie Zhang, Xiangzhong Liu, and Alois Knoll (Technical University of Munich) Abstract Abstract Commercial 3D sensors and vehicle-to-everything (V2X) modules have been widely adopted in autonomous driving systems; however, their outputs are typically available as sparse, post-processed object lists, which are difficult to integrate into mainstream vision-centric 3D multi-object tracking (MOT) pipelines. In this paper, we propose PC-Track, an end-to-end 3D MOT framework based on cross-level fusion (CLF) that incorporates object lists from external black-box sensors as structured priors within a Transformer architecture and jointly fuses them with multi-view image features. To enhance robustness against noisy or delayed priors in real-world deployment, we introduce a prior-conditioned, fusion-aware association mechanism that unifies motion priors and appearance interactions, enabling stable track continuation under a tracking-by-query paradigm. Extensive experiments on the nuScenes benchmark with real and synthetic object lists demonstrate improvements over the CLF tracking baseline and state-of-the-art vision-based 3D trackers. Thursday 0.15 Washington IJCNN Paper IJCNN SS03 Physics-Informed Neural Networks: Advancements and Applications Session Chair: Ciaran Bench (National Physical Laboratory), Michiel Straat (Bielefeld University) SOLIS: Physics-Informed Learning of Interpretable Neural Surrogates for Nonlinear Systems Murat Furkan Mansur and Tufan Kumbasar (Istanbul Technical University) Abstract Abstract Nonlinear system identification must balance physical interpretability with model flexibility. Classical methods yield structured, control-relevant models but rely on rigid parametric forms that often miss complex nonlinearities, whereas Neural ODEs are expressive yet largely black-box. Physics-Informed Neural Networks (PINNs) sit between these extremes, but inverse PINNs typically assume a known governing equation with fixed coefficients, leading to identifiability failures when the true dynamics are unknown or state-dependent. We propose \textbf{SOLIS}, which models unknown dynamics via a \emph{state-conditioned second-order surrogate model} and recasts identification as learning a Quasi-Linear Parameter-Varying (Quasi-LPV) representation, recovering interpretable natural frequency, damping, and gain without presupposing a global equation. SOLIS decouples trajectory reconstruction from parameter estimation and stabilizes training with a cyclic curriculum and \textbf{Local Physics Hints} windowed ridge-regression anchors that mitigate optimization collapse. Experiments on benchmarks show accurate parameter-manifold recovery and coherent physical rollouts from sparse data, including regimes where standard inverse methods fail. Transferable Physics-Informed Representations via Closed-Form Head Adaptation Jian Cheng Wong (Institute of High Performance Computing), Isaac Yin Chung Lai (National University of Singapore), Pao-Hsiung Chiu and Chin Chun Ooi (Institute of High Performance Computing), Abhishek Gupta (Indian Institute of Technology Goa), and Yew-Soon Ong (Nanyang Technological University) Abstract Abstract Physics-informed neural networks (PINNs) have garnered significant interest for their potential in solving partial differential equations (PDEs) that govern a wide range of physical phenomena. By incorporating physical laws into the learning process, PINN models have demonstrated the ability to learn physical outcomes reasonably well. However, current PINN approaches struggle to predict or solve new PDEs effectively when there is a lack of training examples, indicating they do not generalize well to unseen problem instances. In this paper, we present a transferable learning approach for PINNs premised on a fast Pseudoinverse PINN framework (Pi-PINN). Pi-PINN learns a transferable physics-informed representation in a shared embedding space and enables rapid solving of both known and unknown PDE instances via closed-form head adaptation using a least-squares-optimal pseudoinverse under PDE constraints. We further investigate the synergies between data-driven multi-task learning loss and physics-informed loss, providing insights into the design of more performant PINNs. We demonstrate the effectiveness of Pi-PINN on various PDE problems, including Poisson’s equation, Helmholtz equation, and Burgers’ equation, achieving fast and accurate physics-informed solutions without requiring any data for unseen instances. Pi-PINN can produce predictions 100--1000 times faster than a typical PINN, while producing predictions with 10--100 times lower relative error than a typical data-driven model even with only two training samples. Overall, our findings highlight the potential of transferable representations with closed-form head adaptation to enhance the efficiency and generalization of PINNs across PDE families and scientific and engineering applications. Accelerating Thermochemical Equilibrium Calculations with Physics-Informed Neural Networks Klara Meyer (University of Pretoria); Willem Roos (Ex Mente Technologies); Johan Zietsman (Ex Mente Technologies, University of Pretoria); and Anna Bosman (University of Pretoria) Abstract Abstract Estimating thermochemical equilibrium in multi-component systems is essential for materials science and process engineering, yet computational costs scale poorly with system complexity. This work presents a physics-informed neural network (PINN) approach for scalability from binary to quaternary oxide systems in the CaO-MgO-Al2O3-SiO2 space. Systematic evaluation reveals that selective physics-guided feature engineering improves test R² by 8.3%, while embedding the Gibbs-Helmholtz equation as an architectural constraint reduces thermodynamic consistency error by over six orders of magnitude. The proposed two-stage hybrid architecture, comprising a Transformer encoder for phase estimation, and a multi-layer perceptron for property estimation, achieves R² = 0.96 on binary, R² = 0.94 on ternary, and R² = 0.90 on quaternary systems, with inference speeds of ~72 μs/sample (14-140× faster than direct calculation). These results demonstrate that PINNs offer a scalable path toward rapid equilibrium estimation where traditional methods become computationally prohibitive. Physics-Informed Parameter Identification: From Splines to Kolmogorov-Arnold Networks René Schenkendorf (Harz University of Applied Sciences) Abstract Abstract Mathematical models are crucial for advanced process control and digital twins. Accurate parameter identification, however, often involves a trade-off between statistical rigor and computational tractability. Nonlinear Least Squares (NLS) is the benchmark, enforcing physical laws as hard constraints. By contrast, relaxation methods, such as Scientific Machine Learning (SciML), offer flexibility but introduce soft constraints that can obscure parameter identifiability, particularly in dense data regimes where hard constraints typically yield high-precision estimates. In this work, we systematize these perspectives within a unified generalized Tikhonov regularization framework and clarify the structural link between Iterative Principal Differential Analysis (iPDA) and Physics-Informed Kolmogorov-Arnold Networks (PIKANs). Benchmarking five strategies on nonlinear sloshing dynamics reveals a nuanced reality: under dense sampling, NLS achieves superior accuracy, whereas relaxation methods introduce systematic bias in weakly identifiable parameters. However, in the challenging regime of sparse sampling combined with high noise, NLS can converge to poor local minima, while PIKANs demonstrate crucial robustness. These findings caution against uncritical adoption of SciML for well-posed inverse problems while highlighting their essential value under data scarcity, advocating for problem-specific strategy selection. Joint Seismic Inversion and Wavelet Estimation Using Physics-Informed Neural Networks Anthony João Bet and Rafael Santiago (Universidade Federal de Santa Catarina, Departamento de Informática e Estatística); Bruno Barbosa Rodrigues (Petrobras); and Mauro Roisenberg (Universidade Federal de Santa Catarina, Departamento de Informática e Estatística) Abstract Abstract Seismic inversion, used to estimate elastic and petrophysical properties of subsurface regions to identify potential oil and gas reservoirs, is a challenging, ill-posed, and highly nonlinear problem. Neural networks have shown great potential in addressing these challenges due to their ability to model complex nonlinear relationships. However, their application often requires a large amount of training data, which is often a limiting factor in practical applications. In this work, we propose a Physics-Informed Neural Network (PINN) architecture for seismic inversion and wavelet estimation. The seismic wavelet is an essential parameter to be applied in the forward seismic physical model to obtain the seismic traces from impedance values, so determining its correct value is critical. The novelty of our approach lies in the introduction of a network designed to estimate the seismic wavelet, trained jointly with the inversion network. We evaluated the model on the well-known Marmousi2 synthetic dataset and on real poststack seismic data. Results demonstrate that the model outperforms traditional approaches while accurately estimating the wavelet. On the real dataset, the results indicate the framework’s robustness and its ability to generalize to complex geological scenarios, highlighting the potential of PINNs in geophysical applications. Thursday 0.02 Berlin IJCNN Paper, FUZZ-IEEE Position Paper, CEC Late Breaking Paper, CEC Paper, FUZZ J2C Presentation, CEC J2C Presentation, FUZZ-IEEE Paper, CEC Position Paper, IJCNN J2C Presentation, IJCNN Position Paper, IJCNN Late Breaking Paper, FUZZ-IEEE Late Breaking Paper Regular Best Paper Award Self-Attention-Guided Genetic Programming for Dynamic Scheduling: Leveraging BERT for Enhanced Tree-Structured Data Operations Shanshi Mao, Fangfang Zhang, Yi Mei, and Mengjie Zhang (Victoria University of Wellington) Abstract Abstract While Bidirectional Encoder Representations from Trans-formers (BERT) has demonstrated strong performance in learning rich representations from sequential and grid-based inputs like natural language and images, its extension to non-sequential topologies remains an open research question. This study investigates the application of BERT to tree-structured data which presents a significant challenge due to its lack of explicit sequential order and complex topological dependencies. Specifically, we integrate BERT with genetic programming whose classic data representation is tree data structure to solve the dynamic flexible job shop scheduling (DFJSS) problem. The DFJSS problem’s inherent computational complexity and highly dynamic, uncertain nature provide a rigorous testbed for our methodology. Our experiments demonstrate that BERT can effectively capture and integrate the structural information embedded in these tree-based representations. This finding highlights the versatility and adaptability of the self-attention mechanism, extending its utility beyond conventional sequential or grid-based data structures to a broader class of complex and non-sequential topologies. Multi-Fidelity Multi-Objective Optimization of Electric Machines Having Heterogeneous and Blocked Evaluation Times Balija Santoshkumar and Kalyanmoy Deb (Michigan State University) Abstract Abstract Complex engineering design problems often involve multiple objectives and constraints, requiring third-party simulation software based evaluations having heterogeneous computational times. For a time-budget optimization run, a judicious management of high-fidelity evaluations of computationally cheaper and expensive objectives and constraints must be adopted. In this paper, we apply a recently proposed multi-fidelity selection based evolutionary multi-objective optimization (EMO) procedure to optimize an electric machine design problem to demonstrate its effectiveness in solving real-world computationally expensive optimization problems. The problem formulation follows that presented in a foundational study, with appropriate modifications to align with objectives and constraints of the current application. With commercial multi-physics expensive evaluation of electromagnetic and structural analysis softwares, and cheaper mathematical expressions, we demonstrate how the proposed MFE-NSGA-III approach can be efficiently applied to find a number of alternate trade-off designs compared to a conventional surrogate-assisted optimization procedure. The proposed procedure judiciously assigns computationally cheaper and expensive blocked objectives/constraints adaptively to make the overall procedure more computationally effective. How Many Subproblems? A Controlled Study of Decomposition Size in Multi-Policy Multi-Objective Reinforcement Learning Neele Kemper, Jonathan Wurth, Michael Heider, and Jörg Hähner (University of Augsburg) Abstract Abstract Decomposition-based multi-objective reinforcement learning reduces a multi-objective problem to a set of scalar subproblems, each defined by a weight vector on the preference simplex. While this approach has proven effective, the number of weight vectors is typically chosen heuristically without systematic justification. This paper presents a controlled empirical study examining how the number of weight vectors affects Pareto front quality under fixed training budgets. A multi-policy decomposition framework is evaluated with Soft Actor-Critic and Proximal Policy Optimization, training one independent policy per weight vector via linear scalarization across configurations from two to twenty-one weight vectors and six MuJoCo tasks with two and three objectives. A Bayesian temporal ranking model quantifies uncertainty about which configuration is optimal throughout training, and the best-found configurations are compared against state-of-the-art baselines. The results show that the optimal decomposition size fundamentally depends on the algorithmic structure. Off-policy learning favors small decompositions, with four weight vectors serving as a robust default. A shared replay buffer enables information exchange between subproblems and ensures high sample efficiency.On-policy learning operates without replay, requiring larger decompositions that scale with the training budget and produce less pronounced optima. The common default of six weight vectors is suboptimal for both algorithms, and minimal corner weight configurations lead to significant performance degradation. With properly tuned configurations, the proposed implementations outperform all baselines. MerLin: A Discovery Engine for Photonic and Hybrid Quantum Machine Learning Cassandre Notton (Quandela Quantique Inc.); Benjamin Stott (Quandela); Philippe Schoeb (Quandela Quantique Inc.; DIRO and Mila, Université de Montréal); Anthony Walsh (Quandela); Grégoire Leboucher (Quandela, ENS Paris-Saclay); Vincent Espitalier and Vassilis Apostolou (Quandela); Louis-Félix Vigneux (Quandela Quantique Inc.; Département d'informatique, Université de Sherbrooke); and Alexia Salavrakos and Jean Senellart (Quandela) Abstract Abstract Identifying where quantum models may offer practical benefits in near-term quantum machine learning (QML) requires moving beyond isolated algorithmic proposals toward systematic and empirical exploration across models, datasets, and hardware constraints. We introduce MerLin, an open-source framework designed as a discovery engine for photonic and hybrid quantum machine learning. MerLin integrates optimized strong simulation of linear-optical circuits into standard PyTorch and scikit-learn workflows, enabling end-to-end differentiable training of quantum layers. Fuzzy Vertex-Feature Clustering Framework for Predicting Survival of Health Sector Companies And Generative Explanations Lerina Aversano (University of Foggia) and Vincenzo Dentamaro and Felice Franchini (University of Bari "Aldo Moro") Abstract Abstract Accurately forecasting the long-term viability of firms in the highly regulated and competitive health-care sector is critical for investors, managers, and policymakers. Conventional machine-learning methods often struggle to reconcile predictive power with transparency, while standard clustering techniques provide limited insight with high-dimensional, noisy business data. Thursday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 21- Algorithms VI Session Chair: Ruibin Bai (University of Nottingham Ningbo China) Reinforcement Learning for Job-Shop Scheduling via Hybrid GNN Embeddings with Sophisticated State Representations Haohong Xu (National Frontiers Science Center for Industrial Intelligence and Systems Optimization, Northeastern University, China); Qingxin Guo (Key Laboratory of Data Analytics and Optimization for Smart Industry (Northeastern University), Ministry of Education, China); Zhiming Dong (Liaoning Engineering Laboratory of Data Analytics and Optimization for Smart Industry, Northeastern University); and Guodong Zhao (Liaoning Key Laboratory of Manufacturing System and Logistics Optimization, Northeastern University) Abstract Abstract Although neural optimization methods based on priority rule learning and graph isomorphic networks (GINs) have shown promise for the job shop scheduling problem (JSSP), traditional GINs still struggle to effectively capture the dynamic constraints and scheduling relationships when processing disjunctive graphs. This paper proposes an end-to-end algorithmic architecture that integrates a graph attention network (GAT) with a GIN. The architecture leverages the GIN to extract general graph structural features, while introducing the GAT to focus on temporal attributes and dynamic topological changes. The fused network is trained using the proximal policy optimization (PPO) algorithm to achieve adaptive optimization for dynamic job shop scheduling. Experimental results demonstrate that the proposed algorithm outperforms both traditional priority dispatching rules and advanced deep reinforcement learning methods, while also validating the effectiveness of the network fusion. Canny Genetic Programming (CGP) for Edge Detection Daniel van Zyl, Thambo Nyathi, and Nelishia Pillay (University of Pretoria) Abstract Abstract Edge detection is a foundational image processing task. Modern methods are driven by deep learning, which normally requires a large source of labelled training data. Additionally, these methods require large computational resources for training. Conventional methods, although not as accurate as deep learning approaches, have a quick execution time and may not always require labelled data. Amongst the conventional methods, the Canny Edge Detector (CED) is considered to be an industry standard. Despite its efficiency and widespread adoption, the Canny detector has notable limitations. Its gradient-based formulation struggles with semantic boundary localisation. Furthermore, CED is vulnerable to noise interference, which leads it to produce spurious edge results. Critically, Canny’s performance is dependent on manual configuration. The optimal selection of parameters is context-dependent and lacks an automated framework for diverse noise profiles. This study proposes a hybrid approach that combines the Canny Edge Detector and Genetic Programming for edge detection. To evaluate the proposed approach, a subset of images was obtained from the Berkeley Segmentation Dataset. The performance of edge maps evolved by the proposed approach was compared with that of the standard Canny detector. Three metrics, namely recall, precision, and F-measure, were used for comparison. On average across all images, the proposed approach was found to perform significantly better than the standard Canny algorithm. Additionally, the computational training times were found to be within reason. Adaptive Neuro-Evolution for Dynamic Multi-UAV Routing with Load-Dependent Constraints Xiaoyan Gong, Xiaoying YANG, Zhenwei Wang, Fuhua Jia, Siyuan He, Tianxiang Cui, and Ruibin Bai (University of Nottingham, Ningbo, China) Abstract Abstract Energy Capacitated Vehicle Routing Problems with Time Windows (ECVRPTW) are central to modern logistics, particularly UAV-based delivery, owing to the generality of their operational constraints. Existing methods predominantly assume static customer demands, yet real-world operations are inherently dynamic: new service requests emerge during route execution. Conventional static solvers accommodate such changes via customer insertion into ongoing routes—an approach that sacrifices flexibility and often yields suboptimal solutions. While autoregressive reinforcement learning (RL) enables rapid inference for dynamic settings, its probability distribution-based decoding incurs a performance bottleneck when scaling to large instances with complex constraints. To address the uncertainty and generalization requirements of real-time dynamic ECVRPTW, we propose a closed-loop neuro-evolutionary framework built on a hierarchical manager–worker architecture. At the manager level, a deep RL agent trained with Proximal Policy Optimization (PPO) determines, in real time, when and how to re-optimize active routes. At the worker level, an autoregressive pre-trained Transformer constructs an initial solution in a step-by-step manner; this solution is subsequently refined through evolutionary optimization (e.g., NSGA-style operators) to satisfy energy, capacity, and time-window constraints with respect to the current system state. Extensive experiments show that our framework flexibly reorders unserved tasks in response to newly arriving demands, consistently outperforming prior static replanning strategies on benchmark instances and establishing a new state of the art for online routing under uncertainty. An Interactive Multi-Objective Optimization Method for Deriving Measures to Social Issues while Presenting the Predicted Pareto Frontier Toshio Ito, Kota Itakura, Arika Hakoda, and Miwa Ueki (Fujitsu Limited) and Koki Ikeda, Kota Nagakane, and Isao Ono (Institute of Science Tokyo) Abstract Abstract Social issues are complex, and solving them from a single perspective has other negative effects. We set up various Key Performance Indicators (KPIs), in which balanced measures need to be implemented to achieve KPIs. However, it is not easy for decision-makers to find measures for issues while considering multiple perspectives. So, we develop an interactive measure search method that simulates the impact of measures on society through various KPIs, interacts with decision-makers, and evaluates numerous measures to identify those aligned with the decision-makers' values. This method become a type of interactive multi-objective optimization method using evolutionary computation. To find the most preferred measure for decision-makers, this method narrows the search area to the decision-makers' preference region within the Pareto Frontier (PF) while interacting with decision-makers. However, in the early stages of interaction, since only a few optimal solutions calculated by evolutionary computation are available, the overall shape of the PF is not well understood, making it difficult for decision-makers to specify their preference region within the PF. Therefore, we propose a method where we predict the PF from the optimal solutions computed by evolutionary computation, present the predicted PF to decision-makers. The proposed method reduces the number of interactions with decision-makers, enabling faster determination of preferred measures for them. E^2-CTS-DE: Chaotic Thompson Sampling for Operator Selection in Differential Evolution Luis A. Beltran (Universidad de Guadalajara); Daniel F. Zambrano-Gutierrez (Tecnologico de Monterrey); Omar Alvarez, Diego Oliva, Mario A. Navarro, and Javier Galvis-Chacón (Universidad de Guadalajara); and Saul Zapotecas-Martinez (Instituto Nacional de Astrofisica Optica y Electronica) Abstract Abstract Differential Evolution (DE) is highly sensitive to the choice of search operators and to the exploration--exploitation balance induced by their interaction. Many DE variants rely on fixed operator configurations or handcrafted adaptation rules, which can reduce robustness on heterogeneous and deceptive landscapes. This work formulates operator adaptation from a selection hyper-heuristic viewpoint and proposes E\textsuperscript{2}-CTS-DE (Exploration--Exploitation Chaotic Thompson Sampling Differential Evolution), a learning-driven DE variant that adaptively selects mutation--crossover combinations during the search. The method models each mutation--crossover pair as a low-level heuristic and employs a Multi-Armed Bandit scheme driven by Thompson Sampling to perform online selection based on improvement feedback. In addition, the initial population is generated via chaotic maps to enhance diversity and strengthen early exploration. Experiments on a standard suite of continuous black-box benchmark functions show that E\textsuperscript{2}-CTS-DE is competitive and frequently outperforms state-of-the-art DE variants and representative population-based optimizers, achieving robust performance across multimodal, unimodal, composite, and shifted landscapes. Towards an Algorithm-Agnostic Framework for Mitigating Ill-Conditioning in Black-Box Optimization via PCA Whitening Chengjie Zhang (Shandong Key Laboratory of Ubiquitous Intelligent Computing, University of Jinan; Quan Cheng Laboratory); Bo Zhang (Shandong Provincial Key Laboratory of Green and Intelligent Building Materials, University of Jinan); Lin Wang (Shandong Key Laboratory of Ubiquitous Intelligent Computing, University of Jinan; Quan Cheng Laboratory); and Bo Yang (Quan Cheng Laboratory; Shandong Key Laboratory of Ubiquitous Intelligent Computing, University of Jinan) Abstract Abstract Ill-conditioned problems present formidable challenges to search efficiency due to severe scale imbalances. Most existing mitigation strategies rely on embedding adaptive mechanisms within specific algorithms, which often leads to tight algorithmic coupling and restricted generalizability across diverse optimization paradigms. To overcome these limitations, this paper introduces a general, algorithm-agnostic geometric correction framework. By integrating Random Perturbation--Based Sample Augmentation with PCA whitening, the framework fundamentally reparameterizes narrow, steep-walled ``valleys" into isotropic, spherical-like latent spaces that are significantly more amenable to optimization. Designed as a "plug-and-play" external module, the proposed framework enables seamless integration without modifying the core search strategy of the underlying optimizers. Extensive evaluations on ten high-condition-number BBOB benchmark functions demonstrate that the framework consistently and significantly bolsters the performance of multiple representative optimization algorithms on ill-conditioned landscapes, markedly accelerating convergence while enhancing the quality of final solutions. Thursday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 22 - SS19: Evolutionary Computer Vision and Image Processing (ECVIP) Session Chair: Ying Bi (School of Electrical and Information Engineering, Zhengzhou University, China; State Key Laboratory of Intelligent Agricultural Power Equipment) A Novel Genetic Programming Approach for Skeleton-Based Human Action Recognition Youkai Xiao, Mitchell Rogers, Bing Xue, and Mengjie Zhang (Victoria University of Wellington) and Patrice Delmas (The University of Auckland) Abstract Abstract Human action recognition (HAR) plays a vital role in many real-world applications, such as sports. However, illumination changes and complex backgrounds make video-based HAR highly challenging. Skeleton-based HAR (SHAR) alleviates these issues by representing human motion using structured joint coordinates, which are less sensitive to variations in illumination and background. Despite recent progress, popular deep neural network based approaches to SHAR often rely on large-scale labeled data and offer limited interpretability. To address these limitations, we explore genetic programming (GP) as an alternative that can automatically construct discriminative, human-interpretable motion descriptors from skeleton sequences. In this work, we propose a GP-based approach, named SkeletonGP, which extracts informative geometric, motion, statistical, and correlation features via a multi-layer feature extraction framework. The performance of our approach is evaluated on a popular skeleton-based HAR dataset (Penn Action dataset). The results suggest that the proposed SkeletonGP approach achieves competitive performance compared to deep neural network methods, while providing a promising interpretable solution. Furthermore, the ablation study demonstrates the importance of adding a statistical and correlation feature extraction layer on top of the geometric and motion feature extraction layer. Personalized Federated Neural Architecture Search for Edge-Based Monocular Depth Estimation Zhilong He (Hebei University of Technology), Haoyu Zhang (Hangzhou Normal University), Qiqi Liu (Westlake University), Xinjie Wang (Hangzhou Normal University), Junhua Gu (Hebei University of Technology), and Ruyu Liu (Hangzhou Normal University) Abstract Abstract Monocular depth estimation (MDE) is a key component of 3D reconstruction and plays an essential role in applications such as autonomous driving, augmented reality, and robotic navigation. In practical scenarios, MDE models are often required to be deployed on edge devices to support real-time inference. However, deploying MDE models on edge devices faces two major challenges. First, data collected by different devices exhibit significant heterogeneity, while computational and storage resources vary widely across clients, making a unified fixed architecture difficult to meet diverse deployment requirements. Second, existing high-accuracy MDE models usually involve heavy computational costs, which limits their applicability on resource-constrained edge devices. To address these challenges, we propose PTF-EMDE, a personalized federated learning (PFL) framework based on bi-objective evolutionary neural architecture search. By designing a lightweight encoder search space and employing a bi-objective optimization strategy, PTF-EMDE is able to significantly reduce computational complexity while maintaining high estimation accuracy, thereby automatically discovering efficient and client-adaptive model architectures for heterogeneous edge devices. Experimental results on the KITTI dataset demonstrate that PTF-EMDE achieves higher prediction accuracy and lower computational cost than baseline methods, providing an efficient and scalable solution for edge deployment of MDE in privacy-sensitive scenarios. EnergyMamba: Energy Attention Modulated Mamba for Hyperspectral Image Classification Faiq Ahmad, Saad Sohail, Muhammad Usama, and Usman Ghous (FAST); Manuel Mazzara (Innopolis University); Danish Shehzad (New Uzbekistan University); and Muhammad Ahmad (KFUPM) Abstract Abstract Hyperspectral image classification is challenged by extremely high spectral dimensionality, severe inter-band redundancy, and the need to model long-range spectral–spatial dependencies under limited labeled data. Although Transformer-based models provide powerful global context modeling, their quadratic complexity restricts scalability, while recent Mamba-based state-space models, despite their linear complexity, lack explicit mechanisms for adaptive feature discrimination. This paper proposes EnergyMamba, a hybrid architecture that integrates selective state-space modeling with an energy-based attention mechanism to address these limitations in a principled manner. The Mamba branch efficiently captures long-range spectral–spatial dependencies with linear computational complexity, whereas the Energy Attention branch incorporates Hopfield-inspired local energy functions and multi-head energy attention to dynamically reweight features, suppress spectral redundancy, and enhance class separability. By coupling global state-space dynamics with energy-guided associative feature refinement, EnergyMamba provides a discriminative and stable inductive bias for hyperspectral representation learning. Extensive experiments on the Salinas, WHU-Hi-HanChuan, and OHID-1 datasets demonstrate that EnergyMamba achieves state-of-the-art performance, attaining overall accuracies of 99.84\%, 99.36\%, and 94.59\%, respectively, using only 5\% of labeled samples, with particularly strong gains on the challenging OHID-1 dataset. Exploring Shape-based Vision via MultiStage Swarm Optimization with User-Drawn Priors Nordin Zakaria (Universiti Teknologi Petronas) Abstract Abstract An exploratory study of a Computer Vision-by-synthesis approach is presented that maps 2D user-drawn shape priors to contours extracted from an input image or image region. In a preparatory stage, a user observes an object and draws 2D low complexity contours depicting particular views of the object. At deployment time, the drawing is mapped to the input image contours through transformations and deformations. The mapping is framed as an optimization problem solved using |
Monday 0.01 London FUZZ-IEEE Paper FUZZ 1 : Fuzzy data analysis, clustering and classifiers, pattern recognition, bio-informatics Session Chair: Thomas Runkler (Siemens AG, Technische Universtat Munchen) Monday 0.02 Berlin FUZZ-IEEE Paper FUZZ 2: FUZZ-IEEE SS06 Fairness and Trustworthiness in Intelligent Decision Support Systems & Main: Fuzzy web engineering, information retrieval, text mining and social network analysis Session Chair: Luis Martinez (University of Jaén) A Type-2 Fuzzy-Based Feature-Driven Rule Reduction Approach for Demand Forecasting within an Energy Smart Grid System pdfMonday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 1: Algorithms I Session Chair: Pauline Catriona Haddow (NTNU) QiSA: A Parameter-free Quantum-inspired Search Algorithm -- A Preliminary Study on Combinatorial Optimization pdfExploring Island Genetic Algorithms and Unbounded Recognition Regions for Artificial Immune Systems pdfMonday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 2 - Evolutionary Machine Learning I Session Chair: Lukas Sekanina (Brno University of Technology) Monday 0.10 Sydney IEEE CEC (Evolutionary Computation) CEC 3 - Optimization I Session Chair: John Sheppard (Montana State University) Monday 0.11 Cape Town IEEE CEC (Evolutionary Computation) CEC 4 - Related Topics I Session Chair: Sanaz Mostaghim (Otto von Guericke University Magdeburg, Fraunhofer IVI Dresden) Diagnosing Optimizer-Landscape Interaction via Representation in Subset Selection Multitasking: Evidence from Piano Fingering pdfMonday 0.14 Singapore IEEE CEC (Evolutionary Computation) CEC 5 - SS01: Integrating Machine Learning Methods into Evolutionary Optimization Session Chair: Amir Gandomi (University of Technology Sydney) Adaptive Switching between Search and Model Refinement via Explainable Machine Learning in Surrogate-assisted Evolutionary Algorithms pdfOnline electric vehicle charging scheduling with commitment using surrogate-assisted optimization pdfDynamic Multi-Modal Particle Swarm Optimization for Training Neural Network Ensembles Under Concept Drift pdfMonday 0.15 Washington IEEE CEC (Evolutionary Computation) CEC 6-SS10:Evolutionary Computation in Dynamic and Uncertain Environments Session Chair: Michalis Mavrovouniotis (Cyprus University of Technology) Dynamic Multi-Objective Optimization of Integrated Energy Systems in Steel Enterprises via Conditional Variational Autoencoder pdfMulti-Point Search Towards Dynamic Multimodal Optimization with Frequent Solution Changes on their Locations, Moving speeds, and Numbers pdfFrom Environmental Parameters to Pareto Optimal Prediction: LLM-Based Large-Scale Dynamic Multi-Objective Optimization pdfAn Interval-Constrained MultiObjective Evolutionary Algorithm Integrating Interval Constrained Clustering and Transfer Learning pdfMonday Brightlands Foyer IJCNN J2C Presentation, IJCNN Paper, IJCNN Position Paper, IJCNN Late Breaking Paper Poster Presentations IJCNN (1) Session Chair: Thorben Markmann (Bielefeld University), Aleksei Liuliakov (Bielefeld University) STT-BP: Training Kernel-Learnable LIF for Efficient Neuromorphic Vision and Audio Recognition via Spatio-Temporal Tilted Backpropagation pdfProbing the Functional Role of Muscle Synergies in Reinforcement Learning–Based Torso Balance via Neural Reconstruction and Synergy Ablation pdfThe Geometry of Fragility: Benchmarking Robustness of Time Series Explanations via Curvature Maximization pdfSame Accuracy, Different Geometry: Solver-Dependent Representation Learning in Third-Order Neural ODEs pdfReservoir Computing Combined with Optimizable Quantile-Based Subsampling for Small-Sample Time-Series Classification pdfFrom Spatiotemporal Feature Alignment to Transferable Causal Discovery: Process Modeling for Interpretable Analysis of Hot-Rolled Surface Quality Mechanisms pdfLMABC: A Cost-Efficient Large Language Model Multi-Agent Framework for Automated Sequence Labeling in Building Codes pdfO2O-TP: An Offline-to-Online Transformer-Based Reinforcement Learning Approach for Portfolio Management pdfDistributed On-Orbit Privacy Preservation for Satellite Imagery via Lightweight Neural Inpainting pdfEfficient Punctuation Restoration via Weighted Lookahead Scoring Method for Streaming ASR Systems pdfImproving the Robustness of Control of Chaotic Convective Flows with Domain-Informed Reinforcement Learning pdfPATH-NM: Learning Orthogonal N:M Sparse Patterns via Differentiable Transformation for Efficient Heterogeneous MARL pdfExperience Constrained Hierarchical Federated Reinforcement Learning for Large-scale UAV Teams in Hazardous Environments pdfCGSSN: A Coupled Graph Spiking Skeleton Network for Human Action Recognition Using Neuromorphic Vision Sensors pdfKeep the Teacher Non-Fine-Tuned: Knowledge Distillation for Stable Molecular Dynamics with Neural Network Potentials pdf2DGS-Room: Seed-Guided 2D Gaussian Splatting with Geometric Constraints for High-Fidelity Indoor Scene Reconstruction pdfPRISM-Med: A Perception, Routing, Inspection, and Self-Correction Multi-agent Collaboration System for Medical VQA pdfDMR-Seg: Endowing Multimodal Large Language Models with Segmentation Capability via Discretized Mask Representations pdfQ-SARIMA: A Hybrid Quantum–Classical Extension of SARIMA for Time Series Forecasting in Precision Agriculture pdfAdvanced Recombinant Transformer: Instrument-Level Audio Compatibility Scoring with Self-Supervised Acoustic Embeddings pdfCogKT: Cognitive Ability-based Knowledge Tracing towards Interpretable Knowledge Mastery Representation in Personalized Learning pdfReloDepth: Reprojection Loss Filtering for Self-Supervised Depth Estimation in Dynamic Environments pdfMonday 0.01 London FUZZ-IEEE Paper FUZZ 3: FUZZ-IEEE SS11 Software for Soft Computing Session Chair: Giovanni Acampora (University of Naples Federico II), James Keller (University of Missouri) FuzzyLinguistics: A Python Package for Generating, Simplifying, and Comparing Linguistic Summaries of Data pdfMonday 0.04 Brussels IJCNN Paper Neural Learning and Optimization I Session Chair: Jose Principe (University of Florida), Manos Kirtas (Centre for Nanosciences and Nanotechnologies, Aristotle University Of Thessaloniki) Monday 0.05 Paris IJCNN Paper Brain-Inspired and Cognitive Neural Systems Session Chair: Thomas Trappenberg (Dalhousie University), Punit Rathore (Indian Institute of Science) Using Disinhibition versus Direct Control in a Spiking Neural Model of Dopamine-Driven Reinforcement Learning pdfMonday 0.10 Sydney IJCNN Paper LLM Agents and Neural Reasoning Session Chair: Chris Yakopcic (University of Dayton, OH, USA), Snehasis Banerjee (TCS Research) Monday 0.11 Cape Town IJCNN Paper Robust and Adversarial Machine Learning Session Chair: Rafael de Santiago (Universidade Federal de Santa Catarina, Departamento de Informática e Estatística), Pranjala Kolapwar (SGGS Institute of Engineering and Technology) Robustness of Fuzzy ARTMAP to Adversarial Attacks and Progressive Adversarial Training for Streaming Learning pdfMonday 0.14 Singapore IJCNN Paper Scientific ML and Bio-Physical Sensing Session Chair: Eyad Elyan (Robert Gordon University), Mathieu Vandwalle (X-FAB) Hyperspectral Image Classification with Imbalanced Data Based on a Semi-Supervised Learning Algorithm pdfMonday 0.15 Washington IJCNN Paper IJCNN SS17 Generative Foundation Models for Robotics: From Language and Vision to Embodied Action Session Chair: Erdi Sayar (Paderborn University), Alper Yegenoglu (Paderborn University), Van Huyen Dang (Paderborn University), Erdal Kayacan (Paderborn University, Germany) Beyond Geometry: Leveraging Pose and Visual Scene Understanding for Indoor Room Segmentation in Mobile Robots pdfMonday 2.2 Zambezi IJCNN Paper IJCNN SS09 Multimodal Deep Learning in Applications Session Chair: Stefania Tomasiello (University of Salerno), Dawid Połap (Silesian University of Technology, Poland) Data-lightweight particulate matter forecasting using temporal attention and Kolmogorov-Arnold Networks pdfMonday 0.04 Brussels IEEE CEC (Evolutionary Computation), CEC Position Paper CEC 7 - Large Language Models and Neuromorphic EC Session Chair: Mengjie Zhang (Victoria University of Wellington, Centre for Data Science and Artificial Intelligence) Position Paper: Suggestions from Large Language Models about Experimental Settings for Benchmarking Evolutionary Multi-objective Optimization Algorithms pdfMonday 0.10 Sydney IJCNN Position Paper Foundations of Neural and Scientific ML (Position Track) Session Chair: Donald Wunsch (Missouri University of Science and Technology), Alessio Martino (Dept. of AI, Data and Decision Sciences, LUISS University) Monday 0.14 Singapore IJCNN Paper Deep learning for computer vision Session Chair: Alper Yegenoglu (Paderborn University), Stefania Tomasiello (University of Salerno) EEG-MFTNet: An Enhanced EEGNet Architecture with Multi-Scale Temporal Convolutions and Transformer Fusion for Cross-Session Motor Imagery Decoding pdfMonday 0.15 Washington IJCNN Paper IJCNN SS06 Collaborative Learning of Trustworthy Computational Intelligence Systems (CLOTHES 2026) - Third Edition Session Chair: Pietro Ducange (University of Pisa, Italy) Monday 2.18 Mekong IJCNN Paper SS30 Computational Intelligence and AI Applications for Sustainable Energy Management in Smart Grids and Energy Communities Session Chair: Enrico De Santis (University of Rome "La Sapienza") Black-Box Adversarial Attacks on Smart Grid Stability Score Prediction and a Stochastic Quantile Smoothing Defense pdfMonday 2.2 Zambezi IJCNN Paper IJCNN SS26 Brain Machine Intelligence: Models, Systems, and Translational Applications Session Chair: Marwen Belkaid (ETIS UMR 8051, CY Cergy Paris Université, ENSEA, CNRS, F95000 Cergy, France), PANTEA KEIKHOSROKIANI (University of Oulu) Monday 2.1 Volga IJCNN Paper Explainable AI Session Chair: Naoyuki Kubota (Tokyo Metropolitan University), Imen Jdey (REGIM Lab, Sfax University), Simona Casini (University of Pisa) Monday Virtual Room 1 IJCNN Paper SS09 Multimodal Deep Learning in Applications I Session Chair: Fang Yang (Hebei university), Yinfeng Yu (Xinjiang University) Adaptive Global-Local Contrastive Learning and Information Enhancement for Multimodal Entity and Relation Extraction pdfMonday Virtual Room 2 IJCNN Paper SS09 Multimodal Deep Learning in Applications II Session Chair: Zhiyi Zhu (Communication University of China), yu song (East China Normal University) BCMIRec: Behavior Co-occurrence Enhanced Multi-Interest Dual-Graph Learning for Multimodal Recommendation pdfMonday Virtual Room 3 IJCNN Paper SS09 Multimodal Deep Learning in Applications III Session Chair: Haowen Zhu (Southeast University), Alberto Lopez Casanova (Future Connections, R&D Dept.) MIRAGE: Modality-Aware 3D Vision-Language Model for Liver Lesion Diagnosis in MRI via Volumetric-Text Alignment pdfMonday Virtual Room 4 IJCNN Paper SS09 Multimodal Deep Learning in Applications IV Session Chair: mingchao zhang (inner Mongolia University of Technology), Zhikui Chen (Dalian University of Technology) Dustformer: A PM10 Concentration Prediction Model Integrating Spatiotemporal Modeling and Mixture-of-Experts Mechanism pdfD2TNet: A Decoupled Dual-Teacher Framework for Robust Infrared and Visible Image Fusion in Diverse Degraded Scenes pdfCoherentDrive: Conflict-Aware World-State Grounding for Hallucination-Resistant Autonomous Driving Reasoning pdfMonday Virtual Room 5 IJCNN Paper SS09 Multimodal Deep Learning in Applications V Session Chair: Mingyong Li (Chongqing Normal University), Dongxun Jiang (Tongji University) Monday Virtual Room 6 IJCNN Paper SS09 Multimodal Deep Learning in Applications VI Session Chair: Jinhui Yu (Zhejiang Provincial Museum), yiqiang he (zhejiang university of technology) Monday Virtual Room 7 IJCNN Paper SS09 Multimodal Deep Learning in Applications VII Session Chair: 树兰 张 (四川师范大学, College of Computer Science), 斌 王 (East China University of Science and Technology) ChartMVC: A General Multi-View Contrastive Learning Framework for Robust Chart Question Answering pdfA Social Media Emotion Detection Model Based on Implicit Emotion Multi-Hop Memory and Knowledge Embedding pdfMonday Virtual Room 8 IJCNN Paper SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models I Session Chair: Ruwang Jiao (Soochow University), XINGBANG DU (Hokkaido University) Monday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V1 Session Chair: Bin Cao (Hebei University of Technology) Entropy-Guided Maximin-Fitness MOPSO with Simulated Annealing for Multi-Objective Dual-Resource Flexible Job Shop Scheduling pdfA Surrogate-Assisted Evolutionary Algorithm with Geometry-Guided Initialization and Memory-Guided Variation for Expensive Multi-Objective Problems pdfMonday Virtual Room 1 IJCNN Paper SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models II Session Chair: Ruwang Jiao (Soochow University), Lianbo Ma (College of Software, Northeastern University, China; Foshan Graduate School of Innovation, Northeastern University) GFNAS: Towards Gradient-Friendly Neural Network Architecture Search From Sharpness-Aware Perspective pdfMonday Virtual Room 2 IJCNN Paper SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models III Session Chair: Meng Wang (Liaoning Technical University), Zhiguo Hu (Shanxi University) Global-Local nnUNet with Anatomical-Metabolic Statistical Priors for Improved PET/CT based Whole Body Lymphoma Segmentation pdfMonday Virtual Room 3 IJCNN Paper SS43 Computational Intelligence and Software Engineering I Session Chair: Priyaranjan Pattnayak (Oracle Cloud AI), Xiaobing Xiong (Key Laboratory of Cyberspace Security, Ministry of Education; Information Engineering University) CausalVul: Robust Vulnerability Detection via Dual-Granularity Semantic Fusion and Causal Structural Learning pdfSAD-Gen: Structure-Aware Unit Test Generation via Dual-Encoder Fusion and Latent Space Disentanglement pdfMonday Virtual Room 4 IJCNN Paper SS43 Computational Intelligence and Software Engineering II Session Chair: Yunpeng Wang (Taiyuan University of Technology), Chengxiao Zhao (Qilu University of Technology (Shandong Academy of Sciences), Shandong Computer Science Center (National Supercomputer Center in Jinan)) A Vulnerability Detection Method with Semantic-Sensitive Contrastive Learning and Token-Level Graph Representation pdfMonday Virtual Room 5 IJCNN Paper SS43 Computational Intelligence and Software Engineering III Session Chair: Hamed Jelodar (UNB), Yuchuan Chen (Changsha University of Science and Technology, School of Computer Science and Technology) Monday Virtual Room 6 IJCNN Paper SS04 Tiny Machine Learning I Session Chair: Zonglin Yang (Guangdong Police College), Fei Ge (School of Computer Science, Central China Normal University, Wuhan, China) Enabling Memory-efficient Im2win Convolution with Multi-precision Support on GPU CUDA and Tensor Cores pdfMonday Virtual Room 7 IJCNN Paper SS04 Tiny Machine Learning II Session Chair: Chenkai Liao (Hunan University of Technology), jiaping Wang (East China Normal Univerisity) Monday Virtual Room 8 IJCNN Paper SS04 Tiny Machine Learning III Session Chair: Hatem Trigui (Hahn Schickard), BINHUA HUANG (University College Dublin) FedTinyProp: Adaptive Sparse Backpropagation for Efficient Federated Learning on Embedded Devices pdfMonday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V2 Session Chair: Xiangming Jiang (Xidian University) Monday Virtual Room 1 IJCNN Paper SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models I Session Chair: Di Wu (Hebei University of Engineering), Shuaichao Zhang (Southwest University of Science and Technology) Patient-State–Conditioned Pharmacodynamics Modeling With Counterfactual Inference for ADR Risk in Peritoneal Dialysis pdfSCFD: A Self-Calibrated Feature Denoising Framework for Robust Medical Image Classification Under Extreme Class Imbalance pdfMonday Virtual Room 2 IJCNN Paper SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models II Session Chair: Pei Zhou (SiChuanUniversity), Risheng Xie (University of Science and Technology of China, School of Computer Science and Technology) Balance Accuracy and Efficiency: Segment 3D medical images with inter-slice context information guidance pdfHybrid Frequency--Spatial Attention and Arch-Aware Priors for Tooth Detection and FDI Numbering in Dental Images pdfMonday Virtual Room 3 IJCNN Paper SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models III Session Chair: 小雨 刘 (延边大学), yuxin wang (East China Normal University) Class-Conditional Center Alignment for Robust Vision--Language Contrastive Learning in Glaucoma Screening pdfMonday Virtual Room 4 IJCNN Paper SS25 Integrating Large Language Models and Knowledge Graphs I Session Chair: Xinyue Fan (Qilu University of Technology), Fen Zhao (Nanjing Xiaozhuang University) Enhancing Conversational Question Answering through Reinforced Question Reformulation and Prompt Refinement in Children Application pdfEnhancing Knowledge Graph Question Answering through Classified Path Retrieval and Deductive Reasoning pdfMonday Virtual Room 5 IJCNN Paper SS25 Integrating Large Language Models and Knowledge Graphs II Session Chair: 哲平 于 (天津师范大学), 钰琳 张 (延边大学) ISERA-KGC: Enhancing Knowledge Graph Completion with Interactive Semantic Enhancement and Representation Alignment pdfRetrieval, Reasoning, and Generation: A Universal Format Generation Method for Few-Shot Event Argument Extraction pdfMonday Virtual Room 6 IJCNN Paper SS25 Integrating Large Language Models and Knowledge Graphs III Session Chair: Jinxin Liu (Tsinghua University), Zhongtian Bao (Nankai University) REPANA: Reasoning Path Navigated Program Induction for Transferable Reasoning over Heterogeneous Knowledge Bases pdfKG-Hopper: Empowering Compact Open LLMs with Knowledge Graph Reasoning via Reinforcement Learning pdfEnhancing Factual Consistency in Cross-Lingual Dialogue Summarization via Self-Guidance Prompting pdfMonday Virtual Room 7 IJCNN Paper SS39 Computational Audio Intelligence for Perception & Representation I Session Chair: Ravindrakumar Purohit (Dhirubhai Ambani University), Li Xiang (DongHua University) SpecSlice-ViT: Preserving Spectral Integrity via Slicing Transformers for Underwater Acoustic Target Recognition pdfRUCL: Integrating Regularization and Unsupervised Contrastive Learning for Automatic Speech Recognition pdfMonday Virtual Room 8 IJCNN Paper SS39 Computational Audio Intelligence for Perception & Representation II Session Chair: Junbin Zhang (Peking University), Hui Zhang (Southwest University Of Science And Technology) AudioControl: Efficient Audio-to-Image Generation via Global Latent Modulation and Audio–Visual Alignment pdfMonday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V3 Session Chair: Tam Nguyen (Ho Chi Minh City University of Technology) Integrating Priority Structures into NSGA-II for Interval Data-based Multi-Objective Nonlinear Fixed-Cost Transportation Problem pdfMultiform Surrogate-Assisted Evolutionary Algorithm for Expensive Optimization Problems with Mixed Variables pdfMonday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V4 Session Chair: Ishara Hewa Pathiranage (Adelaide University) On the Use of Evolutionary Optimization for the Dynamic Chance Constrained Open-Pit Mine Scheduling Problem pdfMonday Virtual Room 1 IJCNN Paper SS12 XSTASys: Explainability and Security in Trustworthy Artificial Intelligence Systems I Session Chair: Dakai Zhai (Tsinghua University), Xi Zhong (Institute of Computing Technology,Chinese Academy of Sciences; University of Chinese Academy of Sciences) Monday Virtual Room 2 IJCNN Paper SS12 XSTASys: Explainability and Security in Trustworthy Artificial Intelligence Systems II Session Chair: Guangyu Gong (Shandong University), Song Xu (University of Science and Technology of China) PlanGuard: Defending Agents against Indirect Prompt Injection via Planning-based Consistency Verification pdfCross-Domain Vulnerability Detection using LLMs: Knowledge Transfer from Contemporary Software to AI/ML-Specific Software pdfMonday Virtual Room 3 IJCNN Paper SS29 Graph-Based Solutions for Explainable and Efficient AI I Session Chair: Enguang Zuo (Tsinghua University, School of Intelligence Science and Technology), yanqin luo (yunnan university) MRV-GCN: Multi-Relational Semantic Graph Learning with Explainable Subgraph Fusion for Smart Contract Vulnerability Detection pdfSEGCN-AIKC: Trustworthy Herbal Prescription Recommendation via Knowledge-Guided Verification and Correction pdfMonday Virtual Room 4 IJCNN Paper SS29 Graph-Based Solutions for Explainable and Efficient AI II Session Chair: Nannan Hu (Shandong Normal University), YUNQI HAN (Universiti Putra Malaysia) BDSAM: Fusing Graph Attention Networks and Adapters to Address Domain Imbalance in Cross-Domain Segmentation pdfMonday Virtual Room 5 IJCNN Paper SS32 Deep Neural Networks and Generative AI for Multi-Agent Smart Vehicle Perceptron, Learning, Automation and Optimization I Session Chair: Haoqian Song (Institute of Automation, Chinese Academy of Sciences, Beijing, China; Pengcheng Laboratory, Shenzhen, China), Ya Zhang (Southeast University) SARAD: LLM-Based Safety-Aware Hybrid Reinforcement Learning with Collision Prediction for Autonomous Driving pdfA Point Cloud Completion Network Via The Latent Space-driven Two-stage Noise Synthesis And Restoration Strategy pdfMonday Virtual Room 6 IJCNN Paper SS32 Deep Neural Networks and Generative AI for Multi-Agent Smart Vehicle Perceptron, Learning, Automation and Optimization II Session Chair: Huangnan Zheng (Zhejiang University), Xianchang Wang (Shenyang Aerospace University) AeroGraph: Window-Level Large-Graph Association and Uncertainty-Guided Fusion for Multi-Object Tracking in UAV Videos pdfMonday Virtual Room 7 IJCNN Paper SS38 Deep Learning in Computational Biology and Biomedicine: from Biomedical Data to Drug Discovery I Session Chair: Zhe Liu (Jiangnan University), Luojian Xie (East China Normal University) Learning Class-Consistent Metacell Prototypes via Quantization for Few-Shot Single-Cell Annotation pdfPhysicochemically Informed Dual-Conditioned Generative Model of T-Cell Receptor Variable Regions for Cellular Therapy pdfMonday Virtual Room 8 IJCNN Paper SS38 Deep Learning in Computational Biology and Biomedicine: from Biomedical Data to Drug Discovery II Session Chair: Pengwei Hu (Chinese Academy of Sciences, University of Chinese Academy of Sciences), Zhiyuan Chen (Beijing University of Technology; Institute of Automation, Chinese Academy of Sciences) Primate Spatio-Temporal Relational Network for Unified Behavior Quantification in Cage Environments pdfDM-VMUNet: Dual-Context Perception and Multi-Scale Feature Alignment in Vision Mamba for Medical Image Segmentation pdfMonday Virtual Room 1 IJCNN Paper SS10 Trustworthy and Explainable Federated Learning: Towards Security and Privacy Future I Session Chair: Sen Yu (Yunnan University), Hongzhan Ma (Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences) FedPAL: Adaptive Head Aggregation for Personalized Federated Learning Guided by Generative Feature Distributions pdfMonday Virtual Room 2 IJCNN Paper SS10 Trustworthy and Explainable Federated Learning: Towards Security and Privacy Future II Session Chair: 周 荣博 (Xinjiang University), Manel MILI (Faculty of Sciences of Monastir, Medical Technology and Imaging Laboratory LTIM-LR12ES06) Monday Virtual Room 3 IJCNN Paper SS19 Advances in Trustworthy XAI: Novel Methodologies, Benchmarking, and Diverse Data Modality Contexts I Session Chair: Keyuan Wang (Northwest A&F University ), Liangzhou Qu (Shenzhen University) SwinBERT-Fake: Boosting Multimodal Fake News Detection via Hierarchical Visual Modeling and Explainable AI pdfMonday Virtual Room 4 IJCNN Paper SS19 Advances in Trustworthy XAI: Novel Methodologies, Benchmarking, and Diverse Data Modality Contexts II Session Chair: Mufti Mahmud (King Fahd University of Petroleum and Minerals), Yue Guo (南京航空航天大学) Strat-LLM: Stratified Strategy Alignment for LLM-based Stock Trading with Real-time Multi-Source Signals pdfMonday Virtual Room 5 IJCNN Paper SS37 AI in Healthcare: Harnessing Emerging, Generative and Agentic Technologies for Responsible Innovation Session Chair: Matheus Becali Rocha (Universidade Federal do Espírito Santo, Nature Inspired Computing Laboratory), Qishen Chen (Shanghai University) De-biasing multi-question learning in medical visual question answering via semantic-aware natural direct effect subtraction pdfMonday Virtual Room 6 IJCNN Paper SS01 Privacy-Preserving Machine and Deep Learning Session Chair: Wen Yan (Heilongjiang university), Weigang Wu (Sun Yat-sen University) Joint Optimization of Adaptive Gradient Compression and Aggregation for Efficient Asynchronous Federated Learning pdfNoise Aggregation Analysis Driven by Small-Noise Injection: Efficient Membership Inference for Diffusion Models pdfSMRAM: Boosting Class-specific Face Privacy Protection with Stochastic Mask Regularized Adaptive Momentum pdfMonday Virtual Room 7 IJCNN Paper SS03 Physics-Informed Neural Networks: Advancements and Applications Session Chair: Xiaohui Jia (North University of China), Junqi Qu (Florida State University) TPRformer: A Sea Surface Temperature Prediction Method Based on Temporal Periodic Feature Mining and Reconstruction Transformer pdfMonday Virtual Room 8 IJCNN Paper SS21 Novel Networks in Human-Machine Collaboration: Paradigms, Methods, and Applications Session Chair: Haojie Luo (Fudan University), Lian Guo (Huazhong Agricultural University) Agile Imitation: Efficient Robot Skill Learning via Iterative Human Feedback and Data Aggregation pdfMonday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V5 Session Chair: Sunith Bandaru (University of Skövde) A Non-Reductionist, Homeodynamic Simulator Of Ancient-Medicine, Inspired-Artificial Immune Systems For Emergent Intelligence Analysis pdfA Dual-Layer Clustering Guided Genetic Algorithm for Multiobjective Multitype Satellite Observation Scheduling pdfTuesday Virtual Room 1 IJCNN Paper Graph Neural Networks I Session Chair: Xiao Yue (Oakland University), Qian Tao (South China University of Technology) Tuesday Virtual Room 2 IJCNN Paper Graph Neural Networks II Session Chair: Ningyun Chen (BNBU, Zuse School ELIZA), Zhongming Mei (Donghua university) Cross-Stage Reward-Aware Node Injection Attacks against Graph Neural Networks for Fraud Detection pdfTuesday Virtual Room 3 IJCNN Paper Image Restoration and Enhancement I Session Chair: Zhicheng Qian (Fuyang Normal University), Jinao Li (Qilu University of Technology) DCAAFusion: A Diffusion Model-Based Cross-Attention Adaptive Fusion Network for Nighttime Infrared and Visible Image Fusion pdfCDL-FusionNet: A Two-Stage Illumination-Invariant and Luminance-Guided Framework for Infrared and Visible Image Fusion pdfTuesday Virtual Room 4 IJCNN Paper Image Restoration and Enhancement II Session Chair: Chuancheng Fu (Wuhan University of Science and Technology), Xuanchao Lin (Shanghai University) PhyIC-Net: A Robust Unsupervised Framework for Underwater Image Restoration via Intrinsic Consistency pdfTuesday Virtual Room 5 IJCNN Paper Image Segmentation and Dense Prediction I Session Chair: ziyang Tong (Wuhan University of Technology), Xiaohong Jia (Lanzhou Jiaotong University) MEI: Mutual-Enhanced Integration Between Pretrained SAM and Lightweight Model for Medical Image Segmentation pdfTuesday Virtual Room 6 IJCNN Paper Knowledge Graphs and Question Answering I Session Chair: JunWei Yang (East China Normal University), Yi Zhou (Southwest University of Science and Technology) SAMEF-DHNN: Sample-Aware Multimodal Expert Fusion and Dynamic Hard Negatives Network for Multimodal Knowledge Graph Completion pdfQuery-Aware Hierarchical Contrastive Path Learning for Inductive Temporal Knowledge Graph Reasoning pdfTuesday Virtual Room 7 IJCNN Paper LLM Adaptation and Fine-Tuning I Session Chair: Yichen Liu (CASIA), Wentao Hu (University of Science and Technology of China) Tuesday Virtual Room 8 IJCNN Paper LLM Adaptation and Fine-Tuning II Session Chair: Md Zarif Hossain (Florida Atlantic University), 杉杉 陈 (沈阳航空航天大学) Sim-CLIP: Unsupervised Siamese Adversarial Fine-Tuning for Robust and Semantically-Rich Vision-Language Models pdfPre-PEFT Probing: Weight Statistics and Perturbation Robustness for Layer Selection in VLM Vision Encoders pdfBound the Risk, Gate the Bias: A Mechanistic-Driven Sim-to-Real Transfer Method for Rebust Yield Prediction pdfTuesday Virtual Room 1 IJCNN Paper Image Segmentation and Dense Prediction V Session Chair: Shaoguo Cui (Chongqing Normal University), guoqiang zheng (西南科技大学) A Morphology-Aware Deformable Hybrid CNN-Mamba Network for Nuclei Instance Segmentation and Classification pdfSGEA-Net: Stimulus-Guided Gating with Multi-Scale Edge-Enhanced Aggregation for Retinal Vessel Segmentation pdfTuesday Virtual Room 2 IJCNN Paper Image Segmentation and Dense Prediction VI Session Chair: Jin-Chun Piao (Yanbian University), Xiaopeng Liu (Shandong University of Science and Technology) ULSNet: Underwater-Aware Preference–Value Attention with a Boundary-Refined Head for Underwater Instance Segmentation pdfTuesday Virtual Room 3 IJCNN Paper Image Segmentation and Dense Prediction VII Session Chair: zhiyu xiao (Beijing Information Science and Technology University), Yingqi Liang (Guangxi University) FSECrossNet: Frequency–Spatial Enhancement with Cross-Modal Attention for Remote Sensing Semantic Segmentation pdfIDC-YOLO: Efficient Visible-Infrared Object Detection via Illumination-Guided Modulation and Difference-Aware Interaction pdfTuesday Virtual Room 4 IJCNN Paper Image Segmentation and Dense Prediction VIII Session Chair: Chaoli Wang (上海理工大学), Yan Wan (DongHua University) Dual-Domain Feature Enhancement and Boundary-Guided Multi-Branch Attention Network for Polyp Segmentation in colonoscopy images pdfTuesday Virtual Room 5 IJCNN Paper Information Extraction and Text Mining I Session Chair: Boyan Xu (Guangdong University of Technology), Zhiyang Yu (Shandong Normal University) Tuesday Virtual Room 6 IJCNN Paper Information Extraction and Text Mining II Session Chair: Weixin Zuo (Shanghai University of Electric Power, Faculty of Artificial Intelligence), yifan huo (Zhejiang Sci-Tech University) Tuesday Virtual Room 7 IJCNN Paper Information Extraction and Text Mining III Session Chair: Shiao Meng (Tsinghua University), Yuan Gao (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) I2DRE: Intra- and Inter-Pair Interactions based Document-Level Relation Extraction with Binary Hill Loss pdfTuesday Virtual Room 8 IJCNN Paper Information Extraction and Text Mining IV Session Chair: Hao Zhang (School of Computer Science and Engineering, Northeastern University, Shenyang 110819, China), Jipeng Guo (Beijing University of Chemical Technology) Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V6 Session Chair: yilin fang (Wuhan University of Technology) Decoupling Numerical and Structural Parameters: An Empirical Study on Adaptive Genetic Algorithms via Deep Reinforcement Learning for the Large-Scale TSP pdfTKG-DG-DMOEA: Temporal Knowledge Graph-Guided Domain Generalization for Dynamic Multi-Objective Disassembly Line Balancing pdfTuesday Virtual Room 1 IJCNN Paper Knowledge Graphs and Question Answering II Session Chair: Jibing Wu (National University of Defense Technology), Nurmemet Yolwas (Xinjiang University) Few-Shot Temporal Knowledge Graph Completion via Frequency-Based Relation Encoding and Conditional Diffusion pdfDual Quaternion Decoupling and Enhanced Temporal Embeddings for Temporal Knowledge Graph Completion pdfTuesday Virtual Room 2 IJCNN Paper Knowledge Graphs and Question Answering III Session Chair: Xiangfeng Luo (Shanghai University, School of Computer Engineering and Science), Zheng Lin (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Beyond Answer Symbols: Discovering Key Layers of VLMs for Visual Processing in Multiple Choice Question Answering pdfTuesday Virtual Room 3 IJCNN Paper Knowledge Graphs and Question Answering IV Session Chair: Mian Wu (School of Software, Beihang University), wen Zhang (Zhejiang University) MedLTRAG: Co-Augmentation of Knowledge Graphs and Large Language Models for Long-tail Medical Question Answering pdfLLM-SPM:Semantic-Driven Temporal Knowledge Graph Forecasting via Large Language Models and Pattern Mining pdfEnhancing GNN-RAG with Adaptive Subgraph Optimization and Bidirectional Feedback for Robust Knowledge Graph Question Answering pdfTuesday Virtual Room 4 IJCNN Paper Knowledge Graphs and Question Answering V Session Chair: Jiebin Huang (South China University of Technology), yining liu (shandong university) Tuesday Virtual Room 5 IJCNN Paper Knowledge Graphs and Question Answering VI Session Chair: Ziqiong Liu (Southern University of Science and Technology), Yang Liu (Shanxi University) DICE: A Dynamic Subgraph and Interaction–Type Collaborative Enhancement Framework for One-shot Subgraph Reasoning pdfMTAM: Multi-Level Trustworthiness Assessment via Neurosymbolic Synergy for Knowledge Graph Error Detection pdfTuesday Virtual Room 6 IJCNN Paper LLM Adaptation and Fine-Tuning V Session Chair: Wenge Rong (School of Computer Science and Engineering, Beihang University, China; Engineering Research Center of Integration and Application of Digital Learning Technology Ministry of Education,China), baojun tian (Inner Mongolia University of Technology, Inner Mongolia Key Laboratory of Intelligent Perception and System Engineering) HDAF-LLM: Hierarchical Decoupled Adaptive Fusion with Large Language Models for Multi-Domain Recommendation pdfD-Con: Domain-Partitioned Contrastive Learning Framework for Mitigating Domain Interference in Multi-Domain Text Mining pdfTuesday Virtual Room 7 IJCNN Paper LLM Adaptation and Fine-Tuning VI Session Chair: Yixian Kong (Beijing University of Posts and Telecommunications), Shayok Chakraborty (Florida State University) Tuesday Virtual Room 8 IJCNN Paper LLM Adaptation and Fine-Tuning VII Session Chair: Jianwei Xu (Sichuan University), Yicheng Pan (Hangzhou Institute for Advanced Study,UCAS; University of Chinese Academy of Sciences) Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V7 Session Chair: Zhenan He (Sichuan University) Geometric Modeling and Global Optimization for Plant Line Identification in Agricultural Orthomosaics pdfTuesday Virtual Room 1 IJCNN Paper LLM Adaptation and Fine-Tuning VIII Session Chair: zhichao zhang (Northeast Electric Power University), Lindong Wang (Shanghai Jiao Tong University) CREF-LLM: Contrastive Representation Encoding with Frozen Language Models for Few-Shot Wind Power Forecasting pdfTuesday Virtual Room 2 IJCNN Paper LLM Agents and Tool Use V Session Chair: Yang Han (Institute of Software, Chinese Academy of Sciences; Hangzhou Institute for Advanced Study, University of Chinese Academy of Sciences), Yue Liu (Shanghai University) Tuesday Virtual Room 3 IJCNN Paper LLM Agents and Tool Use VI Session Chair: Yue Liu (Shanghai University), Zhilin Zhang (Chongqing Institute of Green and Intelligent Technology,Chinese Academy of Sciences; Chongqing School,University of Chinese Academy Sciences) Task-Tree A-Star: A Hierarchical Framework for Long-Horizon Task Planning with Large Language Models pdfTuesday Virtual Room 4 IJCNN Paper LLM Agents and Tool Use VII Session Chair: Bijia Liu (Independent Researcher), Nirali Sanghvi (Indian Institute of Technology, Roorkee) EEVEE: Ensemble Expertise from Visual-Temporal-Symbolic Embeddings Enhancing Financial Time Series Forecasting pdfTuesday Virtual Room 5 IJCNN Paper LLM Agents and Tool Use VIII Session Chair: Weizhi Kong (Harbin Institute of Technology, Shenzhen), Yinong Chi (Qilu University of Technology; Shandong Provincial Key Laboratory of Industrial Network and Information System Security , Shandong Fundamental Research Center for Computer Science) Tuesday Virtual Room 6 IJCNN Paper LLM Evaluation and Benchmarking III Session Chair: Ponnurangam Kumaraguru (IIIT Hyderabad), Xichen Lin (Nanjing University) Tuesday Virtual Room 7 IJCNN Paper LLM Evaluation and Benchmarking IV Session Chair: Xinhai Chen (National University of Defense Technology), Nan Wang (Lenovo) Tuesday Virtual Room 8 IJCNN Paper LLM Evaluation and Benchmarking V Session Chair: Jie Zhou (Changsha University of Science and Technology, Hunan Provincial Key Laboratory of Mathematical Modeling and Analysis in Engineering), Ruiting Dai (University of Electronic Science and Technology of China) Adaptive State Space Experts: Dynamic Expert Selection with State-Aware Routing for Efficient Sequence Modeling pdfTuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V8 Session Chair: Li Cao (China University of Geosciences) Cost-Efficient Power Orchestration in Heterogeneous EV-Integrated VPPs: A Peak Demand Repair Scheme pdfTE Model: A Transformer Encoder Only Based Model for Prefetch and Replacement in Solid-state Drivers pdfTuesday Virtual Room 1 IJCNN Paper LLM Evaluation and Benchmarking VI Session Chair: Wenpeng Lu (Qilu University of Technology, Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing; Shandong Fundamental Research Center for Computer Science), Zijiao Zhang (Zhengzhou University) LCIC: A Unified Hallucination Detection-and-Mitigation Method Using LLM’s Tail-Pooled Internal Signals pdfTuesday Virtual Room 2 IJCNN Paper LLM Reasoning and Planning III Session Chair: Linna Zhou (Beijing University of Posts and Telecommunications), Jaeeun Jang (Hanwha Systems) Tuesday Virtual Room 3 IJCNN Paper LLM Reasoning and Planning IV Session Chair: Zhiwei Yu (Tsinghua University), Rui Yan (Zhejiang University of Technology) Learning to Reason across Viewpoints: Multi-Viewpoint Spatial Reasoning for Multimodal Large Language Models pdfTuesday Virtual Room 4 IJCNN Paper LLM Reasoning and Planning V Session Chair: Yujia Huo (School of Data Science and Information Engineering, Guizhou Minzu University, China; Guizhou Provincial Key Laboratory of Applied Mathematics and Computing Power & Algorithms, Guizhou, China), wen Zhang (Zhejiang University) Tuesday Virtual Room 5 IJCNN Paper LLM Reasoning and Planning VI Session Chair: Zhe Cui (Beijing University Of Posts and Telecommunications, Beijing Key Laboratory of Network System and Network Culture), a a (none) Tuesday Virtual Room 6 IJCNN Paper Learning Paradigms and Model Efficiency IV Session Chair: Ashish Anand (Indian Institute of Technology Guwahati), Yan Wan (DongHua University) Tuesday Virtual Room 7 IJCNN Paper Learning Paradigms and Model Efficiency V Session Chair: Xufei Zhang (Beijing XingYun Digital Technology Co., Ltd), Cong Hu (Jiangnan University) Interpretable Logical Anomaly Classification via Constraint Decomposition and Instruction Fine-Tuning pdfAnchor Then Align: Anomaly-Aware CLIP Adaptation with Disentangled Text Anchors for Weakly-Supervised VAD pdfTuesday Virtual Room 8 IJCNN Paper Learning Paradigms and Model Efficiency VI Session Chair: Guofeng Zhang (Zhejiang University), Yulin Hu (Southwest University) Tuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V9 Session Chair: Ru Lei (Northwestern Polytechnical University) MODRAL: Two-Phase PPO-Guided ALNS for Multi-Center Home Healthcare Routing with Time-Window Satisfaction pdfTuesday Virtual Room 1 IJCNN Paper Learning Paradigms and Model Efficiency VII Session Chair: Wenge Rong (School of Computer Science and Engineering, Beihang University, China; Engineering Research Center of Integration and Application of Digital Learning Technology Ministry of Education,China), Jia LIU (Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences) Controlling Mutual Information for Generative Few-Shot Data Augmentation via Parametric Distribution Construction pdfTuesday Virtual Room 2 IJCNN Paper Learning with Noisy and Limited Labels I Session Chair: Baixin Li (Northeastern University), Shuohao Li (National University of Defence Technology) Tuesday Virtual Room 3 IJCNN Paper Learning with Noisy and Limited Labels II Session Chair: Jinao Li (Qilu University of Technology), Fei Ding (Peking University) Semi-supervised Label Distribution Learning from Incomplete Annotations via Multi-view Fusion and Reliability-aware Filtering pdfTuesday Virtual Room 4 IJCNN Paper Machine Learning Methods and Applications III Session Chair: Jianbin Jiao (University of Chinese Academy of Sciences), Keji Mao (Zhejiang University of Technology) Deep Covariance Denoising and Attention-Guided On-Grid Classification: A Dual-Stage Network for Robust Underwater DOA Estimation pdfTuesday Virtual Room 5 IJCNN Paper Machine Learning Methods and Applications IV Session Chair: Zhi Wang (Southwest University), SS L (South China University of Technology) Adaptive Dynamic Convolution Module Based Low-Rank Tensor Factorization Framework for Multi-Dimensional Image Recovery pdfTuesday Virtual Room 6 IJCNN Paper Machine Learning Methods and Applications V Session Chair: Zhiming Wang (Imperial College London), Vivek B S (Tata Consultancy Services) Image-set classification using Discriminant Neighborhood Preserving Embedding on Symmetric Positive Definite Manifold pdfTuesday Virtual Room 7 IJCNN Paper Machine Learning Methods and Applications VI Session Chair: Jungang Xu (University of Chinese Academy of Sciences), Min Li (Institute of Information Engineering, Chinese Academy of Sciences,State Key Laboratory of Cyberspace Security Defense; School of Cyber Security, University of Chinese Academy of Sciences) Tuesday Virtual Room 8 IJCNN Paper Machine Learning Methods and Applications VII Session Chair: Dan Qu (Information Engineering University), Zhikui Chen (Dalian University of Technology) CVCAM-K: A Specific Emitter Identification Method with Complex-Valued Convolutional Attention Mamba and KNN pdfPhysio-xLSTM: Unraveling Hemodynamic Dynamics via Dual-Domain KAN-xLSTM Networks for Robust PPG Biometrics pdfA Dynamic-Length Single-Step Sampling Network for Adaptive Inference in Efficient Radio Frequency Fingerprint Identification pdfTuesday Virtual Room 9 IEEE CEC (Evolutionary Computation) CEC V10 Session Chair: Jing-Yu JI (Hong Kong) Tuesday 0.01 London FUZZ-IEEE Paper FUZZ 4: Hybrid systems of computational intelligence techniques Session Chair: Chang-Shing Lee (National University of Tainan) Fuzzy Reliability Redundancy Allocation Problem Using Immigrants-based Multifactorial Evolutionary Algorithm pdfTuesday 0.04 Brussels IJCNN Paper IJCNN SS05 Artificial Intelligence in Healthcare: Leveraging Transformer Models Session Chair: Thorben Markmann (Bielefeld University), Valerie Vaquet (Bielefeld University) Hmsanet: a Hierarchical Multi-scale Attention Network for Precise Retinal Layer Segmentation in Covid-19 Oct Images pdfRG-DermNet: A Multimodal Attention-Based Model with Residual Block Usage for Skin Lesion Classification pdfTuesday 0.05 Paris IJCNN Paper IJCNN SS19 Advances in Trustworthy XAI: Novel Methodologies, Benchmarking, and Diverse Data Modality Contexts Session Chair: Imen Jdey (REGIM Lab, Sfax University), M. Tanveer (Indian Institute of Technology Indore, India) Tuesday 0.10 Sydney IJCNN Paper IJCNN SS04 Tiny Machine Learning Session Chair: Massimo Pavan (Politecnico di Milano) Tiny CNNs or Tiny Transformers? A Systematic Analysis of Latency, Accuracy and Domain Shift Robustness in Extreme-Edge Image Classification pdfTuesday 0.11 Cape Town IJCNN Paper Causal, Probabilistic, and Uncertainty-Aware Learning Session Chair: Varun Sampath Kumar (University of Southern Denmark), Pranab K. Muhuri (South Asian University) Uncertainty Quantification in Radioactivity Level Estimation via Heteroscedastic Count Regression pdfTuesday 0.15 Washington IJCNN Paper IJCNN SS11 Engineering Trust: Ethical, Legal, and Societal Impacts of Computational Intelligence on Human Agency Session Chair: Keeley Crockett (Manchester Metropolitan University, Dalton Building), Robert G. Reynolds (Wayne State University), Tayo Obafemi-Ajayi (Missouri State University, Missouri, USA) From Participants to Co‑Researchers: How AI Literacy Builds Community Trust and Power in Participatory AI Design pdfTuesday 2.18 Mekong IJCNN Paper Human Activity Recognition and Motion Understanding Session Chair: Siyuan Yang (KTH Royal Institute of Technology), Gaurvi Goyal (Maastricht University) Motion-Adaptive Multi-Scale Temporal Modelling with Skeleton-Constrained Spatial Graphs for Efficient 3D Human Pose Estimation pdfBeyond Addition: HDC Binding for Transformers’ Position Encoding applied to Time Series Classification pdfTuesday 2.1 Volga IJCNN Paper IJCNN SS22 Self-organizing Clustering for Continual Learning and its Applications Session Chair: Naoki Masuyama (Osaka Metropolitan University), Yuichiro Toda (Okayama University) Cluster Merging in Adaptive Resonance Theory-based Clustering Guided by Mass and Distance Criteria pdfTuesday 0.01 London IEEE CEC (Evolutionary Computation) CEC 9 - Algorithms II Session Chair: Kalyanmoy Deb (MIchigan State University) Modified GUESS Approach Using Machine Decision-makers for an Efficient Interactive Multi-criterion Decision-Making Procedure pdfAn Evolutionary Method for Joint Optimization of UAV Path and Camera Direction Planning with Optional Viewpoints pdfProjected Hypercube Sampling with Parallel Reference Vectors for Uniform Coverage of Irregular Pareto-Optimal Fronts pdfTuesday 0.02 Berlin IEEE CEC (Evolutionary Computation) CEC 10 - Evolutionary Machine Learning II Session Chair: Efrén Mezura-Montes (University of Veracruz) MODE-based Oblique Decision Tree Induction: Addressing the Accuracy-Complexity Trade-off through a Normalized Complexity Measure pdfTuesday 0.04 Brussels IJCNN Paper Neural Learning and Optimization II Session Chair: Benjamin Redden (Queen's University Belfast), Fabian Hinder (Bielefeld University) Tuesday 0.05 Paris IJCNN Paper Explainable, Fair, and Responsible AI Session Chair: Monowar Bhuyan (Umeå University), M. Tanveer (Indian Institute of Technology Indore, India) FairBound: Boundary-Preserving Fair Downsampling for Accurate and Equitable Imbalanced Classification pdfA Comparative Study of Improved Disparate Impact Remover and Fair Adversarial Learning for Bias Mitigation pdfTuesday 0.10 Sydney IJCNN Paper Edge, Federated, and Privacy-Preserving Learning Session Chair: Run Wang (ETH Zurich), Niklas Melton (Missouri University of Science and Technology) InteFL: Framework for AI-Assisted Design of Trustworthy and Efficient Federated Learning Applications pdfKoalaMamba: A Federated Method with Vision Mamba UNet for Privacy-Preserving Medical Segmentation pdfTuesday 0.11 Cape Town IJCNN Paper Graph and Relational Learning Session Chair: DANIELE ZAMBON (Università della Svizzera italiana), Luca Virgili (Marche Polytechnic University) A self-supervised joint embedding predictive approach for bipartite graphs to solve linear programs pdfGravSpec: Spectral-Gravitational flow on Density Manifolds for Deterministic Overlapping Community Detection pdfTuesday 0.15 Washington IJCNN Paper Trustworthy and Safe Language Models Session Chair: Robert G. Reynolds (Wayne State University), Ihsen Alouani (Queen's University Belfast, Upper Bound) ECRT: Evidence-Centric Cross-Modal Reasoning with Trust-Aware Gating for Multimodal Fake News Detection pdfTuesday 2.1 Volga IJCNN Paper Generative Models and Visual Synthesis Session Chair: Van Huyen Dang (Paderborn University), Dawid Połap (Silesian University of Technology, Poland) An Analysis of Regularization and Fokker–Planck Residuals in Diffusion Models for Image Generation pdfTuesday Brightlands Foyer IJCNN Paper, FUZZ-IEEE Position Paper, CEC Late Breaking Paper, CEC Paper, FUZZ J2C Presentation, CEC J2C Presentation, FUZZ-IEEE Paper, CEC Position Paper, IJCNN J2C Presentation, IJCNN Position Paper, IJCNN Late Breaking Paper, FUZZ-IEEE Late Breaking Paper LBR Posters IJCNN / CEC / FUZZ Session Chair: Maximilian Krentzien (University of Rostock, Institute of Applied Microelectronics and Computer Engineering), Mihail Popescu (University of Missouri) Breaking the Cycle: Stopping Overthinking in Language Models via Cyclic Reasoning Patterns Whova Tag: Poster Presentation Electrical AI Copilot – A Framework for Empowering Power System Engineers Whova Tag: Poster Presentation On the Consistency of Graph Structure Learning in Spatiotemporal Modeling Whova Tag: Poster Presentation Grid-Discretized Multi-Task Transformer for Joint Tropical Cyclone Path and Intensity Forecasting Whova Tag: Poster Presentation Automated Scoliosis Assessment via Dual-Stage YOLOv8 Detection and Attention U-Net Segmentation Whova Tag: Poster Presentation An Evidence-based Arbitration Framework for Zero-shot Medical Image Segmentation Whova Tag: Poster Presentation Successful Adaptation of a Group EEG Generative Model to Individuals Depends on Participant Typicality Whova Tag: Poster Presentation Physiological Time Series Prediction Model Based on Multivariate State Graphs Whova Tag: Poster Presentation Resource Constrained Software Development Paradigm Shift to Neural-Network-Based Energy and Runtime Estimation of Software Execution Whova Tag: Poster Presentation Human-Robot Collaborative Manipulation via Switchable Teleoperation and Autonomous Grasping Whova Tag: Poster Presentation Tuesday 0.01 London FUZZ-IEEE Paper FUZZ 5: FUZZ-IEEE SS03 Information fusion techniques based on aggregation functions, preaggregation functions and their generalizations & Main:Mathematical and theoretical foundations of fuzzy sets, measures and integrals Session Chair: Giancarlo Lucca (Universidade Federal de Pelotas), Graçaliz Dimuro (Universidade Federal do Rio Grande) Power Choquet-Inspired Integral: A Family of Tunable Aggregation Functions with Application to Fuzzy Rule-Based Classification Systems pdfGeneralizing the Antecedent Layer of ANFIS: A Performance Analysis Under Classes of Fuzzy Aggregators pdfTuesday 0.02 Berlin FUZZ-IEEE Paper FUZZ 6 : FUZZ-IEEE SS04 Fuzzy Foundation Models & Main:Fuzzy system Session Chair: Javier Andreu-Perez (University of Essex), Hani Hagras (university of essex) A Neuro-Symbolic System for Interpretable Multimodal Physiological Signals Integration in Human Fatigue Detection pdfTuesday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 11 - Optimization II Session Chair: Frank Neumann (Adelaide University) TC-MFEA: Task Clone-Based Multifactorial Evolutionary Algorithm for Constrained Reliability Redundancy Allocation Problem pdfEvolutionary Optimization of Hospital Bed Allocation to Reduce Mortality from Severe Acute Respiratory Syndrome pdfTuesday 0.10 Sydney IEEE CEC (Evolutionary Computation) CEC 12 - Related Topics II Session Chair: Andy Tyrrell (University of York) Spark-Enabled Binary Particle Swarm Feature Selection Algorithm Evaluated on Genomic Breast Cancer Data pdfTuesday 0.11 Cape Town IEEE CEC (Evolutionary Computation) CEC 13 - Algorithms III Session Chair: Diederick Vermetten (Sorbonne Université, CNRS, LIP6) Turning Parent Degradation into an Advantage: Archive-Directed Mating for Archive-Based Evolutionary Multi-Objective Optimization pdfTuesday 0.15 Washington IEEE CEC (Evolutionary Computation) CEC 14 - Evolutionary Machine Learning III Session Chair: Keiki Takadama (The University of Tokyo) Improving Mutation-Based Evolving Artificial Neural Network via Population Partitioning for Topology and Parameter Mutations pdfTuesday 2.1 Volga IEEE CEC (Evolutionary Computation) CEC 15 - SS17:Computational Intelligence in Power Electrical Engineering Session Chair: Maurizio Repetto (Politecnico di Torino) Systematic Evaluation of Spatial Interpolation Methods for Reconstructing Missing Solar Irradiance Data for Energy Applications pdfGEvTO: A Hybrid Evolutionary Scheme for Large-Scale Multi-objective Binary Topology Optimization of Electromagnetic Devices pdfTuesday Virtual Room 1 IJCNN Paper SS26 Brain Machine Intelligence: Models, Systems, and Translational Applications Session Chair: Dongrui Wu (Huazhong University of Science and Technology), Zhao Boshi (Beijing university of technology; The Center for Excellence in Brain Science and Intelligence Technology, Chinese Academy of Sciences) Unsupervised Alignment using Latent Dynamics in Spiking Neural Networks for Stable and Energy-Efficient iBCI Decoding pdfAn End-To-End Intelligent EEG Bad Channel Detection and Reconstruction System Using Deep Learning and Reservoir Computing pdfTuesday Virtual Room 2 IJCNN Paper SS02 Advances in Machine Learning and Deep Learning for Hyperspectral Image Classification Session Chair: Hongmeng Lu (Xinjiang University), Yuchuan Chen (Changsha University of Science and Technology, School of Computer Science and Technology) CA-DEMAE: Content-Aware Diffusion-Based Masked Autoencoder for Hyperspectral Image Classification pdfDSFMNet: A Dual-Scale Frequency-Modulated Network With a Linear-Complexity Hybrid Architecture for Hyperspectral Image Classification pdfTuesday Virtual Room 3 IJCNN Paper SS20 Deep Edge Intelligence Session Chair: Xujiang Tang (guilin university of electronic technology), Hui Zhang (Tianjin University of Science and Technology) Efficient Pest Detection via Low-Rank Additive Attention and Scale-Dynamic Loss for Deep Edge Intelligence in Smart Agriculture pdfPrismKV: Harmonizing Heterogeneous Compression Strategies through Layer-adaptive Partition Scoring pdfTuesday Virtual Room 4 IJCNN Paper SS30 Computational Intelligence and AI Applications for Sustainable Energy Management in Smart Grids and Energy Communities (2nd ed.) Session Chair: Xiaolu Xu (闽南师范大学), Yuan Qiu (SIT) A Novel Bi-Level Framework for Electric Vehicle Charging Scheduling Using Reinforcement Learning and Transfer Learning pdfTuesday Virtual Room 5 IJCNN Paper SS33 AICS: Artificial Intelligence for Complex Systems Session Chair: zeyu zhang (Inner Mongolia Normal University), Peiwen He (South China Normal University) DualCrossAD: Multivariate Time Series Anomaly Detection with Homogeneous-Heterogeneous Dual-Cross Attention pdfTwo-Level Routed Sparse Mixture-of-Experts with Spatiotemporal Self-Attention for Next POI Recommendation pdfTuesday Virtual Room 6 IJCNN Paper SS34 Safety integration in neural networks for sensor systems in critical applications Session Chair: Miaoji Zheng (Foshan University), Yongtao Zhang (NPIC) RT-CAM Enhanced: A Comparative Study of Multi-Task Brain Tumor Classification Using Explainable Attention Mechanisms pdfPhysDiff: Physics-Guided Diffusion for Manifold-Consistent Data Augmentation in Driving Style Classification pdfTuesday Virtual Room 7 IJCNN Paper SS01 Privacy-Preserving Machine and Deep Learning / SS07 Quantum Machine Learning Algorithms and Applications Session Chair: Wenxia Wang (Information Engineering University), Peiying Xu (Institute of Information Engineering, CAS; School of Cyber Security, University of Chinese Academy of Sciences) Tuesday Virtual Room 8 IJCNN Paper SS03 Physics-Informed Neural Networks: Advancements and Applications / SS08 Automating Model Discovery: Neural Architecture Search in the Era of Large Machine Learning Models Session Chair: Jiyan Qiu (Computer Network Information Center, Chinese Academy of Sciences; University of Chinese Academy of Sciences), Yongwei Tang (Liaoning Technical University) UDCP: Uncertainty Driven Context-prompted nnUNet for Tertiary Lymphoid Structure Segmentation in H&E Images pdfSC-LoRA: Balancing Efficient Fine-tuning and Knowledge Preservation via Subspace-Constrained LoRA pdfLUMF: A Deep Learning Based Accurate and Robust Liver Uptake Measurement Framework for SUV Evaluation in Whole-Body PET/CT Images pdfTuesday Virtual Room 1 IJCNN Paper SS09 Multimodal Deep Learning in Applications / SS21 Novel Networks in Human-Machine Collaboration: Paradigms, Methods, and Applications Session Chair: Anming Dong (Qilu University of Technology (Shandong Academy of Sciences); Shandong Provincial Key Laboratory of Industrial Network and Information System Security, \\ Shandong Fundamental Research Center for Computer Science, Jinan, China), Xiang Li (Wuhan University) Tuesday Virtual Room 2 IJCNN Paper SS13 Neural Networks for Biodiversity / SS24 Advances in Hyperdimensional Computing and Vector Symbolic Architectures Session Chair: jinghao wen (Villanova University), Gabriel Dubus (Muséum National d'Histoire Naturelle; Sebitoli Chimpanzee Project, Great Ape Conservation Project, Sebitoli, Kibale National Park, Fort Portal, Uganda) DeepForestSound: a multi-species automatic detector for passive acoustic monitoring in African tropical forests, a case study in Kibale National Park pdfGraph-Reasoning with Disentangled Attention: A Knowledge-Guided Approach for Fine-Grained Visual Classification pdfTuesday Virtual Room 3 IJCNN Paper SS17 Generative Foundation Models for Robotics: From Language and Vision to Embodied Action / SS26 Brain Machine Intelligence: Models, Systems, and Translational Applications Session Chair: Zhiguo Zhang (Harbin Institute of Technology, Shenzhen, China), Yinfeng Yu (Xinjiang University) Generalizable Audio-Visual Navigation via Binaural Difference Attention and Action Transition Prediction pdfHierarchical Neuro-Semantic Control for UAV Swarm Formation: Bridging LLM Planning and Hard Safety Constraints pdfTuesday Virtual Room 4 IJCNN Paper SS23 NeuroSNN: Neuromorphic Computing and Spiking Neural Networks / SS29 Graph-Based Solutions for Explainable and Efficient AI Session Chair: jia jian (河北工程大学), Jinsheng Xiao (Wuhan University) Tuesday Virtual Room 5 IJCNN Paper SS32 Deep Neural Networks and Generative AI for Multi-Agent Smart Vehicle Perceptron, Learning, Automation and Optimization / SS39 Computational Audio Intelligence for Perception & Representation Session Chair: Zejian Zhou (University of Wyoming), Qing Tian (University of Alabama at Birmingham) Tuesday Virtual Room 6 IJCNN Paper SS36 Human-Centered AI: Multi-modal Agent-based Systems (HCAI) / SS41 Randomization Based Deep and Shallow Learning Algorithms and/or Biomedical Applications Session Chair: zhao wang (Xidian University), Jingyi Lyu (Beijing Normal-Hong Kong Baptist University) Medical Image Classification via Integrated Attention Mechanisms and Dynamic Weighted Loss Function pdfA GAN-Based MRI Denoising Model for Alzheimer’s Disease Using Free Attention and Wasserstein Loss pdfTuesday Virtual Room 7 IJCNN Paper SS38 Deep Learning in Computational Biology and Biomedicine: from Biomedical Data to Drug Discovery / SS48 AI, Law and Regulation Session Chair: Jing Peng (Wuhan University of Technology), Raphaella Revis (UTS) Tuesday Virtual Room 8 IJCNN Paper SS11 Engineering Trust: Ethical, Legal, and Societal Impacts of Computational Intelligence on Human Agency / SS12 XSTASys: Explainability and Security in Trustworthy Artificial Intelligence Systems Session Chair: Ravi Kumar (Florida International University, Rowan University), Ali Karkehabadi (University of California, Davis) Tuesday Virtual Room 1 IJCNN Paper SS14 Neuroevolutionary computation techniques for medical data understanding and pattern analysis / SS15 Agentic Edge Intelligence for the Internet of Things Session Chair: Xingpeng Zhang (Southwest Petroleum University), Ruiqin Bai (College of Artificial Intelligence, Taiyuan University of Technology, Taiyuan, 030024, P.R. China; Postdoctoral Workstation, China Railway 17th Bureau Group Co., Ltd., Taiyuan, 030006, P.R. China This work was supported by the Shanxi Provincial Young Scientists Research Project (202303021212081)) DSIB: Towards Robust Multimodal Survival Prediction via Discrete Self-Distillation Information Bottleneck pdfMMO-GFlowNet: Multi-objective Metahuristic Optimization of GFlowNet Parameters for Molecular Generation pdfA Federated Learning Backdoor Attack Defense Method Based on Frequency Domain Perturbations and Model Repair pdfTuesday Virtual Room 2 IJCNN Paper SS16 Bayesian Neural Networks / SS18 Synergizing Multimodal LLMs, Co-Agent Architectures, and Foundation Models in Manufacturing AI Session Chair: Yun Liu (Sichuan University), Chao Ma (Institute of Information Engineering, CAS; School of Cyber Security,University of Chinese Academy of Sciences) AML-BGNN:Advanced Multi-Level Bayesian Graph Neural Network Framework for Robust Network Security Analysis pdfMCIR-RAG: Multi-Agent Chain-of-Thought Iterative Reflection for Multi-modal Document Understanding pdfTuesday Virtual Room 3 IJCNN Paper SS22 Self-organizing Clustering for Continual Learning and its Applications / SS27 Reservoir Computing for Scalable and Energy-Efficient AI: Theory, Dynamics, and Implementations Session Chair: Shenghao Lyu (Kyoto University), Yi Gong (University Of Electronic Science And Technology Of China) Tuesday Virtual Room 4 IJCNN Paper SS37 AI in Healthcare: Harnessing Emerging, Generative and Agentic Technologies for Responsible Innovation / SS45 Computational Intelligence in Transactive Energy Management and Smart Energy Network (CITESEN) Session Chair: Dongjiao Ge (City University of Macau), Zhiqiang He (Key Laboratory of Universal Wireless Communications, Ministry of Education, Beijing University of Posts and Telecommunications) A Morphology-Aware Contrastive Learning Network for Detecting a Potential Subtype of Knee Osteoarthritis pdfStructDiffSeg: A Structure-Aware Diffusion Framework for Anatomically Consistent Medical Image Segmentation pdfEnhancing Asset Visibility in Smart Energy Systems: A Computational Intelligence Framework for Heterogeneous Infrastructure Inspection pdfTuesday Virtual Room 5 IJCNN Paper SS40 Design Challenges, Methods and Applications in Sustainable Neuromorphic Hardware / SS42 Games / SS46 Computationally Intelligent Techniques in Early Prediction and Detection of Brain Disorders Session Chair: Wannian Xia (School of artificial intelligence, University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences), Tao Hong (Shenyang Aerospace University) A Heterogeneous Feature Fusion Network for EEG Seizure Detection Integrating Spatio-Temporal Channel Networks and Vision Transformer pdfTuesday Virtual Room 6 IJCNN Paper Advances in Computer Vision I Session Chair: Liang Fan (Loughborough University, ai.io), Wei Shi (Hangzhou Dianzi University) Tuesday Virtual Room 7 IJCNN Paper Advances in Computer Vision II Session Chair: Xujian Fang (Hangzhou Dianzi University), Pei Zhou (SiChuanUniversity) AMU-Net: Physics-Aware Adaptive Mamba U-Net for Efficient Multi-Lead Time Precipitation Nowcasting pdfExtending Precipitation Nowcasting Horizons via Spectral Fusion of Radar Observations and Foundation Model Priors pdfHoloEv-Net: Efficient Event-based Action Recognition via Holographic Spatial Embedding and Global Spectral Gating pdfTuesday Virtual Room 8 IJCNN Paper Advances in Computer Vision III Session Chair: Shuo Zhang (Shanghai Dianji University; School of Electronic and Information,), Zheng zhihao (Institute of Information Engineering, Chinese Academy of Sciences; University of Chinese Academy of Sciences) CDSegNet++: Enhancing Diffusion-Based Point Cloud Segmentation with Adaptive Scheduling and Consistency Fusion pdfTuesday Virtual Room 1 IJCNN Paper Advances in Computer Vision IV Session Chair: Tian Lan (Renmin University of China, Center for Applied Statistics and School of Statistics), ZhiJian Fang (Zhejiang Sci-Tech University) Tuesday Virtual Room 2 IJCNN Paper Advances in Computer Vision V Session Chair: pengcheng lu (Sichuan University), Mu Yu (Tongji University) Enhanced Oriented Object Detection in Remote Sensing Using Deformable Quadrilateral Attention and PCA pdfTuesday Virtual Room 3 IJCNN Paper Advances in Machine Learning I Session Chair: 欣 黎 (华中师范大学), Bin Fang (Chongqing university) Tuesday Virtual Room 4 IJCNN Paper Advances in Machine Learning II Session Chair: Jiahui Zhang (Lancaster University), Hongbo Zhao (East China Normal University) Tuesday Virtual Room 5 IJCNN Paper Advances in Machine Learning III Session Chair: Ruifan Li (Beijing University of Posts and Telecommunications), Qicong Wang (Xiamen Univercity) CRoSS: A Continual Robotic Simulation Suite for Scalable Reinforcement Learning with High Task Diversity and Realistic Physics Simulation pdfTuesday Virtual Room 6 IJCNN Paper Advances in Machine Learning IV Session Chair: Hai-Tao Zheng (Tsinghua University), Jian-Yu Li (Nankai University) Tuesday Virtual Room 7 IJCNN Paper Brain-Computer Interfaces and Neural Decoding Session Chair: Guodong Li (Central China Normal University), huan peng (Soochoow University, School Of Computer Science & Technology) FreqTrans-Meta: Subject-Adaptive Cross-Subject SSVEP Decoding via Frequency-Domain Transformer and Meta-Learning pdfTuesday Virtual Room 8 IJCNN Paper Code Generation and Program Synthesis I Session Chair: Bilal FAYE (Sorbonne Paris Nord University), Qiuhong Zhang (Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences) Tuesday Virtual Room 1 IJCNN Paper Computer Vision and Multimodal: Emerging Topics I Session Chair: Shuxiang Song (Guangxi Normal University), Jing Qi (Hebei University) Tuesday Virtual Room 2 IJCNN Paper Computer Vision and Multimodal: Emerging Topics II Session Chair: HaiDi Xu (Zhejiang Sci-Tech University), Cong Liang (Changchun University of Science and Technology) Structure-Constrained GAN: Selective Structure Guided Adversarial Network for Lightweight Image Super-Resolution pdfPC-AGL: Light-Weighted Prior Constrained Adaptive Graph Learning for fMRI Analysis to Diagnose Autism Spectrum Disorders pdfTuesday Virtual Room 3 IJCNN Paper Continual and Incremental Learning I Session Chair: Di Shang (Institute of Automation,Chinese Academy of Sciences), Hechang Chen (Jilin University) DePER: Decoupled Multi-Scheme Prioritized Experience Replay for Sample-Efficient Deep Reinforcement Learning pdfTuesday Virtual Room 4 IJCNN Paper Federated and Distributed Learning I Session Chair: Yixiang Wang (South-Central Minzu University), jingjing fu (clemson university) DFedReweighting: A Unified Framework for Objective-Oriented Reweighting in Decentralized Federated Learning pdfTuesday Virtual Room 5 IJCNN Paper Few-Shot and Meta-Learning Session Chair: Qinghan Wang (Qilu University of Technology (Shandong Academy of Sciences)), Yichao Fu (China) CoDe: Synergizing Contrastive Discriminative Retrieval and Confidence-Aware Fusion for Zero-Shot Image Captioning pdfTuesday Virtual Room 6 IJCNN Paper Generative Priors and Model Inversion Session Chair: Chun Yuan (Tsinghua University), Zixuan Chen (School of Cyber Security and Computer, Hebei University) Tuesday Virtual Room 7 IJCNN Paper Generative Vision and Media Synthesis I Session Chair: Xingzhe Luo (Chongqing University), Paul Henderson (University of Glasgow) Tuesday Virtual Room 8 IJCNN Paper Generative Vision and Media Synthesis II Session Chair: Xin Chen (China Mobile (Hangzhou) Information Technology Co., Ltd.), Shuangquan Lyu (Carnegie Mellong University) Wednesday Virtual Room 1 IJCNN Paper IJCNN Various Tracks I Session Chair: M. Tanveer (Indian Institute of Technology Indore, India), An Zhao (Institute of Artificial Intelligence (TeleAI), China Telecom) FineCAP: Fine-grained Cyclic Augmentation Prompting with Zoom-in Enhancement for Few-shot Learning Whova Tag: Virtual only Wednesday Virtual Room 2 CEC Paper CEC Various Tracks I Session Chair: Mohamed Abouhodaima (University of New South Wales) Bayesian Optimization Framework for Multi-Objective Charging Strategy Design in Lithium-Ion Batteries pdfEnsemble Strategies for Objective Subset Selection: A Decision-Maker-Assisted Framework Using Clustering and Correlation-Based Measures pdfWednesday Virtual Room 3 FUZZ-IEEE Paper FUZZ-IEEE Various Tracks Session Chair: Qian Ma (Nanjing University of Science and Technology), Frank Chung Hoon Rhee (Hanyang University) On The Robustness of Dimensionality Reduction Methods in Selecting Type-1 and Type-2 Fuzzy Membership Functions for High-dimensional Data pdfWednesday Virtual Room 4 CEC Paper CEC Various Tracks II Session Chair: Feby John (Amrita Vishwa Vidyapeetham, Coimbatore) An Evolutionary Algorithm Guided Search of Chained Adversarial Attacks on Generative Adversarial Networks pdfWednesday Virtual Room 1 IJCNN Paper IJCNN Various Tracks II Session Chair: M. Tanveer (Indian Institute of Technology Indore, India), Abhishek Varshney (Indian Institute of Technology Indore) Wednesday Virtual Room 2 IJCNN Paper IJCNN Various Tracks III Session Chair: Stavros Ntalampiras (University of Milan), Ali Muhtaroglu (Oslo Metropolitan University) Low-Cost Real-Time Energy-Efficient Edge Image Classifier Architecture Based on Reservoir Computing With Cellular Automata pdfWednesday Virtual Room 3 IJCNN Paper IJCNN Various Tracks IV Session Chair: Manish Pratap Singh (DYSL-CT, DRDO), Shirin Shujaa (RMIT University) Sharpening Lightweight Models for Generalized Polyp Segmentation: A Boundary Guided Distillation from Foundation Models pdfWednesday Virtual Room 4 IJCNN Paper IJCNN Various Tracks V Session Chair: Chang Wang (National University of Defense Technology), Karim Ali (KING Fahd University of Petroleum & Minerals) Shapley-Enhanced Mean Field Multi-Agent Reinforcement Learning for Multi-UAV Pursuit-Evasion Maneuvering Decision-Making pdfThrough-Wall Respiration State Detection Using Wi-Fi CSI and a Lightweight CNN–LSTM for Edge Deployment pdfGraph-Based Consistency Verification for A Reliable Knowledge-Augmented Question-Answering System pdfWednesday Virtual Room 5 IJCNN Paper IJCNN Various Tracks VI Session Chair: Xiuwen Liu (Florida State University), Shiva Ahir (Stony Brook University, 934-221-1221) Wednesday Virtual Room 6 IJCNN Paper IJCNN Various Tracks VII Session Chair: Yu Liu (University of International Relations), Shuvra Neel Roy (TCS) Local Sensing to Global Hotspots: Cross-Attention Graph Fusion for Congestion-Aware Multi-Agent Path Finding pdfWednesday Virtual Room 7 IJCNN Paper IJCNN Various Tracks VIII Session Chair: Liusha Yang (Shenzhen Technology University), Mengyu Yang (Beijing University of Posts and Telecommunications) Wednesday Virtual Room 8 IJCNN Paper IJCNN Various Tracks IX Session Chair: Linqi Ye (Shanghai University), Junying Chen (South China University of Technology) DiffAFGA-Net: A Diffusion-Based Framework for Cervical Precancerous Cell Classification with Adaptive Fine-Grained Attention pdfPosition Paper: Compressed Models Should Be Robust and Publicly Shared: A Call for Responsible Model Optimization pdfWednesday Virtual Room 1 IJCNN Paper Medical Image Analysis II Session Chair: Yunfeng Liu (Beijing University of Chemical Technology), Song Liu (Qilu University of Technology (Shandong Academy of Sciences)) Toward Reliable Dermatological Medical Report Generation via Knowledge-Enhanced Structured Multi-modal Reasoning pdfWednesday Virtual Room 2 IJCNN Paper Medical Image Analysis III Session Chair: Haifeng Zhao (Anhui University, School of Computer Science and Technology), Junyuan Huang (Guangxi University) MGNet: Multi-Axis Similarity Matching and Grouping Contextual Attention Pyramid Network for Deformable Medical Image Registration pdfWednesday Virtual Room 3 IJCNN Paper Mixture-of-Experts and Modular Networks Session Chair: Gaosheng Sun (AnHui University), Jinyuan Feng (Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences) Tackling Mask Imbalance: Prototypical Mixture-of-Experts with Hardness-Aware Mining for Medical Segmentation pdfPRISM: PRinciple Induced Spectral Mixture of Experts with Test-Time Calibration for Time Series Forecasting pdfWednesday Virtual Room 4 IJCNN Paper Model Compression and Quantization III Session Chair: Dat-Thinh Nguyen (University College Dublin), Yunhan Xing (Beijing Jiaotong University) Wednesday Virtual Room 5 IJCNN Paper Model Compression and Quantization IV Session Chair: Longsheng Zhou (University of Science and Technology of China), Zequan Wang (Northeastern University, China) Wednesday Virtual Room 6 IJCNN Paper Multimodal Representation Learning III Session Chair: shenao peng (Hunan University of Technology), Fengjing Song (Qilu University of Technology) Consistency-Aware Cross-Modal Noisy Correspondence Rectification via Perturbation-Invariant Modeling pdfLG-Cog: Language-Grounded Representation Learning with Cross-Modal Cognitive Priors for Image-Text Retrieval pdfWednesday Virtual Room 7 IJCNN Paper Multimodal Representation Learning IV Session Chair: Gengchen Liu (Qilu University of Technology), He Yuanye (Institute of Information Engineering, Chinese Academy of Sciences) DTPD-MEE: A Dual-path Teacher and Pseudo-label Distillation Framework for Multimodal Event Extraction pdfWednesday Virtual Room 8 IJCNN Paper Multimodal Representation Learning V Session Chair: Jinhui Yi (University of Bonn), Yanwei Yu (Ocean University of China) EPIC: Semantic Inverse Prompting and Evidential Corroboration for Multi-modal Object Re-Identification pdfWednesday Virtual Room 1 IJCNN Paper Multimodal Representation Learning VI Session Chair: Yichen Huang (East China Normal University), Purui Bai (CASIA; School of Artificial Intelligence, University of Chinese Academy of Sciences) ESDN: Event-Centric Semantic Differential Denoising for Long-Form Audio-Visual Video Understanding pdfWednesday Virtual Room 2 IJCNN Paper Multimodal Representation Learning VII Session Chair: PEIFENG LI (Soochow Univeristy, Soochow University), PEIQIANG WANG (Tsinghua University, Shenzhen International Graduate School) Aem45k: a Large-scale Multimodal Dataset and Method for Aesthetic Assessment of Student Artwork in Middle-school Education pdfWednesday Virtual Room 3 IJCNN Paper Multimodal Representation Learning VIII Session Chair: Jia Wang (Xinjiang University), Hui Zhang (Southwest University Of Science And Technology) Wednesday Virtual Room 4 IJCNN Paper Network Security and Fault Diagnosis I Session Chair: Shaily Kabir (University of Nottingham), Chengming Liu (Zhengzhou University, School of Cyber Science and Engineering) Synergizing Laplacian Positional Encodings and Graph Transformers for Robust Network Intrusion Detection pdfUncertainty-Aware Heterogeneous Graph Meta-Learning with Hybrid Attention for IoT Intrusion Detection pdfWednesday Virtual Room 5 IJCNN Paper Network Security and Fault Diagnosis II Session Chair: Gang Shi (Xinjiang University, School of Computer Science and Technology), Zhixin Meng (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Optimal Classifier Alignment via Neural Collapse for Incremental Encrypted Traffic Classification pdfWednesday Virtual Room 6 IJCNN Paper Object Detection and Recognition V Session Chair: Shuailiang Song (Dalian University of Technology), Wenzhu Yang (Hebei University; Machine Vision Engineering Research Center, Hebei University, Baoding, China) A Frequency-Aware and Entropy-Guided Transformer for Real-Time Tiny Object Detection in Remote Sensing pdfWH-Mamba: Wavelet-Enhanced State Space Model with Dynamic Hypergraph Fusion for Robust Traffic Object Detection pdfWednesday Virtual Room 7 IJCNN Paper Object Detection and Recognition VI Session Chair: Keji Mao (Zhejiang University of Technology), Jiong Yu (Xinjiang University, Department of Computer Science and Technology) HPT-SAM2: Adapting Segment Anything Model 2 for Camouflaged Object Detection via Hierarchical Prototypes and Texture Modulation pdfWednesday Virtual Room 8 IJCNN Paper Object Detection and Recognition VII Session Chair: xiao shikang (Hunan University of Technology), Sheng-Chun Yang (Northeast Electric Power University) WMD: A Lightweight YOLO11 Extension for Efficient Complexity Detection in Mechanical Part Inspection pdfPR-DETR for Efficient and Lightweight Object Detection via Weight-sharing Softmask Pruning and Spatial Reduction pdfWednesday Virtual Room 1 IJCNN Paper Object Detection and Recognition VIII Session Chair: Xingpeng Zhang (Southwest Petroleum University), Yunpeng Guo (University of Jinan) Efficient-CANet: Rethinking Cross-Temporal Interaction and Multi-Scale Geometry for Ultra-Lightweight Change Detection pdfFSFENet: An Efficient Frequency-Spatial Feature Extraction Network for Remote Sensing Change Detection pdfWednesday Virtual Room 2 IJCNN Paper Object Tracking and Video Understanding I Session Chair: Hongyu Song (Inner Mongolia University of Technology), Zesen Cai (University of Electronic Science and Technology of China) Wednesday Virtual Room 3 IJCNN Paper Object Tracking and Video Understanding II Session Chair: shiao mi (Tianjin University of Science and Technology), Xun Ruiqi (Innovation Academy for Microsatellites of Chinese Academy of Sciences, University of Chinese Academy of Sciences) Wednesday Virtual Room 4 IJCNN Paper Open-Set and Out-of-Distribution Recognition Session Chair: Rahul Biswas (Indian Institute of Technology Hyderabad, IIT Hyderabad), Cong Hu (Jiangnan University) CALIBER: Benign-Calibrated Dual-Score Rejection for Label-Scarce Open-World Sequence Classification pdfWednesday Virtual Room 5 IJCNN Paper Point Cloud and 3D Vision Session Chair: Ziqi Xi (TongJi University), Ge Li (Peking University) Hybrid Position Encoding and Boundary-Aware Deep Supervision for 3D Point Cloud Semantic Segmentation pdfWednesday Virtual Room 6 IJCNN Paper Recommendation and Information Retrieval I Session Chair: Ziyang Cai (Guangdong University of Technology), Bo Ren (Shanghai University of International Business and Economics) Balancing User Experience and Monetization: An iQoE-Aware Recommendation Framework for Audio Platforms pdfWednesday Virtual Room 7 IJCNN Paper Recommendation and Information Retrieval II Session Chair: Jiwei Qin (Xinjiang University), jiayue wu (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences) Wednesday Virtual Room 8 IJCNN Paper Recommendation and Information Retrieval III Session Chair: Changhong Li (Huazhong University of Science and Technology), Guosheng Kang (Hunan University of Science and Technology) B-STAR: Behavior-aware Sequential Transformer with Adaptive Representations for Multi-behavior Recommendation pdfMSRec: Multimodal Semantic Learning for Cross-modal Recommendation via Multi-curvature Geometric Spaces pdfWednesday Virtual Room 1 IJCNN Paper Recommendation and Information Retrieval IV Session Chair: XiangQin Pang (Sichuan Normal University, College of Computer Science), ZhiJian Fang (Zhejiang Sci-Tech University) Decoupling Preference Aggregation and Interest Evolution via Multi-Agent Reasoning for Next POI Recommendation pdfWednesday Virtual Room 2 IJCNN Paper Reinforcement Learning II Session Chair: Li Chengwei (Institute of Automation Chinese Academy of Sciences; School of Artificial Intelligence,University of Chinese Academy of Sciences), Hongmai Xu (Guangdong University of Technology) Wednesday Virtual Room 3 IJCNN Paper Reinforcement Learning III Session Chair: Zhilin Zhang (Chongqing Institute of Green and Intelligent Technology,Chinese Academy of Sciences; Chongqing School,University of Chinese Academy Sciences), ziming liu (University of Chinese Academy of Sciences, Institute of Software Chinese Academy of Sciences) Wednesday Virtual Room 4 IJCNN Paper Reinforcement Learning IV Session Chair: Zhilin Chen (Fuzhou university), Xing Yu (East China Normal University) RISK-SAC: Risk-Driven Soft Actor-Critic for Safety-Critical Scenario Generation in Autonomous Driving pdfDecoupling Representation Robustness and Behavioral Robustness in End-to-End Driving: A Feature-Space Robustness Module for Actor–Critic Policies pdfWednesday Virtual Room 5 IJCNN Paper Reinforcement Learning V Session Chair: Qixuan Cao (East China Normal University), Tianyou Liu (National University of Defense Technology, College of Intelligence Science and Technology) Verifiability-Aware Training for Efficient Reachability Analysis of Deep Reinforcement Learning Systems pdfRobust UAV Hovering Control under Aggressive Conditions via Bidirectional Thrust and Deep Reinforcement Learning pdfWednesday Virtual Room 6 IJCNN Paper Reinforcement Learning VI Session Chair: Sheng Wang (Guangdong Institute of Intelligence Science and Technology), Huajin Tang (Zhejiang University) Radar Jamming Waveform Design Based on Hierarchical Reinforcement Learning and Cooperative Perception pdfWednesday Virtual Room 7 IJCNN Paper Reinforcement Learning VII Session Chair: Hao Wang (University of Science and Technology of China), Song Liu (Qilu University of Technology (Shandong Academy of Sciences)) Boosting Triple Extraction and Relation Normalization with Trained Retrieval and Policy Optimization pdfWednesday Virtual Room 8 IJCNN Paper Reinforcement Learning VIII Session Chair: Zhengyu Ma (China Mobile Communication Group Tianjin Co., Ltd.; Inspur Group), Jiang Haiqing (Guangzhou City University of Technology) StabTune: A Three-Phase Collaborative Framework for Reliable Microservice Kernel Parameter Optimization pdfWednesday Virtual Room 1 IJCNN Paper Remote Sensing and Earth Observation I Session Chair: Ziming Song (Fuyang Normal University), Wei Wang (Changsha University of Science and Technology) Cascaded Spectral-Spatial Enhancement and Consistency Guided Interaction for Remote Sensing Change Detection pdfSADE-Net: Strip-Aware Difference Enhancement and Multi-Dilation Fusion for Remote Sensing Change Detection pdfWednesday Virtual Room 2 IJCNN Paper Remote Sensing and Earth Observation II Session Chair: Edson Tavares (UESTC), Juan Luo (HUNAN UNIVERSITY) Stereoscopic Dual-Stream Pseudo-Siamese Network for Cloud Top Height Retrieval from Heterogeneous Dual-Satellite Data pdfWednesday Virtual Room 3 IJCNN Paper Remote Sensing and Earth Observation III Session Chair: Chengliang Wang (Chongqing University), Shijie Zhang (Northeastern University) FoumCast: Adaptive Fourier-Meteorological Disentanglement and Diffusion for High-Fidelity Precipitation Nowcasting pdfWednesday Virtual Room 4 IJCNN Paper Retrieval-Augmented Generation III Session Chair: yu pei (Xinjiang University), 宁 张 (ShangHai University) ClauseRouteRAG: A Clause-Level Retrieval-Augmented Generation Framework for 3GPP Telecommunications Standards pdfReaRAG: Knowledge-guided Reasoning Enhances Factuality of Large Reasoning Models with Iterative Retrieval Augmented Generation pdfWednesday Virtual Room 5 IJCNN Paper Retrieval-Augmented Generation IV Session Chair: Qingqiang Wu (Xiamen University, School of Informatics), ShuLin Chen (Korea University) Rethinking the Necessity of Adaptive Retrieval-Augmented Generation through the Lens of Adaptive Listwise Ranking pdfWednesday Virtual Room 6 IJCNN Paper Retrieval-Augmented Generation V Session Chair: Jing Chai (Yunnan University), Ruobing Shang (Soochow University) A Low-Cost Lightweight Large Model Code Generation Optimization Framework Based on Structured Constraint Generation and External Verification pdfTED-RAG: A Tree-Structured Multi-Expert Framework for Domain-Specific Data Synthesis in Retrieval-Augmented Generation pdfWednesday Virtual Room 7 IJCNN Paper Retrieval-Augmented Generation VI Session Chair: SuGe Wang (Shanxi University), Qinghe Li (Zhengzhou University) RPLCD: Training-Free Safety Enhancement for Instruction-Tuned LLMs via Reverse Prompt Layer Contrastive Decoding pdfWednesday Virtual Room 8 IJCNN Paper Robustness and Adversarial Learning III Session Chair: Xu Zhang (Hohai University), chunqi li (Nanjing University of Science and Technology) Gradient Correlation Whitening and Wavelet Packet Decomposition Based Transferable Adversarial Attacks on Vision Transformers pdfWednesday 0.01 London FUZZ-IEEE Paper FUZZ 7: Industry applications; Emerging related topics Session Chair: Autilia Vitiello (University of Naples Federico II) Wednesday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 16 - Optimization III Session Chair: Aldy Gunawan (Singapore Management University) Deep Reinforcement Learning for Stochastic Orienteering Problem with Metaheuristic Enhanced Experience Replay pdfWednesday 0.05 Paris IJCNN Paper Industrial Vision and Visual Quality Inspection Session Chair: Erdi Sayar (Paderborn University), Siamak Mehrkanoon (Utrecht University) Wednesday 0.10 Sydney IJCNN Paper Reliable and Governed LLMs Session Chair: M S MEKALA (Robert Gordon University ), Márcio Basgalupp (UNIFESP) Automatic Generation of Safety-compliant Linear Temporal Logic via Large Language Model: A Self-supervised Framework pdfWednesday 0.11 Cape Town IJCNN Paper Industrial Inspection, Cybersecurity, and Harmful Content Detection Session Chair: Zarka Bashir (IIT Hyderabad, IDRBT), Roberto Corizzo (American) FeatureFox: Sample-Efficient Panoptic Graph Segmentation for Machining Feature Recognition in B-Rep 3D-CAD Models pdfTransformer-Based Intrusion Detection with Feature-Selected Inputs for Cross-Day Generalization and Unseen Attack Detection pdfWednesday 0.14 Singapore IJCNN Paper Remote Sensing and Environmental Intelligence Session Chair: Polat Goktas (Sabanci University), Amirreza Yousefzadeh (University of Twente) Wednesday 0.15 Washington IJCNN Paper Efficient and Resource-Constrained AI I Session Chair: Christos Kyrkou (University of Cyprus; Department of Electrical and Computer Engineering, University of Cyprus), Yanick Christian Tchenko (University of Paris Saclay, University of Evry) Wednesday 2.18 Mekong IJCNN Paper Biomedical Signal Processing Session Chair: Roseline Mary Rozario (University of Wollongong), Susana Vieira (Instituto Superior Técnico) Neural Network Cognitive Feedforward for Habit Change Decision-Making: A systems thinking approach in developing medical software for behavior change and quality assurance pdfWednesday 0.02 Berlin FUZZ-IEEE Paper FUZZ 9 : Fuzzy control and robotics, sensors, fuzzy hardware, fuzzy architectures Session Chair: Uiliam Nelson Lendzion Tomaz Alves (Federal Institute of Education, Science and Technology of Paraná - IFPR) Robust Switched Control with Guaranteed Cost of a Furuta Pendulum Prototype using T-S Fuzzy model pdfWednesday 0.01 London FUZZ-IEEE Position Paper, FUZZ-IEEE Paper FUZZ 8 : FUZZ-IEEE SS05 Fuzzy Federated Learning: Theoretical advances and novel applications (FL-A) & FUZZ Position paper Session Chair: Asier Urio-Larrea (Universidad Pública de Navarra, Spain) A Practical Framework for Federated Fuzzy Clusterwise Regression, Preprocessing and Hyperparameter Tuning pdfWednesday 0.01 London FUZZ-IEEE Paper FUZZ 10 : FUZZ-IEEE SS02 Fuzzy Machine Learning Session Chair: Jie Lu (University of Technology Sydney) Fuzzy Logic-Driven Reward Optimization for Deep Reinforcement Learning in Carbon Emission Trading pdfFuzzy Accuracy Compensates for Label Subjectivity in Classification of Skin Tone Using Wearable Photoplethysmography Signals pdfWednesday 0.02 Berlin IJCNN Paper Neuromorphic and Spiking Neural Networks Session Chair: Toshihisa Tanaka (Tokyo University of Agriculture and Technology), Guangzhi Tang (Maastricht University) Scaling high-dimensional synaptic plasticity on neuromorphic hardware: parameter mapping and vectorized implementation pdfAutomated Model-to-Transistor Design Framework for Memristor Crossbar-Based Spiking Neural Network Architectures pdfKnowledge Distillation at the Edge: Lightweight Deep Neural Networks Deployed on Neuromorphic Hardware pdfSpikeDiffusion: A Fully Spiking Structure-Guided Diffusion Framework for Energy-Efficient Image Generation pdfWednesday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 17 - Related Topics III Session Chair: Carlos Coello Coello (Cinvestav; Faculty of Excellence of the School of Engineering and Sciences, Tecnologico de Monterrey, Monterrey, Mexico) Evolutionary Edge Bundling as a Large-Scale Multi-Objective Optimization Problem: A Comparative Study pdfGraph Models to Forecast Road Traffic Crashes and Resource Optimization on Brazilian Federal Roads pdfWednesday 0.05 Paris IJCNN Paper AI for Cybersecurity, Public Safety, and Surveillance Session Chair: Fabrizio Pittorino (Politecnico di Milano), Israel Efraim de Oliveira (Universidade Federal de Santa Catarina) Wednesday 0.10 Sydney IJCNN Paper Enterprise LLMs, RAG, and Knowledge-Augmented Systems Session Chair: Thiago Oliveira-Santos (UFES, I2CA), Cleber Zanchettin (Universidade Federal de Pernambuco, Northwestern University) Wednesday 0.11 Cape Town IJCNN Paper IJCNN SS31 Systems-Theoretic Approaches to Learning 4.0: From Classical to Quantum Neural Networks Session Chair: Vignesh Narayanan (University of South Carolina), Avimanyu Sahoo (University of Alabama in Huntsville), Krishnan Raghavan (Argonne National Laboratory) Wednesday 0.14 Singapore IJCNN Paper IJCNN SS27 Reservoir Computing for Scalable and Energy-Efficient AI: Theory, Dynamics, and Implementations I Session Chair: Claudio Gallicchio (University of Pisa) As Simple as Possible, but Not Simpler: When Reservoir Computing Rivals Deep Learning in Hydrological Prediction pdfWednesday 0.15 Washington IJCNN Paper Robotics and Autonomous Agents Session Chair: Erdal Kayacan (Paderborn University, Germany), Leon Reznik (Rochester Institute of Technology) End-to-End Learning of Collaborative Loco-Manipulation for Load Transportation using Quadrupeds with Manipulators pdfWednesday 0.01 London FUZZ-IEEE Paper FUZZ 11 : FUZZ-IEEE SS02 Fuzzy Machine Learning Session Chair: Jie Lu (University of Technology Sydney) Interpretable Soil-Robust Bipedal Locomotion on Tilled Soils with Material Curriculum and a Stance-Gated Slip Reward pdfImproving Generalization Performance of Multi-objective Fuzzy Genetics-Based Machine Learning using Synthetic Training Data Generated from High-Performance Black-Box Models pdfWednesday 0.02 Berlin IJCNN Paper Neuromorphic Computing and AI Hardware Systems Session Chair: Guilherme DeSouza (University of Missouri, Vision-Guided and Intelligent Robotics Lab (ViGIR)), Giorgio Morales (University of Caen Normandy) FedMTFI: Feature Importance Based Optimized Multi Teacher Knowledge Distillation in Heterogeneous Federated Learning Environment pdfWednesday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 18 - Algorithms IV Session Chair: Andries Engelbrecht (Stellenbosch University) Beyond Behavioral Sequences: Leveraging Short-Term Memory to Optimize Decision Policies in Anticipatory Learning Classifier System pdfWednesday 0.05 Paris IJCNN Paper AI for Industrial Operations and Planning Session Chair: Ali Minai (University of Cincinnati), Francesco Alesiani (NEC Laboratories Europe GmbH) Preference-Conditioned Dynamic Attention Model for Multi-objective Capacitated Arc Routing Problem pdfPRISM: Deployable Counterfactual Recourse for Metro Maintenance via Physics-Constrained Search and RL Distillation pdfWednesday 0.10 Sydney IJCNN Paper AI for Mobility, Transportation and Infrastructure Systems Session Chair: Yatharth Agarwal (Purdue University), Hanadi Alhamdan (Princess Nourah bint Abdulrahman University, Durham University) MamMA: A Mamba-Based Pedestrian Trajectory Prediction Algorithm Considering Occupancy Map and Pedestrian Awareness States pdfWednesday 0.11 Cape Town IJCNN Paper AI for Healthcare, Medical Imaging, and Biomedical Discovery I Session Chair: Silvia Multari (Ca’Foscari University of Venice), Larissa Zott (Universität der Bundeswehr München) HGDC-Fuse: Clinical Multi-modal Fusion with Heterogeneous Graph and Disease Correlation Learning for Multi-Disease Prediction pdfMLD-Sup: Multi-scale Latent Deep Supervision for Joint Segmentation and Classification of Breast Ultrasound Images pdfWednesday 0.14 Singapore IJCNN Paper AI for Energy and Resource Analytics Session Chair: Enrico De Santis (University of Rome "La Sapienza"), Jean Senellart (Quandela) A Conditional GAN Framework for Joint Generative Modeling of Geological Facies and Acoustic Impedance pdfWednesday 0.15 Washington IJCNN Paper AI for Finance and Economics Session Chair: Aleksei Liuliakov (Bielefeld University), Isaac Amankona Obiri (Quanzhou University of Information Engineering) Fundamental vs. Technical Features for Stock Trend Prediction: A Multi-Horizon Data Stream Learning Analysis pdfRandom Convolution Kernel–Augmented Temporal Convolutional Networks for Volatility-Dominated Financial Time Series pdfAn Enhanced Temporal Graph Network with Retrieval-Augmented Graph for Dynamic Link Prediction in Cryptocurrency Transactions pdfWednesday Virtual Room 1 IJCNN Paper LLM Adaptation and Fine-Tuning III Session Chair: Di Wu (Hebei University of Engineering), Yongqing Wang (School of Computer Science and Technology, Tongji University) Dual-Mechanism Balanced Graph Partitioning: Contrastive Learning for Coarse Partition Generation and Reinforcement Learning for Fine-Tuning pdfRegion Partitioning and Prototype-Query Calibration for Few-Shot Fine-Grained Image Classification pdfWednesday Virtual Room 2 IJCNN Paper LLM Adaptation and Fine-Tuning IV Session Chair: Ken Zhong (Shanghai Jiao Tong University), Long Chen (China West Normal University) A Syntactic Feature Interaction and Cross-Table Relation Search Model for Threat Intelligence Triple Extraction pdfWednesday Virtual Room 3 IJCNN Paper LLM Agents and Tool Use I Session Chair: Junkai Zhang (Institute of Automation, Chinese Academy of Sciences) HDCF: Bi-level Diversity Coordination for Hierarchical Cooperative Multi-Agent Reinforcement Learning pdfWednesday Virtual Room 4 IJCNN Paper LLM Agents and Tool Use II Session Chair: Linqi Ye (Shanghai University), Yibo Chen (PLA Academy of Military Science) Wednesday Virtual Room 5 IJCNN Paper LLM Agents and Tool Use III Session Chair: Lisan Al Amin (University of Maryland, Baltimore County), Zhuoyi Huang (National University of Defense Technology) Wednesday Virtual Room 6 IJCNN Paper LLM Agents and Tool Use IV Session Chair: Gangao Liu (Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences), Yuanjian Zhao (Sichuan University) Graph-RHO: Critical-path-aware Heterogeneous Graph Network for Long-Horizon Flexible Job-Shop Scheduling pdfWednesday Virtual Room 7 IJCNN Paper LLM Evaluation and Benchmarking I Session Chair: Xiuwen Liu (Florida State University), Yufei Zeng (Beijing University of Posts and Telecommunications) TopoAtten: Probing the Internal Topology of Attention for Hallucination Detection in Large Language Models pdfDecoding Emotion in the Deep: A Systematic Study of How LLMs Represent, Retain, and Express Emotion pdfWednesday Virtual Room 8 IJCNN Paper LLM Evaluation and Benchmarking II Session Chair: Boxun Li (North University of China), xueer wang (Xiangtan University) Wednesday Virtual Room 1 IJCNN Paper LLM Reasoning and Planning I Session Chair: Junhong Liang (MBZUAI), kaiyao Tan (sun yat-sen university) Reasoning to Find, Restraining to Fix: An Explainable Chinese Spell Checking Framework via Group Relative Policy Optimization pdfWednesday Virtual Room 2 IJCNN Paper LLM Reasoning and Planning II Session Chair: Huang Shan (Tsinghua University), Jian Chen (Ningxia Research Institute of Transport Sciences) ABEX-RAT: Synergizing Abstractive Augmentation and Adversarial Training for Classification of Occupational Accident Reports pdfWednesday Virtual Room 3 IJCNN Paper Learning Paradigms and Model Efficiency I Session Chair: Chin-Teng Lin (University of Technology Sydney), Huaijun Guang (Dalian Minzu University) Multiscale Spatiotemporal Ensemble Learning for Imbalanced Chiller Fault Diagnosis Using Generative Adversarial Networks pdfBeyond Distribution Matching: Contrastive-Enhanced Diffusion for Imbalanced Tabular Data Augmentation pdfImproving Time Series Generation and Self-Supervised Learning via Instance-Conditioned Contrasting pdfWednesday Virtual Room 4 IJCNN Paper Learning Paradigms and Model Efficiency II Session Chair: Shoji Toyota (Kyushu University), Zitai Kong (Zhejiang university) Stochastic Adaptive Process State Space Model for Portfolio Management Based on Deep Reinforcement Learning pdfWednesday Virtual Room 5 IJCNN Paper Learning Paradigms and Model Efficiency III Session Chair: Fang Wu (Kashgar University), Wei Wang (Yunnan University) Intraday Price-Movement Prediction Using Technical Indicators with Feature Selection and Genetic Algorithm-Based Hyperparameter Optimization pdfWednesday Virtual Room 6 IJCNN Paper Machine Learning Methods and Applications I Session Chair: Wei Guo (Shenyang Aerospace University, School of Computer Science), Guangchao Yang (Chongqing University) Incorporating Dose Level Prior and Texture-Sensitive Contrastive Learning from Frequency Perspective for Multi-dose PET Reconstruction pdfWednesday Virtual Room 7 IJCNN Paper Machine Learning Methods and Applications II Session Chair: Handong Yao (University of Georgia), Ashitabh Misra (University of Illinois at Urbana Champaign) MotiMem: Motion-Aware Approximate Memory for Energy-Efficient Neural Perception in Autonomous Vehicles pdfWednesday Virtual Room 8 IJCNN Paper Medical Image Analysis I Session Chair: Wu chaolin (JiNan University), Dmitrii Kaplun (Saint Petersburg Electrotechnical University "LETI") KGS-UNet: KAN-Gated Skip Connections for Small-Structure Medical Image Segmentation with Limited Data pdfSDG-UNet: Medical Image Segmentation Based on Structure-Guided Dynamic Convolution and Dual-Path Hybrid Attention pdfCALDGSEG: A Medical Image Segmentation Framework via Interactive Feature Fusion and Dynamic Guidance pdfWednesday Virtual Room 1 IJCNN Paper Model Compression and Quantization I Session Chair: DongChen Zhu (Shanghai Institute of Microsystem and Information Technology, Chinese Academy of Sciences), d Liu (Institute of Information Engineering, Chinese Academy of Sciences ; School of Cyber Security, UCAS) Wednesday Virtual Room 2 IJCNN Paper Model Compression and Quantization II Session Chair: Bin Ji (National University of Defense Technology), Saibal Mukhopadhyay (Georgia Institute of Technology) Structured Sparse Deep Nonnegative Matrix Factorization with ℓ_{2,1}-norm and Graph Regularization for Clustering pdfWednesday Virtual Room 3 IJCNN Paper Multimodal Representation Learning I Session Chair: Xing Yu (East China Normal University), qiaoming zhu (Soochow University) Multi-dimensional Hybrid Fact-Checking: Integrating Structured and Unstructured Information to Promote LLMs pdfWednesday Virtual Room 4 IJCNN Paper Multimodal Representation Learning II Session Chair: Peiran Liang (Northwest A&F University), Zhihao Cai (Ocean University of China) CURA-Stereo: Cross-Volume Consistency Uncertainty for Fusion-and-Rectification in Iterative Stereo Matching pdfWednesday Virtual Room 5 IJCNN Paper Neural Architectures and Sequence Models Session Chair: Suhang Qian (Tianjin University), Qizhao Long (Zhejiang University) SHD-Mamba: A Spectral–Temporal Hybrid Mamba with Dual-Suppression Decoding for Satellite Image Time Series pdfWednesday Virtual Room 6 IJCNN Paper Neural Network Foundations and Optimization Session Chair: Jikun Wu (Stellaris AI Limited), yuhao zhang (BeiHang University) Wednesday Virtual Room 7 IJCNN Paper Object Detection and Recognition I Session Chair: xiyu pan (Central South University of Forestry and Technology), Wenzhu Yang (Hebei University; Machine Vision Engineering Research Center, Hebei University, Baoding, China) Wednesday Virtual Room 8 IJCNN Paper Object Detection and Recognition II Session Chair: HaiDi Xu (Zhejiang Sci-Tech University), Tianzhu Xie (De Anza College) Wednesday Virtual Room 1 IJCNN Paper Object Detection and Recognition III Session Chair: Tao Wu (Shanghai Institute of Technology, Faculty of Intelligent Technology), Jin Zhang (Changsha University of Science & Technology) PGA-YOLO: Toward Efficient and Robust Feature Interaction for Defect Detection in Complex Industrial Scenes pdfDS-YOLO-RC: Geometric-Aware Rotated Crack Detection via Over-Parameterized Dilated Convolution and Topology-Preserving Alignment pdfWednesday Virtual Room 2 IJCNN Paper Object Detection and Recognition IV Session Chair: Congyu Liu (Changsha University of Science and Technology), Zhe Wang (East China University of Science and Technology) A Multi-Scale Semantic Feature based Relaxed Alignment Network for Visible-Infrared Person Re-identification pdfSee Fine, Know Clear: Improving Video Object Segmentation via Fine-Grained Matching and MambaVision-Powered Semantics pdfWednesday Virtual Room 3 IJCNN Paper Reinforcement Learning I Session Chair: Xiao Sun (Hefei University of Technology), Zilan Li (Guilin University of Electronic Technology) Hyper-DAG-Based Hierarchical Reinforcement Learning for Unifying Data and Resource Dependencies in Multi-UAV Rescue pdfWednesday Virtual Room 4 IJCNN Paper Retrieval-Augmented Generation I Session Chair: Jundong Liu (Ohio University), Jingyu Wang (Beijing Institute of Technology) URA-NER: A Unified Retrieval-Augmented Framework with Retrieval Alignment and Uncertainty Reduction for Low-Resource NER pdfWednesday Virtual Room 5 IJCNN Paper Retrieval-Augmented Generation II Session Chair: Shiyu Zhu (Beijing Institute of Graphic Commmunication), Jiaxin Wu (Beijing Institute of Graphic Commmunication) Wednesday Virtual Room 6 IJCNN Paper Robustness and Adversarial Learning I Session Chair: Baolin Yan (Institute of Software Chinese Academy of Sciences, University of Chinese Academy of Sciences), Shuaiye Lu (Wuhan University) Wednesday Virtual Room 7 IJCNN Paper Robustness and Adversarial Learning II Session Chair: Jun Li (Capital Normal University), Lixiao Liang (Guangzhou University) Selection-Aware Poisoning: Boosting Clean-Label Backdoor Attacks via Distribution Deviation and Projection Residual Metrics pdfWednesday Virtual Room 8 IJCNN Paper Semi-Supervised and Weakly-Supervised Learning I Session Chair: Donghai Zhai (Southwest Jiaotong University), Xinyu Li (Beijing Jiaotong University) Enhancing Unlabeled Data Efficiency for Semi-Supervised Domain Generalization via Reliability-Guided Energy Prior pdfSABTM: Semi-Supervised Multimodal Emotion Recognition via Adaptive Barlow Twins and Mixture-of-Experts pdfWednesday Virtual Room 1 IJCNN Paper Semi-Supervised and Weakly-Supervised Learning II Session Chair: Weiliang Meng (Chinese Academy of Sciences, Institute of Automation), Zhenming Yuan (Hangzhou Normal University, Hangzhou Hele Tech Co.LTD) Mamba-CAM: Mamba-Driven Class Activation Map Learning for 3D Medical Weakly-Supervised Segmentation pdfWednesday Virtual Room 2 IJCNN Paper Spatiotemporal Learning and Prediction I Session Chair: Ziyi Li (xinjiang university), Zhenli Qian (Shanghai Ocean University) MF-STGCN: Multimodal Fusion Sparse Spatio-Temporal Graph Networks with Physics-Aware Gating for Ocean Eddy Forecasting pdfWednesday Virtual Room 3 IJCNN Paper Time Series Forecasting I Session Chair: Mingshan Du (Chengdu University of Technology), Yiqi Yu (Zhejiang University) Wednesday Virtual Room 4 IJCNN Paper Time Series and Temporal Modeling I Session Chair: Kasun Dewage (university of Central Florida), Jianbin gao (University of Electronic Science and Technology of China) Hybrid Neural-Classical Correction for Frozen Time Series Foundation Models: A Comprehensive Ablation Study on High-Frequency Stock Prediction pdfDecomposing the Time Series Forecasting Pipeline: A Modular Approach for Time Series Representation, Information Extraction, and Projection pdfWednesday Virtual Room 5 IJCNN Paper Vision Domain Adaptation and Generalization I Session Chair: 劭卫 范 (河北工业大学), Zhikui Chen (Dalian University of Technology) Multi-View Structural Consistency Learning for Test-Time Adaptation in Medical Image Segmentation pdfWednesday Virtual Room 6 IJCNN Paper Vision Domain Adaptation and Generalization II Session Chair: Yufan Yi (Huazhong University of Science and Technology), Huiwen Huang (Guilin University of Electronic Technology, School of Computer and Information Security) ID3A: A Multi-Source Domain Adaptation Framework with Identity Disentanglement for Cross-Subject Emotion Recognition pdfWednesday Virtual Room 7 IJCNN Paper Vision-Language Models I Session Chair: YingChao Zeng (Beijing University of Posts and Telecommunications), hongye chen (Nanchang Hangkong University) Wednesday Virtual Room 8 IJCNN Paper Vision-Language Models II Session Chair: Nan Mu (XinJiang university), Arunava Roy (University of Memphis) Causal Inference and Counterfactual Text-Debiasing for Multimodal Aspect-Based Sentiment Analysis pdfThursday Virtual Room 1 IJCNN (Neural Networks) Industrial applications Session Chair: Zhen Zeng (Sun Yat-sen University), Chunling Xi (Google) XSSplore: A Tree-Structured Prompting and Actor-critic Framework for Autonomous Red Teaming XSS Payload Generation pdfThursday Virtual Room 2 IJCNN Paper, IJCNN Late Breaking Paper IJCNN Various Tracks X Session Chair: Patrick Akinwumi (Clemson University) Neuro-Symbolic Reinforcement Learning for Cognitive Digital Twins: A Framework for Lifelong Intelligence Amplification Whova Tag: Poster Presentation Thursday Virtual Room 3 IJCNN (Neural Networks) IJCNN Various Tracks XI Simple Averaging vs. Learned Stacking in K-Fold Ensembles: Evidence from Pore Pressure Prediction pdfPhysics-Guided Multi-Task Learning for Subgrid Scale Turbulence Parameterization: A Comparative Study of Physics Integration Strategies pdfNeuroAPS-Net: Neuro-Anatomically Aware Point Cloud Representation for Efficient Alzheimer's Disease Classification pdfThursday Virtual Room 1 IJCNN Paper Robustness and Adversarial Learning IV Session Chair: Yuanfang Guo (Beihang University), Sk. Subidh Ali (Indian Institute of Technology Bhilai) Generalized Adversarial Overwriting Attack on Deep Learning-Based Watermarking Under Restricted Black-Box Access pdfBlack-box Backdoor Attack on Continual Semi-Supervised Learning for Time-series Internet-of-Things Systems pdfThursday Virtual Room 2 IJCNN Paper Robustness and Adversarial Learning V Session Chair: Jianli Ding (Civil Aviation University of China), Sihan You (National University of Defense Technology) BackFlush: Knowledge-Free Backdoor Detection and Elimination with Watermark Preservation in Large Language Models pdfThursday Virtual Room 3 IJCNN Paper Robustness and Adversarial Learning VI Session Chair: Mustafa Misir (Duke Kunshan University), Yuanbo Xie (Institute of Information Engineering, Chinese Academy of Sciences, China; School of Cyber Security, University of Chinese Academy of Sciences, China) Thursday Virtual Room 4 IJCNN Paper Robustness and Adversarial Learning VII Session Chair: Siya Yao (Zhejiang Gongshang University), Erxiang Wang (ZZU) Buffer-Guided Differentiated Learning for Adversarial Fine-Tuning in Self-Supervised Pretrained Models pdfThursday Virtual Room 5 IJCNN Paper Spatiotemporal Learning and Prediction II Session Chair: Yingchi Mao (Hohai University), Qingzheng Hu (INTI International University, Southwest Petroleum University) Thursday Virtual Room 6 IJCNN Paper Spatiotemporal Learning and Prediction III Session Chair: wei Tang (Civil Aviation University of China), xiao yu (University of Jinan) Seeing Wider, Seeing Longer: OmniGait for Gait Recognition in the Wild with Re-parameterized Spatio-Temporal Large Kernel Convolution pdfSTHP-Net:A Spatio-Temporal Collaborative High-Frequency Purification Network for Intracranial Artery Segmentation in DSA Sequences pdfThursday Virtual Room 7 IJCNN Paper Spatiotemporal Learning and Prediction IV Session Chair: Yuping Wang (Beijing University of Posts and Telecommunications), Haowei Mei (Xi'an Jiaotong University) WaveSFNet: A Wavelet-Based Codec and Spatial–Frequency Dual-Domain Gating Network for Spatiotemporal Prediction pdfThursday Virtual Room 8 IJCNN Paper Speech and Audio Processing Session Chair: Bao Thang Ta (Viettel AI), Xin Dong (Nanjing University of Aeronautics and Astronautics) Thursday Virtual Room 1 IJCNN Paper Text Generation and Multilingual NLP I Session Chair: Jiarui Zhang (Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China; School of Cyber Security, University of Chinese Academy of Sciences, Beijing, China), Haixian Zhang (Sichuan University) Interpretable Semantic Denoising via Neuron-level Functional Separation for Low-resource Multilingual Translation pdfEnhancing Cross-Lingual Embedding Alignment with Additive Keywords for International Trade Product Classification pdfThursday Virtual Room 2 IJCNN Paper Text Generation and Multilingual NLP II Session Chair: 武 庄 (Civil Aviation University of China), Yuyan Zhou (University of Science and Technology of China) Where are next facilities: Emergency planning in non-Euclidean space using graph-based reinforcement learning pdfDACTM: Enhancing Neural Topic Models through Denoising Autoencoders with Intelligent Semantic Masking and Contrastive Learning pdfThursday Virtual Room 3 IJCNN Paper Text Generation and Multilingual NLP III Session Chair: 健 陈 (Shenzhen University), Xianzheng Liu (Central China Normal University) Thursday Virtual Room 4 IJCNN Paper Text Generation and Multilingual NLP IV Session Chair: Zhe Yang (Computer Network Information Center), Shifali Agrahari (Indian Institute of Technology Guwahati, India) Thursday Virtual Room 5 IJCNN Paper Time Series Classification and Anomaly Detection Session Chair: zhang zejun (shenyang aerospace university), Hongbo Zhao (East China Normal University) HRD-Net: Hierarchical Relational Diffusion Network for Anomaly Detection in Multivariate Time-series Data pdfThursday Virtual Room 6 IJCNN Paper Time Series Forecasting II Session Chair: jiannan wan (Fuzhou University, College of Computer and Data Science), Wenhui Hu (Institute of Plasma Physics, Chinese Academy of Sciences) TimeCF: A novel model with adaptive Convolution and Sharpness-Aware Minimization Frequency Domain Loss for long-term time series forecasting pdfThursday Virtual Room 7 IJCNN Paper Time Series Forecasting III Session Chair: Guoqing Tang (University of Science and Technology of China), Yu Wang (Chengdu Institute of Computer Applications, University of Chinese Academy of Sciences) WPKAN-Mixer: Long-Term Time Series Forecasting with Adaptive Function Learning on Wavelet-Decomposed Components pdfThursday Virtual Room 8 IJCNN Paper Time Series Forecasting IV Session Chair: Jiangrong Yang (Beijing Normal University), Dongwei Liu (South China Normal University) Thursday Virtual Room 1 IJCNN Paper Time Series Forecasting V Session Chair: Jiayang Xu (Southern University of Science and Technology), Zenglin Xu (Fudan University) Thursday Virtual Room 2 IJCNN Paper Time Series Forecasting VI Session Chair: Jiameng Chen (Beijing University of Posts and Telecommunications), Hongkai Jiang (Tsinghua University) Stock Price Forecasting Using a Transformer with Time-Interval-Based Attention and Trainable Stock Selection pdfThursday Virtual Room 3 IJCNN Paper Time Series Forecasting VII Session Chair: Hongnian Wang (North Sichuan Medical College), Hongbo Zhao (East China Normal University) DSFormer: Dimension-Segment Transformer with Cross-Stock Correlation Modeling for Stock Prediction pdfThursday Virtual Room 4 IJCNN Paper Time Series Forecasting VIII Session Chair: yiren zhou (sichuan university), Kaixuan Chen (Zhejiang University) Thursday Virtual Room 5 IJCNN Paper Time Series and Temporal Modeling II Session Chair: Mingyue Qin (Shanghai Jiao Tong University), Rui Yan (Zhejiang University of Technology) Fractional-Order Differential Equation-Driven Transformer for Time Series Forecasting via Gauss-Jacobi Quadrature pdfThursday Virtual Room 6 IJCNN Paper Time Series and Temporal Modeling III Session Chair: huanlan yan (tongji university), Shaoqi Tan (University of Electronic Science and Technology of China) Thursday Virtual Room 7 IJCNN Paper Trustworthy and Safe Language Models Session Chair: Jianyuan Ni (Juniata College), Jixuan Guo (Beijing University of Technology, School of Information Science and Technology) Thursday Virtual Room 8 IJCNN Paper Vision Domain Adaptation and Generalization III Session Chair: Junsong Leng (Huazhong University of Science and Technology, School of Artificial Intelligence and Automation), Feifei Zhang (Qilu University of Technology; Shandong Provincial Key Laboratory of Industrial Network and Information System Security, Shandong Fundamental Research Center for Computer Science, Jinan, China) A Source-Free Universal Domain Adaptation Method for Fault Diagnosis Based on Dual-View Disentanglement and Orthogonal Decomposition pdfCross-domain Bearing Fault Diagnosis Based on a Triple Attention Multi-scale Spatiotemporal Convolutional Network pdfThursday Virtual Room 1 IJCNN Paper Vision Domain Adaptation and Generalization IV Session Chair: zhiyong zheng (University of Science and Technology of China; School of AI and Data Science, University of Science and Technology of China), Shayok Chakraborty (Florida State University) Thursday Virtual Room 2 IJCNN Paper Vision-Language Models III Session Chair: Zicheng Wang (Shandong Jianzhu University), Xiangfeng Luo (Shanghai University, School of Computer Engineering and Science) Bridging Naive Semantics through Proxy Injection and Matrix Fusion for Open-Vocabulary Semantic Segmentation pdfLoCoFuse: Locality-Constrained Dual-Branch Inference for Training-Free Underwater Open-Vocabulary Semantic Segmentation pdfThursday Virtual Room 3 IJCNN Paper Vision-Language Models IV Session Chair: An Zhao (Institute of Artificial Intelligence (TeleAI), China Telecom), Ye Shen (Shanghai Jiao Tong University) GLA-MLLM: A Flowchart Understanding Method Integrating Global-Local Awareness With Multimodal Large Language Model Generation pdfThursday Virtual Room 4 IJCNN Paper Vision-Language Models V Session Chair: Haowen Zheng (Central University of Finance and Economics), Jiajie Fan (South China Normal University) Thursday Virtual Room 5 IJCNN Paper Vision-Language Models VI Session Chair: Hao Kong (shanghai university), Ju Wang (Macao Polytechnic University, Faculty of Applied Sciences) Thursday Virtual Room 6 IJCNN Paper Vision-Language Models VII Session Chair: Shuxiang Song (Guangxi Normal University), Zhenjun Tang (Guangxi Normal University) Thursday Virtual Room 7 IJCNN Paper Vision-Language-Action and Robotic Learning I Session Chair: Qian Zhang (National University of Defense Technology), Yanglan Dong (University of Science and Technology of China; SKL of Processors, Institute of Computing Technology, CAS) DWIP: Efficient Depth–Width Integrated Pruning for Vision-Language-Action Models in Robotic Manipulation pdfThursday Virtual Room 8 IJCNN Paper Vision-Language-Action and Robotic Learning II Session Chair: Jian Xue (University of Chinese Academy of Sciences), Weixuan Liu (Shenzhen Institute for Advanced Study, University of Electronic Science and Technology of China) Thursday Virtual Room 1 IJCNN Paper Visual Anomaly and Defect Detection I Session Chair: Chaoli Wang (上海理工大学), Chengming Liu (Zhengzhou University, School of Cyber Science and Engineering) Thursday Virtual Room 2 IJCNN Paper Visual Anomaly and Defect Detection II Session Chair: Mingle Zhou (Qilu University of Technology; Shandong Provincial Key Laboratory of Computing Power Internet and Service Computing, Shandong Fundamental Research Center for Computer Science, Jinan, China.), Lei Duan (Sichuan University) PPRL:Prototype-driven Periodic Representation Learning for Multivariate Time Series Anomaly Detection pdfMSDDS: Multivariate Time Series Anomaly Detection via Multi-Scale Decomposition with Fusion and Dual-Stream Channel Transformer pdfTF-ChebKAN: Time-Frequency Coupled Chebyshev Kolmogorov-Arnold Networks for Multivariate Time Series Anomaly Detection pdfThursday Virtual Room 3 IJCNN Paper Visual Anomaly and Defect Detection III Session Chair: Guitao Cao (Shanghai Key Laboratory of Trustworthy Computing, East China Normal University), Jiong Yu (Xinjiang University, Department of Computer Science and Technology) ReMAP-AD: Unifying Visual Memory Reconstruction and Prompt Learning for Few-Shot Anomaly Detection pdfThursday Virtual Room 4 IJCNN Paper Visual Anomaly and Defect Detection IV Session Chair: Chen Ling (Center for Information Research, Academy of Military Sciences, Beijing 100142, China), Xiaowu Liang (Jinan University) Thursday Virtual Room 5 IJCNN Paper Visual Anomaly and Defect Detection V Session Chair: Hongtao Wang (Department of Computer, North China Electric Power University; Engineering Research Center of Intelligent Computing for Complex Energy Systems, Ministry of Education), Xiaolong Zheng (University of Chinese Academy of Sciences) Thursday Virtual Room 6 IJCNN Paper Visual Anomaly and Defect Detection VI Session Chair: 书一 尚 (Institute of Information Engineering, Chinese Academy of Sciences; School of Cyber Security, University of Chinese Academy of Sciences), Xuanning Liu (Beijing University of Posts and Telecommunications) RESTAD: Residual-Enhanced Spatio-Temporal Modeling with Multi-Scale Attention for Multivariate KPI Anomaly Detection pdfDifferential Contrastive Representation Learning with Frequency-Domain Attention for Anomaly Detection in Aerospace Cyber-Physical Systems pdfThursday Virtual Room 7 IJCNN Paper Visual Anomaly and Defect Detection VII Session Chair: Zewen Wang (Xinjiang University), Na Liu (Inner Mongolia University of Technology) A Weak-Signal-Aware Framework for Subsurface Defect Detection: Mechanisms for Enhancing Low-SCR Hyperbolic Signatures pdfThursday Virtual Room 8 IJCNN Paper Visual Representation Learning Session Chair: Feiyu Chen (Chongqing Normal University), Xiaolin Xiao (South China Normal University) Thursday 0.01 London IJCNN Paper Smart Energy, Grid, and Infrastructure including IJCNN SS35 Computational Intelligence Techniques for Observable Smart Grid and Sustainable Energy Systems Session Chair: Giulia Tanoni (Università Politecnica delle Marche, Ancona, Italy) Two-Stage Prediction Intervals via Spiking Neural Network-Gated State-Specific Residual Models for Solar Power Generation pdfForward--Forward Learning for Imbalanced Tabular Predictive Maintenance on a Real-World Smart-Grid Fault Dataset pdfSimulation of Microgrid Energy Management under Battery Degradation Costs: a PPO-Based Reinforcement Learning Approach pdfThursday 0.02 Berlin IJCNN Paper, FUZZ-IEEE Position Paper, CEC Late Breaking Paper, CEC Paper, FUZZ J2C Presentation, CEC J2C Presentation, FUZZ-IEEE Paper, CEC Position Paper, IJCNN J2C Presentation, IJCNN Position Paper, IJCNN Late Breaking Paper, FUZZ-IEEE Late Breaking Paper Student Best Paper Award Explainable Object Detection in 360° Images Through Fuzzy Logic Systems and Vision-Language Models Integration pdfThursday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 19 - Related Topics IV Session Chair: Juan J. (University of Granada) From Configuration to Evolution: Corporate Social Responsibility Practices and Supply Chain Effects on Environmental Innovation pdfMulti-objective Optimisation of Traffic Light Control for Fast and Safe Traffic Incident Recovery pdfThursday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 20 - Algorithms V Session Chair: Diego Oliva (Universidad de Guadalajara) Automated Algorithm Design of Tailored Metaheuristics for Photovoltaic Parameter Estimation in Single- and Double-Diode Models pdfA Multiobjective Evolutionary Feature Selection Framework for End-Point Molten Steel Temperature Prediction in Ladle Furnace Refining pdfTowards Standardized Evaluation of Feasible Region Identification in Constrained Engineering Design pdfThursday 0.10 Sydney IJCNN Position Paper Reliable, Robust, and Adaptive AI Systems (Position Track) Session Chair: Annabel Latham (Manchester Metropolitan University), Akira Hirose (The University of Tokyo) Position Paper: Post-Solve Robustness in Decision Engines: Feasible Regions and Smoothness Under Perturbations pdfPosition Paper: Uncertainty Quantification in Deep Learning Is Unsatisfactory for Clinical Applications and Complex Decision Making pdfPosition Paper: Just as Humans Need Vaccines, So Do Models: Model Immunization to Combat Falsehoods pdfThursday 0.11 Cape Town IJCNN Paper IJCNN SS07 Quantum Machine Learning Algorithms and Applications I Session Chair: Samuel Yen-Chi Chen (Wells Fargo) Thursday 0.14 Singapore IJCNN Paper Reinforcement Learning, Control, and Autonomous Systems Session Chair: Jen-Tzung Chien (National Yang Ming Chiao Tung University), Jiajie Zhang (Technical University of Munich) A Performance Model for Deadline-Aware Off-Policy Reinforcement Learning with Vectorized Environments pdfReducing Experience-Level Non-Stationarity in Multi-Agent Reinforcement Learning via Policy Constrained Replay. pdfThursday 0.15 Washington IJCNN Paper IJCNN SS03 Physics-Informed Neural Networks: Advancements and Applications Session Chair: Ciaran Bench (National Physical Laboratory), Michiel Straat (Bielefeld University) Thursday 0.02 Berlin IJCNN Paper, FUZZ-IEEE Position Paper, CEC Late Breaking Paper, CEC Paper, FUZZ J2C Presentation, CEC J2C Presentation, FUZZ-IEEE Paper, CEC Position Paper, IJCNN J2C Presentation, IJCNN Position Paper, IJCNN Late Breaking Paper, FUZZ-IEEE Late Breaking Paper Regular Best Paper Award Self-Attention-Guided Genetic Programming for Dynamic Scheduling: Leveraging BERT for Enhanced Tree-Structured Data Operations pdfMulti-Fidelity Multi-Objective Optimization of Electric Machines Having Heterogeneous and Blocked Evaluation Times pdfHow Many Subproblems? A Controlled Study of Decomposition Size in Multi-Policy Multi-Objective Reinforcement Learning pdfThursday 0.04 Brussels IEEE CEC (Evolutionary Computation) CEC 21- Algorithms VI Session Chair: Ruibin Bai (University of Nottingham Ningbo China) Reinforcement Learning for Job-Shop Scheduling via Hybrid GNN Embeddings with Sophisticated State Representations pdfAn Interactive Multi-Objective Optimization Method for Deriving Measures to Social Issues while Presenting the Predicted Pareto Frontier pdfThursday 0.05 Paris IEEE CEC (Evolutionary Computation) CEC 22 - SS19: Evolutionary Computer Vision and Image Processing (ECVIP) Session Chair: Ying Bi (School of Electrical and Information Engineering, Zhengzhou University, China; State Key Laboratory of Intelligent Agricultural Power Equipment) |
