Change search
Link to record
Permanent link

Direct link
Publications (10 of 49) Show all publications
Amiri, M., Viswanathan, V. & Magnússon, S. (2026). Challenger-Based Combinatorial Bandits for Subcarrier Selection in OFDM Systems. In: 2026 IEEE Wireless Communications and Networking Conference (WCNC): . Paper presented at 2026 IEEE Wireless Communications and Networking Conference (WCNC), 13-16 April 2026, Kuala Lumpur, Malaysia..
Open this publication in new window or tab >>Challenger-Based Combinatorial Bandits for Subcarrier Selection in OFDM Systems
2026 (English)In: 2026 IEEE Wireless Communications and Networking Conference (WCNC), 2026Conference paper, Published paper (Refereed)
Abstract [en]

This paper investigates the identification of the top-m user-scheduling sets in multi-user MIMO downlink, which is cast as a combinatorial pure-exploration problem in stochastic linear bandits. Because the action space grows exponentially, exhaustive search is infeasible. We therefore adopt a linear utility model to enable efficient exploration and reliable selection of promising user subsets. We introduce a gap-index framework that maintains a shortlist of current estimates of champion arms (top-m sets) and a rotating shortlist of challenger arms that pose the greatest threat to the champions. This design focuses on measurements that yield the most informative gapindex-based comparisons, resulting in significant reductions in runtime and computation compared to state-of-the-art linear bandit methods, with high identification accuracy. The methodalso exposes a tunable trade-off between speed and accuracy. Simulations on a realistic OFDM downlink show that shortlist driven pure exploration makes online, measurement-efficientsubcarrier selection practical for AI-enabled communication systems.

Series
IEEE Wireless Communications and Networking Conference, ISSN 1525-3511, E-ISSN 1558-2612
National Category
Electrical Engineering, Electronic Engineering, Information Engineering Other Social Sciences Computer Sciences
Identifiers
urn:nbn:se:su:diva-257900 (URN)10.1109/WCNC65185.2026.11555289 (DOI)2-s2.0-105042741769 (Scopus ID)
Conference
2026 IEEE Wireless Communications and Networking Conference (WCNC), 13-16 April 2026, Kuala Lumpur, Malaysia.
Funder
Vinnova, 2024-04058
Available from: 2026-08-04 Created: 2026-08-04 Last updated: 2026-08-17Bibliographically approved
Tian, H., Li, X., Wu, Z. & Magnússon, S. (2026). Communication-efficient online federated composite optimization. Automatica, 183, Article ID 112679.
Open this publication in new window or tab >>Communication-efficient online federated composite optimization
2026 (English)In: Automatica, ISSN 0005-1098, E-ISSN 1873-2836, Vol. 183, article id 112679Article in journal (Refereed) Published
Abstract [en]

Online federated optimization is crucial for sequential decision-making in dynamic environments, yet it often overlooks non-smooth regularizers, which are common in real-world applications such as machine learning and wireless communication. Additionally, communication overhead is an important facet in practice. As such, we focus on addressing the two issues by introducing the online federated composite optimization problem, where the loss function is time-varying and contains a non-smooth regularizer, and employing compressors to reduce communication overhead. An algorithm, named FedOEC is proposed, which simultaneously resolves online optimization problems with non-smooth regularizers and reduces communication overhead by leveraging multi-kernel and compressors efficiently. Through theoretical analysis, FedOEC achieves optimal sublinear regret bound with time-varying step sizes in convex settings, where T represents the number of communication rounds. Finally, numerical experiments confirm the effectiveness of the proposed algorithm.

Keywords
Compression, Federated composite optimization, Multi-kernel learning, Online learning, Regret
National Category
Artificial Intelligence
Identifiers
urn:nbn:se:su:diva-250083 (URN)10.1016/j.automatica.2025.112679 (DOI)001620814000001 ()2-s2.0-105021474415 (Scopus ID)
Available from: 2025-12-05 Created: 2025-12-05 Last updated: 2025-12-05Bibliographically approved
Amiri, M., Avrachenkov, K., El Mimouni, I. & Magnússon, S. (2026). MARBLE: Multi-Armed Restless Bandits in Latent Markovian Environment. In: 2026 European Control Conference (ECC): . Paper presented at 2026 European Control Conference (ECC), 7-10 July 2026, Reykjavík, Iceland. (pp. 956-962). IEEE
Open this publication in new window or tab >>MARBLE: Multi-Armed Restless Bandits in Latent Markovian Environment
2026 (English)In: 2026 European Control Conference (ECC), IEEE, 2026, p. 956-962Conference paper, Published paper (Refereed)
Abstract [en]

Restless Multi-Armed Bandits (RMABs) are powerful models for decision-making under uncertainty, yet classical formulations typically assume fixed dynamics, an assumption often violated in nonstationary environments. We introduce MARBLE (Multi-Armed Restless Bandits in a Latent Markovian Environment), which augments RMABs with a latent Markov state that induces nonstationary behavior. In MARBLE, each arm evolves according to a latent environment state that switches over time, making policy learning substantially more challenging. We further introduce the Markov-Averaged Indexability (MAI) criterion as a relaxed indexability assumption and prove that, despite unobserved regime switches, under the MAI criterion, synchronous Q-learning with Whittle Indices (QWI) converges almost surely to the optimal Q-function and the corresponding Whittle indices. We validate MARBLE on a calibrated simulator-embedded (digital twin) recommender system, where QWI consistently adapts to a shifting latent state and converges to an optimal policy, empirically corroborating our theoretical findings.

Place, publisher, year, edition, pages
IEEE, 2026
National Category
Electrical Engineering, Electronic Engineering, Information Engineering Information Systems, Social aspects
Identifiers
urn:nbn:se:su:diva-258034 (URN)
Conference
2026 European Control Conference (ECC), 7-10 July 2026, Reykjavík, Iceland.
Available from: 2026-08-12 Created: 2026-08-12 Last updated: 2026-08-13Bibliographically approved
Vaishnav, S., Donta, P. K. & Magnússon, S. (2025). Adaptive Budgeted Multi-Armed Bandits for IoT with Dynamic Resource Constraints. In: GLOBECOM 2025 - 2025 IEEE Global Communications Conference: . Paper presented at 2025 IEEE Global Communications Conference (GLOBECOM 2025), Taipei, Taiwan, 8-12 December, 2025 (pp. 4535-4540). Piscataway: IEEE
Open this publication in new window or tab >>Adaptive Budgeted Multi-Armed Bandits for IoT with Dynamic Resource Constraints
2025 (English)In: GLOBECOM 2025 - 2025 IEEE Global Communications Conference, Piscataway: IEEE, 2025, p. 4535-4540Conference paper, Published paper (Refereed)
Abstract [en]

Internet of Things (IoT) systems increasingly operate in environments where devices must respond in real time while managing fluctuating resource constraints, including energy and bandwidth. Yet, current approaches often fall short in addressing scenarios where operational constraints evolve over time. To address these limitations, we propose a novel Budgeted Multi-Armed Bandit framework tailored for IoT applications with dynamic operational limits. Our model introduces a decaying violation budget, which permits limited constraint violations early in the learning process and gradually enforces stricter compliance over time. We present the Budgeted Upper Confidence Bound (UCB) algorithm, which adaptively balances performance optimization and compliance with time-varying constraints. We provide theoretical guarantees showing that Budgeted UCB achieves sublinear regret and logarithmic constraint violations over the learning horizon. Extensive simulations in a wireless communication setting show that our approach achieves faster adaptation and better constraint satisfaction than standard online learning methods. These results highlight the framework’s potential for building adaptive, resource-aware IoT systems.

Place, publisher, year, edition, pages
Piscataway: IEEE, 2025
Series
IEEE Conference on Global Communications (GLOBECOM), E-ISSN 2576-6813
Keywords
Online Learning, Multi-Armed Bandits, Upper Confidence Bound, Dynamic Constraints, Internet of Things
National Category
Information Systems
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-254019 (URN)10.1109/GLOBECOM59602.2025.11432479 (DOI)2-s2.0-105036346401 (Scopus ID)979-8-3315-7781-0 (ISBN)979-8-3315-7782-7 (ISBN)
Conference
2025 IEEE Global Communications Conference (GLOBECOM 2025), Taipei, Taiwan, 8-12 December, 2025
Projects
Digital Futures (project DEMOCRITUS) and the Swedish Research Council (Vetenskapsr˚adet), grant 2024-04058.
Available from: 2026-04-02 Created: 2026-04-02 Last updated: 2026-06-04Bibliographically approved
Beikmohammadi, A., Khirirat, S., Richtárik, P. & Magnússon, S. (2025). Collaborative Value Function Estimation Under Model Mismatch: A Federated Temporal Difference Analysis. In: Rita P. Ribeiro; Bernhard Pfahringer; Nathalie Japkowicz; Pedro Larrañaga; Alípio M. Jorge; Carlos Soares; Pedro H. Abreu; João Gama (Ed.), Machine Learning and Knowledge Discovery in Databases. Research Track: European Conference, ECML PKDD 2025, Porto, Portugal, September 15–19, 2025, Proceedings, Part VI. Paper presented at European Conference, ECML PKDD 2025, Porto, Portugal, September 15–19, 2025. (pp. 41-58). Springer
Open this publication in new window or tab >>Collaborative Value Function Estimation Under Model Mismatch: A Federated Temporal Difference Analysis
2025 (English)In: Machine Learning and Knowledge Discovery in Databases. Research Track: European Conference, ECML PKDD 2025, Porto, Portugal, September 15–19, 2025, Proceedings, Part VI / [ed] Rita P. Ribeiro; Bernhard Pfahringer; Nathalie Japkowicz; Pedro Larrañaga; Alípio M. Jorge; Carlos Soares; Pedro H. Abreu; João Gama, Springer , 2025, p. 41-58Conference paper, Published paper (Refereed)
Abstract [en]

Federated reinforcement learning (FedRL) enables collaborative learning while preserving data privacy by preventing direct data exchange between agents. However, many existing FedRL algorithms assume that all agents operate in identical environments, which is often unrealistic. In real-world applications, such as multi-robot teams, crowdsourced systems, and large-scale sensor networks, each agent may experience slightly different transition dynamics, leading to inherent model mismatches. In this paper, we first establish linear convergence guarantees for single-agent temporal difference learning (TD(0)) in policy evaluation and demonstrate that under a perturbed environment, the agent suffers a systematic bias that prevents accurate estimation of the true value function. This result holds under both i.i.d. and Markovian sampling regimes. We then extend our analysis to the federated TD(0) (FedTD(0)) setting, where multiple agents, each interacting with its own perturbed environment, periodically share value estimates to collaboratively approximate the true value function of a common underlying model. Our theoretical results indicate the impact of model mismatch, network connectivity, and mixing behavior on the convergence of FedTD(0). Empirical experiments corroborate our theoretical gains, highlighting that even moderate levels of information sharing significantly mitigate environment-specific errors.

Place, publisher, year, edition, pages
Springer, 2025
Series
Lecture Notes in Computer Science (LNCS), ISSN 0302-9743, E-ISSN 1611-3349 ; 16018
Keywords
Federated Reinforcement Learning, Model Mismatch in Reinforcement Learning, Temporal Difference Learning, Policy Evaluation
National Category
Computer Sciences
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-248239 (URN)10.1007/978-3-032-06106-5_3 (DOI)2-s2.0-105020022174 (Scopus ID)978-3-032-06106-5 (ISBN)978-3-032-06105-8 (ISBN)
Conference
European Conference, ECML PKDD 2025, Porto, Portugal, September 15–19, 2025.
Available from: 2025-10-20 Created: 2025-10-20 Last updated: 2026-07-16Bibliographically approved
Vaishnav, S., Khirirat, S. & Magnússon, S. (2025). Communication-Adaptive Gradient Sparsification for Federated Learning with Error Compensation. IEEE Internet of Things Journal, 12(2), 1137-1152
Open this publication in new window or tab >>Communication-Adaptive Gradient Sparsification for Federated Learning with Error Compensation
2025 (English)In: IEEE Internet of Things Journal, ISSN 2327-4662, Vol. 12, no 2, p. 1137-1152Article in journal (Other academic) Published
Abstract [en]

Federated learning has emerged as a popular distributed machine-learning paradigm. It involves many rounds of iterative communication between nodes to exchange model parameters. With the increasing complexity of ML tasks, the models can be large, having millions of parameters. Moreover, edge and IoT nodes often have limited energy resources and channel bandwidths. Thus, reducing the communication cost in Federated Learning is a bottleneck problem. This cost could be in terms of energy consumed, delay involved, or amount of data communicated. We propose a communication cost-adaptive model sparsification for Federated Learning with Error Compensation. The central idea is to adapt the sparsification level in run-time by optimizing the ratio between the impact of the communicated model parameters and communication cost. We carry out a detailed convergence analysis to establish the theoretical foundations of the proposed algorithm. We conduct extensive experiments to train both convex and non-convex machine learning models on a standard dataset. We illustrate the efficiency of the proposed algorithm by comparing its performance with three baseline schemes. The performance of the proposed algorithm is validated for two communication models and three cost functions. Simulation results show that the proposed algorithm needs a substantially less amount of communication than the three baseline schemes while achieving the best accuracy and fastest convergence. The results are consistent for all the considered cost models, cost functions, and ML models. Thus, the proposed FL-CATE algorithm can substantially improve the communication efficiency of federated learning, irrespective of the ML tasks, costs, and communication models.

Keywords
Federated learning, Communication efficiency, IoT, Gradient sparsification, Distributed learning
National Category
Computer Sciences
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-235702 (URN)10.1109/JIOT.2024.3490855 (DOI)001395714600019 ()2-s2.0-85208723002 (Scopus ID)
Note

The article is available online under early access area on IEEE Xplore. This article has been accepted for publication in a future issue of this journal, but has not been edited and content may change prior to final publication. It may be cited as an article in a future issue by its Digital Object Identifier.

Available from: 2024-11-19 Created: 2024-11-19 Last updated: 2026-04-13Bibliographically approved
Wang, H., Huang, W., Magnússon, S., Lindgren, T., Chen, C., Wu, J. & Song, Y. (2025). Crowding distance and IGD-driven grey wolf reinforcement learning approach for multi-objective agile earth observation satellite scheduling. International Journal of Digital Earth, 18(1), Article ID 2458024.
Open this publication in new window or tab >>Crowding distance and IGD-driven grey wolf reinforcement learning approach for multi-objective agile earth observation satellite scheduling
Show others...
2025 (English)In: International Journal of Digital Earth, ISSN 1753-8947, E-ISSN 1753-8955, Vol. 18, no 1, article id 2458024Article in journal (Refereed) Published
Abstract [en]

With the rise of low-cost launches, miniaturized space technology, and commercialization, the cost of space missions has dropped, leading to a surge in flexible Earth observation satellites. This increased demand for complex and diverse imaging products requires addressing multi-objective optimization in practice. To this end, we propose a multi-objective agile Earth observation satellite scheduling problem (MOAEOSSP) model and introduce a reinforcement learning-based multi-objective grey wolf optimization (RLMOGWO) algorithm. It aims to maximize observation efficiency while minimizing energy consumption. During population initialization, the algorithm uses chaos mapping and opposition-based learning to enhance diversity and global search, reducing the risk of local optima. It integrates Q-learning into an improved multi-objective grey wolf optimization framework, designing state-action combinations that balance exploration and exploitation. Dynamic parameter adjustments guide position updates, boosting adaptability across different optimization stages. Moreover, the algorithm introduces a reward mechanism based on the crowding distance and inverted generational distance (IGD) to maintain Pareto front diversity and distribution, ensuring a strong multi-objective optimization performance. The experimental results show that the algorithm excels at solving the MOAEOSSP, outperforming competing algorithms across several metrics and demonstrating its effectiveness for complex optimization problems.

Keywords
Earth observation satellite, grey wolf algorithm, multi-objective optimization, Q-learning, reinforcement learning, scheduling
National Category
Computer Sciences
Identifiers
urn:nbn:se:su:diva-240188 (URN)10.1080/17538947.2025.2458024 (DOI)001410804300001 ()2-s2.0-85216608663 (Scopus ID)
Available from: 2025-03-04 Created: 2025-03-04 Last updated: 2025-03-04Bibliographically approved
Makridis, E., Magnússon, S. & Charalambous, T. (2025). Distributed Gradient-Tracking Optimization with Packet-Error Resilience in Unreliable Networks. In: 2025 European Control Conference (ECC): . Paper presented at 23rd European Control Conference (ECC 2025), Thessaloniki, Greece, 24-27 June, 2025 (pp. 635-640). Piscataway: IEEE
Open this publication in new window or tab >>Distributed Gradient-Tracking Optimization with Packet-Error Resilience in Unreliable Networks
2025 (English)In: 2025 European Control Conference (ECC), Piscataway: IEEE, 2025, p. 635-640Conference paper, Published paper (Refereed)
Abstract [en]

In this paper, we address the distributed optimization problem over unreliable error-prone directed networks. We propose a distributed gradient-tracking optimization algorithm (referred to as ARQ-OPT), which exploits packet retransmissions via an Automatic Repeat reQuest (ARQ) error control protocol. Nodes utilize acknowledgement messages transmitted over one-bit error-free channels to trigger retransmissions of packets that were previously received in error. This ensures reliable propagation of information throughout the network, even in the presence of packet errors. We analyze the convergence properties of the proposed algorithm, by augmenting the consensus matrices to align with the retransmission mechanism. Subsequently, we show that by appropriately choosing the maximum number of retransmission attempts, ARQ-OPT can achieve B-step consensus contractivity which allow us to establish asymptotic convergence to the unique optimal solution with probability one. Numerical simulations conducted under various channel conditions validate our findings.

Place, publisher, year, edition, pages
Piscataway: IEEE, 2025
Series
European Control Conference Proceedings, ISSN 2996-8917, E-ISSN 2996-8895
Keywords
ARQ, directed graphs, distributed optimization, gradient tracking, packet-errors, time-varying delays
National Category
Networked, Parallel and Distributed Computing
Identifiers
urn:nbn:se:su:diva-253469 (URN)10.23919/ECC65951.2025.11187185 (DOI)2-s2.0-105030981460 (Scopus ID)978-3-907144-12-1 (ISBN)979-8-3315-0271-3 (ISBN)
Conference
23rd European Control Conference (ECC 2025), Thessaloniki, Greece, 24-27 June, 2025
Available from: 2026-03-13 Created: 2026-03-13 Last updated: 2026-03-13Bibliographically approved
Beikmohammadi, A. & Magnússon, S. (2025). Human-inspired framework to accelerate reinforcement learning. Journal of Supercomputing, 81(12), Article ID 1239.
Open this publication in new window or tab >>Human-inspired framework to accelerate reinforcement learning
2025 (English)In: Journal of Supercomputing, ISSN 0920-8542, E-ISSN 1573-0484, Vol. 81, no 12, article id 1239Article in journal (Refereed) Published
Abstract [en]

Reinforcement learning (RL) is crucial for data science decision-making but suffers from sample inefficiency, particularly in real-world scenarios with costly physical interactions. This paper introduces a novel human-inspired framework to enhance the RL algorithm’s sample efficiency. It achieves this by initially exposing the learning agent to simpler tasks that progressively increase in complexity, ultimately leading to the main task. This method requires no pre-training and involves learning simpler tasks for just one episode. The resulting knowledge can facilitate various transfer learning approaches, such as value and policy transfer, without increasing computational complexity. It can be applied across different goals, environments, and RL algorithms, including value-based, policy-based, tabular, and deep RL methods. Experimental evaluations demonstrate the framework’s effectiveness in enhancing sample efficiency, especially in challenging main tasks, demonstrated through both a simple random walk and more complex optimal control problems with constraints.

Keywords
Deep reinforcement learning, Exploration, Policy optimization, PPO, Sample efficiency
National Category
Human Computer Interaction
Identifiers
urn:nbn:se:su:diva-246714 (URN)10.1007/s11227-025-07737-2 (DOI)001550362100002 ()2-s2.0-105013224205 (Scopus ID)
Available from: 2025-09-11 Created: 2025-09-11 Last updated: 2025-09-11Bibliographically approved
Beikmohammadi, A., Khirirat, S. & Magnússon, S. (2025). On the Convergence of Federated Learning Algorithms Without Data Similarity. IEEE Transactions on Big Data, 11(2), 659-668
Open this publication in new window or tab >>On the Convergence of Federated Learning Algorithms Without Data Similarity
2025 (English)In: IEEE Transactions on Big Data, E-ISSN 2332-7790, Vol. 11, no 2, p. 659-668Article in journal (Refereed) Published
Abstract [en]

Data similarity assumptions have traditionally been relied upon to understand the convergence behaviors of federated learning methods. Unfortunately, this approach often demands fine-tuning step sizes based on the level of data similarity. When data similarity is low, these small step sizes result in an unacceptably slow convergence speed for federated methods. In this paper, we present a novel and unified framework for analyzing the convergence of federated learning algorithms without the need for data similarity conditions. Our analysis centers on an inequality that captures the influence of step sizes on algorithmic convergence performance. By applying our theorems to well-known federated algorithms, we derive precise expressions for three widely used step size schedules: fixed, diminishing, and step-decay step sizes, which are independent of data similarity conditions. Finally, we conduct comprehensive evaluations of the performance of these federated learning algorithms, employing the proposed step size strategies to train deep neural network models on benchmark datasets under varying data similarity conditions. Our findings demonstrate significant improvements in convergence speed and overall performance, marking a substantial advancement in federated learning research.

Keywords
Compression algorithms, federated learning, gradient methods, machine learning
National Category
Computer Sciences
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-232103 (URN)10.1109/TBDATA.2024.3423693 (DOI)001445067800004 ()2-s2.0-105001081071 (Scopus ID)
Available from: 2024-07-24 Created: 2024-07-24 Last updated: 2026-07-16Bibliographically approved
Organisations
Identifiers
ORCID iD: ORCID iD iconorcid.org/0000-0002-6617-8683

Search in DiVA

Show all publications