Multimodal dialogue systems often fail to maintain coherent reasoning over extended conversations and suffer from hallucination due to limited context modeling capabilities.Current approaches struggle with crossmodal ...Multimodal dialogue systems often fail to maintain coherent reasoning over extended conversations and suffer from hallucination due to limited context modeling capabilities.Current approaches struggle with crossmodal alignment,temporal consistency,and robust handling of noisy or incomplete inputs across multiple modalities.We propose Multi Agent-Chain of Thought(CoT),a novel multi-agent chain-of-thought reasoning framework where specialized agents for text,vision,and speech modalities collaboratively construct shared reasoning traces through inter-agent message passing and consensus voting mechanisms.Our architecture incorporates self-reflection modules,conflict resolution protocols,and dynamic rationale alignment to enhance consistency,factual accuracy,and user engagement.The framework employs a hierarchical attention mechanism with cross-modal fusion and implements adaptive reasoning depth based on dialogue complexity.Comprehensive evaluations on Situated Interactive Multi-Modal Conversations(SIMMC)2.0,VisDial v1.0,and newly introduced challenging scenarios demonstrate statistically significant improvements in grounding accuracy(p<0.01),chain-of-thought interpretability,and robustness to adversarial inputs compared to state-of-the-art monolithic transformer baselines and existing multi-agent approaches.展开更多
This paper focuses on the leader-following positive consensus problems of heterogeneous switched multi-agent systems.First,a state-feedback controller with dynamic compensation is introduced to achieve positive consen...This paper focuses on the leader-following positive consensus problems of heterogeneous switched multi-agent systems.First,a state-feedback controller with dynamic compensation is introduced to achieve positive consensus under average dwell time switching.Then sufficient conditions are derived to guarantee the positive consensus.The gain matrices of the control protocol are described using a matrix decomposition approach and the corresponding computational complexity is reduced by resorting to linear programming and co-positive Lyapunov functions.Finally,two numerical examples are provided to illustrate the results obtained.展开更多
With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier...With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier heterogeneous architecture composed of mobile devices,unmanned aerial vehicles(UAVs),and macro base stations(BSs).This scenario typically faces fast channel fading,dynamic computational loads,and energy constraints,whereas classical queuing-theoretic or convex-optimization approaches struggle to yield robust solutions in highly dynamic settings.To address this issue,we formulate a multi-agent Markov decision process(MDP)for an air-ground-fused MEC system,unify link selection,bandwidth/power allocation,and task offloading into a continuous action space and propose a joint scheduling strategy that is based on an improved MATD3 algorithm.The improvements include Alternating Layer Normalization(ALN)in the actor to suppress gradient variance,Residual Orthogonalization(RO)in the critic to reduce the correlation between the twin Q-value estimates,and a dynamic-temperature reward to enable adaptive trade-offs during training.On a multi-user,dual-link simulation platform,we conduct ablation and baseline comparisons.The results reveal that the proposed method has better convergence and stability.Compared with MADDPG,TD3,and DSAC,our algorithm achieves more robust performance across key metrics.展开更多
This paper investigates the consensus tracking control problem for high order nonlinear multi-agent systems subject to non-affine faults,partial measurable states,uncertain control coefficients,and unknown external di...This paper investigates the consensus tracking control problem for high order nonlinear multi-agent systems subject to non-affine faults,partial measurable states,uncertain control coefficients,and unknown external disturbances.Under the directed topology conditions,an observer-based finite-time control strategy based on adaptive backstepping and is proposed,in which a neural network-based state observer is employed to approximate the unmeasurable system state variables.To address the complexity explosion problem associated with the backstepping method,a finite-time command filter is incorporated,with error compensation signals designed to mitigate the filter-induced errors.Additionally,the Butterworth low-pass filter is introduced to avoid the algebraic ring problem in the design of the controller.The finite-time stability of the closed-loop system is rigorously analyzed with the finite-time Lyapunov stability criterion,validating that all closed-loop signals of the system remain bounded within a finite time.Finally,the effectiveness of the proposed control strategy is verified through a simulation example.展开更多
This paper mainly focuses on the velocity-constrained consensus problem of discrete-time heterogeneous multi-agent systems with nonconvex constraints and arbitrarily switching topologies,where each agent has first-ord...This paper mainly focuses on the velocity-constrained consensus problem of discrete-time heterogeneous multi-agent systems with nonconvex constraints and arbitrarily switching topologies,where each agent has first-order or second-order dynamics.To solve this problem,a distributed algorithm is proposed based on a contraction operator.By employing the properties of the stochastic matrix,it is shown that all agents’position states could converge to a common point and second-order agents’velocity states could remain in corresponding nonconvex constraint sets and converge to zero as long as the joint communication topology has one directed spanning tree.Finally,the numerical simulation results are provided to verify the effectiveness of the proposed algorithms.展开更多
The Internet of Unmanned Aerial Vehicles(I-UAVs)is expected to execute latency-sensitive tasks,but limited by co-channel interference and malicious jamming.In the face of unknown prior environmental knowledge,defendin...The Internet of Unmanned Aerial Vehicles(I-UAVs)is expected to execute latency-sensitive tasks,but limited by co-channel interference and malicious jamming.In the face of unknown prior environmental knowledge,defending against jamming and interference through spectrum allocation becomes challenging,especially when each UAV pair makes decisions independently.In this paper,we propose a cooperative multi-agent reinforcement learning(MARL)-based anti-jamming framework for I-UAVs,enabling UAV pairs to learn their own policies cooperatively.Specifically,we first model the problem as a modelfree multi-agent Markov decision process(MAMDP)to maximize the long-term expected system throughput.Then,for improving the exploration of the optimal policy,we resort to optimizing a MARL objective function with a mutual-information(MI)regularizer between states and actions,which can dynamically assign the probability for actions frequently used by the optimal policy.Next,through sharing their current channel selections and local learning experience(their soft Q-values),the UAV pairs can learn their own policies cooperatively relying on only preceding observed information and predicting others’actions.Our simulation results show that for both sweep jamming and Markov jamming patterns,the proposed scheme outperforms the benchmarkers in terms of throughput,convergence and stability for different numbers of jammers,channels and UAV pairs.展开更多
This paper investigates the observer-based prescribed-time time-varying output formation-containment(PT-TV-OFC)control problem for heterogeneous multi-agent systems in which the different agents have different state d...This paper investigates the observer-based prescribed-time time-varying output formation-containment(PT-TV-OFC)control problem for heterogeneous multi-agent systems in which the different agents have different state dimensions.The system comprises one tracking leader,multiple formation leaders,and followers,where two types of leaders are used to generate a reference trajectory for movement and achieve specific formation,respectively.Firstly,a prescribed-time dynamics observer is constructed for the formation leaders to estimate the tracking leader's dynamic model and state.On this basis,a prescribed-time control protocol is designed for the formation leaders to achieve time-varying output formation.Then,a prescribed-time convex hull observer is designed for the followers to estimate information regarding the convex hull formed by the formation leaders.Using the estimated convex hull information,a prescribed-time containment control protocol is designed to ensure the followers converge into the convex hull.Furthermore,using Lyapunov stability theory,the stability of systems is proved in detail,which implies that the heterogeneous multi-agent systems can achieve PT-TV-OFC control.Finally,numerical simulations validate the feasibility of the theoretical results.展开更多
This paper delves into the problem of optimal placement conditions for a group of agents collaboratively localizing a target using range-only or bearing-only measurements.The challenge in this study stems from the unc...This paper delves into the problem of optimal placement conditions for a group of agents collaboratively localizing a target using range-only or bearing-only measurements.The challenge in this study stems from the uncertainty associated with the positions of the agents,which may experience drift or disturbances during the target localization process.Initially,we derive the Cramer-Rao lower bound(CRLB)of the target position as the primary analytical metric.Subsequently,we establish the necessary and sufficient conditions for the optimal placement of agents.Based on these conditions,we analyze the maximal allowable agent position error for an expected mean squared error(MSE),providing valuable guidance for the selection of agent positioning sensors.The analytical findings are further validated through simulation experiments.展开更多
Policy training against diverse opponents remains a challenge when using Multi-Agent Reinforcement Learning(MARL)in multiple Unmanned Combat Aerial Vehicle(UCAV)air combat scenarios.In view of this,this paper proposes...Policy training against diverse opponents remains a challenge when using Multi-Agent Reinforcement Learning(MARL)in multiple Unmanned Combat Aerial Vehicle(UCAV)air combat scenarios.In view of this,this paper proposes a novel Dominant and Non-dominant strategy sample selection(DoNot)mechanism and a Local Observation Enhanced Multi-Agent Proximal Policy Optimization(LOE-MAPPO)algorithm to train the multi-UCAV air combat policy and improve its generalization.Specifically,the LOE-MAPPO algorithm adopts a mixed state that concatenates the global state and individual agent's local observation to enable efficient value function learning in multi-UCAV air combat.The DoNot mechanism classifies opponents into dominant or non-dominant strategy opponents,and samples from easier to more challenging opponents to form an adaptive training curriculum.Empirical results demonstrate that the proposed LOE-MAPPO algorithm outperforms baseline MARL algorithms in multi-UCAV air combat scenarios,and the DoNot mechanism leads to stronger policy generalization when facing diverse opponents.The results pave the way for the fast generation of cooperative strategies for air combat agents with MARLalgorithms.展开更多
This paper addresses the time-varying formation-containment(FC) problem for nonholonomic multi-agent systems with a desired trajectory constraint, where only the leaders can acquire information about the desired traje...This paper addresses the time-varying formation-containment(FC) problem for nonholonomic multi-agent systems with a desired trajectory constraint, where only the leaders can acquire information about the desired trajectory. Input the fixed time-varying formation template to the leader and start executing, this process also needs to track the desired trajectory, and the follower needs to converge to the convex hull that the leader crosses. Firstly, the dynamic models of nonholonomic systems are linearized to second-order dynamics. Then, based on the desired trajectory and formation template, the FC control protocols are proposed. Sufficient conditions to achieve FC are introduced and an algorithm is proposed to resolve the control parameters by solving an algebraic Riccati equation. The system is demonstrated to achieve FC, with the average position and velocity of the leaders converging asymptotically to the desired trajectory. Finally, the theoretical achievements are verified in simulations by a multi-agent system composed of virtual human individuals.展开更多
An Interval Type-2(IT-2)fuzzy controller design approach is proposed in this research to simultaneously achievemultiple control objectives inNonlinearMulti-Agent Systems(NMASs),including formation,containment,and coll...An Interval Type-2(IT-2)fuzzy controller design approach is proposed in this research to simultaneously achievemultiple control objectives inNonlinearMulti-Agent Systems(NMASs),including formation,containment,and collision avoidance.However,inherent nonlinearities and uncertainties present in practical control systems contribute to the challenge of achieving precise control performance.Based on the IT-2 Takagi-Sugeno Fuzzy Model(T-SFM),the fuzzy control approach can offer a more effective solution for NMASs facing uncertainties.Unlike existing control methods for NMASs,the Formation and Containment(F-and-C)control problem with collision avoidance capability under uncertainties based on the IT-2 T-SFM is discussed for the first time.Moreover,an IT-2 fuzzy tracking control approach is proposed to solve the formation task for leaders in NMASs without requiring communication.This control scheme makes the design process of the IT-2 fuzzy Formation Controller(FC)more straightforward and effective.According to the communication interaction protocol,the IT-2 Containment Controller(CC)design approach is proposed for followers to ensure convergence into the region defined by the leaders.Leveraging the IT-2 T-SFM representation,the analysis methods developed for linear Multi-Agent Systems(MASs)are successfully extended to perform containment analysis without requiring the additional assumptions imposed in existing research.Notably,the IT-2 fuzzy tracking controller can also be applied in collision avoidance situations to track the desired trajectories calculated by the avoidance algorithm under the Artificial Potential Field(APF).Benefiting from the combination of vortex and source APFs,the leaders can properly adjust the system dynamics to prevent potential collision risk.Integrating the fuzzy theory and APFs avoidance algorithm,an IT-2 fuzzy controller design approach is proposed to achieve the F-and-C purposewhile ensuring collision avoidance capability.Finally,amulti-ship simulation is conducted to validate the feasibility and effectiveness of the designed IT-2 fuzzy controller.展开更多
Unmanned Aerial Vehicles(UAVs)have demonstrated significant potential as Aerial Base Stations(A-BSs)for providing data services to Ground Users(GUs),attributed to their flexibility,cost-effectiveness,and high likeliho...Unmanned Aerial Vehicles(UAVs)have demonstrated significant potential as Aerial Base Stations(A-BSs)for providing data services to Ground Users(GUs),attributed to their flexibility,cost-effectiveness,and high likelihood of establishing line-of-sight links.In this article,we formulate the joint power and trajectory optimization problem for a multi-UAV assisted wireless network with no-fly zones constrained,aiming at maximizing the Accumulated Service Data(ASD)of UAVs and minimizing the Average End Age of Information(AEAoI)of GUs.Specifically,this paper proposes the Multi-Agent worst-case Soft Actor Critic(MA-wcSAC)algorithm with a distributional safety-critic.The simulation results demonstrate that,compared to the Multi-Agent Soft Actor Critic(MA-SAC)algorithm,the proposed algorithm exhibits comparable data service performance while reducing security risks by at least 30%at different risk levels.展开更多
With the aid of multi-agent based modeling approach to complex systems, the hierarchy simulation models of carrier-based aircraft catapult launch are developed. Ocean, carrier, aircraft, and atmosphere are treated as ...With the aid of multi-agent based modeling approach to complex systems, the hierarchy simulation models of carrier-based aircraft catapult launch are developed. Ocean, carrier, aircraft, and atmosphere are treated as aggregation agents, the detailed components like catapult, landing gears, and disturbances are considered as meta-agents, which belong to their aggregation agent. Thus, the model with two layers is formed i.e. the aggregation agent layer and the meta-agent layer. The information communication among all agents is described. The meta-agents within one aggregation agent communicate with each other directly by information sharing, but the meta-agents, which belong to different aggregation agents exchange their information through the aggregation layer first, and then perceive it from the sharing environment, that is the aggregation agent. Thus, not only the hierarchy model is built, but also the environment perceived by each agent is specified. Meanwhile, the problem of balancing the independency of agent and the resource consumption brought by real-time communication within multi-agent system (MAS) is resolved. Each agent involved in carrier-based aircraft catapult launch is depicted, with considering the interaction within disturbed atmospheric environment and multiple motion bodies including carrier, aircraft, and landing gears. The models of reactive agents among them are derived based on tensors, and the perceived messages and inner frameworks of each agent are characterized. Finally, some results of a simulation instance are given. The simulation and modeling of dynamic system based on multi-agent system is of benefit to express physical concepts and logical hierarchy clearly and precisely. The system model can easily draw in kinds of other agents to achieve a precise simulation of more complex system. This modeling technique makes the complex integral dynamic equations of multibodies decompose into parallel operations of single agent, and it is convenient to expand, maintain, and reuse the program codes.展开更多
The similarities and differences between the container terminal logistics system(CTLS)and the Harvard-architecture computer system are compared in terms of organization and architecture.The mapping relation and the mo...The similarities and differences between the container terminal logistics system(CTLS)and the Harvard-architecture computer system are compared in terms of organization and architecture.The mapping relation and the modeling framework of the CTLS are presented based on multi-agent,and the successful algorithms in the computer domain are applied to the modeling framework,such as the dynamic priority and multilevel feedback scheduling algorithm.In addition,a model and simulation on a certain quay at Shanghai harbor is built up on the AnyLogic platform to support the decision-making of terminal on service cost.It validates the feasibility and creditability of the above systematic methodology.展开更多
文摘Multimodal dialogue systems often fail to maintain coherent reasoning over extended conversations and suffer from hallucination due to limited context modeling capabilities.Current approaches struggle with crossmodal alignment,temporal consistency,and robust handling of noisy or incomplete inputs across multiple modalities.We propose Multi Agent-Chain of Thought(CoT),a novel multi-agent chain-of-thought reasoning framework where specialized agents for text,vision,and speech modalities collaboratively construct shared reasoning traces through inter-agent message passing and consensus voting mechanisms.Our architecture incorporates self-reflection modules,conflict resolution protocols,and dynamic rationale alignment to enhance consistency,factual accuracy,and user engagement.The framework employs a hierarchical attention mechanism with cross-modal fusion and implements adaptive reasoning depth based on dialogue complexity.Comprehensive evaluations on Situated Interactive Multi-Modal Conversations(SIMMC)2.0,VisDial v1.0,and newly introduced challenging scenarios demonstrate statistically significant improvements in grounding accuracy(p<0.01),chain-of-thought interpretability,and robustness to adversarial inputs compared to state-of-the-art monolithic transformer baselines and existing multi-agent approaches.
基金supported by the National Natural Science Foundation of China(62463007,62463005)the Natural Science Foundation of Hainan Province(625RC710,625MS047)+1 种基金the System Control and Information Processing Education Ministry Key Laboratory Open Funding,China(Scip20240119)the Science Research Funding of Hainan University,China(KYQD(ZR)22180,KYQD(ZR)23180).
文摘This paper focuses on the leader-following positive consensus problems of heterogeneous switched multi-agent systems.First,a state-feedback controller with dynamic compensation is introduced to achieve positive consensus under average dwell time switching.Then sufficient conditions are derived to guarantee the positive consensus.The gain matrices of the control protocol are described using a matrix decomposition approach and the corresponding computational complexity is reduced by resorting to linear programming and co-positive Lyapunov functions.Finally,two numerical examples are provided to illustrate the results obtained.
文摘With the advent of sixth-generation mobile communications(6G),space-air-ground integrated networks have become mainstream.This paper focuses on collaborative scheduling for mobile edge computing(MEC)under a three-tier heterogeneous architecture composed of mobile devices,unmanned aerial vehicles(UAVs),and macro base stations(BSs).This scenario typically faces fast channel fading,dynamic computational loads,and energy constraints,whereas classical queuing-theoretic or convex-optimization approaches struggle to yield robust solutions in highly dynamic settings.To address this issue,we formulate a multi-agent Markov decision process(MDP)for an air-ground-fused MEC system,unify link selection,bandwidth/power allocation,and task offloading into a continuous action space and propose a joint scheduling strategy that is based on an improved MATD3 algorithm.The improvements include Alternating Layer Normalization(ALN)in the actor to suppress gradient variance,Residual Orthogonalization(RO)in the critic to reduce the correlation between the twin Q-value estimates,and a dynamic-temperature reward to enable adaptive trade-offs during training.On a multi-user,dual-link simulation platform,we conduct ablation and baseline comparisons.The results reveal that the proposed method has better convergence and stability.Compared with MADDPG,TD3,and DSAC,our algorithm achieves more robust performance across key metrics.
基金supported in part by the Beijing Natural Science Foundation under Grant 4252050in part by the National Science Fund for Distinguished Young Scholars under Grant 62425304in part by the Basic Science Center Programs of NSFC under Grant 62088101.
文摘This paper investigates the consensus tracking control problem for high order nonlinear multi-agent systems subject to non-affine faults,partial measurable states,uncertain control coefficients,and unknown external disturbances.Under the directed topology conditions,an observer-based finite-time control strategy based on adaptive backstepping and is proposed,in which a neural network-based state observer is employed to approximate the unmeasurable system state variables.To address the complexity explosion problem associated with the backstepping method,a finite-time command filter is incorporated,with error compensation signals designed to mitigate the filter-induced errors.Additionally,the Butterworth low-pass filter is introduced to avoid the algebraic ring problem in the design of the controller.The finite-time stability of the closed-loop system is rigorously analyzed with the finite-time Lyapunov stability criterion,validating that all closed-loop signals of the system remain bounded within a finite time.Finally,the effectiveness of the proposed control strategy is verified through a simulation example.
基金2024 Jiangsu Province Youth Science and Technology Talent Support Project2024 Yancheng Key Research and Development Plan(Social Development)projects,“Research and Application of Multi Agent Offline Distributed Trust Perception Virtual Wireless Sensor Network Algorithm”and“Research and Application of a New Type of Fishery Ship Safety Production Monitoring Equipment”。
文摘This paper mainly focuses on the velocity-constrained consensus problem of discrete-time heterogeneous multi-agent systems with nonconvex constraints and arbitrarily switching topologies,where each agent has first-order or second-order dynamics.To solve this problem,a distributed algorithm is proposed based on a contraction operator.By employing the properties of the stochastic matrix,it is shown that all agents’position states could converge to a common point and second-order agents’velocity states could remain in corresponding nonconvex constraint sets and converge to zero as long as the joint communication topology has one directed spanning tree.Finally,the numerical simulation results are provided to verify the effectiveness of the proposed algorithms.
基金supported in part by the National Natural Science Foundation of China under Grants 62001225,62071236,62071234 and U22A2002in part by the Major Science and Technology plan of Hainan Province under Grant ZDKJ2021022+1 种基金in part by the Scientific Research Fund Project of Hainan University under Grant KYQD(ZR)-21008in part by the Key Technologies R&D Program of Jiangsu(Prospective and Key Technologies for Industry)under Grants BE2023022 and BE2023022-2.
文摘The Internet of Unmanned Aerial Vehicles(I-UAVs)is expected to execute latency-sensitive tasks,but limited by co-channel interference and malicious jamming.In the face of unknown prior environmental knowledge,defending against jamming and interference through spectrum allocation becomes challenging,especially when each UAV pair makes decisions independently.In this paper,we propose a cooperative multi-agent reinforcement learning(MARL)-based anti-jamming framework for I-UAVs,enabling UAV pairs to learn their own policies cooperatively.Specifically,we first model the problem as a modelfree multi-agent Markov decision process(MAMDP)to maximize the long-term expected system throughput.Then,for improving the exploration of the optimal policy,we resort to optimizing a MARL objective function with a mutual-information(MI)regularizer between states and actions,which can dynamically assign the probability for actions frequently used by the optimal policy.Next,through sharing their current channel selections and local learning experience(their soft Q-values),the UAV pairs can learn their own policies cooperatively relying on only preceding observed information and predicting others’actions.Our simulation results show that for both sweep jamming and Markov jamming patterns,the proposed scheme outperforms the benchmarkers in terms of throughput,convergence and stability for different numbers of jammers,channels and UAV pairs.
基金supported in part by the National Natural Science Foundation of China(Grant Nos.62473135 and 62173121)。
文摘This paper investigates the observer-based prescribed-time time-varying output formation-containment(PT-TV-OFC)control problem for heterogeneous multi-agent systems in which the different agents have different state dimensions.The system comprises one tracking leader,multiple formation leaders,and followers,where two types of leaders are used to generate a reference trajectory for movement and achieve specific formation,respectively.Firstly,a prescribed-time dynamics observer is constructed for the formation leaders to estimate the tracking leader's dynamic model and state.On this basis,a prescribed-time control protocol is designed for the formation leaders to achieve time-varying output formation.Then,a prescribed-time convex hull observer is designed for the followers to estimate information regarding the convex hull formed by the formation leaders.Using the estimated convex hull information,a prescribed-time containment control protocol is designed to ensure the followers converge into the convex hull.Furthermore,using Lyapunov stability theory,the stability of systems is proved in detail,which implies that the heterogeneous multi-agent systems can achieve PT-TV-OFC control.Finally,numerical simulations validate the feasibility of the theoretical results.
文摘This paper delves into the problem of optimal placement conditions for a group of agents collaboratively localizing a target using range-only or bearing-only measurements.The challenge in this study stems from the uncertainty associated with the positions of the agents,which may experience drift or disturbances during the target localization process.Initially,we derive the Cramer-Rao lower bound(CRLB)of the target position as the primary analytical metric.Subsequently,we establish the necessary and sufficient conditions for the optimal placement of agents.Based on these conditions,we analyze the maximal allowable agent position error for an expected mean squared error(MSE),providing valuable guidance for the selection of agent positioning sensors.The analytical findings are further validated through simulation experiments.
文摘Policy training against diverse opponents remains a challenge when using Multi-Agent Reinforcement Learning(MARL)in multiple Unmanned Combat Aerial Vehicle(UCAV)air combat scenarios.In view of this,this paper proposes a novel Dominant and Non-dominant strategy sample selection(DoNot)mechanism and a Local Observation Enhanced Multi-Agent Proximal Policy Optimization(LOE-MAPPO)algorithm to train the multi-UCAV air combat policy and improve its generalization.Specifically,the LOE-MAPPO algorithm adopts a mixed state that concatenates the global state and individual agent's local observation to enable efficient value function learning in multi-UCAV air combat.The DoNot mechanism classifies opponents into dominant or non-dominant strategy opponents,and samples from easier to more challenging opponents to form an adaptive training curriculum.Empirical results demonstrate that the proposed LOE-MAPPO algorithm outperforms baseline MARL algorithms in multi-UCAV air combat scenarios,and the DoNot mechanism leads to stronger policy generalization when facing diverse opponents.The results pave the way for the fast generation of cooperative strategies for air combat agents with MARLalgorithms.
文摘This paper addresses the time-varying formation-containment(FC) problem for nonholonomic multi-agent systems with a desired trajectory constraint, where only the leaders can acquire information about the desired trajectory. Input the fixed time-varying formation template to the leader and start executing, this process also needs to track the desired trajectory, and the follower needs to converge to the convex hull that the leader crosses. Firstly, the dynamic models of nonholonomic systems are linearized to second-order dynamics. Then, based on the desired trajectory and formation template, the FC control protocols are proposed. Sufficient conditions to achieve FC are introduced and an algorithm is proposed to resolve the control parameters by solving an algebraic Riccati equation. The system is demonstrated to achieve FC, with the average position and velocity of the leaders converging asymptotically to the desired trajectory. Finally, the theoretical achievements are verified in simulations by a multi-agent system composed of virtual human individuals.
基金founded by the National Science and Technology Council of the Republic of China under contract NSTC113-2221-E-019-032.
文摘An Interval Type-2(IT-2)fuzzy controller design approach is proposed in this research to simultaneously achievemultiple control objectives inNonlinearMulti-Agent Systems(NMASs),including formation,containment,and collision avoidance.However,inherent nonlinearities and uncertainties present in practical control systems contribute to the challenge of achieving precise control performance.Based on the IT-2 Takagi-Sugeno Fuzzy Model(T-SFM),the fuzzy control approach can offer a more effective solution for NMASs facing uncertainties.Unlike existing control methods for NMASs,the Formation and Containment(F-and-C)control problem with collision avoidance capability under uncertainties based on the IT-2 T-SFM is discussed for the first time.Moreover,an IT-2 fuzzy tracking control approach is proposed to solve the formation task for leaders in NMASs without requiring communication.This control scheme makes the design process of the IT-2 fuzzy Formation Controller(FC)more straightforward and effective.According to the communication interaction protocol,the IT-2 Containment Controller(CC)design approach is proposed for followers to ensure convergence into the region defined by the leaders.Leveraging the IT-2 T-SFM representation,the analysis methods developed for linear Multi-Agent Systems(MASs)are successfully extended to perform containment analysis without requiring the additional assumptions imposed in existing research.Notably,the IT-2 fuzzy tracking controller can also be applied in collision avoidance situations to track the desired trajectories calculated by the avoidance algorithm under the Artificial Potential Field(APF).Benefiting from the combination of vortex and source APFs,the leaders can properly adjust the system dynamics to prevent potential collision risk.Integrating the fuzzy theory and APFs avoidance algorithm,an IT-2 fuzzy controller design approach is proposed to achieve the F-and-C purposewhile ensuring collision avoidance capability.Finally,amulti-ship simulation is conducted to validate the feasibility and effectiveness of the designed IT-2 fuzzy controller.
基金supported in part by the National Natural Science Foundation of China(Nos.62371369 and 62376204)the National Key R&D Program of China(No.2022YFC3301300).
文摘Unmanned Aerial Vehicles(UAVs)have demonstrated significant potential as Aerial Base Stations(A-BSs)for providing data services to Ground Users(GUs),attributed to their flexibility,cost-effectiveness,and high likelihood of establishing line-of-sight links.In this article,we formulate the joint power and trajectory optimization problem for a multi-UAV assisted wireless network with no-fly zones constrained,aiming at maximizing the Accumulated Service Data(ASD)of UAVs and minimizing the Average End Age of Information(AEAoI)of GUs.Specifically,this paper proposes the Multi-Agent worst-case Soft Actor Critic(MA-wcSAC)algorithm with a distributional safety-critic.The simulation results demonstrate that,compared to the Multi-Agent Soft Actor Critic(MA-SAC)algorithm,the proposed algorithm exhibits comparable data service performance while reducing security risks by at least 30%at different risk levels.
基金Aeronautical Science Foundation of China (2006ZA51004)
文摘With the aid of multi-agent based modeling approach to complex systems, the hierarchy simulation models of carrier-based aircraft catapult launch are developed. Ocean, carrier, aircraft, and atmosphere are treated as aggregation agents, the detailed components like catapult, landing gears, and disturbances are considered as meta-agents, which belong to their aggregation agent. Thus, the model with two layers is formed i.e. the aggregation agent layer and the meta-agent layer. The information communication among all agents is described. The meta-agents within one aggregation agent communicate with each other directly by information sharing, but the meta-agents, which belong to different aggregation agents exchange their information through the aggregation layer first, and then perceive it from the sharing environment, that is the aggregation agent. Thus, not only the hierarchy model is built, but also the environment perceived by each agent is specified. Meanwhile, the problem of balancing the independency of agent and the resource consumption brought by real-time communication within multi-agent system (MAS) is resolved. Each agent involved in carrier-based aircraft catapult launch is depicted, with considering the interaction within disturbed atmospheric environment and multiple motion bodies including carrier, aircraft, and landing gears. The models of reactive agents among them are derived based on tensors, and the perceived messages and inner frameworks of each agent are characterized. Finally, some results of a simulation instance are given. The simulation and modeling of dynamic system based on multi-agent system is of benefit to express physical concepts and logical hierarchy clearly and precisely. The system model can easily draw in kinds of other agents to achieve a precise simulation of more complex system. This modeling technique makes the complex integral dynamic equations of multibodies decompose into parallel operations of single agent, and it is convenient to expand, maintain, and reuse the program codes.
基金The National Key Technology R&D Program of China during the 11th Five-Year Plan Period(No.2006BAH02A06)
文摘The similarities and differences between the container terminal logistics system(CTLS)and the Harvard-architecture computer system are compared in terms of organization and architecture.The mapping relation and the modeling framework of the CTLS are presented based on multi-agent,and the successful algorithms in the computer domain are applied to the modeling framework,such as the dynamic priority and multilevel feedback scheduling algorithm.In addition,a model and simulation on a certain quay at Shanghai harbor is built up on the AnyLogic platform to support the decision-making of terminal on service cost.It validates the feasibility and creditability of the above systematic methodology.