Revise the JFM research workflow documentation to clarify project structure, decision-making processes, and tool routing. Remove outdated figures and images that are no longer relevant to the current manuscript. Ensure that the workflow aligns with the latest research objectives and safety protocols.
394 lines
68 KiB
TeX
394 lines
68 KiB
TeX
\documentclass[lineno]{JFM-FLM_Au}
|
|
|
|
\usepackage{amsmath,amssymb,bm,booktabs}
|
|
\newcommand{\dd}{\mathrm{d}}
|
|
\newcommand{\vect}[1]{\boldsymbol{#1}}
|
|
\newcommand{\draftfigure}[3]{%
|
|
\begin{figure}
|
|
\centering
|
|
\fbox{\parbox[c][0.22\textheight][c]{0.88\textwidth}{}}
|
|
\caption{#2}
|
|
\label{#3}
|
|
\end{figure}
|
|
}
|
|
|
|
\lefttitle{Y. Wang, F. Ren, X. Wang, H. Bai, H. Tang and S. Chen}
|
|
\righttitle{Symbolic control laws for hydrodynamic cloaking and illusion}
|
|
|
|
\title{Interpretable control laws for active hydrodynamic cloaking and illusion of the fluidic pinball}
|
|
|
|
\author{Yanqi Wang\aff{1,3}, Feng Ren\aff{2}, Xu Wang\aff{1}, Honglei Bai\aff{4}, Hui Tang\aff{1}, \and Shiyi Chen\aff{3} }
|
|
|
|
\affiliation{\aff{1}Department of Mechanical Engineering, The Hong Kong Polytechnic University, Hong Kong SAR, PR China
|
|
\aff{2}School of Marine Science and Technology, Northwestern Polytechnical University, Xi'an 710072, PR China
|
|
\aff{3}Eastern Institute for Advanced Study, Eastern Institute of Technology, Ningbo 315200, PR China
|
|
\aff{4}School of Aeronautics and Astronautics, Sun Yat-Sen University, Shenzhen Campus, Shenzhen 518107, PR China}
|
|
|
|
\corresau{Hui Tang, \email{h.tang@polyu.edu.hk}}
|
|
|
|
\begin{document}
|
|
\maketitle
|
|
|
|
\begin{abstract}
|
|
Hydrodynamic cloaking and illusion require active control of the wake transported beyond a body, rather than optimization of an integrated force alone. We investigate these declared-reference objectives using a fluidic pinball of three independently rotating cylinders. A proximal-policy-optimization controller maps the forces on the cylinders and velocities from three downstream probes to rotational commands, while registered velocity fields provide an independent measure of the resulting wake. For periodic K\'arm\'an cloaking, the controlled mean and phase-resolved fields approach the target substantially more closely than the uncontrolled zero-rotation wake. Post-hoc symbolic regression reveals an executable control law organized around persistent counter-rotation of the rear cylinders, supplemented by a smaller time-dependent correction. Closed-loop deletion tests identify the persistent rear rotation as the dominant tested component. A complementary four-flow analysis shows that fixing the cylinders at the mean learned action accounts for 93.42\% of the measured reduction in scalar mean-field error. The remaining time-dependent field organization is examined separately through action-correlated correction-field modes, linking low-rank wake structure to front, rear-symmetric and rear-antisymmetric actuation coordinates. The framework also accommodates an event-relative Vortex case and a non-zero-target Illusion case, demonstrating how the control problem changes with its temporal coordinate and prescribed wake signature. A separate steady-flow analysis connects persistent rear rotation to a low-deficit endpoint and tests its relation to a prescribed-circulation outer flow. Together, these results show how learned rotational control can be evaluated against declared wake references and distilled into a compact, physically readable K\'arm\'an control structure.
|
|
\end{abstract}
|
|
|
|
\begin{keywords}
|
|
hydrodynamic cloaking, flow control, symbolic regression, deep reinforcement learning, fluidic pinball
|
|
\end{keywords}
|
|
|
|
\section{Introduction}
|
|
\label{sec:introduction}
|
|
|
|
A hydrodynamic cloak seeks to make the flow around a body approach a declared background, while a hydrodynamic illusion seeks a declared non-zero substitute signature. Transformation optics established complementary routes to controlling linear fields through material design \citep{PendryEtAl2006Controlling,Leonhardt2006OpticalMapping}. Active source-based cloaking has also been formulated for specified linear equations \citep{Miller2006PerfectActiveCloaking,VasquezEtAl2009ActiveExterior}. In a distinct role, active constructions have retargeted exterior fields towards non-zero illusion signatures \citep{MaEtAl2013ActiveCloaking,LinEtAl2021ActiveAcousticIllusions}. These branches distinguish redirecting a field, generating compensation and prescribing a substitute signature. They assume linear governing relations and prescribed material or source responses, whereas a viscous wake evolves with the actuation and body-generated vorticity; their role here is conceptual rather than predictive. A multilayer hydrodynamic cloak in free flow has been reported \citep{ChenEtAl2024MultilayerHydrodynamicCloak}. Separated finite-Reynolds-number flow nevertheless poses a distinct problem even relative to active hydrodynamic-cloak formulations \citep{UrzhumovSmith2012ActiveCloaks}: nonlinear advection couples forcing to the evolving wake, no-slip bodies generate shear layers and separation, and viscosity and confinement alter the response. The attainable match depends on the declared target and comparison, especially when coherent shedding carries information beyond local measurements. We therefore consider controlled approach to a specified background or substitute wake, assessed beyond the control observations, rather than universal invisibility or exact cancellation in this setting.
|
|
|
|
The desired outcome is an extended Navier--Stokes field, whereas actuation is confined to compact boundaries. The fluidic pinball---three independently rotating cylinders in channel flow---is a useful model because a small multi-input plant generates rich nonlinear wakes and admits established open- and closed-loop control studies \citep{DengEtAl2020PinballBifurcations,CornejoMacedaEtAl2021PinballControl,RaibaudoEtAl2020PinballControl}. Transient and post-transient pinball-force dynamics have been treated with reduced modelling \citep{DengEtAl2021PinballForceModel}, supporting integrated load as a useful intermediary between local rotation and the wake. Localized actuation can reorganize bluff-body shedding, and deep reinforcement learning has been applied to pinball force control \citep{FengWang2010SyntheticJetCylinder,FengEtAl2023PinballForceControl}. Conventional active flow control commonly asks how actuation changes shedding or integrated load. The present cloak/illusion objective instead adds a separately acquired reference and an extended-field test: neither scalar drag reduction nor force control establishes that the downstream velocity field approaches that reference. The resulting problem is whether three rotation commands can use force-mediated feedback to approach a declared hydrodynamic signature while leaving the downstream field independently testable.
|
|
|
|
The difficulty increases when the reference contains periodic shedding, an incoming event or a non-zero target wake. Deep reinforcement learning has enabled closed-loop decisions in nonlinear flows under limited observations \citep{RabaultEtAl2019DRLFlowControl,ParisEtAl2021RobustFlowControl}, including rotating-cylinder actuation in towing experiments and simulations \citep{FanEtAl2020BluffBodyRL}. Related bluff-body work has framed learned active control as hydrodynamic stealth \citep{RenEtAl2021HydrodynamicStealthDRL}, although its observer metric is not the target-relative field comparison used here. Unlike a scalar optimization objective, reference following must distinguish what the controller observes from what is later evaluated. The present policy receives the two force components of each cylinder together with sparse downstream velocities. Integrated forces report body--flow interaction, while remote measurements sample the emerging wake; neither implies full observability. Sparse-probe utility is flow- and task-dependent, including sensitivity to Reynolds number and blockage \citep{LiZhang2022ConfinedWakeRL}. This interface motivates three tests: whether the policy realizes its objective, whether the velocity field approaches the target beyond the reward and sensor locations, and whether recurrent action and field structures can be reduced to forms that remain checkable in closed-loop execution.
|
|
|
|
These tests define a layered analysis. Learned execution is first assessed within each case's own configuration, target roles and temporal coordinate. The periodic K\'arm\'an cloak is then evaluated independently by registering target, controlled and physical-zero velocity fields, while temporal signal comparison remains complementary rather than a substitute for field agreement. On the action side, symbolic regression compresses the policy into a candidate control law tested by closed-loop deployment, following the principle that readable learned feedback should be judged in execution \citep{GautierEtAl2015ClosedLoop,CornejoMacedaEtAl2021PinballControl}. Section~\ref{sec:karman-law} distinguishes this post-hoc policy-map regression from direct genetic-programming control and sparse dynamical-system identification \citep{LiEtAl2017GeneticProgrammingControl,BruntonEtAl2016SINDy}; PySR is the symbolic-regression implementation \citep{Cranmer2023PySR}, not evidence that the resulting expression controls the flow. This distinction prevents a compact expression from being promoted from representation to explanation before deployment. On the field side, the persistent mean response is separated from a centred residual, and action-correlated correction-field modes describe organization without assigning causality. Sparse-to-field and reduced-order precedents motivate separating online information from fuller-state interpretation \citep{LoiseauEtAl2018SparseROM,LiEtAl2022FlowEstimation}. The action and field analyses answer different questions and do not validate one another.
|
|
|
|
We therefore ask whether force-plus-sparse-observation rotational control can approach declared cloak and illusion references under separately specified channel-flow formulations, and how far the periodic K\'arm\'an cloak can be evaluated and interpreted. Rotational feedback was executed for the periodic K\'arm\'an cloak, the non-zero Illusion target and the event-relative Vortex case; only K\'arm\'an supports a target-relative field-proximity claim, while Illusion and Vortex provide execution and morphology evidence. For Erase, the available target and comparator roles do not yet permit the same performance comparison. Registered target, controlled and physical-zero fields show improved target proximity for the periodic K\'arm\'an cloak in the parabolic-inflow/no-slip-wall channel configuration. Closed-loop action-law tests further identify a K\'arm\'an control skeleton: dominant persistent rear-cylinder rotation accompanied by a smaller dynamic correction, with the associated field organization described by action-correlated modes. This deeper account is possible for K\'arm\'an because policy execution can be followed through registered field comparison, executable law reduction and organized mean and residual fields. The separate steady analysis instead shows that the corresponding inviscid closure fails, and does not extend the learned periodic result. The next section defines the configurations, actuation, observations, target roles and independent evaluation criteria used to keep these comparisons distinct.
|
|
|
|
\section{Problem formulation, control framework and evaluation}
|
|
\label{sec:formulation}
|
|
|
|
\subsection{Flow configurations and numerical method}
|
|
\label{subsec:configurations}
|
|
The controlled body is a two-dimensional fluidic pinball comprising three circular cylinders of diameter $D$. A front cylinder is followed by upper and lower rear cylinders, and all three rotate independently. Streamwise and transverse coordinates are $(x,y)$, and the computational body order is $(F,U,L)=(\text{front},\text{rear}_{y+},\text{rear}_{y-})$. Lateral confinement bounds cross-stream motion and provides a repeatable environment in which rotation can alter the integrated forces and wake. Two channel configurations are used for different parts of the study. The parabolic-inflow/no-slip-wall channel supplies the periodic K\'arm\'an field comparison and the detailed action- and field-side analyses. The uniform-inflow/free-slip-wall channel supplies the broader controlled cases and the separate steady analysis. Their results are interpreted within their own physical and numerical settings rather than combined as realizations of one plant. A case-dependent upstream disturbance or target body defines the requested wake, with its dimensions specified for that case.
|
|
|
|
Velocities are normalized by a configuration-specific reference speed $U_{ref}$, lengths by $D$, and time by $D/U_{ref}$. The article-wide Reynolds number is
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
Re_D &= \frac{U_{ref}D}{\nu},\\
|
|
U_{max} &\simeq 1.5U_{ref}
|
|
\quad\text{for parabolic inflow}.
|
|
\end{aligned}
|
|
\end{equation}
|
|
where $\nu$ is the kinematic viscosity. For uniform inflow, $U_{ref}$ is the inlet speed; for parabolic inflow, it is the bulk/mean inlet speed. This distinction matters when comparing inherited case labels: a label based on $2D$ can be halved only if its velocity scale is also $U_{ref}$. Historical parabolic records do not always identify $U_0$ as centreline or bulk/mean speed, so no physical Reynolds number is inferred from such labels here. Both configurations are advanced with a D2Q9 lattice-Boltzmann method using multiple-relaxation-time collision and curved moving-body treatment. Inlet, lateral-wall and outlet implementations follow the stated configuration. The detailed K\'arm\'an analyses use data from the parabolic/no-slip realization, whereas the control formulation described next is implemented in the uniform/free-slip configuration; no numerical equivalence between the two implementations is assumed.
|
|
|
|
\subsection{Rotational actuation, observations and PPO training}
|
|
\label{subsec:ppo-framework}
|
|
For a radius vector $(r_x,r_y)$ measured from a cylinder centre, angular velocity $\omega$ imposes the wall velocity
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
(U_w,V_w) &= (-\omega r_y,\omega r_x),\\
|
|
s_i &= \frac{\omega_iR}{U_{ref}}.
|
|
\end{aligned}
|
|
\end{equation}
|
|
where $R=D/2$, positive $\omega$ is counter-clockwise, and $s_i$ is the nondimensional surface-speed command for cylinder $i\in\{F,U,L\}$. The historical symbolic-regression records use $\alpha_i=\omega_i/U_0$ instead. The two descriptions are related by $s_i=\alpha_iR(U_0/U_{ref})$ and coincide as $s_i=\alpha_iR$ only when $U_0=U_{ref}$. Thus $\alpha_i$ is not dimensionless unless a radius-one convention is also specified. In the canonical control formulation, the policy commands rotation without a preset steady-derived bias, and successive commands are smoothed before application. This choice defines the computational interface; it is not a comparison of bias strategies or a map to an experimental actuator. Prior pinball work has treated force dynamics with reduced models and used learned control for force objectives \citep{DengEtAl2021PinballForceModel,FengEtAl2023PinballForceControl}. This motivates integrated force as a compact actuation-coupled signal, without making it a downstream-field measure.
|
|
|
|
The observation vector contains the two force components of each cylinder and both velocity components from three circular wake probes. The probes are centred $10D$ downstream of the signature source at $(y-y_0)/D=0,\pm2$ and have radius $0.25D$. Each probe contributes area-averaged streamwise and transverse velocity. Forces and wake velocities describe complementary aspects of the controlled flow: periodic loads can carry phase information, while sparse wake measurements provide downstream coordinates \citep{NairEtAl2021PhaseFlowControl,MarraEtAl2024ActuationManifold}. The centre probe is intended to sample coherent-shedding phase, and the off-centre pair to sample lateral deflection and alternating cross-wake motion. Comparable downstream sensing has been used to characterize and control fluidic-pinball wakes \citep{RaibaudoEtAl2020PinballControl,CornejoMacedaEtAl2021PinballControl,LiEtAl2022FlowEstimation}; sparse wake information is also physically accessible in other wake-sensing settings \citep{BeemTriantafyllou2015WakeSensing}. These precedents motivate the measurement classes, not the exact present geometry.
|
|
|
|
Sensor placement is inseparable from its use and from the time scales carried by the chosen observable. In periodic wakes, a phase-bearing signal can encode dominant shedding dynamics in the regime for which its model is constructed \citep{GongEtAl2020VortexEstimation}. For feedback, however, moving a sensor downstream can increase developed-wake information while also increasing convective lag, so a location that is effective for estimation need not be best for closed-loop control \citep{JinEtAl2022SensorActuatorPlacement}. Sparse instantaneous measurements can be augmented with temporal histories, but such lifting increases input dimension and redundancy \citep{WangEtAl2024DynamicFeatureDRL}. These results motivate retaining spatially distributed and temporally resolved information, but do not choose the present geometry. The three-probe arrangement is consequently a compact wake monitor rather than a claim of optimal placement or full observability. Solver averaging, fixed physical pre-scaling and online observation whitening are applied as distinct operations. Because whitening parameters form part of the policy input map, evaluation uses the normalizer saved with the selected policy.
|
|
|
|
PPO updates a stochastic policy by maximizing a clipped surrogate objective relative to the preceding policy \citep{SchulmanEtAl2017PPO}. Clipping limits the incentive for a large policy change, but does not guarantee convergence or successful control. The canonical network has two hidden layers of 64 units with sinusoidal activation. Training uses 2048-step rollouts, batches of 64, ten update epochs, discount factor $0.995$ and learning rate $3\times10^{-4}$. These settings specify the reusable canonical training procedure. Reproducing an individual trained policy additionally requires its selected checkpoint, matching normalizer, target signals and run-specific objective settings; the canonical defaults alone do not determine that policy.
|
|
|
|
\subsection{Target construction and online objective}
|
|
\label{subsec:targets-objective}
|
|
Every comparison distinguishes how its signals and fields are acquired. The target or reference is the wake requested by the case; controlled denotes the trajectory generated with the policy; physical zero or uncontrolled denotes an independently acquired zero-rotation flow; and constant denotes a separately acquired fixed-actuation comparator where one exists. A cloak target is a declared background flow, whereas an illusion target is a separately generated non-zero wake, including a target-body acquisition when required. The target trajectory and the physical-zero trajectory therefore play different roles even when both contain no policy actuation. Target-body size and upstream geometry are specified case by case so that the requested signature is not silently transferred between configurations.
|
|
|
|
During control, the policy minimizes discrepancies in the same compact information space that it observes. We write the online objective as
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
\mathcal{J}_{online}
|
|
={}& w_F\,\mathcal{D}_F(\mathbf F,\mathbf F_t)\\
|
|
&+w_S\,\mathcal{D}_S(\mathbf q,\mathbf q_t).
|
|
\end{aligned}
|
|
\end{equation}
|
|
where $\mathbf F$ collects the selected body-force components, $\mathbf q$ contains the six probe velocities, and subscript $t$ denotes the corresponding target signal. The force and sensor terms allow rotation to respond both to the body's integrated interaction with the flow and to the developing wake. Their channel aggregation, normalization, orientation and weights are selected for each case. Because this objective samples forces and three wake regions rather than the complete velocity field, reward improvement establishes signal-space execution only. Field matching is tested separately.
|
|
|
|
\subsection{Independent signal- and field-level evaluation}
|
|
\label{subsec:evaluation}
|
|
Dynamic time warping (DTW) uses dynamic programming to find a minimum-cost monotone path through the pairwise discrepancies of two sampled sequences \citep{SakoeChiba1978DynamicProgramming,Muller2015FundamentalsMusicProcessing}. Declared endpoint, window and step constraints permit limited local time reparameterization without reducing the comparison to a single global phase shift. For each result, we state whether DTW is a distance or higher-is-better similarity, whether channels are native or normalized, and the time window, scaling, aggregation and path constraints. These conventions are project-specific rather than supplied by the generic DTW references. Because excessive warping can hide timing errors, DTW is not interpreted as a physical delay, causal relation, observability test or field-equivalence measure.
|
|
|
|
Field agreement is evaluated over a registered downstream rectangle,
|
|
\begin{equation}
|
|
\Omega=\left\{(x,y):
|
|
\begin{aligned}
|
|
-6 &\leq \frac{x-x_s}{D} \leq 14,\\
|
|
\left|\frac{y-y_0}{D}\right| &\leq 5
|
|
\end{aligned}
|
|
\right\}.
|
|
\end{equation}
|
|
where $x_s$ is the sensor-plane streamwise location and $y_0$ is the channel centreline. For $r\in\{\text{controlled},\text{physical zero}\}$, the complete-cycle mean-field error is
|
|
\begin{equation}
|
|
\begin{split}
|
|
E_{mean,r}=\Biggl[\frac{1}{|\Omega|}\int_{\Omega}
|
|
&\frac{\|\overline{\mathbf u}_r
|
|
-\overline{\mathbf u}_t\|_2^2}{U_{ref}^2}\\
|
|
&\,\mathrm d\Omega\Biggr]^{1/2}.
|
|
\end{split}
|
|
\end{equation}
|
|
The fixed $20D\times10D$ region is registered to the sensor plane and centreline so that each role is compared over the same downstream extent. Saved fields do not preserve solver-exact masks, however, so the comparison uses a geometry guard rather than claiming an exact common-fluid mask. A comparator-relative reduction is formed only when target, controlled and physical-zero fields share the same acquisition definition, region, weighting, normalization and averaging window.
|
|
|
|
The same spatial error is evaluated at eight canonical phase slots for the periodic K\'arm\'an case. Each role is phased independently from its own centre-probe transverse velocity, and each slot uses the nearest saved snapshot. The resulting phase diagnostic samples one canonical cycle; it is not a repeated-cycle ensemble or an unrestricted phase optimization. Other cases require different temporal coordinates: Vortex uses offsets from a declared event, Erase admits a mean-field comparison only when its roles are resolved, and the steady analysis uses a late field or profile. These quantities are not pooled because they answer different physical questions. DTW measures temporal signatures close to the controller's information space, whereas mean and phase errors interrogate the extended velocity field. Either can improve without the other, making their separation central to the evaluation.
|
|
|
|
\subsection{Compact CFD qualification}
|
|
\label{subsec:cfd-qualification}
|
|
The numerical method is assessed through selected observables in two CelerisLab benchmark configurations. In the rotating-cylinder Kan99b K2 case, Strouhal number, mean drag and force-fluctuation amplitudes lie within their prescribed bands. The mean-lift sign differs because the CFD and comparison conventions use opposite force directions, so the overall K2 comparison remains a \emph{partial pass} despite the conventional origin of that sign difference. In the confined-cylinder Sah04 S2 case, the Strouhal number meets its prescribed band. Only these two completed simulations enter the qualification. They support the named frequency and force observables under their benchmark conditions, but do not assess PPO training, target construction, symbolic regression, correction-field interpretation or hydrodynamic cloaking.
|
|
|
|
With the configurations, controller interface, target roles and independent measures defined, the following sections compare the controlled cases without combining results across configurations or treating signal and field agreement as interchangeable.
|
|
|
|
\draftfigure{Two configuration schematics showing inlet and wall conditions, fluidic-pinball body order and rotation sign, the case-dependent disturbance or target body, three downstream two-component probes, and the registered evaluation region.}{Planned control configurations and measurements. The parabolic-inflow/no-slip-wall and uniform-inflow/free-slip-wall channels will be shown separately, together with body order $(F,U,L)$, counter-clockwise-positive rotation, and the three circular probes at $x/D=10$, $(y-y_0)/D=0,\pm2$ with radius $0.25D$. The schematic defines geometry, boundaries and observations only; it does not imply configuration equivalence, optimal sensor placement or full observability.}{fig:configurations-sensors}
|
|
|
|
\section{From steady cloaking to wake retargeting}
|
|
\label{sec:progression}
|
|
|
|
Rotational control is considered through four comparisons of increasing temporal and target complexity: an approximately steady uniform profile, a periodic K\'arm\'an street, an isolated-vortex event and a non-zero wake target. Because the configuration, clock and measure change between cases, their ordering describes a progression of control tasks rather than a common performance ranking. The steady case asks whether a simple downstream profile can be recovered. K\'arm\'an cloaking then makes the desired wake periodic, the Vortex case replaces phase by an event-relative coordinate, and Illusion changes the desired wake itself.
|
|
|
|
The steady case uses the uniform-inflow/free-slip-wall channel configuration. At $x/D=10$, the controlled endpoint follows the uniform target more closely than the stationary pinball in both velocity components: the streamwise profile approaches the uniform level and the transverse departure is reduced. This comparison concerns one downstream profile at one endpoint, not a periodic mean or full-field identity. It establishes neither optimality nor stability or mechanism. The actuation magnitude, settling behaviour and physical interpretation are treated with the steady-flow evidence later; here the result provides the simplest instance in which rotation moves measured velocity components towards the target.
|
|
|
|
Periodic K\'arm\'an cloaking provides the most complete comparison. In the parabolic-inflow/no-slip-wall channel configuration, the target, frozen-policy controlled and physical-zero roles define the alternating-street problem. For both the complete-cycle mean and the separate eight-phase comparison, the controlled downstream velocity field is closer to the target than physical zero, with a dimensionless error reduction of approximately 82\% in each case. The velocity errors are normalized by the simulation speed $U_0$ used in these data. Its relation to article-wide $U_{ref}$ remains unresolved, but no conversion is needed for either relative reduction. Reward and visual resemblance do not establish the field result; Section~\ref{sec:karman-results} separates the mean, phase, signal and action comparisons in detail.
|
|
|
|
The Vortex case removes periodic phase and instead compares the target, controlled and physical-zero flow morphology at one event-relative coordinate. At that instant the three fields describe how the wake is organized relative to the passing vortex. Without a field-error scalar, one snapshot cannot establish event evolution or improvement over physical zero. The comparison is therefore morphological: it extends the progression from a repeating wake to an isolated interaction while leaving quantitative event response to the cross-case analysis.
|
|
|
|
Illusion changes the desired field rather than only its temporal coordinate. In the uniform-inflow/free-slip-wall configuration, one acquisition contains the $0.75L_0$ non-zero target, the seed-43 PPO-controlled trajectory and the physical-zero trajectory under the same sampling and phase procedure. The policy was executed toward that non-zero target, and the controlled morphology at the common literal phase is consistent with wake retargeting. Without a scalar field comparison, this morphology does not establish approach to the target or improvement over physical zero. It also neither constructs a target action nor revives the negative Illusion symbolic-regression result.
|
|
|
|
The progression therefore changes one principal feature at a time: the target evolves from a uniform background to a periodic street, the temporal coordinate changes from phase to an isolated event, and the target finally becomes a different non-zero wake. K\'arm\'an cloaking receives detailed treatment because it combines repeated learning histories with target, controlled and physical-zero fields, sparse downstream signals and applied rotations. This broader set of comparisons does not make K\'arm\'an representative of the other tasks. It instead permits two bounded questions: over which tested K\'arm\'an conditions were high-performing policies found, and how closely did the controlled wake approach its target?
|
|
|
|
\draftfigure{A K\'arm\'an-dominant progression from the steady profile to periodic target/controlled/physical-zero fields, with one event-relative Vortex comparison and one non-zero-target Illusion comparison.}{Planned progression of controlled wake definitions. The K\'arm\'an comparison will dominate the composition; the Vortex panels will show target, controlled and physical-zero morphology at one stated event offset, and the Illusion panels will show execution toward the retained non-zero target. Scales and acquisition coordinates will be stated within each comparison. Only the K\'arm\'an panels support a target-relative field-error result; the changed-task panels are morphology comparisons without a pooled efficacy measure.}{fig:capability-progression}
|
|
|
|
\section{K\'arm\'an-wake cloaking}
|
|
\label{sec:karman-results}
|
|
|
|
\subsection{Learning across tested K\'arm\'an conditions}
|
|
\label{subsec:karman-learning}
|
|
|
|
The uniform-inflow/free-slip-wall series samples $Re_D=30$, 50, 100 and 200 and disturbance-radius ratios 0.75, 1.0, 1.5 and 2.0. These discrete points delimit the tested conditions; they do not define a continuous Reynolds-number or geometry law. Only the $Re_D=50$ reference condition has five stochastic training realizations, whereas each other condition has one deterministic policy demonstration. The two outer disturbance-radius cases also used a learning rate of $10^{-4}$ instead of $3\times10^{-4}$, so their differences cannot be attributed to geometry alone. Reward and six-sensor dynamic-time-warping (DTW) similarity are reported as separate, higher-is-better quantities: reward measures the learned objective, whereas DTW measures similarity of the selected downstream signals. Neither is treated as a fitted trend or combined score.
|
|
|
|
The condition series compares one deterministic policy at each tested condition, whereas the learning histories show how five stochastic searches developed at the reference condition. They are not paired before-and-after observations. At $Re_D=50$, all five 500-iteration histories progress from low reward to high-performing checkpoints. Their best rewards are 0.9266, 0.9302, 0.9160, 0.9222 and 0.9412, reached at iterations 471, 315, 360, 465 and 447, respectively. High-performance policy discovery therefore recurred in all five tested realizations rather than depending on a single selected run.
|
|
|
|
The recurrence applies to these five searches only: it provides no uncertainty for other seeds or hyperparameters and establishes neither asymptotic convergence nor closed-loop stability. The deterministic policy used for the condition comparison is one deployment drawn from this set, not a sixth realization or a continuation of training reward. Moreover, high reward establishes success under the online objective, which combines force and sparse downstream-signal terms; it does not by itself establish proximity of the downstream velocity field.
|
|
|
|
That distinction also separates the condition survey from the field analysis below. The survey uses the uniform-inflow/free-slip-wall configuration, whereas the detailed reference wake uses a parabolic-inflow/no-slip-wall configuration. Their numerical case labels therefore do not supply a common physical Reynolds number. The field comparison instead asks directly whether the controlled wake approaches its target in the velocity field and downstream signals.
|
|
|
|
\subsection{Fields and trajectories in the reference wake}
|
|
\label{subsec:karman-fields}
|
|
|
|
In the parabolic-inflow/no-slip-wall reference wake, the target, frozen-PPO controlled and physical-zero roles define the field comparison. The controlled wake reproduces the target's alternating organization more closely than the physical-zero pinball. These roles are separate trajectories in the same physical configuration, not paired realizations of one trajectory. Section~\ref{sec:karman-law} uses these same target and physical-zero fields to compare symbolic-regression control, but its controlled trajectory is different from the PPO-controlled trajectory here. Coincident target and physical-zero values therefore come from shared comparator fields, not independent agreement between the two controllers.
|
|
|
|
The complete-cycle mean tests the persistent velocity organization after periodic variation has been averaged out. With velocity normalized by the simulation speed $U_0$ used in these data, its error is $E_{mean}^{(U_0)}=0.085771$ for the controlled wake and $0.475206$ for physical zero, giving a zero-relative reduction of 0.819508. Thus the mean controlled field lies substantially closer to the target over the registered $20D\times10D$ downstream region. The result applies to this downstream region, which lies beyond the solid geometry; without a saved solver-exact fluid mask it is not a claim of pointwise identity throughout the domain.
|
|
|
|
The eight-slot diagnostic retains periodic organization instead of averaging it away. It gives $E_{phase8}^{(U_0)}=0.112680$ for controlled and $0.630878$ for physical zero, a zero-relative reduction of 0.821392. Each slot is the nearest saved snapshot after the target, controlled and physical-zero trajectories have been phased independently; the slots are not repeated-cycle ensembles and no global phase minimization is applied. The mean and phase comparisons give the same target-relative ordering but are not replicate measurements or uncertainty estimates. Both use $U_0$, so their relative reductions are dimensionless even though the relation between $U_0$ and article-wide $U_{ref}$ remains unresolved.
|
|
|
|
Sparse downstream observations provide a complementary trajectory-level comparison. Target-normalized six-channel DTW similarity is 0.932899 for controlled and 0.625127 for physical zero; unlike spatial error, a larger value denotes closer signal evolution. DTW does not imply full-field equivalence, but its ordering is consistent with the smaller controlled spatial errors. Over samples 96--145, corresponding to $tU_0/D=38.4$--58.0 in the simulation scaling and approximately three target cycles, the centre-sensor trajectory shows the repeating downstream response while the three applied PPO rotations show the simultaneous periodic actuation. The target has no action trajectory, so no target-action surrogate is inferred.
|
|
|
|
Together, the condition survey and learning histories show repeated policy discovery over the tested training set, while the reference-wake analysis shows target-relative proximity in mean fields, phase-resolved fields and sparse signals. These results imply neither mechanism nor a universal optimal law. Similar numbers in the later symbolic-regression or correction-field analyses do not provide validation because those calculations use different controlled trajectories and measures. The remaining control-side question is whether the periodic action produced by the neural policy can be represented by a compact executable law.
|
|
|
|
\draftfigure{Tested K\'arm\'an condition points, five reference-condition learning histories, and the parabolic/no-slip target, PPO-controlled and physical-zero field and signal comparisons.}{Planned K\'arm\'an learning and field evidence. Separate panels will report deterministic reward and six-sensor DTW similarity over the finite uniform/free-slip condition series, the five stochastic learning histories at $Re_D=50$, and target-relative mean, phase and sparse-signal comparisons for the distinct parabolic/no-slip reference wake. The configurations and estimands will remain separate; points do not define a continuous parameter law, and reward or DTW does not establish field agreement.}{fig:karman-learning-fields}
|
|
|
|
\section{A compact control law for K\'arm\'an cloaking}
|
|
\label{sec:karman-law}
|
|
|
|
The compressed action has a simple physical organization: the rear cylinders maintain persistent counter-rotation, rear-lift feedback modulates this background motion, and the front cylinder receives a weaker drag-asymmetry correction. The question is whether this organization can replace the PPO policy with an executable observation-to-action surrogate for the same periodic K\'arm\'an cloak in the parabolic-inflow/no-slip-wall channel configuration.
|
|
|
|
Direct symbolic controller search evaluates explicit feedback expressions in the control loop, as in genetic-programming studies of separation and vibration control \citep{GautierEtAl2015ClosedLoop,DebienEtAl2016GeneticProgrammingRamp,RenEtAl2019MachineLearningVIV}. Comparisons of compact genetic-programming laws and neural policies also separate readability from executed behaviour \citep{CastellanosEtAl2022FewSensorMLControl}. By contrast, sparse identification of nonlinear dynamics (SINDy) identifies parsimonious dynamical equations from data \citep{BruntonEtAl2016SINDy}, while SINDy-RL learns sparse dynamics and reward models jointly with a policy \citep{ZolmanEtAl2025SINDyRL}. Both address different mathematical objects from a fixed-policy map for the present controlled wake problem.
|
|
|
|
Here symbolic regression was applied after PPO training to recorded state--action trajectories, making the expression a fixed-policy map rather than a directly searched controller, identified wake equation or jointly learned reinforcement-learning model. PySR supplies the search method and software identity \citep{Cranmer2023PySR}, not evidence of control. Each post-step state at index $i$ was paired with the PPO action at $i+1$. Agreement on PPO-visited states is an offline diagnostic; after deployment, the surrogate chooses its own actions and visits a different trajectory. Closed-loop execution against the target is therefore the relevant test, and the cited methods do not validate the present map or its performance.
|
|
|
|
Before specifying the coefficients, we impose a reflection organization consistent with the centreline-symmetric geometry and cloaking objective. For a measured state $x$, $Gx$ denotes its centreline-reflected state: upper and lower sensor values are exchanged, streamwise velocity is unchanged, transverse velocity changes sign, front drag is unchanged, front lift changes sign, and the upper and lower force pairs are exchanged with drag even and lift odd. For actions ordered as front, upper rear and lower rear, the corresponding action map is
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
G_\alpha(\alpha_F,\alpha_U,\alpha_L)
|
|
&=(-\alpha_F,-\alpha_L,-\alpha_U),\\
|
|
\alpha_F(x)&=\frac{h_F(x)-h_F(Gx)}{2},\\
|
|
\alpha_U(x)&=h_R(x),\\
|
|
\alpha_L(x)&=-h_R(Gx).
|
|
\end{aligned}
|
|
\end{equation}
|
|
The front action is odd, while the rear actions share one mirrored generator. This construction replaces three separately parameterized heads with an odd front head and one rear generator. The reflection map was imposed on the surrogate after the PPO trajectories had been collected; it is not evidence that the PPO policy or controlled plant is reflection-equivariant.
|
|
|
|
The law uses signed native solver forces in the streamwise and transverse coordinate directions, so positive lift is the positive-$y$ force. For cylinder $i$, the fitted inputs are
|
|
\begin{equation*}
|
|
\begin{aligned}
|
|
C_{d,i}&=\frac{2f_{x,i}}{\rho U_0^2D}, &
|
|
C_{l,i}&=\frac{2f_{y,i}}{\rho U_0^2D},
|
|
\end{aligned}
|
|
\end{equation*}
|
|
with $\rho=1$ in the retained lattice data. Thus
|
|
\begin{equation*}
|
|
\begin{aligned}
|
|
C_{d,\mathrm{rear},a}&=\frac{C_{d,U}-C_{d,L}}{2}, &
|
|
C_{l,\mathrm{rear},s}&=\frac{C_{l,U}+C_{l,L}}{2}.
|
|
\end{aligned}
|
|
\end{equation*}
|
|
Within this imposed action space, the executable map is
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
\alpha_F&=\operatorname{odd}
|
|
(-0.381391\,C_{d,\mathrm{rear},a}),\\
|
|
\alpha_U&=1.307782\,C_{l,\mathrm{rear},s}-3.431209,\\
|
|
\alpha_L&=-\alpha_U(Gx).
|
|
\end{aligned}
|
|
\end{equation}
|
|
where $\operatorname{odd}$ denotes $f(x)\mapsto[f(x)-f(Gx)]/2$. The constant $-3.431209$ sets the persistent upper-rear rotation and, through reflection, the opposite lower-rear rotation. Rear-mean lift adjusts this pair, while rear drag asymmetry supplies the front correction. These terms describe the surrogate action, not a governing law for the fluid.
|
|
|
|
The historical output is $\alpha=\omega/U_0$, and deployment applies $\omega=\alpha U_0$; hence $\alpha$ has inverse-length units in dimensional notation. The corresponding surface-speed command is $s=\omega R/U_{ref}=\alpha R(U_0/U_{ref})$. It cannot be reduced to $s=\alpha R$ because the historical parabolic source does not establish $U_0=U_{ref}$.
|
|
|
|
Closed-loop execution tests whether the compact action remains useful away from PPO-visited states. This deployment test is also the selection boundary: expression simplicity and offline action fit can nominate a readable map, but only the executed surrogate determines the trajectory on which its control performance is measured. No external controller-search or sparse-identification result supplies that evidence for this case.
|
|
|
|
In the reference run, 480 control intervals were used to establish the trajectory, followed by 160 recorded post-step boundaries spanning eight complete cycles. The mean native dynamic-time-warping (DTW) similarity was $0.942357$, where higher is better. This rolling comparison uses a 30-boundary window: one circular lag is applied before the six channel scores are averaged. The lag aligns signals for comparison and is not a physical delay.
|
|
|
|
The same run also supplies two spatial comparisons with the same-case target and physical-zero trajectories from the shared acquisition lineage. Both errors are two-component velocity RMS differences over the geometry-guarded downstream region, normalized by the historical inlet scale $U_0$: $E_{\mathrm{mean}}$ compares complete-cycle arithmetic mean fields, whereas $E_{\mathrm{phase8}}$ is the RMS over eight independently phase-referenced nearest-boundary snapshots. The values were $E_{\mathrm{mean}}=0.125404$ for SR and $0.475206$ for physical zero, and $E_{\mathrm{phase8}}=0.156040$ for SR and $0.630878$ for physical zero. The surrogate-controlled wake is therefore closer to the target under both measures. Each phase entry is the nearest saved snapshot on a phase coordinate determined separately for that flow, and the region has no stored solver-exact mask. These SR-run values answer a different question from the independently recorded PPO-controlled field errors in Section~\ref{sec:karman-results} and are not repeat measurements of them.
|
|
|
|
Term deletion asks which parts of the executed map matter most among the tested variants. In each of four source cases, the full map and a deletion variant were run for the same 40 control steps. Six-channel native DTW similarity was evaluated after one shared circular-lag alignment, and $\Delta S=S_{deleted}-S_{parent}$. Deleting the rear constant gave $-0.106537$, $-0.093869$, $-0.178525$ and $-0.191082$; deleting rear-lift feedback gave $-0.028755$, $-0.041660$, $-0.061363$ and $-0.048948$; deleting the front correction gave $-0.006520$, $+0.000579$, $-0.005179$ and $-0.025194$.
|
|
|
|
Removing the persistent rear rotation produces the largest degradation in every tested case. Removing rear-lift feedback has a smaller but non-zero effect, while removing the front correction is weak or mixed over this window. A deletion changes the actions and hence the states subsequently visited, so this ranking does not identify causal or necessary terms in the flow-control process. It establishes a tested action hierarchy and motivates a separate field question: how do the persistent reference action and the remaining time-dependent organization appear in the mean and centred velocity fields?
|
|
|
|
\draftfigure{Closed-loop symbolic-controller action traces and target/SR/physical-zero wake fields, accompanied by a fixed-window parent-relative term-deletion plot.}{Planned compact K\'arm\'an law and deployment evidence in the parabolic-inflow/no-slip-wall configuration. The action traces will show persistent opposite rear rotation and smaller modulation; the field panels will compare the exact target, SR-controlled and physical-zero roles; and deletion will report $\Delta S=S_{deleted}-S_{parent}$ over the stated 40-step window. The deletion ranks tested action terms but does not establish causality, necessity, stability or transfer.}{fig:karman-sr-law}
|
|
|
|
\section{Mean-field correction and action-correlated organization}
|
|
\label{sec:field-ccd}
|
|
|
|
The field analysis begins with four flows: the target $T$, physical zero $0$, constant control $C$ fixed at the mean action of the PPO policy, and PPO-controlled flow $D$. For the periodic K\'arm\'an cloak in the parabolic-inflow/no-slip-wall configuration, each mean averages the same 360 post-step boundaries, indices $[480,840)$. The comparison covers $34\leq x/D\leq54$, $|y/D|\leq5$, on the 80,200 fluid points shared by all four simulations. With coordinate weights $w_k$ and both velocity components, the target error of flow $r$ is
|
|
\begin{equation}
|
|
E_r=\left[
|
|
\frac{\displaystyle\sum_{k\in K}w_k
|
|
\|\overline{\boldsymbol{u}}_r(k)
|
|
-\overline{\boldsymbol{u}}_T(k)\|_2^2}
|
|
{\displaystyle\sum_{k\in K}w_k}
|
|
\right]^{1/2}.
|
|
\end{equation}
|
|
Here $\boldsymbol{u}_r(k)=(u_{x,r}(k),u_{y,r}(k))$ is the two-component velocity at point $k$, and $\overline{\boldsymbol{u}}_T$ is the target mean, not an instantaneous field. This four-flow comparison uses different data, averaging windows, masks and weighting from the PPO-controlled comparison in Section~\ref{sec:karman-results} and the SR comparison in Section~\ref{sec:karman-law}; the resulting numbers are not repeat measurements.
|
|
|
|
The streamwise means give the first physical picture. Across much of the wake region, $\overline u_{x,0}-\overline u_{x,T}$ is negative: the physical-zero pinball has a mean streamwise deficit relative to the target. The difference $\overline u_{x,D}-\overline u_{x,0}$ is largely opposite in sign and shows where PPO control adds or removes streamwise velocity relative to physical zero. This pattern is compatible with adding downstream velocity where the physical-zero flow produces a deficit, but it does not establish momentum restoration. These fields show only the streamwise component; the errors below use both velocity components.
|
|
|
|
The exact errors are $E_0=0.4694352273$, $E_C=0.1110140739$ and $E_D=0.0857787860$. Their ordered differences give
|
|
\begin{equation}
|
|
\frac{E_0-E_C}{E_0-E_D}=93.42\%,\qquad
|
|
\frac{E_C-E_D}{E_0-E_D}=6.58\%.
|
|
\end{equation}
|
|
Thus constant control accounts for $93.42\%$ of the measured reduction from the physical-zero error to the PPO-controlled error, and the change from constant control to PPO control accounts for the remaining $6.58\%$. These percentages divide changes in one scalar distance from the target mean. They are neither energy fractions nor measurements of the norm of the residual defined next, and they do not assign shares to terms in the SR law.
|
|
|
|
The time-dependent comparison requires a velocity field rather than a difference between scalar errors. Mean-field descriptions of natural and actuated cylinder wakes use shift modes to represent changing means and their coupling with fluctuations \citep{TadmorEtAl2010CylinderMeanField}. That context motivates keeping the mean correction and fluctuation organization as distinct objects here; it does not identify the present constant-control subtraction with a shift-mode model. The PPO-controlled run contains 19 cycles resolved into ten phase bins. A separate 19-cycle constant-control run supplies the two-component phase field $\widehat{\boldsymbol{u}}^C_b(k)$. After subtracting each flow's own mean, the centred velocity residual is
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
\boldsymbol{u}^{\mathrm{res}}_{c,b}(k)
|
|
={}&[\boldsymbol{u}^D_{c,b}(k)-\overline{\boldsymbol{u}}_D(k)]\\
|
|
&-[\widehat{\boldsymbol{u}}^C_b(k)-\overline{\boldsymbol{u}}_C(k)],\\
|
|
&c=0,\ldots,18,\qquad b=0,\ldots,9.
|
|
\end{aligned}
|
|
\end{equation}
|
|
Every PPO-controlled cycle is compared with the constant-control ensemble template at phase bin $b$; individual cycles from the two runs are not paired. No cross-run lag, phase wrapping or interpolation is introduced. The residual is therefore a separately centred comparison, not a matched counterfactual response, and its magnitude is not the 6.58\% scalar share.
|
|
|
|
Canonical correlation decomposition (CCD) asks whether this residual varies with the executed actions. The actions are the same-boundary, counter-clockwise-positive cylinder speeds in native solver units. From front, upper and lower actions $(a_F,a_U,a_L)$, the analysis forms $a_F$, the rear-symmetric coordinate $(a_U+a_L)/\sqrt{2}$ and the rear-antisymmetric coordinate $(a_U-a_L)/\sqrt{2}$. It removes the measured mean action and then centres the samples, without rescaling or whitening them. A rank-3 decomposition relates these three action coordinates to the centred velocity residual. Its field vectors are termed \emph{action-correlated correction-field modes}. Their singular strengths measure cross-correlation, not field energy, and each mode may be multiplied by $-1$ without changing the decomposition. The modes describe action association rather than a causal response of the flow. This placement has a broader correlation-oriented ancestry: extended POD relates fields to correlated events, observable-ranked decompositions order velocity structures by their linear observability in another signal, and point--field correlations have been used for descriptive field reconstruction \citep{Boree2003ExtendedPOD,JordanEtAl2007ObservableJetModes,DiscettiEtAl2018CorrelatedFieldEstimation}. Sparse sensor-to-field modelling provides a related observable-associated context \citep{LoiseauEtAl2018SparseROM}. These precedents motivate the distinction from energy ranking; they do not validate CCD, import another decomposition, or establish causality.
|
|
|
|
Proper orthogonal decomposition (POD) provides the field-only comparison \citep{SchlegelEtAl2012LeastOrder}. It orders velocity-field variance under the same weights without using the action coordinates. Here POD and CCD use exactly the same centred residual snapshots, rank, weights and spatial region. Their rank-3 field subspaces are essentially the same. CCD therefore does not provide a superior or uniquely physical velocity basis; it adds coordinates that associate the shared low-rank field organization with front, rear-symmetric and rear-antisymmetric motion.
|
|
|
|
Together, the mean fields show that constant rear-cylinder rotation accounts for most of the measured scalar error reduction. A separately centred velocity residual has descriptive action-correlated organization. This relationship does not identify SR terms with CCD coordinates or independently confirm the action analysis. It applies only to this K\'arm\'an cloak, these four flows and the stated spatial and temporal comparison.
|
|
|
|
\draftfigure{Four role means and signed streamwise differences with the scalar error ledger, followed by the separately centred residual construction and representative action-correlated phase organization.}{Planned mean-field and correction-field analysis for the periodic K\'arm\'an cloak. Target, physical-zero, constant and PPO-controlled means will precede the two-component scalar error accounting, for which constant control accounts for 93.42\% of the measured zero-to-PPO reduction. Separate panels will define the non-paired centred residual and its action-correlated correction-field modes. The scalar shares are not residual norms or energy fractions, and the modal association is descriptive and noncausal.}{fig:mean-field-ccd}
|
|
|
|
\section{Extensions under changed temporal and target contracts}
|
|
\label{sec:extensions}
|
|
|
|
The periodic K\'arm\'an case provides a reference for changing the control problem, not a numerical or interpretive template for the cases considered here. The incoming vortex replaces periodic phase by event-relative time; Illusion replaces background matching by a declared non-zero target; and Erase compares with a clean-flow target while an upstream disturbance remains in the PPO-controlled and physical-zero flows. The policy maps its current force-plus-sparse-observation state to cylinder actions in each case, but the clocks, target fields and physical comparisons differ. Field morphology can therefore be interpreted only within each case.
|
|
|
|
In the central Taylor-vortex scene, the target, PPO-controlled and physical-zero flows are sampled at offsets relative to the detected incoming event, together with the action history. The offset indexes the event, not a periodic phase or a calibrated physical time. At offset $+10$, the three vorticity fields show distinct event-relative morphologies. Their common event coordinate places the incoming structure, its passage around the pinball and the downstream wake in one spatial comparison, so differences can be assigned to role without treating the offset as a phase. No scalar comparison is reported; only morphology is considered. Because the policy acts from the current observation state, the event motivates a hypothesis of reactive control from local information for a comparable-scale disturbance. This single scene demonstrates neither control across vortex families, amplitudes or offsets nor observability or physical-flow memorylessness.
|
|
|
|
Illusion changes the desired wake signature rather than the clock. In the uniform-inflow/free-slip-wall configuration, the retained $0.75$ target-size, seed-43 policy was executed toward a declared non-zero target. Unlike background matching, that target contains a wake pattern, and the PPO-controlled morphology is consistent with retargeting toward it. No scalar change in target-relative error is reported. The historical negative Illusion symbolic-regression result remains specific to that method: this PPO execution supplies neither a positive Illusion symbolic law nor a transfer of the K\'arm\'an surrogate.
|
|
|
|
Erase changes both the physical scene and the temporal comparison. In the parabolic-inflow/no-slip-wall configuration, its clean target is one late field, whereas the PPO-controlled and physical-zero flows contain the upstream disturbance and contribute eight phase-referenced fields. The clean, PPO-controlled and physical-zero morphologies are visible, but one late target and an eight-snapshot set are not the same temporal object. The historical rolling reward compares an evolving same-rollout history rather than the clean target field. No scalar field evaluation is reported; averaging the retained snapshots would first remove phase-dependent variation, whereas aggregating snapshot-wise errors would retain that variation in the error. These operations define different estimands, neither of which is used here.
|
|
|
|
Together, the fields show PPO execution under changed temporal or target definitions. For Illusion, the bounded result is execution toward a declared non-zero target with morphology consistent with retargeting; Vortex and Erase remain morphology-only comparisons, and reactive-control or target-erasure interpretations remain hypotheses. The distinct configurations and temporal comparisons admit neither a common efficacy measure nor transfer of the K\'arm\'an SR/CCD interpretation. A separately obtained steady endpoint can therefore provide only non-equivalent physical context.
|
|
|
|
\draftfigure{Target, PPO-controlled and physical-zero fields for one stated Vortex event offset and the retained non-zero-target Illusion acquisition, with acquisition-specific coordinates and scales.}{Planned changed-task morphology. Vortex fields will be compared at a stated event-relative offset, never a periodic phase, and Illusion fields will show policy execution toward the declared non-zero target. These panels describe role identity and morphology only: no scalar efficacy, K\'arm\'an-law transfer, common metric or causal response is inferred.}{fig:changed-task-morphology}
|
|
|
|
\section{Steady control as physical context and a theory boundary}
|
|
\label{sec:steady-theory}
|
|
|
|
Steady rear-cylinder rotation was the historical discovery seed, and the periodic K\'arm\'an analysis later identified persistent rear counter-rotation as the dominant tested symbolic-regression term, with a smaller dynamic correction. Those findings belong to the parabolic-inflow/no-slip-wall configuration. Here a separately acquired uniform-inflow/free-slip-wall endpoint at $Re_D=50$ supplies a distinct physical comparison. It is not a continuation, amplitude match, replication or validation of the periodic control chain, and no cross-configuration arithmetic is meaningful.
|
|
|
|
With body order $(\text{front},\text{rear}_{y+},\text{rear}_{y-})$, the accepted action is $[0,+\Omega,-\Omega]$ and the surface-speed ratio is $s=\Omega R/U_\infty$. At $x/D=10$, define the two-component profile error by
|
|
\begin{equation*}
|
|
\begin{aligned}
|
|
E_{\infty,\mathrm{vector}}
|
|
&(\boldsymbol u_{\mathrm{ctl}},\boldsymbol u_{\mathrm{in}})\\
|
|
&=\frac{1}{U_\infty}\max_{y\,\in\,\mathcal C}
|
|
\left|\boldsymbol u_{\mathrm{ctl}}(y)
|
|
-\boldsymbol u_{\mathrm{in}}(y)\right|,
|
|
\end{aligned}
|
|
\end{equation*}
|
|
where $\mathcal C$ is the common-fluid cross-section and $|\cdot|$ is the Euclidean norm of the streamwise and transverse velocity differences. Thus this $L^\infty$ profile metric uses a cross-sectional maximum, not quadrature weighting. The settling-qualified endpoint at $s=3.55$ has $E_{\infty,\mathrm{vector}}=0.02317$. Stationary flow is an unsteady reference, while the reverse endpoint at $s=5.2$ differs in both sign and amplitude. The accepted velocity field is therefore a distinct low-deficit endpoint, not an optimum, a stable branch or a universal law.
|
|
|
|
The wall law $(U_w,V_w)=(-\omega r_y,\omega r_x)$ fixes the tangential surface velocity: for the accepted rear-cylinder signs, both gap-facing cardinal points move downstream and both outer cardinal points move upstream. Relative to the stationary and reverse flows, the accepted vorticity and streamline fields show a different inter-cylinder and near-wake organization. These endpoint observations motivate a viscous hypothesis: imposed wall motion may alter wall shear and vorticity production, shear-layer development, inter-cylinder transport, separation and recirculation, and hence downstream mean transport. The sequence is not a demonstrated mechanism because no amplitude-matched intervention, complete same-control-volume momentum or mechanical-energy budget, force attribution, or stability analysis isolates its links.
|
|
|
|
Rotating-cylinder studies make this separation necessary. Wake modification by auxiliary rotating cylinders varies with Reynolds number, gap, rotation rate and dimensionality \citep{Mittal2001RotatingControlCylinders,ChanEtAl2011CounterRotatingCylinders}; even a strongly suppressed viscous wake that resembles a potential doublet does not thereby become an inviscid solution \citep{ChanEtAl2011CounterRotatingCylinders}. Viscous rotating-cylinder flows can also occupy multiple steady or cyclic branches, including bistable regimes \citep{SierraEtAl2020RotatingCylinderBifurcations}, while low-Reynolds-number two-cylinder analysis requires matched inner Stokes and outer Oseen descriptions \citep{Watson1996TwoRotatingCylinders}. These results motivate an outer reference as a deliberately restricted test of streamline organization, not as a parameter transfer, branch identification or mechanism for the present endpoint.
|
|
|
|
The rearranged outer streamlines raise a specific closure question: can the finite-Re organization be explained by the circulation degrees of freedom that remain after an inviscid flow satisfies impermeability? To test that proposition without assigning viscous physics to the outer model, consider the irrotational family \citep{Crowdy2006MultipleCylinders}
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
\boldsymbol{u}(\boldsymbol{x})
|
|
&=\boldsymbol{u}_{N}(\boldsymbol{x})
|
|
+\sum_{j=1}^{3}\Gamma_j\boldsymbol{h}_j(\boldsymbol{x}),\\
|
|
\nabla\!\cdot\boldsymbol{h}_j
|
|
&=\nabla\!\times\boldsymbol{h}_j=0,
|
|
\qquad \boldsymbol{h}_j\!\cdot\boldsymbol{n}=0,\\
|
|
\oint_{B_k}\boldsymbol{h}_j\!\cdot\mathrm{d}\boldsymbol{\ell}
|
|
&=\delta_{jk}.
|
|
\end{aligned}
|
|
\end{equation}
|
|
Here $\boldsymbol{u}_N$ is the zero-period Neumann solution and the counter-clockwise circulations $\Gamma_j$ are prescribed periods spanning the harmonic freedom left by impermeability; cylinder spin does not select them. This freedom reorganizes the outer streamlines, but it cannot impose rotating no slip; viscous rotating-cylinder theory retains that tangential boundary condition under its own low-Reynolds-number assumptions \citep{UedaEtAl2003RotatingCylinders}. For one cylinder,
|
|
\begin{equation}
|
|
\begin{aligned}
|
|
u_\theta(R,\theta)
|
|
&=-2U_\infty\sin\theta+\frac{\Gamma}{2\pi R}\\
|
|
&\neq R\Omega
|
|
\quad\hbox{for all }\theta\hbox{ when }U_\infty\neq0.
|
|
\end{aligned}
|
|
\end{equation}
|
|
For the normalized rear-opposite basis, define the dimensionless coefficient $g$ by
|
|
\begin{equation*}
|
|
\frac{(\Gamma_1,\Gamma_2,\Gamma_3)}{U_\infty D}
|
|
=\frac{g(0,1,-1)}{\sqrt{2}},
|
|
\end{equation*}
|
|
with positive circulation geometrically counter-clockwise. The outer strip objective selects $g_{\mathrm{opt}}=6.148$, whereas the descriptive projection of the accepted finite-Re endpoint gives $g_{\mathrm{fit}}=-12.20$. The no-slip incompatibility and opposite sign reject wall-spin-to-circulation closure: the outer family organizes geometry but does not explain the viscous endpoint.
|
|
|
|
The periodic K\'arm\'an fields and actions contain a persistent rear-rotation reference plus a smaller correction under their own configuration. The separate steady velocity field adds a low-deficit endpoint and motivates the viscous sequence above, but does not establish its links. Distinguishing those links requires amplitude-matched interventions, held-out endpoint and convergence tests, wall-vorticity flux and separation diagnostics, and a closed same-control-volume momentum/mechanical-energy budget. Until those measurements are available, the one-circle no-slip incompatibility and the opposite signs of $g_{\mathrm{opt}}$ and $g_{\mathrm{fit}}$ leave the inviscid circulation closure rejected.
|
|
|
|
\draftfigure{Stationary, accepted and reverse steady endpoint fields and profiles, together with prescribed-circulation outer-reference streamlines and the opposite-sign closure comparison.}{Planned steady endpoint and theory boundary in the uniform-inflow/free-slip-wall configuration. The finite-$Re_D$ panels will distinguish the settling-qualified accepted endpoint from the unsteady stationary reference and the non-amplitude-matched reverse endpoint. Outer-reference panels will compare $g=0$, $g_{\mathrm{opt}}$ and the descriptive fitted $g$; no-slip incompatibility and the opposite signs of $g_{\mathrm{opt}}$ and $g_{\mathrm{fit}}$ reject the proposed wall-spin-to-circulation closure. The composition provides context, not a viscous mechanism or cross-configuration validation.}{fig:steady-theory-boundary}
|
|
|
|
\section{Qualitative open-loop experiment}
|
|
\label{sec:experiment}
|
|
% AUTHOR-OWNED PLACEHOLDER: replace only after provenance review; this is not a reported result.
|
|
\textit{[Author-supplied qualitative open-loop experiment text will be inserted here after provenance review.]}
|
|
|
|
\draftfigure{Author-supplied apparatus schematic and, only after provenance review, selected experimental and separately generated CFD single-frame views for named open-loop conditions.}{Planned qualitative open-loop experiment figure. The apparatus and command/acquisition path will be identified from author-supplied records. Any retained experimental and CFD images will be labelled as separate, unregistered, acquisition-specific single-frame views for qualitative morphology only. No temporal change, velocity, force, field matching, quantitative agreement or CFD validation will be inferred.}{fig:open-loop-experiment}
|
|
|
|
\section{Discussion and conclusions}
|
|
\label{sec:discussion-conclusions}
|
|
|
|
\subsection{Discussion}
|
|
\label{subsec:discussion}
|
|
|
|
Three independently rotating cylinders can alter wakes defined by different temporal and target conditions. In the uniform-inflow/free-slip-wall configuration, repeated K\'arm\'an training runs produced high-performing policies across the tested realizations. In the separately documented parabolic-inflow/no-slip-wall configuration, the controlled K\'arm\'an velocity field is closer to its declared target than physical zero in both complete-cycle mean and phase-resolved comparisons. The first result concerns repeatable policy discovery under the online objective; the second concerns the downstream field under a different configuration. Together they locate the periodic K\'arm\'an cloak as the case in which learned rotation is accompanied by a direct target-relative field comparison, without treating the two configurations as one continuous data set.
|
|
|
|
Persistent counter-rotation of the rear cylinders is the clearest recurring feature of the periodic control. Removing that term degrades the compact surrogate more than removing the tested modulations. In the mean velocity field, the physical-zero pinball leaves a mean streamwise deficit relative to the target, while the controlled-minus-physical-zero field is largely opposite in sign. Holding the cylinders at the mean DRL action produces a constant-control field that accounts for most of the measured scalar mean-field error reduction; the remaining constant-to-DRL scalar increment is distinct from the centred velocity residual. That residual subtracts separately centred DRL and constant-control phase fields without cycle pairing, and its action-correlated modes describe phase organization rather than its size or cause. The surrogate ranking, scalar error increments and residual modes are therefore compatible but non-equivalent: they support persistent rear-cylinder rotation as an organizing component, not a unique decomposition or mechanism.
|
|
|
|
Changing the task changes what can be inferred. For the incoming-vortex case, the policy was executed on an event-relative clock and the retained fields show target, controlled and physical-zero morphology at the selected event offset, without a scalar efficacy result. For Illusion, the policy was executed toward a declared non-zero wake target, and the retained role fields are consistent with that bounded morphology comparison; no scalar target-relative improvement is claimed. The separate steady case adds physical context: prescribed rear-cylinder counter-rotation accompanies a low-deficit mean-flow endpoint, while the wall kinematics point toward viscous shear-layer and transport effects. The prescribed-circulation outer flow cannot select circulation from cylinder spin, and its opposite sign rejects the proposed inviscid closure. Persistent rotation is therefore a useful kinematic coordinate across the stated cases, but not a transferred control law or causal explanation.
|
|
|
|
\subsection{Conclusion}
|
|
\label{subsec:conclusion}
|
|
|
|
Rotational control was executed for periodic, event-relative and non-zero-target wake definitions. The periodic K\'arm\'an cloak provides the strongest field-level result: in the parabolic-inflow/no-slip-wall configuration, its controlled registered velocity field is closer to the declared target than physical zero. The Vortex and Illusion cases extend the tested task definitions through event-relative execution and bounded non-zero-target morphology, without supplying a common efficacy measure.
|
|
|
|
For the periodic K\'arm\'an case, persistent rear-cylinder counter-rotation links the dominant tested surrogate term to the constant-control mean field, while the centred velocity residual describes time-dependent organization separately. This physical reading is deliberately limited: action ranking, scalar mean-field increments and residual modes answer different questions. Hydrodynamic cloak and illusion here denote control relative to a declared reference for specified observations and comparisons, not universal invisibility, configuration-independent transfer or an unrestricted physical law.
|
|
|
|
\section*{Acknowledgements}
|
|
The authors acknowledge the computational resources and discussions that supported this work.
|
|
|
|
\section*{Funding}
|
|
Funding information will be supplied before submission.
|
|
|
|
\section*{Declaration of interests}
|
|
The authors report no conflict of interest.
|
|
|
|
\section*{Author ORCIDs}
|
|
Author ORCID information will be supplied before submission.
|
|
|
|
\bibliographystyle{jfm}
|
|
\bibliography{../ROUND5_BOTTOM_UP/R5_REFERENCES}
|
|
|
|
\end{document}
|