refactor(workflow): update JFM research workflow and remove obsolete figures

Revise the JFM research workflow documentation to clarify project structure, decision-making processes, and tool routing. Remove outdated figures and images that are no longer relevant to the current manuscript. Ensure that the workflow aligns with the latest research objectives and safety protocols.
This commit is contained in:
Frank14f
2026-08-11 21:25:33 +08:00
parent 0ade812864
commit 1a8b994a03
101 changed files with 5757 additions and 23827 deletions
+291 -569
View File
@@ -1,28 +1,28 @@
\documentclass[lineno]{JFM-FLM_Au}
\usepackage{amsmath,amssymb,bm,booktabs}
\newcommand{\Rey}{\mathit{Re}}
\newcommand{\dd}{\mathrm{d}}
\newcommand{\vect}[1]{\boldsymbol{#1}}
\newcommand{\draftfigure}[3]{%
\begin{figure}
\centering
\fbox{\parbox[c][0.22\textheight][c]{0.88\textwidth}{\centering\textbf{Figure placeholder}\\[0.6em]#1}}
\fbox{\parbox[c][0.22\textheight][c]{0.88\textwidth}{}}
\caption{#2}
\label{#3}
\end{figure}
}
\lefttitle{Y. Wang, F. Ren, S. Chen and H. Tang}
\righttitle{Symbolic feedback laws for hydrodynamic cloaking and illusion}
\lefttitle{Y. Wang, F. Ren, X. Wang, H. Bai, H. Tang and S. Chen}
\righttitle{Symbolic control laws for hydrodynamic cloaking and illusion}
\title{Interpretable feedback laws for active hydrodynamic cloaking and illusion of the fluidic pinball}
\title{Interpretable control laws for active hydrodynamic cloaking and illusion of the fluidic pinball}
\author{Yanqi Wang\aff{1,3}, Feng Ren\aff{2}, Shiyi Chen\aff{3} \and Hui Tang\aff{1}}
\author{Yanqi Wang\aff{1,3}, Feng Ren\aff{2}, Xu Wang\aff{1}, Honglei Bai\aff{4}, Hui Tang\aff{1}, \and Shiyi Chen\aff{3} }
\affiliation{\aff{1}Department of Mechanical Engineering, The Hong Kong Polytechnic University, Hong Kong, PR China
\affiliation{\aff{1}Department of Mechanical Engineering, The Hong Kong Polytechnic University, Hong Kong SAR, PR China
\aff{2}School of Marine Science and Technology, Northwestern Polytechnical University, Xi'an 710072, PR China
\aff{3}Eastern Institute for Advanced Study, Eastern Institute of Technology, Ningbo 315200, PR China}
\aff{3}Eastern Institute for Advanced Study, Eastern Institute of Technology, Ningbo 315200, PR China
\aff{4}School of Aeronautics and Astronautics, Sun Yat-Sen University, Shenzhen Campus, Shenzhen 518107, PR China}
\corresau{Hui Tang, \email{h.tang@polyu.edu.hk}}
@@ -30,7 +30,7 @@
\maketitle
\begin{abstract}
Hydrodynamic cloaking and illusion require a controller to manipulate not only an integrated force but also the information transported by a nonlinear wake. We study these observer-centred control objectives using the fluidic pinball, a compact array of three independently rotating cylinders. Proximal-policy-optimization controllers are first trained from the forces on the cylinders and three downstream velocity probes. We then use symbolic regression to determine whether the successful neural policies contain a compact and transferable feedback structure. Candidate expressions are constrained by the reflection symmetry of the pinball and are accepted only after deployment in closed-loop lattice-Boltzmann simulations; one-step regression accuracy alone is shown to be an unreliable selection criterion. For K\'arm\'an-street cloaking, a joint expression extracted across diameter-based Reynolds numbers $\Rey_D=25$--$200$ separates into an approximately steady counter-rotation of the rear cylinders and dynamic front-cylinder rate--lift feedback. The symbolic controller obtains downstream similarities of $0.847$, $0.888$, $0.845$ and $0.806$, respectively, and transfers without refitting to a Lamb vortex dipole and a Taylor vortex monopole with similarities of $0.949$ and $0.905$. Illusion of target-cylinder wakes admits a common target-error law near target diameters $0.75D$--$1.0D$, but transfer deteriorates as the target scale departs from the naturally controllable wake. At $1.5D$, the neural action contains a high-frequency modulation that is not represented by the present instantaneous feature library, exposing a clear symbolic-model regime boundary. A compact correction-field analysis provides independent spatial support for the actuator interpretation. The results show that deep reinforcement learning can serve as a policy-discovery stage from which a physically interpretable cloak backbone is distilled, while also identifying the conditions under which target-dependent temporal structure must be retained.
Hydrodynamic cloaking and illusion require active control of the wake transported beyond a body, rather than optimization of an integrated force alone. We investigate these declared-reference objectives using a fluidic pinball of three independently rotating cylinders. A proximal-policy-optimization controller maps the forces on the cylinders and velocities from three downstream probes to rotational commands, while registered velocity fields provide an independent measure of the resulting wake. For periodic K\'arm\'an cloaking, the controlled mean and phase-resolved fields approach the target substantially more closely than the uncontrolled zero-rotation wake. Post-hoc symbolic regression reveals an executable control law organized around persistent counter-rotation of the rear cylinders, supplemented by a smaller time-dependent correction. Closed-loop deletion tests identify the persistent rear rotation as the dominant tested component. A complementary four-flow analysis shows that fixing the cylinders at the mean learned action accounts for 93.42\% of the measured reduction in scalar mean-field error. The remaining time-dependent field organization is examined separately through action-correlated correction-field modes, linking low-rank wake structure to front, rear-symmetric and rear-antisymmetric actuation coordinates. The framework also accommodates an event-relative Vortex case and a non-zero-target Illusion case, demonstrating how the control problem changes with its temporal coordinate and prescribed wake signature. A separate steady-flow analysis connects persistent rear rotation to a low-deficit endpoint and tests its relation to a prescribed-circulation outer flow. Together, these results show how learned rotational control can be evaluated against declared wake references and distilled into a compact, physically readable K\'arm\'an control structure.
\end{abstract}
\begin{keywords}
@@ -40,632 +40,354 @@ hydrodynamic cloaking, flow control, symbolic regression, deep reinforcement lea
\section{Introduction}
\label{sec:introduction}
A bluff body leaves a spatially and temporally extended signature in the surrounding fluid. The signature includes a mean velocity deficit, separated shear layers, coherent vortices, pressure fluctuations and characteristic frequencies. These quantities are not merely consequences of the force on the body: they are also information from which a downstream observer may infer the presence, size or state of an upstream object. Active flow control may therefore be formulated not only as drag reduction, lift control or suppression of vortex-induced vibration, but also as control of the information transmitted by the wake.
A hydrodynamic cloak seeks to make the flow around a body approach a declared background, while a hydrodynamic illusion seeks a declared non-zero substitute signature. Transformation optics established complementary routes to controlling linear fields through material design \citep{PendryEtAl2006Controlling,Leonhardt2006OpticalMapping}. Active source-based cloaking has also been formulated for specified linear equations \citep{Miller2006PerfectActiveCloaking,VasquezEtAl2009ActiveExterior}. In a distinct role, active constructions have retargeted exterior fields towards non-zero illusion signatures \citep{MaEtAl2013ActiveCloaking,LinEtAl2021ActiveAcousticIllusions}. These branches distinguish redirecting a field, generating compensation and prescribing a substitute signature. They assume linear governing relations and prescribed material or source responses, whereas a viscous wake evolves with the actuation and body-generated vorticity; their role here is conceptual rather than predictive. A multilayer hydrodynamic cloak in free flow has been reported \citep{ChenEtAl2024MultilayerHydrodynamicCloak}. Separated finite-Reynolds-number flow nevertheless poses a distinct problem even relative to active hydrodynamic-cloak formulations \citep{UrzhumovSmith2012ActiveCloaks}: nonlinear advection couples forcing to the evolving wake, no-slip bodies generate shear layers and separation, and viscosity and confinement alter the response. The attainable match depends on the declared target and comparison, especially when coherent shedding carries information beyond local measurements. We therefore consider controlled approach to a specified background or substitute wake, assessed beyond the control observations, rather than universal invisibility or exact cancellation in this setting.
This viewpoint gives two related objectives. \emph{Hydrodynamic cloaking} requires the downstream flow to approach the background that would have reached the observer in the absence of the controlled body. For a steady background this reference is an undisturbed channel flow; for an incoming vortex street or isolated vortex it is the corresponding disturbance propagated without the body. \emph{Hydrodynamic illusion} instead requires the downstream flow to reproduce the non-zero signature of a different target object. The body remains present and exchanges momentum with the fluid in both problems. Cloaking is consequently not defined as disappearance of the local near-body disturbance, and illusion is not defined by force matching alone. The relevant criterion is the field or signal perceived outside the immediate actuation region.
The desired outcome is an extended Navier--Stokes field, whereas actuation is confined to compact boundaries. The fluidic pinball---three independently rotating cylinders in channel flow---is a useful model because a small multi-input plant generates rich nonlinear wakes and admits established open- and closed-loop control studies \citep{DengEtAl2020PinballBifurcations,CornejoMacedaEtAl2021PinballControl,RaibaudoEtAl2020PinballControl}. Transient and post-transient pinball-force dynamics have been treated with reduced modelling \citep{DengEtAl2021PinballForceModel}, supporting integrated load as a useful intermediary between local rotation and the wake. Localized actuation can reorganize bluff-body shedding, and deep reinforcement learning has been applied to pinball force control \citep{FengWang2010SyntheticJetCylinder,FengEtAl2023PinballForceControl}. Conventional active flow control commonly asks how actuation changes shedding or integrated load. The present cloak/illusion objective instead adds a separately acquired reference and an extended-field test: neither scalar drag reduction nor force control establishes that the downstream velocity field approaches that reference. The resulting problem is whether three rotation commands can use force-mediated feedback to approach a declared hydrodynamic signature while leaving the downstream field independently testable.
Nonlinear separated wakes make these objectives difficult. Instability, mean-flow deformation, vortex formation and convection are coupled, and the observer responds only after a finite propagation delay. A controller may reduce drag while leaving a readily detectable wake, or suppress a dominant frequency while producing an incorrect mean deficit. Illusion is more restrictive still because frequency, phase, amplitude and spatial organisation must be retuned towards a non-zero target.
The difficulty increases when the reference contains periodic shedding, an incoming event or a non-zero target wake. Deep reinforcement learning has enabled closed-loop decisions in nonlinear flows under limited observations \citep{RabaultEtAl2019DRLFlowControl,ParisEtAl2021RobustFlowControl}, including rotating-cylinder actuation in towing experiments and simulations \citep{FanEtAl2020BluffBodyRL}. Related bluff-body work has framed learned active control as hydrodynamic stealth \citep{RenEtAl2021HydrodynamicStealthDRL}, although its observer metric is not the target-relative field comparison used here. Unlike a scalar optimization objective, reference following must distinguish what the controller observes from what is later evaluated. The present policy receives the two force components of each cylinder together with sparse downstream velocities. Integrated forces report body--flow interaction, while remote measurements sample the emerging wake; neither implies full observability. Sparse-probe utility is flow- and task-dependent, including sensitivity to Reynolds number and blockage \citep{LiZhang2022ConfinedWakeRL}. This interface motivates three tests: whether the policy realizes its objective, whether the velocity field approaches the target beyond the reward and sensor locations, and whether recurrent action and field structures can be reduced to forms that remain checkable in closed-loop execution.
Deep reinforcement learning (DRL) is attractive for this setting because it can discover feedback strategies directly through interaction with the flow, without requiring a tractable model of the controlled Navier--Stokes dynamics. Its principal limitation is interpretability. A successful neural policy does not by itself reveal which measurements are essential, whether the actions contain a reusable physical mechanism, or why a policy transfers between incident disturbances. This limitation is especially important for hydrodynamic perception control: a high scalar reward establishes performance only in the chosen observer space and does not identify the mechanism that produces the downstream correction.
These tests define a layered analysis. Learned execution is first assessed within each case's own configuration, target roles and temporal coordinate. The periodic K\'arm\'an cloak is then evaluated independently by registering target, controlled and physical-zero velocity fields, while temporal signal comparison remains complementary rather than a substitute for field agreement. On the action side, symbolic regression compresses the policy into a candidate control law tested by closed-loop deployment, following the principle that readable learned feedback should be judged in execution \citep{GautierEtAl2015ClosedLoop,CornejoMacedaEtAl2021PinballControl}. Section~\ref{sec:karman-law} distinguishes this post-hoc policy-map regression from direct genetic-programming control and sparse dynamical-system identification \citep{LiEtAl2017GeneticProgrammingControl,BruntonEtAl2016SINDy}; PySR is the symbolic-regression implementation \citep{Cranmer2023PySR}, not evidence that the resulting expression controls the flow. This distinction prevents a compact expression from being promoted from representation to explanation before deployment. On the field side, the persistent mean response is separated from a centred residual, and action-correlated correction-field modes describe organization without assigning causality. Sparse-to-field and reduced-order precedents motivate separating online information from fuller-state interpretation \citep{LoiseauEtAl2018SparseROM,LiEtAl2022FlowEstimation}. The action and field analyses answer different questions and do not validate one another.
The present study uses the fluidic pinball as a controlled laboratory for this question. Three equal circular cylinders form an equilateral triangle with one cylinder facing the incident flow. Independent rotation provides a compact three-input actuator capable of changing stagnation points, gap flow, base bleed, separation and shedding phase. The uncontrolled pinball also undergoes Hopf, pitchfork and higher-frequency transitions as the Reynolds number increases. It is therefore simple enough for repeated high-fidelity simulation but rich enough to expose nonlinear, multi-attractor control dynamics.
We therefore ask whether force-plus-sparse-observation rotational control can approach declared cloak and illusion references under separately specified channel-flow formulations, and how far the periodic K\'arm\'an cloak can be evaluated and interpreted. Rotational feedback was executed for the periodic K\'arm\'an cloak, the non-zero Illusion target and the event-relative Vortex case; only K\'arm\'an supports a target-relative field-proximity claim, while Illusion and Vortex provide execution and morphology evidence. For Erase, the available target and comparator roles do not yet permit the same performance comparison. Registered target, controlled and physical-zero fields show improved target proximity for the periodic K\'arm\'an cloak in the parabolic-inflow/no-slip-wall channel configuration. Closed-loop action-law tests further identify a K\'arm\'an control skeleton: dominant persistent rear-cylinder rotation accompanied by a smaller dynamic correction, with the associated field organization described by action-correlated modes. This deeper account is possible for K\'arm\'an because policy execution can be followed through registered field comparison, executable law reduction and organized mean and residual fields. The separate steady analysis instead shows that the corresponding inviscid closure fails, and does not extend the learned periodic result. The next section defines the configurations, actuation, observations, target roles and independent evaluation criteria used to keep these comparisons distinct.
The central hypothesis is that successful cloaking is organised by a reusable actuator-side compensation mechanism. The pinball creates a blockage-induced velocity deficit and an additional force signature. The learned policy compensates the mean deficit through rear-cylinder counter-rotation and regulates the residual phase and antisymmetric force through dynamic front-cylinder actuation. We test this hypothesis by distilling the observation-to-action map with symbolic regression (SR). Crucially, symbolic expressions are judged in closed-loop computational fluid dynamics (CFD), not only by their ability to fit actions recorded from the DRL trajectory. A compact correction-field decomposition is then used as an independent check of the spatial interpretation.
\section{Problem formulation, control framework and evaluation}
\label{sec:formulation}
The principal contributions are as follows. First, cloaking and illusion are formulated as observer-centred wake-matching problems for a nonlinear bluff-body flow. Secondly, a symmetry-constrained SR procedure is developed for post hoc distillation of a trained PPO policy, with closed-loop CFD as the decisive model-selection test. Thirdly, a common K\'arm\'an-cloak law is identified across $\Rey_D=25$--$200$ and transferred without refitting to periodic and transient incident disturbances. Fourthly, illusion is shown to reuse this compensation backbone only over a finite target-scale range; high-frequency neural actuation at larger mismatch reveals the information missing from the present symbolic state.
\subsection{Flow configurations and numerical method}
\label{subsec:configurations}
The controlled body is a two-dimensional fluidic pinball comprising three circular cylinders of diameter $D$. A front cylinder is followed by upper and lower rear cylinders, and all three rotate independently. Streamwise and transverse coordinates are $(x,y)$, and the computational body order is $(F,U,L)=(\text{front},\text{rear}_{y+},\text{rear}_{y-})$. Lateral confinement bounds cross-stream motion and provides a repeatable environment in which rotation can alter the integrated forces and wake. Two channel configurations are used for different parts of the study. The parabolic-inflow/no-slip-wall channel supplies the periodic K\'arm\'an field comparison and the detailed action- and field-side analyses. The uniform-inflow/free-slip-wall channel supplies the broader controlled cases and the separate steady analysis. Their results are interpreted within their own physical and numerical settings rather than combined as realizations of one plant. A case-dependent upstream disturbance or target body defines the requested wake, with its dimensions specified for that case.
The paper is organised as follows. Section~\ref{sec:problem} defines the physical configuration and the Reynolds-number convention. Section~\ref{sec:methods} describes the numerical environment, DRL policy and SR protocol. Section~\ref{sec:srresults} presents the symbolic controllers and their closed-loop transfer. Sections~\ref{sec:oid} and~\ref{sec:ccd} provide independent spatial and spatio-temporal tests of the mechanism. Conclusions are given in \S~\ref{sec:conclusion}.
\subsection{Relation to interpretable flow control and system identification}
The aim of extracting an analytical policy has precedents in machine-learning control, sparse system identification and reduced-order modelling, but the present task differs from each of these in a consequential way. Genetic-programming control searches directly over mathematical expressions and has exposed physically meaningful feedback in separated flows \citep{Gautier2015,Li2017}. Such an approach makes interpretability native to the optimisation, but repeated evaluation of candidate laws in a large CFD environment is costly. Here PPO first explores a flexible policy class, after which symbolic regression asks whether the successful behaviour occupies a much smaller algebraic class. This ordering treats the neural controller as a discovery instrument rather than as the final scientific result.
Sparse identification of nonlinear dynamics (SINDy) provides a second point of reference \citep{Brunton2016,Loiseau2018Galerkin}. SINDy and constrained sparse Galerkin regression identify parsimonious evolution equations from a prescribed library. They have established two lessons that are directly relevant here: structural constraints can matter more than small changes in residual error, and a model must be assessed by its long-time dynamics rather than its derivative fit alone. In the present work, however, the principal pipeline is PySR applied to the policy map, not SINDy applied to the Navier--Stokes state. SINDy is used only as a historical sparse-regression reference and a cross-diagnostic in the difficult $1.5D$ illusion case. There is no DANTE pipeline in this study, and no claim is made that the discovered expressions are governing equations.
The distinction between policy imitation and closed-loop equivalence is central. Supervised policy distillation usually minimises an action discrepancy over states visited by the teacher. In a convectively delayed flow, replacing the teacher changes those states. A formula can accurately interpolate the PPO trajectory yet introduce a small phase bias that compounds until the downstream signal no longer matches the reference. Conversely, a low-$R^2$ formula can preserve the stabilising or tracking geometry that matters in closed loop while ignoring high-variance action details. We therefore place closed-loop CFD inside the selection loop conceptually, even though the computational pipeline performs fitting and validation as separate reproducible stages.
The fluidic pinball is especially suitable for this test. Its unforced dynamics pass through successive Hopf and symmetry-breaking bifurcations, and its force observables obey symmetry-constrained low-order structure \citep{Deng2020,Loiseau2018Galerkin}. Rotation of the downstream pair can create base bleed, while front-cylinder rotation can supply phase-sensitive forcing; related roles emerged in gradient-enriched machine-learning control \citep{CornejoMaceda2021}. These prior results provide physical hypotheses, not labels supplied to PySR. A credible symbolic law should recover actuator roles compatible with this dynamical organisation and should remain useful beyond one recorded waveform.
Finally, the term \emph{cloaking} is used here in an observer-centred and explicitly limited sense. Active hydrodynamic metamaterial studies show that cancellation of a body-induced disturbance can be designed over restricted regimes, while also exposing stability limitations at finite Reynolds number \citep{Urzhumov2012}. Our sparse-probe similarity measures whether selected downstream signals are recovered; it is not evidence of pointwise invisibility throughout the domain. Full-field correction analyses reduce, but do not eliminate, this observer dependence. Likewise, illusion denotes reproduction of selected force and wake signatures of a target cylinder, not a claim that every external measurement would classify the pinball as that cylinder.
\section{Flow configuration and control objectives}
\label{sec:problem}
\subsection{Fluidic pinball}
The controlled body consists of three circular cylinders of diameter $D$. Their centres form an equilateral triangle of side $1.5D$, leaving a gap of $0.5D$ between adjacent surfaces. The upstream cylinder is denoted by $F$ and the downstream upper and lower cylinders by $T$ and $B$. The three angular velocities form the actuation vector
Velocities are normalized by a configuration-specific reference speed $U_{ref}$, lengths by $D$, and time by $D/U_{ref}$. The article-wide Reynolds number is
\begin{equation}
\vect{\omega}(t)=(\omega_F,\omega_T,\omega_B)^{\mathsf T},
\begin{aligned}
Re_D &= \frac{U_{ref}D}{\nu},\\
U_{max} &\simeq 1.5U_{ref}
\quad\text{for parabolic inflow}.
\end{aligned}
\end{equation}
The Legacy solver stores the applied rotational command in velocity units as $\omega_i$; the SR pipeline therefore defines $\alpha_i=\omega_i/U_0$. This project-specific variable should not be confused with a dimensional angular frequency. The corresponding no-slip wall velocity is prescribed by the solver convention below.
On cylinder $i$, the no-slip surface velocity is
where $\nu$ is the kinematic viscosity. For uniform inflow, $U_{ref}$ is the inlet speed; for parabolic inflow, it is the bulk/mean inlet speed. This distinction matters when comparing inherited case labels: a label based on $2D$ can be halved only if its velocity scale is also $U_{ref}$. Historical parabolic records do not always identify $U_0$ as centreline or bulk/mean speed, so no physical Reynolds number is inferred from such labels here. Both configurations are advanced with a D2Q9 lattice-Boltzmann method using multiple-relaxation-time collision and curved moving-body treatment. Inlet, lateral-wall and outlet implementations follow the stated configuration. The detailed K\'arm\'an analyses use data from the parabolic/no-slip realization, whereas the control formulation described next is implemented in the uniform/free-slip configuration; no numerical equivalence between the two implementations is assumed.
\subsection{Rotational actuation, observations and PPO training}
\label{subsec:ppo-framework}
For a radius vector $(r_x,r_y)$ measured from a cylinder centre, angular velocity $\omega$ imposes the wall velocity
\begin{equation}
\vect{u}_{w,i}=\omega_i R
\begin{pmatrix}-\sin\theta_i\\ \cos\theta_i\end{pmatrix},
\qquad R=D/2.
\begin{aligned}
(U_w,V_w) &= (-\omega r_y,\omega r_x),\\
s_i &= \frac{\omega_iR}{U_{ref}}.
\end{aligned}
\end{equation}
Three downstream probes measure both velocity components. The force on each pinball cylinder is also available to the controller.
where $R=D/2$, positive $\omega$ is counter-clockwise, and $s_i$ is the nondimensional surface-speed command for cylinder $i\in\{F,U,L\}$. The historical symbolic-regression records use $\alpha_i=\omega_i/U_0$ instead. The two descriptions are related by $s_i=\alpha_iR(U_0/U_{ref})$ and coincide as $s_i=\alpha_iR$ only when $U_0=U_{ref}$. Thus $\alpha_i$ is not dimensionless unless a radius-one convention is also specified. In the canonical control formulation, the policy commands rotation without a preset steady-derived bias, and successive commands are smoothed before application. This choice defines the computational interface; it is not a comparison of bias strategies or a map to an experimental actuator. Prior pinball work has treated force dynamics with reduced models and used learned control for force objectives \citep{DengEtAl2021PinballForceModel,FengEtAl2023PinballForceControl}. This motivates integrated force as a compact actuation-coupled signal, without making it a downstream-field measure.
\draftfigure{Insert the fluidic-pinball geometry, disturbance generator and three downstream probes.}{Computational arrangement and observer-centred control objectives. The three pinball cylinders rotate independently; the reference trajectory is recorded without the pinball for cloaking and with a separate target cylinder for illusion.}{fig:configuration}
The observation vector contains the two force components of each cylinder and both velocity components from three circular wake probes. The probes are centred $10D$ downstream of the signature source at $(y-y_0)/D=0,\pm2$ and have radius $0.25D$. Each probe contributes area-averaged streamwise and transverse velocity. Forces and wake velocities describe complementary aspects of the controlled flow: periodic loads can carry phase information, while sparse wake measurements provide downstream coordinates \citep{NairEtAl2021PhaseFlowControl,MarraEtAl2024ActuationManifold}. The centre probe is intended to sample coherent-shedding phase, and the off-centre pair to sample lateral deflection and alternating cross-wake motion. Comparable downstream sensing has been used to characterize and control fluidic-pinball wakes \citep{RaibaudoEtAl2020PinballControl,CornejoMacedaEtAl2021PinballControl,LiEtAl2022FlowEstimation}; sparse wake information is also physically accessible in other wake-sensing settings \citep{BeemTriantafyllou2015WakeSensing}. These precedents motivate the measurement classes, not the exact present geometry.
The incompressible flow satisfies
\begin{align}
\nabla\boldsymbol{\cdot}\vect{u}&=0,\\
\rho_f\left(\frac{\partial\vect{u}}{\partial t}
+\vect{u}\boldsymbol{\cdot}\nabla\vect{u}\right)
&=-\nabla p+\mu\nabla^2\vect{u}+\vect{f}_e,
\end{align}
where $\vect{f}_e$ represents the force density used to impose the moving curved boundaries.
Sensor placement is inseparable from its use and from the time scales carried by the chosen observable. In periodic wakes, a phase-bearing signal can encode dominant shedding dynamics in the regime for which its model is constructed \citep{GongEtAl2020VortexEstimation}. For feedback, however, moving a sensor downstream can increase developed-wake information while also increasing convective lag, so a location that is effective for estimation need not be best for closed-loop control \citep{JinEtAl2022SensorActuatorPlacement}. Sparse instantaneous measurements can be augmented with temporal histories, but such lifting increases input dimension and redundancy \citep{WangEtAl2024DynamicFeatureDRL}. These results motivate retaining spatially distributed and temporally resolved information, but do not choose the present geometry. The three-probe arrangement is consequently a compact wake monitor rather than a claim of optimal placement or full observability. Solver averaging, fixed physical pre-scaling and online observation whitening are applied as distinct operations. Because whitening parameters form part of the policy input map, evaluation uses the normalizer saved with the selected policy.
\subsection{Reynolds-number convention}
\label{sec:reconv}
PPO updates a stochastic policy by maximizing a clipped surrogate objective relative to the preceding policy \citep{SchulmanEtAl2017PPO}. Clipping limits the incentive for a large policy change, but does not guarantee convergence or successful control. The canonical network has two hidden layers of 64 units with sinusoidal activation. Training uses 2048-step rollouts, batches of 64, ten update epochs, discount factor $0.995$ and learning rate $3\times10^{-4}$. These settings specify the reusable canonical training procedure. Reproducing an individual trained policy additionally requires its selected checkpoint, matching normalizer, target signals and run-specific objective settings; the canonical defaults alone do not determine that policy.
Throughout this paper, the physical Reynolds number follows the fluidic-pinball convention and is based on the diameter of one pinball cylinder,
\subsection{Target construction and online objective}
\label{subsec:targets-objective}
Every comparison distinguishes how its signals and fields are acquired. The target or reference is the wake requested by the case; controlled denotes the trajectory generated with the policy; physical zero or uncontrolled denotes an independently acquired zero-rotation flow; and constant denotes a separately acquired fixed-actuation comparator where one exists. A cloak target is a declared background flow, whereas an illusion target is a separately generated non-zero wake, including a target-body acquisition when required. The target trajectory and the physical-zero trajectory therefore play different roles even when both contain no policy actuation. Target-body size and upstream geometry are specified case by case so that the requested signature is not silently transferred between configurations.
During control, the policy minimizes discrepancies in the same compact information space that it observes. We write the online objective as
\begin{equation}
\Rey_D=\frac{U_0D}{\nu}.
\label{eq:red}
\begin{aligned}
\mathcal{J}_{online}
={}& w_F\,\mathcal{D}_F(\mathbf F,\mathbf F_t)\\
&+w_S\,\mathcal{D}_S(\mathbf q,\mathbf q_t).
\end{aligned}
\end{equation}
The project-level solver and several historical file names use a different reference length, $2D$, and therefore store
where $\mathbf F$ collects the selected body-force components, $\mathbf q$ contains the six probe velocities, and subscript $t$ denotes the corresponding target signal. The force and sensor terms allow rotation to respond both to the body's integrated interaction with the flow and to the developing wake. Their channel aggregation, normalization, orientation and weights are selected for each case. Because this objective samples forces and three wake regions rather than the complete velocity field, reward improvement establishes signal-space execution only. Field matching is tested separately.
\subsection{Independent signal- and field-level evaluation}
\label{subsec:evaluation}
Dynamic time warping (DTW) uses dynamic programming to find a minimum-cost monotone path through the pairwise discrepancies of two sampled sequences \citep{SakoeChiba1978DynamicProgramming,Muller2015FundamentalsMusicProcessing}. Declared endpoint, window and step constraints permit limited local time reparameterization without reducing the comparison to a single global phase shift. For each result, we state whether DTW is a distance or higher-is-better similarity, whether channels are native or normalized, and the time window, scaling, aggregation and path constraints. These conventions are project-specific rather than supplied by the generic DTW references. Because excessive warping can hide timing errors, DTW is not interpreted as a physical delay, causal relation, observability test or field-equivalence measure.
Field agreement is evaluated over a registered downstream rectangle,
\begin{equation}
\Rey_{\mathrm{code}}=\frac{U_0(2D)}{\nu}=2\Rey_D.
\label{eq:recode}
\Omega=\left\{(x,y):
\begin{aligned}
-6 &\leq \frac{x-x_s}{D} \leq 14,\\
\left|\frac{y-y_0}{D}\right| &\leq 5
\end{aligned}
\right\}.
\end{equation}
These quantities are not interchangeable. All Reynolds numbers reported in the text, tables and conclusions below are $\Rey_D$ unless the code-level quantity is explicitly identified.
The mechanism dataset used for the present SR study was generated with the Legacy configuration on a $1280\times512$ lattice, with $D=20$ lattice units, $U_0=0.01$, a parabolic inlet and no-slip upper and lower walls. A newer V5 training environment exists on a $2000\times600$ lattice with a uniform inlet and free-slip walls, but it is not mixed with the numerical coefficients reported here. The distinction is necessary because the two environments have different plant dynamics, action decoders and observation normalisation. The present expressions and performance values refer to the Legacy analysis dataset.
\subsection{Cloaking and illusion}
Let $\vect{q}_{\mathrm{in}}$ denote the incident reference flow, $\vect{q}_{\mathrm{blk}}$ the flow with a fixed pinball and $\vect{q}_{\mathrm{ctl}}$ the controlled flow. Cloaking seeks
where $x_s$ is the sensor-plane streamwise location and $y_0$ is the channel centreline. For $r\in\{\text{controlled},\text{physical zero}\}$, the complete-cycle mean-field error is
\begin{equation}
\mathcal{H}\vect{q}_{\mathrm{ctl}}(t)
\simeq \mathcal{H}\vect{q}_{\mathrm{in}}(t),
\label{eq:cloakobj}
\begin{split}
E_{mean,r}=\Biggl[\frac{1}{|\Omega|}\int_{\Omega}
&\frac{\|\overline{\mathbf u}_r
-\overline{\mathbf u}_t\|_2^2}{U_{ref}^2}\\
&\,\mathrm d\Omega\Biggr]^{1/2}.
\end{split}
\end{equation}
where $\mathcal{H}$ is the downstream observation operator. For K\'arm\'an cloaking, $\vect{q}_{\mathrm{in}}$ contains the street generated by an upstream cylinder but no pinball. For the Lamb and Taylor cases it contains an isolated dipole or monopole-like event.
The fixed $20D\times10D$ region is registered to the sensor plane and centreline so that each role is compared over the same downstream extent. Saved fields do not preserve solver-exact masks, however, so the comparison uses a geometry guard rather than claiming an exact common-fluid mask. A comparator-relative reduction is formed only when target, controlled and physical-zero fields share the same acquisition definition, region, weighting, normalization and averaging window.
For illusion, a separate target cylinder generates $\vect{q}_{\mathrm{tar}}$, and the objective becomes
The same spatial error is evaluated at eight canonical phase slots for the periodic K\'arm\'an case. Each role is phased independently from its own centre-probe transverse velocity, and each slot uses the nearest saved snapshot. The resulting phase diagnostic samples one canonical cycle; it is not a repeated-cycle ensemble or an unrestricted phase optimization. Other cases require different temporal coordinates: Vortex uses offsets from a declared event, Erase admits a mean-field comparison only when its roles are resolved, and the steady analysis uses a late field or profile. These quantities are not pooled because they answer different physical questions. DTW measures temporal signatures close to the controller's information space, whereas mean and phase errors interrogate the extended velocity field. Either can improve without the other, making their separation central to the evaluation.
\subsection{Compact CFD qualification}
\label{subsec:cfd-qualification}
The numerical method is assessed through selected observables in two CelerisLab benchmark configurations. In the rotating-cylinder Kan99b K2 case, Strouhal number, mean drag and force-fluctuation amplitudes lie within their prescribed bands. The mean-lift sign differs because the CFD and comparison conventions use opposite force directions, so the overall K2 comparison remains a \emph{partial pass} despite the conventional origin of that sign difference. In the confined-cylinder Sah04 S2 case, the Strouhal number meets its prescribed band. Only these two completed simulations enter the qualification. They support the named frequency and force observables under their benchmark conditions, but do not assess PPO training, target construction, symbolic regression, correction-field interpretation or hydrodynamic cloaking.
With the configurations, controller interface, target roles and independent measures defined, the following sections compare the controlled cases without combining results across configurations or treating signal and field agreement as interchangeable.
\draftfigure{Two configuration schematics showing inlet and wall conditions, fluidic-pinball body order and rotation sign, the case-dependent disturbance or target body, three downstream two-component probes, and the registered evaluation region.}{Planned control configurations and measurements. The parabolic-inflow/no-slip-wall and uniform-inflow/free-slip-wall channels will be shown separately, together with body order $(F,U,L)$, counter-clockwise-positive rotation, and the three circular probes at $x/D=10$, $(y-y_0)/D=0,\pm2$ with radius $0.25D$. The schematic defines geometry, boundaries and observations only; it does not imply configuration equivalence, optimal sensor placement or full observability.}{fig:configurations-sensors}
\section{From steady cloaking to wake retargeting}
\label{sec:progression}
Rotational control is considered through four comparisons of increasing temporal and target complexity: an approximately steady uniform profile, a periodic K\'arm\'an street, an isolated-vortex event and a non-zero wake target. Because the configuration, clock and measure change between cases, their ordering describes a progression of control tasks rather than a common performance ranking. The steady case asks whether a simple downstream profile can be recovered. K\'arm\'an cloaking then makes the desired wake periodic, the Vortex case replaces phase by an event-relative coordinate, and Illusion changes the desired wake itself.
The steady case uses the uniform-inflow/free-slip-wall channel configuration. At $x/D=10$, the controlled endpoint follows the uniform target more closely than the stationary pinball in both velocity components: the streamwise profile approaches the uniform level and the transverse departure is reduced. This comparison concerns one downstream profile at one endpoint, not a periodic mean or full-field identity. It establishes neither optimality nor stability or mechanism. The actuation magnitude, settling behaviour and physical interpretation are treated with the steady-flow evidence later; here the result provides the simplest instance in which rotation moves measured velocity components towards the target.
Periodic K\'arm\'an cloaking provides the most complete comparison. In the parabolic-inflow/no-slip-wall channel configuration, the target, frozen-policy controlled and physical-zero roles define the alternating-street problem. For both the complete-cycle mean and the separate eight-phase comparison, the controlled downstream velocity field is closer to the target than physical zero, with a dimensionless error reduction of approximately 82\% in each case. The velocity errors are normalized by the simulation speed $U_0$ used in these data. Its relation to article-wide $U_{ref}$ remains unresolved, but no conversion is needed for either relative reduction. Reward and visual resemblance do not establish the field result; Section~\ref{sec:karman-results} separates the mean, phase, signal and action comparisons in detail.
The Vortex case removes periodic phase and instead compares the target, controlled and physical-zero flow morphology at one event-relative coordinate. At that instant the three fields describe how the wake is organized relative to the passing vortex. Without a field-error scalar, one snapshot cannot establish event evolution or improvement over physical zero. The comparison is therefore morphological: it extends the progression from a repeating wake to an isolated interaction while leaving quantitative event response to the cross-case analysis.
Illusion changes the desired field rather than only its temporal coordinate. In the uniform-inflow/free-slip-wall configuration, one acquisition contains the $0.75L_0$ non-zero target, the seed-43 PPO-controlled trajectory and the physical-zero trajectory under the same sampling and phase procedure. The policy was executed toward that non-zero target, and the controlled morphology at the common literal phase is consistent with wake retargeting. Without a scalar field comparison, this morphology does not establish approach to the target or improvement over physical zero. It also neither constructs a target action nor revives the negative Illusion symbolic-regression result.
The progression therefore changes one principal feature at a time: the target evolves from a uniform background to a periodic street, the temporal coordinate changes from phase to an isolated event, and the target finally becomes a different non-zero wake. K\'arm\'an cloaking receives detailed treatment because it combines repeated learning histories with target, controlled and physical-zero fields, sparse downstream signals and applied rotations. This broader set of comparisons does not make K\'arm\'an representative of the other tasks. It instead permits two bounded questions: over which tested K\'arm\'an conditions were high-performing policies found, and how closely did the controlled wake approach its target?
\draftfigure{A K\'arm\'an-dominant progression from the steady profile to periodic target/controlled/physical-zero fields, with one event-relative Vortex comparison and one non-zero-target Illusion comparison.}{Planned progression of controlled wake definitions. The K\'arm\'an comparison will dominate the composition; the Vortex panels will show target, controlled and physical-zero morphology at one stated event offset, and the Illusion panels will show execution toward the retained non-zero target. Scales and acquisition coordinates will be stated within each comparison. Only the K\'arm\'an panels support a target-relative field-error result; the changed-task panels are morphology comparisons without a pooled efficacy measure.}{fig:capability-progression}
\section{K\'arm\'an-wake cloaking}
\label{sec:karman-results}
\subsection{Learning across tested K\'arm\'an conditions}
\label{subsec:karman-learning}
The uniform-inflow/free-slip-wall series samples $Re_D=30$, 50, 100 and 200 and disturbance-radius ratios 0.75, 1.0, 1.5 and 2.0. These discrete points delimit the tested conditions; they do not define a continuous Reynolds-number or geometry law. Only the $Re_D=50$ reference condition has five stochastic training realizations, whereas each other condition has one deterministic policy demonstration. The two outer disturbance-radius cases also used a learning rate of $10^{-4}$ instead of $3\times10^{-4}$, so their differences cannot be attributed to geometry alone. Reward and six-sensor dynamic-time-warping (DTW) similarity are reported as separate, higher-is-better quantities: reward measures the learned objective, whereas DTW measures similarity of the selected downstream signals. Neither is treated as a fitted trend or combined score.
The condition series compares one deterministic policy at each tested condition, whereas the learning histories show how five stochastic searches developed at the reference condition. They are not paired before-and-after observations. At $Re_D=50$, all five 500-iteration histories progress from low reward to high-performing checkpoints. Their best rewards are 0.9266, 0.9302, 0.9160, 0.9222 and 0.9412, reached at iterations 471, 315, 360, 465 and 447, respectively. High-performance policy discovery therefore recurred in all five tested realizations rather than depending on a single selected run.
The recurrence applies to these five searches only: it provides no uncertainty for other seeds or hyperparameters and establishes neither asymptotic convergence nor closed-loop stability. The deterministic policy used for the condition comparison is one deployment drawn from this set, not a sixth realization or a continuation of training reward. Moreover, high reward establishes success under the online objective, which combines force and sparse downstream-signal terms; it does not by itself establish proximity of the downstream velocity field.
That distinction also separates the condition survey from the field analysis below. The survey uses the uniform-inflow/free-slip-wall configuration, whereas the detailed reference wake uses a parabolic-inflow/no-slip-wall configuration. Their numerical case labels therefore do not supply a common physical Reynolds number. The field comparison instead asks directly whether the controlled wake approaches its target in the velocity field and downstream signals.
\subsection{Fields and trajectories in the reference wake}
\label{subsec:karman-fields}
In the parabolic-inflow/no-slip-wall reference wake, the target, frozen-PPO controlled and physical-zero roles define the field comparison. The controlled wake reproduces the target's alternating organization more closely than the physical-zero pinball. These roles are separate trajectories in the same physical configuration, not paired realizations of one trajectory. Section~\ref{sec:karman-law} uses these same target and physical-zero fields to compare symbolic-regression control, but its controlled trajectory is different from the PPO-controlled trajectory here. Coincident target and physical-zero values therefore come from shared comparator fields, not independent agreement between the two controllers.
The complete-cycle mean tests the persistent velocity organization after periodic variation has been averaged out. With velocity normalized by the simulation speed $U_0$ used in these data, its error is $E_{mean}^{(U_0)}=0.085771$ for the controlled wake and $0.475206$ for physical zero, giving a zero-relative reduction of 0.819508. Thus the mean controlled field lies substantially closer to the target over the registered $20D\times10D$ downstream region. The result applies to this downstream region, which lies beyond the solid geometry; without a saved solver-exact fluid mask it is not a claim of pointwise identity throughout the domain.
The eight-slot diagnostic retains periodic organization instead of averaging it away. It gives $E_{phase8}^{(U_0)}=0.112680$ for controlled and $0.630878$ for physical zero, a zero-relative reduction of 0.821392. Each slot is the nearest saved snapshot after the target, controlled and physical-zero trajectories have been phased independently; the slots are not repeated-cycle ensembles and no global phase minimization is applied. The mean and phase comparisons give the same target-relative ordering but are not replicate measurements or uncertainty estimates. Both use $U_0$, so their relative reductions are dimensionless even though the relation between $U_0$ and article-wide $U_{ref}$ remains unresolved.
Sparse downstream observations provide a complementary trajectory-level comparison. Target-normalized six-channel DTW similarity is 0.932899 for controlled and 0.625127 for physical zero; unlike spatial error, a larger value denotes closer signal evolution. DTW does not imply full-field equivalence, but its ordering is consistent with the smaller controlled spatial errors. Over samples 96--145, corresponding to $tU_0/D=38.4$--58.0 in the simulation scaling and approximately three target cycles, the centre-sensor trajectory shows the repeating downstream response while the three applied PPO rotations show the simultaneous periodic actuation. The target has no action trajectory, so no target-action surrogate is inferred.
Together, the condition survey and learning histories show repeated policy discovery over the tested training set, while the reference-wake analysis shows target-relative proximity in mean fields, phase-resolved fields and sparse signals. These results imply neither mechanism nor a universal optimal law. Similar numbers in the later symbolic-regression or correction-field analyses do not provide validation because those calculations use different controlled trajectories and measures. The remaining control-side question is whether the periodic action produced by the neural policy can be represented by a compact executable law.
\draftfigure{Tested K\'arm\'an condition points, five reference-condition learning histories, and the parabolic/no-slip target, PPO-controlled and physical-zero field and signal comparisons.}{Planned K\'arm\'an learning and field evidence. Separate panels will report deterministic reward and six-sensor DTW similarity over the finite uniform/free-slip condition series, the five stochastic learning histories at $Re_D=50$, and target-relative mean, phase and sparse-signal comparisons for the distinct parabolic/no-slip reference wake. The configurations and estimands will remain separate; points do not define a continuous parameter law, and reward or DTW does not establish field agreement.}{fig:karman-learning-fields}
\section{A compact control law for K\'arm\'an cloaking}
\label{sec:karman-law}
The compressed action has a simple physical organization: the rear cylinders maintain persistent counter-rotation, rear-lift feedback modulates this background motion, and the front cylinder receives a weaker drag-asymmetry correction. The question is whether this organization can replace the PPO policy with an executable observation-to-action surrogate for the same periodic K\'arm\'an cloak in the parabolic-inflow/no-slip-wall channel configuration.
Direct symbolic controller search evaluates explicit feedback expressions in the control loop, as in genetic-programming studies of separation and vibration control \citep{GautierEtAl2015ClosedLoop,DebienEtAl2016GeneticProgrammingRamp,RenEtAl2019MachineLearningVIV}. Comparisons of compact genetic-programming laws and neural policies also separate readability from executed behaviour \citep{CastellanosEtAl2022FewSensorMLControl}. By contrast, sparse identification of nonlinear dynamics (SINDy) identifies parsimonious dynamical equations from data \citep{BruntonEtAl2016SINDy}, while SINDy-RL learns sparse dynamics and reward models jointly with a policy \citep{ZolmanEtAl2025SINDyRL}. Both address different mathematical objects from a fixed-policy map for the present controlled wake problem.
Here symbolic regression was applied after PPO training to recorded state--action trajectories, making the expression a fixed-policy map rather than a directly searched controller, identified wake equation or jointly learned reinforcement-learning model. PySR supplies the search method and software identity \citep{Cranmer2023PySR}, not evidence of control. Each post-step state at index $i$ was paired with the PPO action at $i+1$. Agreement on PPO-visited states is an offline diagnostic; after deployment, the surrogate chooses its own actions and visits a different trajectory. Closed-loop execution against the target is therefore the relevant test, and the cited methods do not validate the present map or its performance.
Before specifying the coefficients, we impose a reflection organization consistent with the centreline-symmetric geometry and cloaking objective. For a measured state $x$, $Gx$ denotes its centreline-reflected state: upper and lower sensor values are exchanged, streamwise velocity is unchanged, transverse velocity changes sign, front drag is unchanged, front lift changes sign, and the upper and lower force pairs are exchanged with drag even and lift odd. For actions ordered as front, upper rear and lower rear, the corresponding action map is
\begin{equation}
\mathcal{H}\vect{q}_{\mathrm{ctl}}(t)
\simeq \mathcal{H}\vect{q}_{\mathrm{tar}}(t).
\label{eq:illusionobj}
\begin{aligned}
G_\alpha(\alpha_F,\alpha_U,\alpha_L)
&=(-\alpha_F,-\alpha_L,-\alpha_U),\\
\alpha_F(x)&=\frac{h_F(x)-h_F(Gx)}{2},\\
\alpha_U(x)&=h_R(x),\\
\alpha_L(x)&=-h_R(Gx).
\end{aligned}
\end{equation}
The target is non-zero and changes with target diameter. Sensor agreement, force response and phase-aligned fields are evaluated separately so that no single reward component is equated with complete invisibility.
The front action is odd, while the rear actions share one mirrored generator. This construction replaces three separately parameterized heads with an odd front head and one rear generator. The reflection map was imposed on the surrogate after the PPO trajectories had been collected; it is not evidence that the PPO policy or controlled plant is reflection-equivariant.
\section{Numerical and data-driven methods}
\label{sec:methods}
\subsection{Lattice-Boltzmann environment}
The flow is advanced with a GPU-accelerated multiple-relaxation-time lattice-Boltzmann method on a D2Q9 lattice. The populations obey
The law uses signed native solver forces in the streamwise and transverse coordinate directions, so positive lift is the positive-$y$ force. For cylinder $i$, the fitted inputs are
\begin{equation*}
\begin{aligned}
C_{d,i}&=\frac{2f_{x,i}}{\rho U_0^2D}, &
C_{l,i}&=\frac{2f_{y,i}}{\rho U_0^2D},
\end{aligned}
\end{equation*}
with $\rho=1$ in the retained lattice data. Thus
\begin{equation*}
\begin{aligned}
C_{d,\mathrm{rear},a}&=\frac{C_{d,U}-C_{d,L}}{2}, &
C_{l,\mathrm{rear},s}&=\frac{C_{l,U}+C_{l,L}}{2}.
\end{aligned}
\end{equation*}
Within this imposed action space, the executable map is
\begin{equation}
f_i(\vect{x}+\vect{c}_i\Delta t,t+\Delta t)-f_i(\vect{x},t)
=-\left[\vect{M}^{-1}\vect{S}
(\vect{m}-\vect{m}^{\mathrm{eq}})\right]_i.
\begin{aligned}
\alpha_F&=\operatorname{odd}
(-0.381391\,C_{d,\mathrm{rear},a}),\\
\alpha_U&=1.307782\,C_{l,\mathrm{rear},s}-3.431209,\\
\alpha_L&=-\alpha_U(Gx).
\end{aligned}
\end{equation}
Curved ghost-node interpolation imposes the rotating no-slip cylinder boundaries. The low lattice velocity $U_0=0.01$ limits compressibility effects. Cylinder rotation, force integration and sensor sampling are coupled directly to the Python control environment, allowing each policy decision to be followed by a prescribed number of CFD steps without restarting the solver.
where $\operatorname{odd}$ denotes $f(x)\mapsto[f(x)-f(Gx)]/2$. The constant $-3.431209$ sets the persistent upper-rear rotation and, through reflection, the opposite lower-rear rotation. Rear-mean lift adjusts this pair, while rear drag asymmetry supplies the front correction. These terms describe the surrogate action, not a governing law for the fluid.
\subsection{Observation, action and reward}
The historical output is $\alpha=\omega/U_0$, and deployment applies $\omega=\alpha U_0$; hence $\alpha$ has inverse-length units in dimensional notation. The corresponding surface-speed command is $s=\omega R/U_{ref}=\alpha R(U_0/U_{ref})$. It cannot be reduced to $s=\alpha R$ because the historical parabolic source does not establish $U_0=U_{ref}$.
For cloaking, the observation contains six force components and six downstream velocity components,
Closed-loop execution tests whether the compact action remains useful away from PPO-visited states. This deployment test is also the selection boundary: expression simplicity and offline action fit can nominate a readable map, but only the executed surrogate determines the trajectory on which its control performance is measured. No external controller-search or sparse-identification result supplies that evidence for this case.
In the reference run, 480 control intervals were used to establish the trajectory, followed by 160 recorded post-step boundaries spanning eight complete cycles. The mean native dynamic-time-warping (DTW) similarity was $0.942357$, where higher is better. This rolling comparison uses a 30-boundary window: one circular lag is applied before the six channel scores are averaged. The lag aligns signals for comparison and is not a physical delay.
The same run also supplies two spatial comparisons with the same-case target and physical-zero trajectories from the shared acquisition lineage. Both errors are two-component velocity RMS differences over the geometry-guarded downstream region, normalized by the historical inlet scale $U_0$: $E_{\mathrm{mean}}$ compares complete-cycle arithmetic mean fields, whereas $E_{\mathrm{phase8}}$ is the RMS over eight independently phase-referenced nearest-boundary snapshots. The values were $E_{\mathrm{mean}}=0.125404$ for SR and $0.475206$ for physical zero, and $E_{\mathrm{phase8}}=0.156040$ for SR and $0.630878$ for physical zero. The surrogate-controlled wake is therefore closer to the target under both measures. Each phase entry is the nearest saved snapshot on a phase coordinate determined separately for that flow, and the region has no stored solver-exact mask. These SR-run values answer a different question from the independently recorded PPO-controlled field errors in Section~\ref{sec:karman-results} and are not repeat measurements of them.
Term deletion asks which parts of the executed map matter most among the tested variants. In each of four source cases, the full map and a deletion variant were run for the same 40 control steps. Six-channel native DTW similarity was evaluated after one shared circular-lag alignment, and $\Delta S=S_{deleted}-S_{parent}$. Deleting the rear constant gave $-0.106537$, $-0.093869$, $-0.178525$ and $-0.191082$; deleting rear-lift feedback gave $-0.028755$, $-0.041660$, $-0.061363$ and $-0.048948$; deleting the front correction gave $-0.006520$, $+0.000579$, $-0.005179$ and $-0.025194$.
Removing the persistent rear rotation produces the largest degradation in every tested case. Removing rear-lift feedback has a smaller but non-zero effect, while removing the front correction is weak or mixed over this window. A deletion changes the actions and hence the states subsequently visited, so this ranking does not identify causal or necessary terms in the flow-control process. It establishes a tested action hierarchy and motivates a separate field question: how do the persistent reference action and the remaining time-dependent organization appear in the mean and centred velocity fields?
\draftfigure{Closed-loop symbolic-controller action traces and target/SR/physical-zero wake fields, accompanied by a fixed-window parent-relative term-deletion plot.}{Planned compact K\'arm\'an law and deployment evidence in the parabolic-inflow/no-slip-wall configuration. The action traces will show persistent opposite rear rotation and smaller modulation; the field panels will compare the exact target, SR-controlled and physical-zero roles; and deletion will report $\Delta S=S_{deleted}-S_{parent}$ over the stated 40-step window. The deletion ranks tested action terms but does not establish causality, necessity, stability or transfer.}{fig:karman-sr-law}
\section{Mean-field correction and action-correlated organization}
\label{sec:field-ccd}
The field analysis begins with four flows: the target $T$, physical zero $0$, constant control $C$ fixed at the mean action of the PPO policy, and PPO-controlled flow $D$. For the periodic K\'arm\'an cloak in the parabolic-inflow/no-slip-wall configuration, each mean averages the same 360 post-step boundaries, indices $[480,840)$. The comparison covers $34\leq x/D\leq54$, $|y/D|\leq5$, on the 80,200 fluid points shared by all four simulations. With coordinate weights $w_k$ and both velocity components, the target error of flow $r$ is
\begin{equation}
\vect{o}_t=
(F_{F,x},F_{F,y},F_{T,x},F_{T,y},F_{B,x},F_{B,y},
u_1,v_1,u_2,v_2,u_3,v_3)^{\mathsf T},
\label{eq:obs}
E_r=\left[
\frac{\displaystyle\sum_{k\in K}w_k
\|\overline{\boldsymbol{u}}_r(k)
-\overline{\boldsymbol{u}}_T(k)\|_2^2}
{\displaystyle\sum_{k\in K}w_k}
\right]^{1/2}.
\end{equation}
with scene-specific physical scaling. Illusion augments this state with the target drag and lift reconstructed at the current target phase. The PPO output $\vect{a}_t\in[-1,1]^3$ is converted to physical rotation using the action scale and bias of the corresponding Legacy scene. SR is fitted only after undoing this decoder.
Here $\boldsymbol{u}_r(k)=(u_{x,r}(k),u_{y,r}(k))$ is the two-component velocity at point $k$, and $\overline{\boldsymbol{u}}_T$ is the target mean, not an instantaneous field. This four-flow comparison uses different data, averaging windows, masks and weighting from the PPO-controlled comparison in Section~\ref{sec:karman-results} and the SR comparison in Section~\ref{sec:karman-law}; the resulting numbers are not repeat measurements.
The reward combines force and downstream-signal objectives. For cloak, Gaussian force scores encourage small additional mean drag and lift,
The streamwise means give the first physical picture. Across much of the wake region, $\overline u_{x,0}-\overline u_{x,T}$ is negative: the physical-zero pinball has a mean streamwise deficit relative to the target. The difference $\overline u_{x,D}-\overline u_{x,0}$ is largely opposite in sign and shows where PPO control adds or removes streamwise velocity relative to physical zero. This pattern is compatible with adding downstream velocity where the physical-zero flow produces a deficit, but it does not establish momentum restoration. These fields show only the streamwise component; the errors below use both velocity components.
The exact errors are $E_0=0.4694352273$, $E_C=0.1110140739$ and $E_D=0.0857787860$. Their ordered differences give
\begin{equation}
r_D=\exp(-K_D C_D^2),\qquad
r_L=\exp(-K_L C_L^2),
\frac{E_0-E_C}{E_0-E_D}=93.42\%,\qquad
\frac{E_C-E_D}{E_0-E_D}=6.58\%.
\end{equation}
while a signal score compares the six probe channels with the reference. For illusion, $C_D$ and $C_L$ are replaced by the differences between the pinball assembly and the target-cylinder forces.
Thus constant control accounts for $93.42\%$ of the measured reduction from the physical-zero error to the PPO-controlled error, and the change from constant control to PPO control accounts for the remaining $6.58\%$. These percentages divide changes in one scalar distance from the target mean. They are neither energy fractions nor measurements of the norm of the residual defined next, and they do not assign shares to terms in the SR law.
Temporal similarity is based on dynamic time warping (DTW). For vector sequences $X=\{\vect{x}_i\}$ and $Y=\{\vect{y}_j\}$, the accumulated cost is
The time-dependent comparison requires a velocity field rather than a difference between scalar errors. Mean-field descriptions of natural and actuated cylinder wakes use shift modes to represent changing means and their coupling with fluctuations \citep{TadmorEtAl2010CylinderMeanField}. That context motivates keeping the mean correction and fluctuation organization as distinct objects here; it does not identify the present constant-control subtraction with a shift-mode model. The PPO-controlled run contains 19 cycles resolved into ten phase bins. A separate 19-cycle constant-control run supplies the two-component phase field $\widehat{\boldsymbol{u}}^C_b(k)$. After subtracting each flow's own mean, the centred velocity residual is
\begin{equation}
D(i,j)=\lVert\vect{x}_i-\vect{y}_j\rVert_2+
\min\{D(i-1,j-1),D(i-1,j),D(i,j-1)\}.
\begin{aligned}
\boldsymbol{u}^{\mathrm{res}}_{c,b}(k)
={}&[\boldsymbol{u}^D_{c,b}(k)-\overline{\boldsymbol{u}}_D(k)]\\
&-[\widehat{\boldsymbol{u}}^C_b(k)-\overline{\boldsymbol{u}}_C(k)],\\
&c=0,\ldots,18,\qquad b=0,\ldots,9.
\end{aligned}
\end{equation}
The distance is normalised by sequence length and a scene-specific fluctuation scale, and then mapped to a bounded similarity. DTW accommodates modest phase and convection-time offsets, but the full-field comparisons are additionally phase aligned.
Every PPO-controlled cycle is compared with the constant-control ensemble template at phase bin $b$; individual cycles from the two runs are not paired. No cross-run lag, phase wrapping or interpolation is introduced. The residual is therefore a separately centred comparison, not a matched counterfactual response, and its magnitude is not the 6.58\% scalar share.
\subsection{PPO policy discovery}
Canonical correlation decomposition (CCD) asks whether this residual varies with the executed actions. The actions are the same-boundary, counter-clockwise-positive cylinder speeds in native solver units. From front, upper and lower actions $(a_F,a_U,a_L)$, the analysis forms $a_F$, the rear-symmetric coordinate $(a_U+a_L)/\sqrt{2}$ and the rear-antisymmetric coordinate $(a_U-a_L)/\sqrt{2}$. It removes the measured mean action and then centres the samples, without rescaling or whitening them. A rank-3 decomposition relates these three action coordinates to the centred velocity residual. Its field vectors are termed \emph{action-correlated correction-field modes}. Their singular strengths measure cross-correlation, not field energy, and each mode may be multiplied by $-1$ without changing the decomposition. The modes describe action association rather than a causal response of the flow. This placement has a broader correlation-oriented ancestry: extended POD relates fields to correlated events, observable-ranked decompositions order velocity structures by their linear observability in another signal, and point--field correlations have been used for descriptive field reconstruction \citep{Boree2003ExtendedPOD,JordanEtAl2007ObservableJetModes,DiscettiEtAl2018CorrelatedFieldEstimation}. Sparse sensor-to-field modelling provides a related observable-associated context \citep{LoiseauEtAl2018SparseROM}. These precedents motivate the distinction from energy ranking; they do not validate CCD, import another decomposition, or establish causality.
The actor and critic contain two hidden layers of width 64 with sinusoidal activation,
Proper orthogonal decomposition (POD) provides the field-only comparison \citep{SchlegelEtAl2012LeastOrder}. It orders velocity-field variance under the same weights without using the action coordinates. Here POD and CCD use exactly the same centred residual snapshots, rank, weights and spatial region. Their rank-3 field subspaces are essentially the same. CCD therefore does not provide a superior or uniquely physical velocity basis; it adds coordinates that associate the shared low-rank field organization with front, rear-symmetric and rear-antisymmetric motion.
Together, the mean fields show that constant rear-cylinder rotation accounts for most of the measured scalar error reduction. A separately centred velocity residual has descriptive action-correlated organization. This relationship does not identify SR terms with CCD coordinates or independently confirm the action analysis. It applies only to this K\'arm\'an cloak, these four flows and the stated spatial and temporal comparison.
\draftfigure{Four role means and signed streamwise differences with the scalar error ledger, followed by the separately centred residual construction and representative action-correlated phase organization.}{Planned mean-field and correction-field analysis for the periodic K\'arm\'an cloak. Target, physical-zero, constant and PPO-controlled means will precede the two-component scalar error accounting, for which constant control accounts for 93.42\% of the measured zero-to-PPO reduction. Separate panels will define the non-paired centred residual and its action-correlated correction-field modes. The scalar shares are not residual norms or energy fractions, and the modal association is descriptive and noncausal.}{fig:mean-field-ccd}
\section{Extensions under changed temporal and target contracts}
\label{sec:extensions}
The periodic K\'arm\'an case provides a reference for changing the control problem, not a numerical or interpretive template for the cases considered here. The incoming vortex replaces periodic phase by event-relative time; Illusion replaces background matching by a declared non-zero target; and Erase compares with a clean-flow target while an upstream disturbance remains in the PPO-controlled and physical-zero flows. The policy maps its current force-plus-sparse-observation state to cylinder actions in each case, but the clocks, target fields and physical comparisons differ. Field morphology can therefore be interpreted only within each case.
In the central Taylor-vortex scene, the target, PPO-controlled and physical-zero flows are sampled at offsets relative to the detected incoming event, together with the action history. The offset indexes the event, not a periodic phase or a calibrated physical time. At offset $+10$, the three vorticity fields show distinct event-relative morphologies. Their common event coordinate places the incoming structure, its passage around the pinball and the downstream wake in one spatial comparison, so differences can be assigned to role without treating the offset as a phase. No scalar comparison is reported; only morphology is considered. Because the policy acts from the current observation state, the event motivates a hypothesis of reactive control from local information for a comparable-scale disturbance. This single scene demonstrates neither control across vortex families, amplitudes or offsets nor observability or physical-flow memorylessness.
Illusion changes the desired wake signature rather than the clock. In the uniform-inflow/free-slip-wall configuration, the retained $0.75$ target-size, seed-43 policy was executed toward a declared non-zero target. Unlike background matching, that target contains a wake pattern, and the PPO-controlled morphology is consistent with retargeting toward it. No scalar change in target-relative error is reported. The historical negative Illusion symbolic-regression result remains specific to that method: this PPO execution supplies neither a positive Illusion symbolic law nor a transfer of the K\'arm\'an surrogate.
Erase changes both the physical scene and the temporal comparison. In the parabolic-inflow/no-slip-wall configuration, its clean target is one late field, whereas the PPO-controlled and physical-zero flows contain the upstream disturbance and contribute eight phase-referenced fields. The clean, PPO-controlled and physical-zero morphologies are visible, but one late target and an eight-snapshot set are not the same temporal object. The historical rolling reward compares an evolving same-rollout history rather than the clean target field. No scalar field evaluation is reported; averaging the retained snapshots would first remove phase-dependent variation, whereas aggregating snapshot-wise errors would retain that variation in the error. These operations define different estimands, neither of which is used here.
Together, the fields show PPO execution under changed temporal or target definitions. For Illusion, the bounded result is execution toward a declared non-zero target with morphology consistent with retargeting; Vortex and Erase remain morphology-only comparisons, and reactive-control or target-erasure interpretations remain hypotheses. The distinct configurations and temporal comparisons admit neither a common efficacy measure nor transfer of the K\'arm\'an SR/CCD interpretation. A separately obtained steady endpoint can therefore provide only non-equivalent physical context.
\draftfigure{Target, PPO-controlled and physical-zero fields for one stated Vortex event offset and the retained non-zero-target Illusion acquisition, with acquisition-specific coordinates and scales.}{Planned changed-task morphology. Vortex fields will be compared at a stated event-relative offset, never a periodic phase, and Illusion fields will show policy execution toward the declared non-zero target. These panels describe role identity and morphology only: no scalar efficacy, K\'arm\'an-law transfer, common metric or causal response is inferred.}{fig:changed-task-morphology}
\section{Steady control as physical context and a theory boundary}
\label{sec:steady-theory}
Steady rear-cylinder rotation was the historical discovery seed, and the periodic K\'arm\'an analysis later identified persistent rear counter-rotation as the dominant tested symbolic-regression term, with a smaller dynamic correction. Those findings belong to the parabolic-inflow/no-slip-wall configuration. Here a separately acquired uniform-inflow/free-slip-wall endpoint at $Re_D=50$ supplies a distinct physical comparison. It is not a continuation, amplitude match, replication or validation of the periodic control chain, and no cross-configuration arithmetic is meaningful.
With body order $(\text{front},\text{rear}_{y+},\text{rear}_{y-})$, the accepted action is $[0,+\Omega,-\Omega]$ and the surface-speed ratio is $s=\Omega R/U_\infty$. At $x/D=10$, define the two-component profile error by
\begin{equation*}
\begin{aligned}
E_{\infty,\mathrm{vector}}
&(\boldsymbol u_{\mathrm{ctl}},\boldsymbol u_{\mathrm{in}})\\
&=\frac{1}{U_\infty}\max_{y\,\in\,\mathcal C}
\left|\boldsymbol u_{\mathrm{ctl}}(y)
-\boldsymbol u_{\mathrm{in}}(y)\right|,
\end{aligned}
\end{equation*}
where $\mathcal C$ is the common-fluid cross-section and $|\cdot|$ is the Euclidean norm of the streamwise and transverse velocity differences. Thus this $L^\infty$ profile metric uses a cross-sectional maximum, not quadrature weighting. The settling-qualified endpoint at $s=3.55$ has $E_{\infty,\mathrm{vector}}=0.02317$. Stationary flow is an unsteady reference, while the reverse endpoint at $s=5.2$ differs in both sign and amplitude. The accepted velocity field is therefore a distinct low-deficit endpoint, not an optimum, a stable branch or a universal law.
The wall law $(U_w,V_w)=(-\omega r_y,\omega r_x)$ fixes the tangential surface velocity: for the accepted rear-cylinder signs, both gap-facing cardinal points move downstream and both outer cardinal points move upstream. Relative to the stationary and reverse flows, the accepted vorticity and streamline fields show a different inter-cylinder and near-wake organization. These endpoint observations motivate a viscous hypothesis: imposed wall motion may alter wall shear and vorticity production, shear-layer development, inter-cylinder transport, separation and recirculation, and hence downstream mean transport. The sequence is not a demonstrated mechanism because no amplitude-matched intervention, complete same-control-volume momentum or mechanical-energy budget, force attribution, or stability analysis isolates its links.
Rotating-cylinder studies make this separation necessary. Wake modification by auxiliary rotating cylinders varies with Reynolds number, gap, rotation rate and dimensionality \citep{Mittal2001RotatingControlCylinders,ChanEtAl2011CounterRotatingCylinders}; even a strongly suppressed viscous wake that resembles a potential doublet does not thereby become an inviscid solution \citep{ChanEtAl2011CounterRotatingCylinders}. Viscous rotating-cylinder flows can also occupy multiple steady or cyclic branches, including bistable regimes \citep{SierraEtAl2020RotatingCylinderBifurcations}, while low-Reynolds-number two-cylinder analysis requires matched inner Stokes and outer Oseen descriptions \citep{Watson1996TwoRotatingCylinders}. These results motivate an outer reference as a deliberately restricted test of streamline organization, not as a parameter transfer, branch identification or mechanism for the present endpoint.
The rearranged outer streamlines raise a specific closure question: can the finite-Re organization be explained by the circulation degrees of freedom that remain after an inviscid flow satisfies impermeability? To test that proposition without assigning viscous physics to the outer model, consider the irrotational family \citep{Crowdy2006MultipleCylinders}
\begin{equation}
\vect{h}_1=\sin(\vect{W}_1\vect{o}+\vect{b}_1),\qquad
\vect{h}_2=\sin(\vect{W}_2\vect{h}_1+\vect{b}_2).
\begin{aligned}
\boldsymbol{u}(\boldsymbol{x})
&=\boldsymbol{u}_{N}(\boldsymbol{x})
+\sum_{j=1}^{3}\Gamma_j\boldsymbol{h}_j(\boldsymbol{x}),\\
\nabla\!\cdot\boldsymbol{h}_j
&=\nabla\!\times\boldsymbol{h}_j=0,
\qquad \boldsymbol{h}_j\!\cdot\boldsymbol{n}=0,\\
\oint_{B_k}\boldsymbol{h}_j\!\cdot\mathrm{d}\boldsymbol{\ell}
&=\delta_{jk}.
\end{aligned}
\end{equation}
The policy is trained with the clipped PPO objective
Here $\boldsymbol{u}_N$ is the zero-period Neumann solution and the counter-clockwise circulations $\Gamma_j$ are prescribed periods spanning the harmonic freedom left by impermeability; cylinder spin does not select them. This freedom reorganizes the outer streamlines, but it cannot impose rotating no slip; viscous rotating-cylinder theory retains that tangential boundary condition under its own low-Reynolds-number assumptions \citep{UedaEtAl2003RotatingCylinders}. For one cylinder,
\begin{equation}
\mathcal{L}^{\mathrm{clip}}=
\mathbb{E}_t\left[
\min\left(\rho_t\widehat A_t,
\operatorname{clip}(\rho_t,1-\epsilon,1+\epsilon)\widehat A_t\right)
\right].
\begin{aligned}
u_\theta(R,\theta)
&=-2U_\infty\sin\theta+\frac{\Gamma}{2\pi R}\\
&\neq R\Omega
\quad\hbox{for all }\theta\hbox{ when }U_\infty\neq0.
\end{aligned}
\end{equation}
The neural policy is used as a discovery mechanism. Its deterministic trajectories supply the force, sensor and physical-action records used for SR; the policy weights are not inspected to infer the mechanism.
For the normalized rear-opposite basis, define the dimensionless coefficient $g$ by
\begin{equation*}
\frac{(\Gamma_1,\Gamma_2,\Gamma_3)}{U_\infty D}
=\frac{g(0,1,-1)}{\sqrt{2}},
\end{equation*}
with positive circulation geometrically counter-clockwise. The outer strip objective selects $g_{\mathrm{opt}}=6.148$, whereas the descriptive projection of the accepted finite-Re endpoint gives $g_{\mathrm{fit}}=-12.20$. The no-slip incompatibility and opposite sign reject wall-spin-to-circulation closure: the outer family organizes geometry but does not explain the viscous endpoint.
\subsection{Symbolic-regression target and feature libraries}
\label{sec:srmethod}
The periodic K\'arm\'an fields and actions contain a persistent rear-rotation reference plus a smaller correction under their own configuration. The separate steady velocity field adds a low-deficit endpoint and motivates the viscous sequence above, but does not establish its links. Distinguishing those links requires amplitude-matched interventions, held-out endpoint and convergence tests, wall-vorticity flux and separation diagnostics, and a closed same-control-volume momentum/mechanical-energy budget. Until those measurements are available, the one-circle no-slip incompatibility and the opposite signs of $g_{\mathrm{opt}}$ and $g_{\mathrm{fit}}$ leave the inviscid circulation closure rejected.
The SR target is the dimensionless applied actuation
\begin{equation}
\vect{\alpha}_t=\frac{\vect{\omega}_t}{U_0}
= (\alpha_F,\alpha_T,\alpha_B)^{\mathsf T}.
\label{eq:srtarget}
\end{equation}
Converting the stored PPO action to $\vect{\alpha}$ before regression is essential: otherwise the fitted coefficients combine policy physics with scene-dependent decoder scale and bias.
\draftfigure{Stationary, accepted and reverse steady endpoint fields and profiles, together with prescribed-circulation outer-reference streamlines and the opposite-sign closure comparison.}{Planned steady endpoint and theory boundary in the uniform-inflow/free-slip-wall configuration. The finite-$Re_D$ panels will distinguish the settling-qualified accepted endpoint from the unsteady stationary reference and the non-amplitude-matched reverse endpoint. Outer-reference panels will compare $g=0$, $g_{\mathrm{opt}}$ and the descriptive fitted $g$; no-slip incompatibility and the opposite signs of $g_{\mathrm{opt}}$ and $g_{\mathrm{fit}}$ reject the proposed wall-spin-to-circulation closure. The composition provides context, not a viscous mechanism or cross-configuration validation.}{fig:steady-theory-boundary}
Candidate variables include total drag and lift,
\begin{equation}
C_{d,\mathrm{tot}}=\sum_i C_{d,i},\qquad
C_{l,\mathrm{tot}}=\sum_i C_{l,i},\qquad
C_{d,\mathrm{rear}}=C_{d,T}+C_{d,B},
\end{equation}
selected velocity probes, target-relative force errors for illusion, and finite-time rates. For a generic signal $g$,
\begin{equation}
\dot g(t)\simeq\frac{g(t)-g(t-\Delta t_c)}{\Delta t_c},
\label{eq:rate}
\end{equation}
where $\Delta t_c$ is the physical control interval rather than one raw LBM step. The joint K\'arm\'an library also includes viscosity and selected action-rate variables to represent changes in phase dynamics across $\Rey_D$.
\section{Qualitative open-loop experiment}
\label{sec:experiment}
% AUTHOR-OWNED PLACEHOLDER: replace only after provenance review; this is not a reported result.
\textit{[Author-supplied qualitative open-loop experiment text will be inserted here after provenance review.]}
The search uses PySR, whose evolutionary search alternates mutation, crossover, simplification and numerical coefficient optimisation over populations of expression trees \citep{Cranmer2023}. The binary operator set is restricted to $\{+,-,\times,\div\}$ and the unary set to a square operation. Protected numerical evaluation rejects non-finite candidates. This deliberately conservative grammar avoids introducing trigonometric or transcendental functions merely because they can interpolate a finite trajectory. Expression complexity is the weighted node count; constants and variables have unit cost, while compound operations incur their tree cost. The Pareto front is ranked by loss improvement per added unit of complexity, but its nominal winner is not accepted automatically.
\draftfigure{Author-supplied apparatus schematic and, only after provenance review, selected experimental and separately generated CFD single-frame views for named open-loop conditions.}{Planned qualitative open-loop experiment figure. The apparatus and command/acquisition path will be identified from author-supplied records. Any retained experimental and CFD images will be labelled as separate, unregistered, acquisition-specific single-frame views for qualitative morphology only. No temporal change, velocity, force, field matching, quantitative agreement or CFD validation will be inferred.}{fig:open-loop-experiment}
Four nested feature classes are considered: instantaneous force and sensor variables; phase-state variables with temporal derivatives; a cross-$\Rey_D$ physics-rate library including viscosity and action rates; and an illusion library including target errors. The purpose is not to maximise library size but to determine the smallest closed-loop state that preserves the policy behaviour. Table~\ref{tab:features} records the physical meaning and reflection parity of the features used in the reported searches.
\section{Discussion and conclusions}
\label{sec:discussion-conclusions}
\begin{table}
\centering
\caption{Symbolic-regression features. ``Even'' and ``odd'' denote parity under centreline reflection; paired quantities are evaluated after exchanging upper and lower channels. Rates use the physical control interval $\Delta t_c$.}
\label{tab:features}
\begin{tabular}{p{0.19\textwidth}p{0.43\textwidth}p{0.12\textwidth}p{0.16\textwidth}}
\toprule
symbol & definition or role & parity & library\\
\midrule
$u_m,u_a,u_c$ & streamwise velocities at the three downstream probes & paired/even & all\\
$v_a$ & selected transverse probe velocity & odd & cloak\\
$C_{d,\mathrm{tot}}$ & total pinball drag coefficient & even & all\\
$C_{d,\mathrm{rear}}$ & sum of rear-cylinder drag coefficients & even & all\\
$C_{l,\mathrm{tot}}$ & total pinball lift coefficient & odd & all\\
$C_{l,\mathrm{diff}}$ & symmetry-adapted rear lift difference & parity-adapted & cloak\\
$\dot a_F,\dot a_T,\dot a_B$ & decoded physical-action rates & transformed & cross-$\Rey_D$\\
$\mu g$ & viscosity-weighted feature $g$ & parity of $g$ & cross-$\Rey_D$\\
$C_{d,\mathrm{err}},C_{l,\mathrm{err}}$ & pinball-minus-target force errors & even/odd & illusion\\
$\dot u_a,\dot C_l,\dot C_{d,\mathrm{err}},\dot C_{l,\mathrm{err}}$ & finite-time phase and error rates & inherited & phase/illusion\\
\bottomrule
\end{tabular}
\end{table}
\subsection{Discussion}
\label{subsec:discussion}
\subsection{Trajectory collection and decoder audit}
Three independently rotating cylinders can alter wakes defined by different temporal and target conditions. In the uniform-inflow/free-slip-wall configuration, repeated K\'arm\'an training runs produced high-performing policies across the tested realizations. In the separately documented parabolic-inflow/no-slip-wall configuration, the controlled K\'arm\'an velocity field is closer to its declared target than physical zero in both complete-cycle mean and phase-resolved comparisons. The first result concerns repeatable policy discovery under the online objective; the second concerns the downstream field under a different configuration. Together they locate the periodic K\'arm\'an cloak as the case in which learned rotation is accompanied by a direct target-relative field comparison, without treating the two configurations as one continuous data set.
The fitting records are deterministic PPO rollouts collected after training. At every control decision the archive stores the six force components, six probe velocities and three normalised policy outputs. Illusion records additionally contain the target force phase used by the environment. The four K\'arm\'an datasets are stored under the code-level labels $\Rey_{\mathrm{code}}=50$, 100, 200 and 400 and therefore cover $\Rey_D=25$, 50, 100 and 200; the joint illusion fit uses the $0.75D$ and $1.0D$ targets. Lamb and Taylor records are held out from the K\'arm\'an fitting and are used only for transfer evaluation.
Persistent counter-rotation of the rear cylinders is the clearest recurring feature of the periodic control. Removing that term degrades the compact surrogate more than removing the tested modulations. In the mean velocity field, the physical-zero pinball leaves a mean streamwise deficit relative to the target, while the controlled-minus-physical-zero field is largely opposite in sign. Holding the cylinders at the mean DRL action produces a constant-control field that accounts for most of the measured scalar mean-field error reduction; the remaining constant-to-DRL scalar increment is distinct from the centred velocity residual. That residual subtracts separately centred DRL and constant-control phase fields without cycle pairing, and its action-correlated modes describe phase organization rather than its size or cause. The surrogate ranking, scalar error increments and residual modes are therefore compatible but non-equivalent: they support persistent rear-cylinder rotation as an organizing component, not a unique decomposition or mechanism.
Undoing the action decoder is a necessary audit step rather than a change of units made for presentation. In the Legacy environment the normalised action $a_i\in[-1,1]$ is mapped according to
\begin{equation}
\alpha_i=s a_i+b_i,
\label{eq:decoder}
\end{equation}
where $s=8$ with biases $(0,-4,4)$ for K\'arm\'an, $s=8$ with $(0,-2,2)$ for illusion, and $s=4$ with $(0,-4,4)$ for the transient-vortex scenes. The precise sign and ordering follow the Legacy cylinder convention. Regression on $a_i$ would make a constant rear rotation appear smaller or disappear into the decoder bias; equation~\eqref{eq:decoder} is therefore inverted before constructing the targets and re-applied consistently during deployment.
Changing the task changes what can be inferred. For the incoming-vortex case, the policy was executed on an event-relative clock and the retained fields show target, controlled and physical-zero morphology at the selected event offset, without a scalar efficacy result. For Illusion, the policy was executed toward a declared non-zero wake target, and the retained role fields are consistent with that bounded morphology comparison; no scalar target-relative improvement is claimed. The separate steady case adds physical context: prescribed rear-cylinder counter-rotation accompanies a low-deficit mean-flow endpoint, while the wall kinematics point toward viscous shear-layer and transport effects. The prescribed-circulation outer flow cannot select circulation from cylinder spin, and its opposite sign rejects the proposed inviscid closure. Persistent rotation is therefore a useful kinematic coordinate across the stated cases, but not a transferred control law or causal explanation.
The control interval is also part of the identified model. K\'arm\'an and transient-vortex scenes normally use 800 LBM steps per decision, while the $0.75D$, $1.0D$ and $1.5D$ illusion scenes use 400, 600 and 800 steps, respectively. A derivative in the symbolic library is a finite difference over this decision interval, not over one lattice update. Thus changing the sampling interval changes both the information available to feedback and the numerical value of a rate feature. The dedicated $\Rey_D=200$ ($\Rey_{\mathrm{code}}=400$) sensitivity test halves the interval to 400 steps and doubles the number of validation decisions so that the physical horizon remains comparable.
\subsection{Conclusion}
\label{subsec:conclusion}
\subsection{PySR fitting and model-selection protocol}
Rotational control was executed for periodic, event-relative and non-zero-target wake definitions. The periodic K\'arm\'an cloak provides the strongest field-level result: in the parabolic-inflow/no-slip-wall configuration, its controlled registered velocity field is closer to the declared target than physical zero. The Vortex and Illusion cases extend the tested task definitions through event-relative execution and bounded non-zero-target morphology, without supplying a common efficacy measure.
For each channel, samples from the relevant scenes are concatenated after scene-consistent feature construction. Front and shared-rear heads are searched separately. Candidate expressions are first screened for finite values and action bounds, then evaluated on withheld trajectory segments. Search depth and population evolution produce a family of loss--complexity compromises rather than one privileged expression. We retain candidates that improve materially over a constant baseline without adding structurally redundant terms.
Model selection then proceeds through four filters. The first is offline action prediction, reported with $R^2$ only as a diagnostic. The second is a symbolic audit of dimensions, decoder convention and $\mathbb{Z}_2$ equivariance. The third is a deployment audit: every term must remain active when the expression generates its own actions. The fourth and decisive filter is closed-loop CFD similarity over a horizon long enough for actuation information to convect from the pinball to the observer. Force histories and boundedness are inspected alongside similarity. This hierarchy is designed to reject formulas that exploit correlations peculiar to the teacher trajectory.
The pipeline is reproducible in four stages: deterministic PPO inference, PySR fitting, CFD validation and figure generation. It should not be confused with an end-to-end symbolic reinforcement-learning algorithm. PPO and PySR are trained sequentially, and validation is an independent CFD rollout. This separation permits the neural and symbolic controllers to be compared from the same flow state with the same normalisation, actuator smoothing and observation schedule.
\subsection{Reflection symmetry}
Reflection about the centreline exchanges $T$ and $B$ and changes the sign of transverse quantities and rotation. The action transforms as
\begin{equation}
\mathcal{G}_a
\begin{bmatrix}\alpha_F\\\alpha_T\\\alpha_B\end{bmatrix}
=
\begin{bmatrix}-\alpha_F\\-\alpha_B\\-\alpha_T\end{bmatrix}.
\label{eq:gsym}
\end{equation}
One front relation and one shared rear relation are fitted; the lower-cylinder law is generated by evaluating the shared relation on the reflected state. This removes degeneracy caused by the strong correlation of the rear actions and prevents symmetry-related actuators from being interpreted as independent mechanisms.
\subsection{Closed-loop model selection}
\label{sec:clselection}
Each candidate expression is subjected to three tests. First, its one-step action fit is measured on held-out PPO data. Secondly, its dimensions, decoder convention, deployment state and reflection symmetry are audited. Thirdly, the PPO decoder is replaced by the symbolic expression and the controller is deployed in CFD from the same initial condition and with the same sampling and actuation filtering.
The third test is decisive because a policy surrogate changes the state distribution on which all later actions are evaluated. In the present trajectories, a full-lag expression reached a one-step $R^2\simeq0.94$ but only $0.62$ closed-loop similarity. A lower-dimensional expression with $R^2\simeq0.32$ reached approximately $0.75$ similarity. Regression accuracy alone would therefore have selected the inferior controller. Terms that are correlated on the PPO trajectory but inactive after deployment are also removed. In particular, an early rear-action-rate term vanished because the deployed rear action was approximately constant.
\section{Symbolic feedback laws and closed-loop performance}
\label{sec:srresults}
\subsection{A joint law for K\'arm\'an cloaking}
Scene-specific fits at four Reynolds numbers used different but correlated force and phase variables. Literal equality of the expressions was therefore not used as evidence of a shared mechanism. A stronger test was performed by fitting all four datasets jointly and deploying one algebraic structure in every scene.
The raw joint PySR candidate retained in the formula archive is
\begin{equation}
\alpha_F^{\mathrm{raw}}=0.38505\,\dot a_B+\dot a_F-14.95165\,\mu C_{l,\mathrm{tot}}.
\label{eq:rawfrontlaw}
\end{equation}
It has a nominal offline $R^2$ of unity on the highly correlated joint trajectory. That number is not interpreted as perfect identification. When equation~\eqref{eq:rawfrontlaw} is deployed with the shared rear head, the rear action is constant, hence $\dot a_B=0$ after initialisation. The first term is therefore a teacher-trajectory correlate, not an independently exercised feedback route. Retaining it in the paper's mechanism statement would falsely suggest dynamic rear feedback.
Because the shared rear head is constant, $\dot a_B$ becomes inactive after initialisation. The reduced relation used only for mechanism interpretation is
\begin{equation}
\boxed{\alpha_F=\dot a_F-14.952\,\mu C_{l,\mathrm{tot}}},
\label{eq:frontlaw}
\end{equation}
and the rear relation reduces to
\begin{equation}
\boxed{\alpha_T\simeq3.414,\qquad \alpha_B\simeq-3.414.}
\label{eq:rearlaw}
\end{equation}
Here $\dot a_F$ is the backward-looking front action-rate feature formed with the scene control interval, and $\mu$ is the viscosity variable used in the joint library. The numerical coefficients depend on the Legacy normalisation, control interval and action convention and are not proposed as universal constants. The closed-loop values in table~\ref{tab:crossre} validate the archived raw expression~\eqref{eq:rawfrontlaw}; the reduced equation~\eqref{eq:frontlaw} has not yet been assigned a separate validation record and must not be read as an independently tested replacement.
Equations~\eqref{eq:frontlaw}--\eqref{eq:rearlaw} separate the controller into a dynamic front contribution and a persistent antisymmetric rear contribution. The rear pair modifies the gap flow and supplies mean streamwise momentum to compensate the pinball velocity deficit. The front cylinder responds to the phase-rate and antisymmetric force state, regulating the residual lift and wake phase.
\subsection{Cross-$\Rey_D$ validation}
Table~\ref{tab:crossre} reports closed-loop performance. One symbolic expression is deployed without changing its algebraic structure.
\begin{table}
\centering
\caption{Closed-loop K\'arm\'an-cloak similarity. All Reynolds numbers are based on one pinball-cylinder diameter.}
\label{tab:crossre}
\begin{tabular}{cccc}
\toprule
$\Rey_D$ & PPO & joint SR & comment\\
\midrule
25 & 0.961 & 0.847 & successful transfer, below PPO\\
50 & 0.954 & 0.888 & closest agreement with PPO\\
100 & 0.884 & 0.845 & robust intermediate-$\Rey_D$ transfer\\
200 & 0.795 & 0.806 & 0.819 with shorter control interval\\
\bottomrule
\end{tabular}
\end{table}
The symbolic controller remains effective over the full range. At $\Rey_D=200$ ($\Rey_{\mathrm{code}}=400$), reducing the control interval improves similarity from $0.806$ to $0.819$, indicating that temporal resolution contributes to the high-$\Rey_D$ degradation. This observation does not establish sampling as the only cause: the PPO baseline is also weaker, and the wake contains shorter coupled time scales.
The important result is structural rather than numerical equality with PPO. A joint expression fitted to combined data retains most of the closed-loop performance across a factor of eight in $\Rey_D$. Reynolds-number dependence enters through the rate construction and viscosity-weighted lift feedback instead of through four unrelated policies.
\draftfigure{Insert the SR cross-$\Rey_D$ performance figure and representative PPO/SR action traces.}{Closed-loop validation of the joint cloak law. The figure should compare PPO, symbolic and uncontrolled similarity at each $\Rey_D$ and show that the rear actions are approximately constant while the front action carries the dynamic response.}{fig:sr-crossre}
\subsection{Transfer from periodic to transient disturbances}
The same law was applied without refitting to two transient-vortex cases. The Lamb dipole and Taylor monopole differ qualitatively from a periodic street and are not labelled in the symbolic expression. Table~\ref{tab:vortextransfer} shows that the SR law remains competitive with PPO.
\begin{table}
\centering
\caption{Transfer of the joint K\'arm\'an SR law to transient disturbances.}
\label{tab:vortextransfer}
\begin{tabular}{lcc}
\toprule
incident disturbance & PPO & transferred SR\\
\midrule
Lamb vortex dipole & 0.942 & 0.949\\
Taylor vortex monopole & 0.916 & 0.905\\
\bottomrule
\end{tabular}
\end{table}
The transfer is difficult to explain by waveform memorisation because the law contains neither a disturbance identifier nor an explicit target waveform. It uses local force and action-phase information that remains available when the incident field changes. The Taylor result is slightly below PPO and therefore also defines the scope of the claim: the relation is a reusable cloak backbone, not a universally optimal transient controller.
\subsection{Physical interpretation}
The joint relation may be written schematically as
\begin{equation}
\vect{\alpha}(t)\simeq
\underbrace{(0,\alpha_0,-\alpha_0)^{\mathsf T}}_{\text{mean rear compensation}}
+\underbrace{(\alpha_F(t),0,0)^{\mathsf T}}_{\text{front phase--force regulation}}.
\label{eq:mechanism}
\end{equation}
Opposite rear rotation accelerates the gap flow and changes the base-bleed jet, providing a mean correction to the blocked wake. The front cylinder carries most of the time-dependent action and changes the instantaneous antisymmetric state. This decomposition is consistent with known actuator roles in the fluidic pinball, but the present objective differs from conventional drag reduction or stabilisation: the corrected wake must retain an imposed incident event at a downstream observer.
The physical conclusion is deliberately narrower than the algebraic expression. The transferable finding is the division into persistent rear compensation and dynamic front rate--lift feedback. The coefficient in \eqref{eq:frontlaw} is tied to the dataset and is not a new nondimensional fluid constant.
\subsection{Illusion as target-dependent retuning}
Illusion adds a non-zero target state. Compact scene-specific laws are obtained near the natural controllable scale. For target diameter $0.75D$,
\begin{equation}
\alpha_F=-0.169(C_{l,\mathrm{tot}}+\dot C_{l,\mathrm{tot}})-1.240,
\label{eq:ill075}
\end{equation}
whereas for $1.0D$,
\begin{equation}
\alpha_F=0.0123(\dot u_a+u_a+26.5).
\label{eq:ill100}
\end{equation}
These formulas use different local coordinates around different target manifolds but both retain low-dimensional force or phase feedback.
A joint law fitted to the $0.75D$ and $1.0D$ targets is
\begin{align}
\alpha_F={}&C_{d,\mathrm{tot}}-(C_{d,\mathrm{err}}+5.428)
+0.00978(\dot u_a+u_a),
\label{eq:illfront}\\
\alpha_T={}&0.535\left[C_{d,\mathrm{err}}
-(C_{d,\mathrm{rear}}-C_{l,\mathrm{err}})\right]+2.782,
\label{eq:illrear}
\end{align}
with the lower action generated by reflection. It reaches full-horizon similarities of $0.982$ and $0.958$ on the two fitted target diameters in the current validation JSON files; the corresponding tail-window values are $0.807$ and $0.886$. The distinction shows that a high aggregate score need not imply uniform long-time retention. Thus moderate illusion retains mean-flow correction, phase information and force regulation, but expresses them relative to target error.
\begin{table}
\centering
\caption{Cross-diameter deployment of the joint illusion law.}
\label{tab:illusion}
\begin{tabular}{ccl}
\toprule
target diameter & joint-SR similarity & interpretation\\
\midrule
$0.5D$ & 0.854 & partial transfer\\
$0.6D$ & 0.939 & strong transfer\\
$0.75D$ & $\approx0.98$ & fitted near-native regime\\
$1.0D$ & 0.958 & fitted near-native regime\\
$1.2D$ & 0.849 & beginning of degradation\\
$1.5D$ & -- & not represented by present library\\
$2.0D$ & 0.675 & weak transfer\\
\bottomrule
\end{tabular}
\end{table}
Table~\ref{tab:illusion} defines a finite generalisation range. The shared law is effective while the target wake remains near the natural controllable state, but it degrades as the target scale changes.
\draftfigure{Insert the illusion diameter-generalisation curve. Add the $1.5D$ PPO spectrum and autocorrelation only after the original trajectory has been restored and the diagnostics regenerated.}{Target-scale dependence of the symbolic illusion controller. Near-native targets admit a compact target-error law, whereas no useful symbolic closed-loop controller has yet been validated for $1.5D$ with the present feature library.}{fig:illusion-regimes}
The result supports
\begin{equation}
\text{illusion}=\text{shared compensation}+\text{target-specific retuning},
\end{equation}
not a universal formula for all target cylinders. The individual formulas in equations~\eqref{eq:ill075} and~\eqref{eq:ill100} are useful local descriptions, whereas equations~\eqref{eq:illfront}--\eqref{eq:illrear} test whether a common target-error coordinate survives across scenes. The latter is the stronger generalisation test even when an individual fit has a better offline score.
The coefficient signs in the joint front law deserve explicit attention. The archived PySR expression is
\begin{equation}
C_{d,\mathrm{tot}}-(C_{d,\mathrm{err}}+5.4277)
-(-0.0097839)(\dot u_a+u_a),
\end{equation}
so its final phase term is positive after resolving the double negative. Writing the unresolved archive syntax alongside the simplified mathematical law prevents a transcription error between the machine-readable formula and the manuscript.
Cross-diameter deployment is performed without introducing target diameter as an input. Consequently, the degradation curve measures extrapolation of a law whose only target information is supplied through the instantaneous target-error features. A separate diameter-marker regression can fit a diameter-dependent expression, but it answers an easier question and is not the principal result reported here. The $0.8D$ validation, omitted from the compact table for readability, gives a similarity of $0.908$ and follows the same near-training-range trend.
\subsection{The $1.5D$ regime boundary}
The $1.5D$ case is a failure of the present symbolic representation rather than a failure of neural control. The stored PPO baseline remains effective, whereas PySR returns a degenerate lag-copy relation with nominal one-step $R^2=1$ and no useful canonical symbolic closed-loop validation. Exploratory analysis indicates rapid periodic modulation, but the previously quoted frequency ratio, lag-two autocorrelation and feature correlations cannot presently be recomputed because the underlying trajectory NPZ is absent from the archived analysis tree. Those numerical diagnostics are therefore withheld pending restoration of the raw trajectory.
Both PySR and sparse linear identification then favour lagged-action proxies. Such expressions predict one step but do not identify a mechanism: the lagged action stands in for a rapidly switching internal phase that is absent from the feature library. The appropriate response is not to report the proxy as a physical law. The case instead shows that an explicit oscillator, a sufficient delay embedding or a target-scale variable is required. SR is therefore useful both when it succeeds and when it exposes a missing state coordinate.
\section{Observable-inferred decomposition of the active correction}
\label{sec:oid}
Symbolic regression answers which measured quantities are mapped to which rotations, but an action-space law does not identify the spatial structures responsible for force and downstream matching. We therefore analyse the active correction field independently. Proper orthogonal decomposition (POD) provides an energy-ranked basis, while observable-inferred decomposition (OID) rotates a retained POD subspace towards covariance with a selected output, following the broader observable-correlation viewpoint of extended POD \citep{Boree2003}. This distinction is important because a low-energy near-body structure can dominate integrated force, whereas an energetic convected structure can dominate a downstream probe.
\subsection{Correction fields and OID construction}
For aligned vectorised correction snapshots $\vect{x}_t$, let
\begin{equation}
\vect{X}=\vect{U}\vect{\Sigma}\vect{V}^{\mathsf T},\qquad
\vect{z}_t=\vect{U}_r^{\mathsf T}\vect{x}_t.
\end{equation}
For an observable $\vect{y}_t$, the reduced cross-covariance is
\begin{equation}
\vect{C}_{zy}=N^{-1}\sum_t\vect{z}_t\vect{y}_t^{\mathsf T}
=\vect{U}_y\vect{\Sigma}_y\vect{V}_y^{\mathsf T},
\end{equation}
and the OID field directions are $\vect{\psi}^{(y)}_k=\vect{U}_r\vect{u}_{y,k}$. Force, downstream-signature, suppression and action observables are treated separately while sharing the underlying correction subspace. OID is therefore not a latent-state reconstruction of PPO and is not used to modify the SR law.
Five POD modes capture $99.97\%$ of steady-cloak correction energy, $99.90\%$ for K\'arm\'an cloak, $99.93\%$ and $99.91\%$ for the $0.75D$ and $1.0D$ illusions, and $97.90\%$ for $1.5D$. This compactness concerns the active change, not the full multi-scene flow. The lower value at $1.5D$ is consistent with a correction distributed over additional scales.
\begin{table}
\centering
\caption{Two-coordinate held-out prediction from observable-informed and energy-ranked coordinates. Negative POD $R^2$ means that the two-coordinate predictor is worse than the held-out mean.}
\label{tab:oidpred}
\begin{tabular}{llrr}
\toprule
scene & observable & OID $R^2$ & POD $R^2$\\
\midrule
K\'arm\'an cloak & force & 0.750 & 0.418\\
illusion $0.75D$ & force & 0.435 & -2.426\\
illusion $0.75D$ & signature & 0.661 & -0.034\\
illusion $1.0D$ & force & 0.671 & -0.237\\
illusion $1.0D$ & signature & 0.586 & -0.160\\
illusion $1.5D$ & force & 0.640 & 0.264\\
illusion $1.5D$ & signature & 0.315 & 0.060\\
\bottomrule
\end{tabular}
\end{table}
Table~\ref{tab:oidpred} shows that observable ranking, rather than extra state dimension, produces the predictive advantage. The K\'arm\'an force result is particularly clear. For the corresponding delayed signature error, successful cloaking leaves very little target variance; an $R^2$ close to zero is consequently not evidence that the correction is absent. Metrics that normalise by target variance become ill-conditioned precisely when the error is nearly eliminated.
\subsection{Task-dependent force--signature geometry}
The leading force and signature directions vary systematically with objective. In steady cloak, the signed force--suppression overlap is $+0.763$: reducing natural fluctuations and regulating force rely on related near-wake changes. The full-field fluctuation RMS decreases by $99.43\%$, lift RMS by $83.3\%$, and recirculation area by $38.5\%$, while recirculation length changes by only $3.2\%$. The correction narrows and quietens the wake rather than simply deleting its streamwise recirculation extent.
In K\'arm\'an cloak, the leading force--signature overlap is approximately $-0.034$ and remains close to zero over the tested convective delays. This near-orthogonality is physically plausible: the controller must regulate additional pinball force while preserving, not suppressing, the incoming street. For illusion the signed overlaps are $-0.082$, $-0.495$ and $-0.932$ for $0.75D$, $1.0D$ and $1.5D$. The $0.75D$ value is rank-sensitive and supports only partial separation; the latter two are more stable. A negative overlap is an orientation within the retained correction subspace, not proof of dynamically independent channels.
Raw observations predict K\'arm\'an PPO action with $R^2=0.956$, whereas two force-OID coordinates explain only about $22.5\%$ of action variance. This contrast reinforces the division of labour: SR approximates the observation-to-action map, while OID identifies output-relevant structures in the action-induced field. Neither should be substituted for the other.
\draftfigure{Insert the selected OID field-analysis panels after final figure assembly.}{Observable-informed correction modes and held-out prediction. Force relevance is concentrated in the body-connected near wake, whereas signature relevance shifts towards the convected correction near the observer. Error bars should show retained-rank or leave-one-cycle-out sensitivity where available.}{fig:oid}
\section{Lagged canonical-correlation decomposition and propagation}
\label{sec:ccd}
OID is instantaneous. Canonical-correlation decomposition (CCD) augments the observable with delays to distinguish the source correction near the cylinders from its downstream descendant. For delays $\{\tau_q\}_{q=1}^Q$, define
\begin{equation}
\vect{P}_t=[\vect{p}(t+\tau_1),\ldots,\vect{p}(t+\tau_Q)]
\end{equation}
and
\begin{equation}
\vect{C}_{PZ}=N^{-1}\sum_t\vect{P}_t^{\mathsf T}\vect{z}_t
=\vect{R}\vect{\Sigma}\vect{W}^{\mathsf T}.
\end{equation}
The field directions are $\vect{\psi}^{\mathrm{CCD}}_k=\vect{U}_r\vect{w}_k$. Periodic cases use leave-one-cycle-out tests; the steady case is evaluated with suppression and recirculation diagnostics. Delay embedding is not a claim of causality by itself, but the known ordering of cylinder actuation, near-wake response and sensor arrival gives the correlation a physically constrained interpretation.
\subsection{Common cloak correction and spatial zones}
Across steady, K\'arm\'an, Lamb and Taylor cloak, the dominant correction is a positive streamwise-velocity increment behind the pinball, with dipole-like velocity and concentrated vorticity changes around the rotating cylinders. In the common analysis normalisation, correction-field RMS values are $0.196$, $0.397$, $0.146$ and $0.188$, respectively. Similarity of structure does not mean equality of amplitude: the periodic K\'arm\'an task requires the largest correction because it must remove pinball distortion while transmitting an incoming street.
The field is examined in overlapping near-body, body-wake and sensor zones. Force-correlated directions are strongest around the surfaces, separated shear layers and body-connected rolled-up vorticity, consistent with impulse-based interpretations of instantaneous force. Delayed signature directions extend farther downstream as the active correction is convected and deformed. The zones are diagnostic masks, not independent subsystems.
Action-CCD associates the leading correction with the persistent rear rotation; higher directions carry the dynamic front contribution. Force-CCD emphasises the body-connected source, whereas signature-CCD emphasises the transported consequence. This yields the physically testable ordering
\begin{equation}
\text{rotation}\rightarrow\text{near-body correction}
\rightarrow\text{wake transport}\rightarrow\text{observer signal},
\end{equation}
although the present evidence is entirely numerical and no experimental measurements are included.
\subsection{Illusion target-correction comparison}
For illusion, the active correction can be compared with the correction required to transform the blocked pinball field into the target-cylinder field. At retained rank six, leading-mode absolute overlaps are $0.383$, $0.926$ and $0.922$ for target diameters $0.75D$, $1.0D$ and $1.5D$; at rank ten they become $0.320$, $0.684$ and $0.661$. The $1.0D$ target provides the clearest robust evidence of a shared leading correction. The $0.75D$ match is partial and rank-sensitive. At $1.5D$, a target-like leading spatial direction coexists with higher-order and temporal disagreement, consistent with the PPO high-frequency action and the failure of the instantaneous SR library.
These results refine, rather than simply repeat, the SR interpretation. Mean deficit compensation can remain spatially recognisable even when the temporal policy needed to realise it cannot be represented by the current symbolic coordinates. A strong leading-mode overlap is therefore not sufficient evidence for a transferable feedback law.
\draftfigure{Insert phase-aligned active corrections and lagged CCD modes after assembling the final SR, OID and CCD panels.}{Propagation of the active correction. Columns should distinguish action-, force- and delayed-signature-informed structures, with near-body, body-wake and sensor zones marked. Lamb and Taylor panels test whether the positive streamwise cloak correction survives a change from periodic to transient incidence.}{fig:ccd}
\section{Integrated mechanism and limitations}
\label{sec:field}
SR maps measurements to actions but does not show where those actions modify the flow. We therefore use the active correction
\begin{equation}
\Delta\vect{q}_{\mathrm{ctl}}
=\vect{q}_{\mathrm{ctl}}-\vect{q}_{\mathrm{blk}}
\label{eq:dqctl}
\end{equation}
rather than decomposing the raw controlled field. For illusion, the correction required to transform the blocked pinball wake into the target is
\begin{equation}
\Delta\vect{q}_{\mathrm{tar}}
=\vect{q}_{\mathrm{tar}}-\vect{q}_{\mathrm{blk}}.
\end{equation}
This subtraction separates the active change from the incident disturbance and natural pinball wake.
Five POD modes capture $99.97\%$ of the steady-cloak correction energy, $99.90\%$ for K\'arm\'an cloak, $99.93\%$ and $99.91\%$ for the $0.75D$ and $1.0D$ illusions, and $97.90\%$ for the $1.5D$ illusion. The actuator therefore modifies the flow within a compact subspace even though the complete flow is not low-dimensional over all scenes.
The leading spatial correction shared by steady, K\'arm\'an, Lamb and Taylor cloak is a positive streamwise-velocity increment behind the pinball, accompanied by near-body dipole-like velocity and vorticity changes. This is the field counterpart of the rear counter-rotation in \eqref{eq:rearlaw}: it compensates the pinball-induced velocity deficit. Time-dependent near-body changes are consistent with the front phase--force term in \eqref{eq:frontlaw}.
Observable-informed projections further show why force matching and downstream matching must not be conflated. Two observable-informed coordinates predict K\'arm\'an force with $R^2=0.750$, compared with $0.418$ for two energy-ranked POD coordinates. For illusion, the same two-coordinate comparison gives force/signature $R^2$ values of $0.435/0.661$, $0.671/0.586$ and $0.640/0.315$ for $0.75D$, $1.0D$ and $1.5D$, respectively; the corresponding POD values are substantially lower. The force-relevant structure is concentrated in the body-connected near wake, whereas the signature-relevant structure is its convected descendant near the observer.
\draftfigure{Insert the phase-aligned active correction fields for steady, K\'arm\'an, Lamb and Taylor cloak, together with the leading force- and signature-informed modes.}{Spatial consistency of the symbolic mechanism. Rear-cylinder compensation produces a common downstream streamwise increment, while force- and signature-relevant projections emphasise the near-body source and convected descendant, respectively.}{fig:correction-fields}
These field results are not used to fit the symbolic law. They provide an independent consistency check for the mechanism chain
\begin{equation}
\text{measurements}
\rightarrow\text{compact actuation}
\rightarrow\text{near-body correction}
\rightarrow\text{body-wake transport}
\rightarrow\text{downstream signature}.
\label{eq:chain}
\end{equation}
They also explain why a common actuator mechanism can have different force and signature projections in steady suppression, incident-street preservation and target-wake synthesis.
\section{Discussion}
\label{sec:discussion}
\subsection{Why closed-loop validation changes the SR problem}
Policy distillation differs fundamentally from ordinary supervised regression. The target data are generated by a controller that shapes its own state distribution. Once the symbolic surrogate replaces PPO, each action changes all future observations. High one-step accuracy can preserve correlations that are incidental to the PPO trajectory while accumulating a systematic phase error in deployment. Closed-loop CFD must therefore be treated as part of symbolic model selection rather than as a final illustration.
This criterion also limits algebraic over-interpretation. Correlated forces, sensor values and action lags can yield multiple formulas with similar offline error. Reflection symmetry, dimensional construction, deployment activity and transfer performance are used to select a canonical representative. The identified mechanism is the robust actuator-role decomposition, not the uniqueness of every algebraic term.
\subsection{Scope of the common cloak backbone}
The strongest evidence for a reusable mechanism is the combination of cross-$\Rey_D$ and cross-disturbance transfer. The joint law retains substantial performance from $\Rey_D=25$ to 200 and transfers from a periodic street to isolated vortices. This suggests that the controller responds primarily to the additional distortion introduced by the pinball. The incident field changes the timing and amplitude of the response, but rear momentum compensation and front phase--force regulation remain useful.
The result does not imply complete invariance. Sampling sensitivity increases at high $\Rey_D$, the Taylor transient remains imperfect, and the formula has been evaluated on the state manifolds reached by the trained policies. Arbitrary initial conditions, different sensor locations or different actuator limits may require a new rate coordinate or refitting.
\subsection{Why illusion has a regime boundary}
Cloaking cancels the pinball contribution relative to an incident background; illusion must additionally create the difference between the corrected pinball wake and a target wake. When target and natural controllable scales are close, target-error feedback retunes the cloak backbone. As the target diameter grows, the controller must reorganise frequency, amplitude and spatial scale more strongly. The $1.5D$ policy appears to use a rapid internal phase state, and the $2.0D$ transfer is weak. The symbolic failure is consequently physically informative: the control state requires a temporal coordinate that cannot be reconstructed from the present instantaneous force and probe library.
\subsection{What the symbolic law establishes}
The evidence supports three levels of claim, which should not be collapsed. At the empirical level, one algebraic structure provides useful closed-loop K\'arm\'an cloaking over a factor of eight in $\Rey_D$ and transfers to two transient disturbances. At the mechanistic level, the deployment audit and field correction support a division between mean rear compensation and dynamic front regulation. At the universal level, however, the available evidence is insufficient: coefficients depend on the Legacy plant, and neither three-dimensional flow nor alternative actuator geometry has been tested. The paper's main contribution is therefore a transferable mechanism within a documented numerical family, not a universal constitutive law for hydrodynamic cloaks.
The raw-versus-canonical distinction in equations~\eqref{eq:rawfrontlaw} and~\eqref{eq:frontlaw} illustrates why symbolic interpretability requires intervention. An expression tree does not become a mechanism merely by being short. The raw $\dot a_B$ term is mathematically legitimate on the PPO trajectory and improves neither deployment activity nor physical explanation once the rear head becomes constant. Removing it is analogous to testing a putative causal input by changing the operating distribution: the symbolic controller itself supplies that intervention. This is also why the nominal $R^2=1$ attached to the archived joint formula should not dominate interpretation.
The implicit appearance of $\dot a_F$ warrants similar caution. Written literally, equation~\eqref{eq:frontlaw} contains the derivative of the output being computed. In implementation it is an available finite difference from the previous applied action, so the law is a discrete dynamic controller rather than an algebraic loop. If $k$ indexes control decisions,
\begin{equation}
\dot a_F^k=\frac{a_F^{k-1}-a_F^{k-2}}{\Delta t_c}
\end{equation}
is formed from stored applied actions before $\alpha_F^k$ is evaluated. This convention must accompany any reproduction because using the newly requested action in the difference would define a different controller. The rate term acts as a compact phase coordinate, but it is not asserted to equal a material derivative or a state derivative of the fluid.
\subsection{Implications for sensor and controller design}
The extracted feature set suggests a hierarchy for future controller design. Mean rear compensation requires little dynamic information and could plausibly be supplied as a calibrated operating-point schedule. The front cylinder requires an antisymmetric force measurement and a temporal coordinate. The downstream probes remain important to PPO training and reward evaluation, but their limited appearance in the canonical K\'arm\'an law suggests that the distilled controller can use local force feedback once the compensation operating point has been discovered. This is a hypothesis about the tested closed-loop manifold, not a demonstration that probes can be removed under noise, drift or unseen incidence.
The illusion law differs because target errors enter directly. A target wake is not specified by one scalar diameter in the principal joint formula; it is represented by phase-dependent force errors and a downstream phase coordinate. This design allows the same expression to interpolate moderate target changes, but it also explains why extrapolation fails. As target scale grows, the available error coordinates may no longer resolve whether a mismatch should be corrected by changing mean momentum, shedding phase or oscillation frequency. The degenerate lag-copy fit and absence of a useful symbolic closed-loop result reveal this ambiguity; the detailed spectrum should be restored from the original trajectory before publication.
A practical successor to the present library would introduce a minimal oscillator state rather than indiscriminately adding lags. Two quadrature variables estimated from a narrow-band force or probe signal could encode phase without using the previous action as a proxy. Alternatively, a delay-coordinate state could be selected by observability and closed-loop tests. Both choices would increase complexity and should be justified by improvement at $1.5D$ without degrading the compact laws at $0.75D$ and $1.0D$. The present negative result supplies a concrete benchmark for that extension.
\subsection{Numerical provenance and reproducibility}
All mechanism coefficients and validation values in this article originate from the Legacy $1280\times512$ parabolic-inlet, no-slip-wall dataset. This provenance is repeated because a superficially similar V5 environment uses $2000\times600$ nodes, uniform inflow and free-slip walls. V5 may be appropriate for new training studies, but mixing its trajectories with the Legacy regression would conflate changes in blockage, wall boundary layers, action maps and normalisation. A future cross-plant study should treat V5 as an external validation domain and should report whether coefficients are refitted, not silently pool the two datasets.
The computational record contains machine-readable scene definitions, formula files, deterministic PPO trajectories and closed-loop validation summaries. Formula JSON files retain feature names and archive syntax; validation JSON files retain scene, controller mode, number of decisions and similarity. The scene registry is the source of truth when summary documents disagree. One such example is the $1.0D$ illusion similarity: the registry records $0.958$ for the current canonical validation, while an earlier results summary reports approximately $0.970$. We have retained approximate wording in the narrative where archival summaries differ and recommend regenerating the final table from a frozen registry immediately before submission.
Figure placeholders intentionally compile without copied graphics. The suggested paths point to generated SR plots in the analysis directory, but the files have not been assumed to exist beside the Overleaf source. Before submission, each selected PDF should be copied into a manuscript figure directory, its plotted Reynolds-number labels should be audited against the $\Rey_D$ convention, and any caption language such as ``universal'' should be weakened to ``shared over the tested cases''. OID and CCD panels require separate assembly from their field outputs and uncertainty checks.
\subsection{Limitations}
The study is two-dimensional and uses numerical trajectories. No experimental results are included, and the simulations should not be described as experimental validation. Three-dimensional instability, measurement noise, actuator bandwidth and laboratory delay may change the extracted coefficients and possibly the preferred features. The reported SR coefficients belong to the Legacy solver configuration and should not be transplanted directly to the newer V5 plant. The correction-field records contain a limited number of independent cycles, and the larger-target cases require longer trajectories and explicit frequency-state features.
Similarity is defined in a sparse downstream observer space and is not equivalent to pointwise cancellation everywhere. Full-field correction and force results reduce this ambiguity but do not establish invisibility to every possible observer. Finally, symbolic expressions are empirical closed-loop surrogates, not governing equations for the Navier--Stokes dynamics.
\subsection{Comparison of cloak and illusion identification problems}
Cloak and illusion produce superficially similar regression tables but pose different identification problems. For cloak, the desired observer signal is inherited from the incident flow. The controller should remove the incremental distortion created by the pinball, so force and phase variables can be interpreted relative to a recurring compensation task. This is why a formula trained on a periodic street can remain useful for isolated vortices: the incident waveform changes, but the additional blockage and body-connected distortion retain common actuator-side features.
For illusion, the reference itself changes. The error is generated jointly by the pinball dynamics and a target cylinder that is absent from the controlled domain. A target-force harmonic reconstruction supplies phase information during training, but the controlled flow can depart from that phase manifold in deployment. The joint formula succeeds near $0.75D$--$1.0D$ because target errors remain informative local coordinates there. Beyond this range, equal instantaneous errors can demand different actions depending on target phase and frequency history. The cross-diameter decline is thus not merely ordinary extrapolation in a scalar parameter; it is loss of state identifiability in the chosen feature space.
This difference also clarifies why a diameter marker is not a complete remedy. A marker can tell PySR which target family is requested, but it does not supply the missing oscillator phase at $1.5D$. Conversely, a phase coordinate without target scale may distinguish switching direction while failing to choose the required mean correction. A robust large-range illusion controller will probably require both a target descriptor and a dynamic phase state. Establishing the minimal pair is a natural next symbolic-identification problem.
\subsection{Why the field analyses remain secondary to SR}
OID and CCD strengthen the mechanism claim but do not replace the central policy result. A spatial mode can recur across cases because geometrically similar actuators create similar local disturbances, even if the feedback rules that time those disturbances are unrelated. The cross-$\Rey_D$ and cross-disturbance closed-loop deployments directly test the feedback relation; field decompositions then ask whether its inferred actuator roles have a consistent consequence. This ordering avoids using visually similar vorticity panels as proof of controller equivalence.
Conversely, SR alone could mistake a compact correlation for physics. The positive downstream streamwise correction is the expected spatial consequence of rear counter-rotation compensating a blocked mean wake. The body-connected force modes and convected signature modes explain why the same action can influence two reward channels at different locations and times. The agreement is therefore triangulation: policy transfer, deployment activity and field organisation support one another while retaining distinct uncertainties.
The field analyses also constrain the language of invisibility. Sparse-sensor similarity can be high even when correction energy remains near the cylinders, and OID deliberately selects structures relevant to one observable. CCD shows transport towards the sensor zone, not cancellation for every possible observer. A full observability claim would require dense multi-component measurements over a prescribed exterior region and robustness to alternative observer placement. Those tests are beyond the current dataset.
\subsection{Prospects for closed-loop symbolic discovery}
The present workflow uses CFD validation after PySR search because evaluating every evolutionary candidate in CFD would be prohibitive. A future method could incorporate closed-loop information more efficiently through staged screening. Offline Pareto fronts could first be filtered by symmetry, boundedness and local sensitivity. A differentiable or data-driven short-horizon surrogate could then reject candidates with obvious phase drift, after which only a small set would enter full CFD. Importantly, the final acceptance criterion should remain the original solver, since surrogate agreement on the teacher distribution recreates the problem identified here.
Multiple independent PPO seeds would provide another valuable axis. If different neural policies converge to the same symbolic actuator roles but different coefficients or correlated coordinates, the shared structure would be stronger evidence of mechanism. If they instead produce distinct successful laws, the result would reveal policy multiplicity: several feedback strategies may generate the same observer signature. Such multiplicity is scientifically relevant and should not be hidden by selecting one best seed.
Noise and delay studies should likewise be performed at the symbolic level. Compact formulas can be more transparent but less forgiving than neural policies with implicit smoothing. Rate features amplify measurement noise, and the $\dot a_F$ coordinate depends on actuator telemetry and timing. Low-pass differentiation, state observers or explicit oscillator coordinates could improve robustness, but each changes the identified law. Reporting the filtering operator and its delay would be as important as reporting the coefficients.
Finally, the symbolic law offers a tractable object for stability and sensitivity analysis. Around a periodic controlled orbit one could linearise the discrete controller together with a reduced flow response and examine Floquet multipliers as the control interval changes. Such an analysis could determine whether the SI400 improvement at $\Rey_D=200$ ($\Rey_{\mathrm{code}}=400$) results from increased phase margin or simply finer waveform tracking. The current data establish sensitivity but do not resolve that mechanism, so this remains a proposed extension rather than a conclusion.
\section{Conclusions}
\label{sec:conclusion}
DRL policies for hydrodynamic cloaking and illusion of the fluidic pinball have been distilled into explicit feedback laws and tested in closed-loop CFD. The main conclusions are:
\begin{enumerate}
\item One-step action fit is insufficient for policy distillation. A model with $R^2\simeq0.94$ can perform substantially worse in closed loop than a lower-dimensional model with $R^2\simeq0.32$. Closed-loop CFD, symmetry and deployment relevance must be included in model selection.
\item K\'arm\'an cloaking across $\Rey_D=25$--200 admits a common actuator-side structure: approximately steady and opposite rear-cylinder rotation supplies mean wake compensation, while front-cylinder rate--lift feedback regulates phase and residual force.
\item The same symbolic law transfers without refitting to Lamb and Taylor transient vortices, supporting a reusable cloak backbone rather than memorisation of a periodic waveform.
\item Illusion near target diameters $0.75D$--$1.0D$ reuses the compensation mechanism with target-error feedback. Generalisation deteriorates as target scale departs from this regime.
\item The $1.5D$ target reveals a high-frequency policy state absent from the current feature library. This regime requires an oscillator, delay embedding or explicit target-frequency coordinate rather than a lagged-action proxy.
\item The active correction field is strongly low-dimensional and contains the downstream streamwise increment predicted by the symbolic rear-cylinder mechanism. Force and downstream signature are different observable projections of this common correction.
\end{enumerate}
The broader result is that DRL and symbolic regression play complementary roles. DRL discovers high-performing strategies in a nonlinear delayed environment; symmetry-constrained SR identifies the smallest transferable feedback relation; and field analysis tests whether that relation has a consistent spatial consequence. This combination turns hydrodynamic cloaking from a black-box control demonstration into a testable mechanism for active wake-information management.
\section*{Supplementary data}
Supplementary material will include the scene definitions, symbolic expressions, closed-loop validation records and correction-field datasets.
For the periodic K\'arm\'an case, persistent rear-cylinder counter-rotation links the dominant tested surrogate term to the constant-control mean field, while the centred velocity residual describes time-dependent organization separately. This physical reading is deliberately limited: action ranking, scalar mean-field increments and residual modes answer different questions. Hydrodynamic cloak and illusion here denote control relative to a declared reference for specified observations and comparisons, not universal invisibility, configuration-independent transfer or an unrestricted physical law.
\section*{Acknowledgements}
The authors acknowledge the computational resources and discussions that supported this work.
\section*{Funding}
Funding information will be added in the final manuscript.
Funding information will be supplied before submission.
\section*{Declaration of interests}
The authors report no conflict of interest.
\section*{Author ORCIDs}
Author ORCID information will be added before submission.
Author ORCID information will be supplied before submission.
\bibliographystyle{jfm}
\bibliography{jfm}
\bibliography{../ROUND5_BOTTOM_UP/R5_REFERENCES}
\end{document}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
Binary file not shown.

After

Width:  |  Height:  |  Size: 816 KiB

Binary file not shown.

Before

Width:  |  Height:  |  Size: 226 KiB

After

Width:  |  Height:  |  Size: 226 KiB

Binary file not shown.
Binary file not shown.

After

Width:  |  Height:  |  Size: 689 KiB