Search arXivSearch

arXiv subjects

Ban

Publications and source records attributed to Ban.

11 recordsLinked to original sources

Directional Influence Function: Estimating Training Data Influence in Constrained Learning

As constrained learning becomes increasingly common, models are trained under explicit feasibility requirements to enforce fairness, safety, robustness, regulariza- tion, and physics or logic constraints. Understanding how training samples in- fluence the model solution (e.g., learned parameters) is crucial for interpretability and robustness. The classical influence function (IF) estimates sample contribu- tions via local sensitivity analysis, measuring how the solution changes when a specific training sample is perturbed or removed. However, IF becomes unreli- able in constrained settings: data perturbations can reshape both the objective and the feasible region, leading to estimates that violate feasibility. In response, we propose the Directional Influence Function (DIF), a novel estimator that explicitly incorporates these constraints into influence estimation. DIF formulates the opti- mality conditions of constrained learning as a variational inequality (VI) and ana- lyzes how perturbing training data affects this VI. We validate DIF on constrained linear regression and demonstrate that it recovers leave-one-out retraining results, whereas IF and penalty-based IF exhibit significant bias. We further apply DIF to fairness-constrained CNNs, where DIF accurately predicts test loss changes under data removal and aligns closely with actual retraining. Our results establish DIF as an efficient and reliable tool for data attribution in constrained learning.

cs.LG

Decorrelating the Future: Joint Frequency Domain Learning for Spatio-temporal Forecasting

Standard direct forecasting models typically rely on point-wise objectives such as Mean Squared Error, which fail to capture the complex spatio-temporal dependencies inherent in graph-structured signals. While recent frequency-domain approaches such as FreDF mitigate temporal autocorrelation, they often overlook spatial and cross spatio-temporal interactions. To address this limitation, we propose FreST Loss, a frequency-enhanced spatio-temporal training objective that extends supervision to the joint spatio-temporal spectrum. By leveraging the Joint Fourier Transform (JFT), FreST Loss aligns model predictions with ground truth in a unified spectral domain, effectively decorrelating complex dependencies across both space and time. Theoretical analysis shows that this formulation reduces estimation bias associated with time-domain training objectives. Extensive experiments on six real-world datasets demonstrate that FreST Loss is model-agnostic and consistently improves state-of-the-art baselines by better capturing holistic spatio-temporal dynamics.

cs.LG

Machine Unlearning of Traffic State Estimation and Prediction

Data-driven traffic state estimation and prediction (TSEP) relies heavily on data sources that contain sensitive information. While the abundance of data has fueled significant breakthroughs, particularly in machine learning-based methods, it also raises concerns regarding privacy, cybersecurity, and data freshness. These issues can erode public trust in intelligent transportation systems. Recently, regulations have introduced the "right to be forgotten", allowing users to request the removal of their private data from models. As machine learning models can remember old data, simply removing it from back-end databases is insufficient in such systems. To address these challenges, this study introduces a novel learning paradigm for TSEP-Machine Unlearning TSEP-which enables a trained TSEP model to selectively forget privacy-sensitive, poisoned, or outdated data. By empowering models to "unlearn," we aim to enhance the trustworthiness and reliability of data-driven traffic TSEP.

cs.LG

Model-Targeted Data Poisoning Attacks against ITS Applications with Provable Convergence

The growing reliance of intelligent systems on data makes the systems vulnerable to data poisoning attacks. Such attacks could compromise machine learning or deep learning models by disrupting the input data. Previous studies on data poisoning attacks are subject to specific assumptions, and limited attention is given to learning models with general (equality and inequality) constraints or lacking differentiability. Such learning models are common in practice, especially in Intelligent Transportation Systems (ITS) that involve physical or domain knowledge as specific model constraints. Motivated by ITS applications, this paper formulates a model-target data poisoning attack as a bi-level optimization problem with a constrained lower-level problem, aiming to induce the model solution toward a target solution specified by the adversary by modifying the training data incrementally. As the gradient-based methods fail to solve this optimization problem, we propose to study the Lipschitz continuity property of the model solution, enabling us to calculate the semi-derivative, a one-sided directional derivative, of the solution over data. We leverage semi-derivative descent to solve the bi-level optimization problem, and establish the convergence conditions of the method to any attainable target model. The model and solution method are illustrated with a simulation of a poisoning attack on the lane change detection using SVM.

math.OC

Discovering Car-following Dynamics from Trajectory Data through Deep Learning

This study aims to discover the governing mathematical expressions of car-following dynamics from trajectory data directly using deep learning techniques. We propose an expression exploration framework based on deep symbolic regression (DSR) integrated with a variable intersection selection (VIS) method to find variable combinations that encourage interpretable and parsimonious mathematical expressions. In the exploration learning process, two penalty terms are added to improve the reward function: (i) a complexity penalty to regulate the complexity of the explored expressions to be parsimonious, and (ii) a variable interaction penalty to encourage the expression exploration to focus on variable combinations that can best describe the data. We show the performance of the proposed method to learn several car-following dynamics models and discuss its limitations and future research directions.

cs.LG

Minimum-Delay Opportunity Charging Scheduling for Electric Buses

Transit agencies that operate battery-electric buses must carefully manage fast-charging infrastructure to extend daily bus range without degrading on-time performance. To support this need, we propose a mixed-integer linear programming model to schedule opportunity charging that minimizes the amount of departure delay in all trips served by electric buses. Our novel approach directly tracks queuing at chargers in order to set and propagate departure delays. Allowing but minimizing delays makes it possible to optimize performance when delays due to traffic conditions and charging needs are inevitable, in contrast with existing methods that require charging to occur during scheduled layover time. To solve the model, we develop two algorithms based on decomposition. The first is an exact solution method based on Combinatorial Benders (CB) decomposition, which avoids directly enumerating the model's logic-based "big M" constraints and their inevitable computational challenges. The second, inspired by the CB approach but more efficient, is a polynomial-time heuristic based on linear programming that we call 3S. Computational experiments on both a simple notional transit network and the real bus system of King County, Washington, USA demonstrate the performance of both methods. The 3S method appears particularly promising for creating good charging schedules quickly at real-world scale.

math.OC

Joint Estimation of Multi-phase Traffic Demands at Signalized Intersections Based on Connected Vehicle Trajectories

Accurate traffic demand estimation is critical for the dynamic evaluation and optimization of signalized intersections. Existing studies based on connected vehicle (CV) data are designed for a single phase only and have not sufficiently studied the real-time traffic demand estimation for oversaturated traffic conditions. Therefore, this study proposes a cycle-by-cycle multi-phase traffic demand joint estimation method at signalized intersections based on CV data that considers both undersaturated and oversaturated traffic conditions. First, a joint weighted likelihood function of traffic demands for multiple phases is derived given real-time observed CV trajectories, which considers the initial queue and relaxes the first-in-first-out assumption by treating each queued CV as an independent observation. Then, the sample size of the historical CVs is used to derive a joint prior distribution of traffic demands. Ultimately, a joint estimation method based on the maximum a posteriori (i.e., the JO-MAP method) is developed for cycle-based multi-phase traffic demand estimation. The proposed method is evaluated using both simulation and empirical data. Simulation results indicate that the proposed method can produce reliable estimates under different penetration rates, arrival patterns, and traffic demands. The feature of joint estimation makes our method less demanding for the penetration rate of CVs and the consideration of prior distribution can significantly improve the estimation accuracy. Empirical results show that the proposed method achieves accurate cycle-based traffic demand estimation with a MAPE of 12.73%, outperforming the other four methods.

math.ST

Cumulative Flow Diagram-Based Fixed-Time Signal Timing Optimization at Isolated Intersections Using Connected Vehicle Trajectory Data

Time-dependent fixed-time control is a cost-effective control method that is widely employed at signalized intersections in numerous countries. Existing optimization models rely on traditional delay models with specific assumptions regarding vehicle arrivals. Recent advancements in intelligent mobility have led to development of high-resolution trajectory data of connected vehicles (CVs), thereby providing opportunities for improving fixed-time signal control. Taking advantage of CV trajectories, this study proposes a cumulative flow diagram (CFD)-based signal timing optimization method for fixed-time signal control at isolated intersections, which includes a CFD model and a multi-objective optimization model. The CFD model is formulated to profile the time-dependent vehicle arrival and departure processes under varying signal timing plans, where the intersection demand is estimated based on a weighted maximum likelihood estimation method. Then, a CFD-based multi-objective optimization model is proposed for both undersaturated and oversaturated traffic conditions. The primary objective is to minimize the exceeded queue dissipation time, whereas the secondary objective is to minimize the average delay at the intersection. Considering the data-driven property of the CFD model, a bi-level particle swarm optimization-based algorithm is then specially designed to solve the optimal cycle length (and the reference point if it is considered) and green ratios separately. The proposed method is evaluated based on simulation data and compared with Synchro. The results indicate that the proposed method outperforms Synchro under various traffic conditions in terms of average delay and queue since the time-dependent vehicle arrivals during the cycle are considered in a CV trajectory data-driven way.

math.OC

The Effects of the COVID-19 Pandemic on Transportation Systems in New York City and Seattle, USA

This paper continues to highlight trends in mobility and sociability in New York City (NYC), and supplements them with similar data from Seattle, WA, two of the cities most affected by COVID-19 in the U.S. Seattle may be further along in its recovery from the pandemic and ensuing lockdown than NYC, and may offer some insights into how travel patterns change. Finally, some preliminary findings from cities in China are discussed, two months following the lifting of their lockdowns, to offer a glimpse further into the future of recovery.

physics.soc-ph

Commuting Service Platform: Concept and Analysis

We propose and investigate the concept of commuting service platforms (CSP) that leverage emerging mobility services to provide commuting services and connect directly commuters (employees) and their worksites (employers). By applying the two-sided market analysis framework, we show under what conditions a CSP may present the two-sidedness. Both the monopoly and duopoly CSPs are then analyzed. We showhowthe price allocation, i.e., the prices charged to commuters and worksites, can impact the participation and profit of the CSPs. We also add demand constraints to the duopoly model so that the participation rates ofworksites and employees are (almost) the same. With demand constraints, the competition between the two CSPs becomes less intense in general. Discussions are presented on how the results and findings in this paper may help build CSP in practice and how to develop new, CSP-based travel demand management strategies.

econ.GN

Extracting Trips from Multi-Sourced Data for Mobility Pattern Analysis: An App-Based Data Example

Passively-generated data, such as GPS data and cellular data, bring tremendous opportunities for human mobility analysis and transportation applications. Since their primary purposes are often non-transportation related, the passively-generated data need to be processed to extract trips. Most existing trip extraction methods rely on data that are generated via a single positioning technology such as GPS or triangulation through cellular towers (thereby called single-sourced data), and methods to extract trips from data generated via multiple positioning technologies (or, multi-sourced data) are absent. And yet, multi-sourced data are now increasingly common. Generated using multiple technologies (e.g., GPS, cellular network- and WiFi-based), multi-sourced data contain high variances in their temporal and spatial properties. In this study, we propose a 'Divide, Conquer and Integrate' (DCI) framework to extract trips from multi-sourced data. We evaluate the proposed framework by applying it to an app-based data, which is multi-sourced and has high variances in both location accuracy and observation interval (i.e. time interval between two consecutive observations). On a manually labeled sample of the app-based data, the framework outperforms the state-of-the-art SVM model that is designed for GPS data. The effectiveness of the framework is also illustrated by consistent mobility patterns obtained from the app-based data and an externally collected household travel survey data for the same region and the same period.

stat.AP