Search arXiv⌕ Search

arXiv subjects

Anatoly A. Krasnovsky

Publications and source records attributed to Anatoly A. Krasnovsky.

7 recordsLinked to original sources

Emergence-as-Code as a Foundation for Reliable Self-Governance

Local adaptations can preserve component health while changing whether a system meets its reliability requirements. The system-level consequence depends on the interaction affected and its role in the user journey. Emergence-as-Code (EmaC) makes this relationship part of an executable, evidence-reconciled CompositeSLO assessment. A declared journey obligation persists while discovery maintains a hypothesis of its runtime realization. An inferred operator-to-role binding selects the measurements used in the calculation; its evidential status determines whether the binding supports a numerical assessment. Each result retains the obligation and versioned hypothesis as a basis for governance decisions. A PetClinic proof of concept recovered the exact operator-state/edge-binding change in all 20 randomized treatments, with no false change in 20 controls. EmaC and a manual dynamic composite matched disjoint semantic outcomes in all 40 conditions. A model with frozen suppression knowledge overestimated required-history availability by 0.01, placing treatments above a 0.995 target despite observed availability of 0.99. Ambiguous and contradictory evidence produced UNASSESSABLE. The research programme evaluates sustained binding maintenance, calibrated uncertainty, and the decision value of this requirement-level account of local adaptation.

cs.SE↗

Stochastic Connectivity as the Foundation of a Runtime Model for Microservice Availability Analysis

Microservice availability is commonly assessed by fault injection and chaos experiments, but such experiments are costly, operationally risky, and difficult to repeat for every architectural change. Distributed tracing and deployment metadata provide cheaper evidence, yet they usually remain descriptive: they show which services interacted, not what endpoint-level availability property follows. This paper proposes a formal runtime availability model based on stochastic connectivity for resilience-oriented analysis of microservice endpoints. It treats endpoint availability under explicit fault scenarios as a measurable facet of microservice resilience, combining a typed service-dependency graph, a replication map, a probability measure over node and edge states, and request-specific success predicates. Its semantics separates computational failures of service replicas from communication failures of logical dependencies, showing that replication cannot compensate for bottleneck dependencies. The model can be reconstructed from traces and deployment artifacts, parameterized for architectural what-if analysis, and analyzed by Monte Carlo simulation before or alongside fault injection. We define the model, its trace-to-model construction, elementary semantic properties, and a synthetic adequacy study. The study matches closed-form oracle cases within sampling error and exposes boundaries caused by edge bottlenecks, correlated failures, missing traces, and time-dependent failures.

cs.SE↗

Toward Operationalizing Rasmussen: Drift Observability on the Simplex for Evolving Systems

Software operations increasingly rely on SLOs, traces, deployment specifications, and change events, yet dashboards and thresholding practices often expose share-like operational signals as separate scalar panels or baseline distances. This can create false alarms under benign redistribution and miss movement toward policy boundaries. Rasmussen's dynamic safety model motivates drift under competing pressures, but operationalizing it for software is difficult because relevant state variables (remaining margin, engineering effort, and risk/impact) are often compositional and their parts evolve. We formulate an automated, artifact-derived drift-monitor design that maps changing software artifacts into a stable compositional monitoring state: it extracts a current part inventory and policy constraints, maps telemetry to a positive composition, stabilizes splits, merges, and renames through lineage-aware canonical groups, and analyzes boundary-directed drift in log-ratio coordinates. The proposed monitor would report drift direction, step-to-boundary, balance-level attribution, and model-health indicators under architectural churn. We specify the approach, identify its zero/noise/lineage assumptions, and report a reproducible synthetic sanity check of boundary-aware drift and controlled part churn.

eess.SY↗

Evaluating Asynchronous Semantics in Trace-Discovered Resilience Models: A Case Study on the OpenTelemetry Demo

While distributed tracing and chaos engineering are becoming standard for microservices, resilience models remain largely manual and bespoke. We revisit a trace-discovered connectivity model that derives a service dependency graph from traces and uses Monte Carlo simulation to estimate endpoint availability under fail-stop service failures. Compared to earlier work, we (i) derive the graph directly from raw OpenTelemetry traces, (ii) attach endpoint-specific success predicates, and (iii) add a simple asynchronous semantics that treats Kafka edges as non-blocking for immediate HTTP success. We apply this model to the OpenTelemetry Demo ("Astronomy Shop") using a GitHub Actions workflow that discovers the graph, runs simulations, and executes chaos experiments that randomly kill microservices in a Docker Compose deployment. Across the studied failure fractions, the model reproduces the overall availability degradation curve, while asynchronous semantics for Kafka edges change predicted availabilities by at most about 10^(-5) (0.001 percentage points). This null result suggests that for immediate HTTP availability in this case study, explicitly modeling asynchronous dependencies is not warranted, and a simpler connectivity-only model is sufficient.

cs.SE↗

Model Discovery and Graph Simulation: A Lightweight Gateway to Chaos Engineering

Chaos engineering reveals resilience risks but is expensive and operationally risky to run broadly and often. Model-based analyses can estimate dependability, yet in practice they are tricky to build and keep current because models are typically handcrafted. We claim that a simple connectivity-only topological model - just the service-dependency graph plus replica counts - can provide fast, low-risk availability estimates under fail-stop faults. To make this claim practical without hand-built models, we introduce model discovery: an automated step that can run in CI/CD or as an observability-platform capability, synthesizing an explicit, analyzable model from artifacts teams already have (e.g., distributed traces, service-mesh telemetry, configs/manifests) - providing an accessible gateway for teams to begin resilience testing. As a proof by instance on the DeathStarBench Social Network, we extract the dependency graph from Jaeger and estimate availability across two deployment modes and five failure rates. The discovered model closely tracks live fault-injection results; with replication, median error at mid-range failure rates is near zero, while no-replication shows signed biases consistent with excluded mechanisms. These results create two opportunities: first, to triage and reduce the scope of expensive chaos experiments in advance, and second, to generate real-time signals on the system's resilience posture as its topology evolves, preserving live validation for the most critical or ambiguous scenarios.

cs.SE↗

Measuring Uncertainty in Transformer Circuits with Effective Information Consistency

Mechanistic interpretability has identified functional subgraphs within large language models (LLMs), known as Transformer Circuits (TCs), that appear to implement specific algorithms. Yet we lack a formal, single-pass way to quantify when an active circuit is behaving coherently and thus likely trustworthy. Building on prior systems-theoretic proposals, we specialize a sheaf/cohomology and causal emergence perspective to TCs and introduce the Effective-Information Consistency Score (EICS). EICS combines (i) a normalized sheaf inconsistency computed from local Jacobians and activations, with (ii) a Gaussian EI proxy for circuit-level causal emergence derived from the same forward state. The construction is white-box, single-pass, and makes units explicit so that the score is dimensionless. We further provide practical guidance on score interpretation, computational overhead (with fast and exact modes), and a toy sanity-check analysis. Empirical validation on LLM tasks is deferred.

cs.LG↗

Sheaf-Theoretic Causal Emergence for Resilience Analysis in Distributed Systems

Distributed systems often exhibit emergent behaviors that impact their resilience (Franz-Kaiser et al., 2020; Adilson E. Motter, 2002; Jianxi Gao, 2016). This paper presents a theoretical framework combining attributed graph models, flow-on-graph simulation, and sheaf-theoretic causal emergence analysis to evaluate system resilience. We model a distributed system as a graph with attributes (capturing component state and connections) and use sheaf theory to formalize how local interactions compose into global states. A flow simulation on this graph propagates functional loads and failures. To assess resilience, we apply the concept of causal emergence, quantifying whether macro-level dynamics (coarse-grained groupings) exhibit stronger causal efficacy (via effective information) than micro-level dynamics. The novelty lies in uniting sheaf-based formalization with causal metrics to identify emergent resilient structures. We discuss limitless potential applications (illustrated by microservices, neural networks, and power grids) and outline future steps toward implementing this framework (Lake et al., 2015).

eess.SY↗