Testing framework and monitoring system for the ATLAS EventIndex

EPJ Web of Conferences EDP Sciences 295 (2024) 01047-01047

Authors:

Elizaveta Cherepanova, Elizabeth J Gallas, Fedor Prokoshin, Miguel Villaplana Pérez

Abstract:

The ATLAS EventIndex is a global catalogue of the events collected, processed or generated by the ATLAS experiment. The system was upgraded in advance of LHC Run 3, with a migration of the Run 1 and Run 2 data from HDFS MapFiles to HBase tables with a Phoenix interface. Two frameworks for testing functionality and performance of the new system have been developed. There are two types of tests running. First, the functional test that must check the correct functioning of the import chain. These tests run event picking over a random set of recently imported data to see if the data have been imported correctly, and can be accessed by both the CLI and the PanDA client. The second, the performance test, generates event lookup queries on sets of the EventIndex data and measures the response times. These tests enable studies of the response time dependence on the amount of requested data, and data sample type and size. Both types of tests run regularly on the existing system. The results of the regular tests as well as the statuses of the main EventIndex subsystems (services health, loaders status, filesystem usage, etc.) are sent to InfluxDB in JSON format via HTTP requests and are displayed on Grafana monitoring dashboards. In case (part of) the system misbehaves or becomes unresponsive, alarms are raised by the monitoring system

The ATLAS Alarm Helper

EPJ Web of Conferences EDP Sciences 295 (2024) 02014-02014

Authors:

Florian Haslbeck, Carlos Solans Sánchez, Ignacio Asensi Tortajada, Daniela Bortoletto, André Rummler, Gustavo A Uribe

Abstract:

The Detector Safety System is the last line of defence to protect the ATLAS detector against abnormal and potentially even unforeseen situations. It is designed to return the detector to a safe state based on predefined actions triggered by alarms which are triggered on their part by specific sets of conditions. Every alarm whether it results in an action taken or not is followed up by the operations team that assesses the criticality, takes countermeasures and identifies the point of failure. From experience abnormal situations can result either from faults or from side effects of planned interventions which were either not properly identified despite the mandatory planning and review or where a mistake during execution occurred. In many cases there are multiple interventions ongoing simultaneously in order to profit from shutdown periods. The rapid analysis of alarms while the incident is ongoing is often complicated due to the complexity of the ATLAS detector and its infrastructure and the large number of responsible groups and experts. A new Alarm Helper tool was designed to assist the operation team, particularly the operator in the control room responsible for infrastructure and safety (SLIMOS – Shift Leader in Matters of Safety), by providing real-time information about ongoing interventions and the possible related causes of failure. The new tool will combine historical events, documentation, and limited knowledge about ongoing interventions. It extends the Expert System which visualizes and simulates infrastructure inter-dependencies and allows to trace faults or alarms to a list of potential points of failure. The new tool also proposes which experts should be contacted in the particular circumstances.

Towards a new conditions data infrastructure in ATLAS

EPJ Web of Conferences EDP Sciences 295 (2024) 01013-01013

Authors:

Evgeny Alexandrov, Luca Canali, Davide Costanzo, Andrea Formica, Elizabeth J Gallas, Mikhail Mineev, Nurcan Ozturk, Shaun Roe, Vakho Tsulaia, Marcelo Vogel

Abstract:

The ATLAS experiment is preparing a major change in the conditions data infrastructure in view of LHC Run 4. In this paper we describe the ongoing changes in the database architecture which have been implemented for Run 3, and describe the motivations and the on-going developments for the deployment of a new system (called CREST for Conditions Representational State Transfer, as a reference to REST architectures). The main goal is to set up a parallel infrastructure for full scale testing before the end of Run 3

Improving topological cluster reconstruction using calorimeter cell timing in ATLAS

European Physical Journal C Springer Nature 84:5 (2024) 455

A search for R-parity-violating supersymmetry in final states containing many jets in pp collisions at s = 13 TeV with the ATLAS detector

Journal of High Energy Physics Springer 2024:5 (2024) 3

Authors:

G Aad, B Abbott, K Abeling, NJ Abicht, SH Abidi, A Aboulhorma, H Abramowicz, H Abreu, Y Abulaiti, BS Acharya, C Adam Bourdarios, L Adamczyk, SV Addepalli, MJ Addison, J Adelman, A Adiguzel, T Adye, AA Affolder, Y Afik, MN Agaras, J Agarwala, A Aggarwal, C Agheorghiesei, A Ahmad

Abstract:

A search for R-parity-violating supersymmetry in final states with high jet multiplicity is presented. The search uses 140 fb−1 of proton-proton collision data at s = 13 TeV collected by the ATLAS experiment during Run 2 of the Large Hadron Collider. The results are interpreted in the context of R-parity-violating supersymmetry models that feature prompt gluino-pair production decaying directly to three jets each or decaying to two jets and a neutralino which subsequently decays promptly to three jets. No significant excess over the Standard Model expectation is observed and exclusion limits at the 95% confidence level are extracted. Gluinos with masses up to 1800 GeV are excluded when decaying directly to three jets. In the cascade scenario, gluinos with masses up to 2340 GeV are excluded for a neutralino with mass up to 1250 GeV.