Publications

Preprint arXiv 2026

COMPASS: A Unified Decision-Intelligence System for Navigating Performance Trade-off in HPC

Ankur Lahiry, Banooqa Banday, Yugesh Bhattarai, Tanzima Z Islam, and Mohammad Zaeed

Decision Intelligence HPC Explainability Uncertainty Ranking

Decision-intelligence framework for HPC configuration optimization using uncertainty-aware ranking and explainable recommendations.

Key Contribution: Introduced an uncertainty-aware decision-intelligence framework for HPC with up to 100× faster training and 80× faster inference than generative baselines.

  • Scaled to 1.3B samples (126 GB) for high-volume HPC traces.
  • Balanced performance, cost, and reliability in a single recommendation pipeline.
  • Delivered explainable outputs with uncertainty-aware ranking for system decisions.
Workshop Euro-par 2025

A Distributed Framework for Causal Modeling of Performance Variability in GPU Traces

Ankur Lahiry, Ayush Pokharel, Banooqa Banday, Seth Ockerman, Amal Gueroudji, Mohammad Zaeed, Tanzima Z Islam, and Line Pouchard

GPU Causal Inference HPC Performance Variability

Distributed causal analysis framework for GPU traces that identifies root causes of performance variability in HPC workloads.

Key Contribution: Framed large-scale GPU performance variability analysis as a distributed causal inference problem for root-cause discovery.

  • Modeled performance variability in GPU traces with causal structure rather than correlation alone.
  • Scaled analysis across large trace collections for HPC workflows.
  • Targeted root-cause identification for recurring performance anomalies.
Poster SSDBM 2025

Scalable GPU Performance Variability Analysis Framework

Ankur Lahiry, Ayush Pokharel, Seth Ockerman, Amal Gueroudji, Line Pouchard, and Tanzima Z Islam

GPU Distributed Systems HPC Trace Analysis

Distributed GPU trace-analysis framework that reduces memory overhead and improves scalability for large HPC log pipelines.

Key Contribution: Improved GPU trace-analysis scalability by 67% through distributed partitioning and parallel processing.

  • Reduced analysis time for repeated GPU trace workloads.
  • Lowered memory overhead on large trace datasets.
  • Enabled faster detection of stalls, variability, and bottlenecks.
Preprint arXiv 2025

WANDER: An Explainable Decision-Support Framework for HPC

Ankur Lahiry, Banooqa Banday, Yugesh Bhattarai, and Tanzima Z Islam

Explainable AI HPC Anomaly Detection Graph Learning

Explainable HPC decision-support framework that turns complex system logs into interpretable signals for early anomaly detection.

Key Contribution: Unified graph-based modeling and explainable AI for earlier and more interpretable diagnosis of HPC performance anomalies.

  • Converted complex HPC logs into actionable diagnostic signals.
  • Improved interpretability for anomaly detection decisions.
  • Supported earlier diagnosis of unusual large-scale system behavior.
Research Paper COMPSAC'24 2024

Reimagine Application Performance as a Graph: Novel Graph-Based Method for Performance Anomaly Classification in High-Performance Computing

Chase Phelps, Ankur Lahiry, Tanzima Z Islam, and Line C Pouchard

Graph Learning HPC Anomaly Detection Performance Analytics

Graph-based application-performance representation for anomaly classification in HPC using graph neural networks.

Key Contribution: Showed that graph-structured performance modeling improves anomaly classification in large-scale HPC systems.

  • Represented tasks and resources as structured graph relationships.
  • Applied graph neural networks to HPC anomaly classification.
  • Improved expressiveness over flat performance feature views.
Research Paper ICMLA'23 2023

Novel Representation Learning Technique Using Graphs for Performance Analytics

Tarek Ramadan, Ankur Lahiry, and Tanzima Z Islam

Graph Learning Performance Prediction HPC Representation Learning

Graph representation learning method that transforms tabular HPC performance data into graphs for performance prediction.

Key Contribution: Converted tabular HPC performance data into graph structure to improve representation learning for regression tasks.

  • Mapped tabular performance records into graph form for richer dependencies.
  • Applied graph neural network reasoning to performance analytics.
  • Targeted execution-time prediction and related regression problems.
Poster SC'23 2023

Graph Based Anomaly Detection in Chimbuko: Feasible or Fallible?

Chase Phelps, Ankur Lahiry, Tanzima Z Islam, and Christopher Kelly

HPC Anomaly Detection Graph Learning Performance Analytics

Evaluation of graph-based anomaly detection for Chimbuko performance analytics in large-scale supercomputing applications.

Key Contribution: Assessed the feasibility of graph-based deep learning for anomaly classification inside the Chimbuko HPC analytics framework.

  • Tested graph-based anomaly classification in an operational HPC analytics pipeline.
  • Focused on monitoring efficiency in large-scale supercomputing applications.
  • Surfaced practical trade-offs of graph methods for real HPC anomaly workloads.