| BuildStreaM Pipeline Architecture and API Enhancements |
Enhanced pipeline architecture and API capabilities with resume and retry functionality, pipeline decomposition, dynamic child pipeline generation, image group lifecycle tracking, manual cleanup operations, and PowerScale S3 backend support. For more details, see Path D: BuildStreaM Automated Deployment. |
| BMC Discovery via Dell OpenManage Enterprise |
Automated BMC discovery via Dell OpenManage Enterprise (OME) with paginated API queries, automatic extraction of service tags and iDRAC details, Scalable Unit extraction, timestamped file generation, and OME group mapping. For more details, see BMC Discovery. |
| Multi-Subnet DHCP for Rack-Based Provisioning |
Multi-subnet DHCP configuration for rack-based network provisioning with per-rack /24 subnet assignment, CoreDHCP multi-subnet configuration generation, and CoreDNS forward and reverse zone generation. For more details, see Configure Multi-Subnet DHCP. |
| CoreDNS-Based Hostname Resolution for Slurm and MPI |
Dynamic DNS resolution powered by coresmd replacing static /etc/hosts file management with automatic hostname resolution, real-time inventory updates from OpenCHAMI SMD, cloud-init based /etc/resolv.conf configuration, and Kubernetes CoreDNS forwarding. For more details, see Configure Cluster DNS. |
| Vector Telemetry Pipeline for Data Routing |
Vector high-performance data pipeline for collecting, transforming, and routing telemetry data from LDMS and OME sources to VictoriaMetrics and VictoriaLogs with dedicated write-buffer components. For more details, see Telemetry Architecture. |
| PowerScale Telemetry for Storage Monitoring |
PowerScale Telemetry for comprehensive storage observability collecting storage performance metrics and logs from Dell PowerScale storage nodes with CSM Metrics, OpenTelemetry Collector, and CSI Driver integration. For more details, see Configure PowerScale Telemetry. |
| UFM Telemetry to VictoriaMetrics |
UFM (Unified Fabric Manager) telemetry collection for InfiniBand fabric monitoring through vmagent scraping with secure HTTPS, TLS certificate validation, and dual-destination forwarding to local and remote VictoriaMetrics clusters. For more details, see Configure UFM (NVIDIA Unified Fabric Manager) Telemetry. |
| VAST Storage Telemetry Integration |
VAST storage telemetry integration through VMagent scraping of VAST Prometheus endpoints and VLAgent syslog log collection with secure HTTPS, TLS authentication, and dual-destination forwarding. For more details, see Configure VAST (VAST Storage) Telemetry. |
| External Log Aggregation to VictoriaLogs |
Centralized log collection from external sources including network devices, storage systems, and fabric managers through VLAgent with syslog (plaintext/TLS) and HTTP forwarding support, TLS certificate validation, and JSON Lines format ingestion. For more details, see Collect Logs from External Clients to VictoriaLogs. |
| Configurable Pod Replicas for vmagent and vlagent |
Configurable replica counts for vmagent and vlagent pods with default value of 2 replicas each, providing improved availability and scalability for telemetry data collection and log aggregation. For more details, see Telemetry Configuration. |
| Minimal OS Functional Groups |
Minimal OS functional groups (os_x86_64 and os_aarch64) providing a clean operating system baseline designed specifically for downstream platform software installation. For more details, see Create a Mapping File. |
| Unattended OS Installation via iDRAC Virtual Media |
Automated operating system deployment on iDRAC-enabled servers by mounting the OS image as virtual media, removing the need for manual ISO handling and enabling unattended provisioning workflows. For more details, see Unattended OS Installation. |
| Custom Cloud-Init for Post-Boot Mounts and Scripts |
Added support for injecting custom cloud-init directives during node provisioning to enable boot-time node customization without modifying platform-managed templates. Administrators can apply configurations globally to all nodes (common) or to specific functional groups (groups). Supported cloud-init directives include write_files for creating or modifying files and runcmd for executing setup commands during node initialization. For more information, see Configure Additional Cloud-Init. |
| NVIDIA DCGM and CUDA Toolkit Provisioning for Slurm GPU Nodes |
End-to-end automated GPU readiness for Slurm clusters with NVIDIA driver installation, CUDA toolkit distribution to shared cluster storage, and DCGM setup during stateless node provisioning. For more details, see Set Up Slurm. |
| NVIDIA HPC SDK Provisioning for Slurm Clusters |
Cluster-wide deployment of NVIDIA HPC SDK (nvhpc) for Slurm compiler and compute nodes with single installation on compiler node, NFS sharing, and bind mount distribution. For more details, see NVIDIA HPC SDK Setup. |
| Vast Repo and Vast Client Installation |
Vast NFS client installation streamlined by building Vast repository from source, hosting RPMs on HTTP server, configuring repository, and automatic installation during provisioning when InfiniBand NIC is present. For more details, see Configure VAST Storage. |
| One-Shot Combined Log Extraction for Debugging |
One-shot log collection playbook for gathering cluster logs from Kubernetes and Slurm nodes with full and curated support collection modes, log collection from all node types, and timestamped tar.gz bundle output. For more details, see Log Management. |
| ETCD on Local Disk Support for Kubernetes Service Cluster |
ETCD deployment on local disk instead of NFS for Kubernetes service cluster with configurable etcd_on_local_disk setting in omnia_config.yml, automatic disk selection prioritizing BOSS cards (BOSS-N1/N2) with fallback to SSD/SATA disks, /var/lib/etcd mount point, support for pre-configured RAID 1 or RAID 10 on BOSS cards, and minimum 20 GB disk space recommendation. For more details, see Configure Kubernetes HA. |
| Unattended OS Installation via iDRAC Virtual Media |
Automated bare-metal OS installation using iDRAC Virtual Media with NFS-based Kickstart, one server at a time. Supports aarch64 nodes via install_os_arm_node.yml orchestrator and generic x86_64 nodes via install_os.yml. For more details, see Unattended OS Installation via iDRAC. |