abstract blue wave
abstract blue wave

Intelligence Case Study

DRIVE: Dynamic Resource Integration for Versatile Execution

of imagery processed in 17 minutes
more model setup tests per minute than three traditional HPC servers
hours per day of untapped compute, reclaimed

Across agencies and use cases, artificial intelligence and machine learning (AI/ML) workflows have driven an increase in demand for compute resources. At the same time, budgets are contracting – particularly in the domains of cloud and on-premise computing – and cost efficiency has become even more of a priority.

This increased demand and supply chain constraints have caused the expensive GPUs required to power high-performance computing (HPC) environments that many AI/ML projects need more difficult to procure. Tasks enabled by these machines, like data-intensive processing or complex model tuning, are precisely the technologies that have the potential to make the most transformative impacts on missions.

It’s against this backdrop that mission partners proactively look for ways to help customers manage resources, talent and costs in creative and efficient ways. This was the impetus for the Dynamic Resource Integration for Versatile Execution (DRIVE) proof of concept GDIT recently completed alongside a long-time intelligence customer. This customer had thousands of thick client workstations with powerful GPUs that were underutilized – meaning, they had the capacity to operate 24 hours per day but were only being used during traditional working hours.

Our team came to the customer with an idea: What if we could leverage their existing hardware in its idle state to expand processing power? We could do it by automatically detecting when a workstation with a GPU was idle and then registering that workstation to become part of a distributed computing network.

As an example, on an image inference task, a sensor would deposit a large amount of raw imagery into an environment. The workload would be split into hundreds of jobs, each with a different set of images to work on. Dozens of idle workstation GPUs would run the jobs independently and then send annotated images back to a central location. The otherwise dormant machines would expand the processing power and complete the task faster, with minimal new hardware requirements.

The team came up with the DRIVE concept while expanding technology capabilities for the customer. We saw the requirements for GPU-equipped workstations along with the demand for AI/ML resources, and an opportunity to leverage those machines for compute-intensive tasks. In our view, this would greatly enhance the customer’s resources, allowing them to meet their mission with greater efficiency and speed. We believed in the idea enough to pilot it with our own resources, showing the customer the art of the possible before deploying to their environment.

In practice, the DRIVE concept also helps tackle another major problem confronting customers today: forecasting infrastructure while expanding use of AI/ML and HPC. Customers know they’ll need additional resources but find it hard to predict how much they will need and when. Using existing resources and hardware can give them the capacity to perform more AI/ML initiatives while also fine tuning their demand forecasts for the future.

Initial results of the DRIVE pilot are promising. For a computer vision inference project, the customer processed 8TB of imagery in just 17 minutes, using 500 machines. On a model optimization exercise, 400 DRIVE machines tested more than ten times as many model setups per minute compared to three traditional datacenter HPC servers, enabling the team to find the best-performing model much faster.

The DRIVE program has zero user impacts because it only leverages logged-out workstations, terminating workloads immediately upon user login. This allows for consistent login times between DRIVE and non-DRIVE workstations, making them indistinguishable from one another. A security-focused architecture ensures sensitive customer data is not inadvertently exposed to the user, and exemption lists and “neighborhood” limits enable the customer to manage power demand and network bandwidth.

DRIVE is an unconventional HPC managed service, but it is exactly the kind of creative and responsive model that today’s technological and economic landscapes require. This brand of thinking helps customers tackle mission problems more efficiently. And, it's the right thing to do for the customer and for the mission.


Learn more about how GDIT collaborates with customers across a variety of HPC use cases.