Resource Consumption and Development Effort in GPU-Optimized Industrial Anomaly Detection Inference: A Comparison between PyTorch and a C++/CUDA Pipeline