Well, admittedly I have no knowledge of Lightroom's internals, but from an end-user's perspective the bottleneck in the Develop Module has always been rendering—especially when making local adjustments, zooming to 1:1, and performing other drawing operations that are computationally expensive. Again, just an impression, but LR otherwise seems to have enough CPU cycles to do what it needs to do on a modern multi-core machine: when it blocks, it almost always seems to be when it is updating the display. I presume the Adobe developers use DTrace and something similar on Windows to determine for what functions greater parallelization would be useful.
I do not agree - perhaps we differ in the definition of rendering. Absolutely, the bottleneck is in processing the image. However, I believe that it is in processing the image prior to display, not actually mapping the processed image to the display. On the other hand, the sensitivity of speed to screen resolution
does suggest that the simply displaying it is a major bottleneck. I've always found that difficult to accept, given the speed with which other software can display an existing in-memory image.
I would suggest that LightRoom may use an internal representation that is fairly easy to process, but difficult to display, but the processing required for local adjustments has clearly been a bottleneck.
For general-purpose processing, OpenGL is rarely a good choice. OpenACC is easy to use, but is not as widely supported as I'd like. CUDA is Nvidia-specific, but provides a reasonable programming environment. OpenCL runs on almost anything, but had rather primitive development tools the last time I looked. (That was several years ago, and is probably no longer the case. I
hope OpenCL has a good development environment and toolchain by now.) I would expect any of these to be preferable to do the sort of processing needed to apply adjustments to an image.