Not only that, but I also want cloud movement to match that of the water movement (4 minutes of water movement vs 30 seconds of cloud movement, if bracketing 3 stops, just doesn't work for me).
What you say is interesting. I think you don't need to renounce to having the same movement for both water and clouds even if not using the grad filter. Imagine you use the same aperture and ISO in the following 3 methods:
METHOD 1 (grad filter)
1 shot 4 min with 3 stops ND grad filter (well exposed water and clouds)
METHOD 2 (blend method)
1 shot 4 min (4 min water movement and blown clouds)
+
1 shot 30 secs (underexposed water and 30 secs clouds movement)
METHOD 3 (grad filter simulation)
8 consecutive shots 30 seconds each (water will be 3 stops underexposed with respect to method 1 and none of them will have blown clouds).
Method 3 provides even more flexibility than Method 1. We average the 8 shots in 8 layers in PS (this can even be done linearly using a gamma=1 profile to perfectly mimic camera photon count) and push exposure for the water with a
tone preserving curve and custom layer mask. If the 30 secs shots were taken right one after another you will not notice any time gap (times between shots will be negligible vs 30 secs), so the averaging will look 100% the same as a single 4 min shot both for the sky and clouds.
Regarding IQ, Method 3 will collect the same amount of photons for the water in the 8 shots of 30 secs as in the single 4 min shot with the grad filter in Method 1, so SNR in the water will be the similar (we could do more precise calculations here). And SNR will even be better for the clouds because you will collect the photons rejected by the grad filter in this area.
If the camera provides Liveview, the mirror can stay locked up all the time so the loss of sharpness because of misalignement will be minimised (only the shutter vibration remains).
Regards