That significantly depends on the HDR compositing method used. Most methods also have to account for things like ghosting and use a weighted blend in the transition zone going from one exposure to the next. When blending with exposures that differ more than 2 stops, the increased noise contribution of the lower exposure will show, depending on the HDR compositing or exposure blending strategy.
Thanks to sensor linearity, the transition zone can be made negligible as long as the relative exposures between adjacent shots are properly calculated. Unluckily most programs just look at EXIF data to find out the relative exposure, or worse, allow the user to set a EV value, therefore making progressive transitions necessary to prevent visible gaps. Funily the most accurate source for relative exposure calculation is not the metadata nor the user, but the image data itself.
This HDR composite was done with 0 transition zones. Just correcting adjacent shots by 2EV (as the EXIF data indicated and an ingenuous user would suggest) resulted in a visible exposure gap, but an accurate relative exposure calculation of 1.93EV (just based on image pixels) provided a seamless result:

The relative exposure histograms explain why (look at the 1.93EV calculated in the first histogram):

With a precise calculation, strongly fragmented non-progressive fusion maps such as this one can be used:

Without any risk of visible exposure gaps in the output image:

Regarding the number of shots, with current sensors there is no reason to bracket less than 3EV apart to prevent noise gaps. I prefer much more a {0, +3, +6} bracket than the classic {0, +2, +4}. Assuming both series shared the same ETTR'ed 0EV shot, the +6 shot would display about 4 times less noise than the +4 shot in the deep shadows (where read noise prevails at 6dB/EV). Unfortunately one cannot do the first bracketing in most Canons without touching the camera or using some external triggering. Bravo for Canon!.
Regards