A question and a couple of thoughts, not scientifically tested at all:
Question: how far are you downsampling and still sensing a difference?
1) 2x linear (4x in pixel count) is the minimum for the JPEG or TIFF or whatever to have even one pixel of each color in the collection used to produce each output pixel, going the lower pixel count image has it's pixel count reduced by les than 4x, I can see why it might have less sharpeners.
2) In practice, demosaicing algorithms use a weighted average of data from many nearby pixels, not just the nearest ones of the needed color, so could it be that the "footprint" of each output pixel covers date for, say three or four photosites in each direction, which would mean that only when one goes past about 6x to 8x linear downsampling (36x to 64x pixel count reduction) is the smearing due to this process completely gone?
3) Downsampling can increase the SNR of the output pixels, so could downsampling from more sensor pixels give more local contrast and thus perceived sharpness in some cases?
P.S. in the experiment you describe, of a 12MP RGB output file format (so R, G, B values at each of 12 million locations) from a 12MP Bayer CFA camera (6 million green, 3 million each red and blue) and a 24MP camera (12 million green, 6 million each or red and blue) it would not at all surprise me that the latter gives more resolution, because the 12MP RGB output file format can hold more information than either sensor delivers, and specifically more luminosity information than the 12MP sensor gives, since that is based mostly or entirely on values from green pixels. Also, in each case, the output luminosity values rely on raw data from locations other than that of the output pixel: even the 12 million green pixels of the 24MP sensor are not at the same places as those of the output file.