Hello Guillermo,
Lots of questions to which I only have vague comments to, I am learning as you are. I'll give you my matrix perspective so you can give us the NN's.
With current hardware, I think we can assume linearity of the photodiodes/pixels/Raw data to, say, 1 stop below clipping. Assuming perfect Spectral Sensitivity Functions of the image acquisition hardware (i.e. one exact matrix multiplication away from Cone Fundamentals and therefore CIE Color Matching Functions), the system should then be perfectly linear, and in theory all we would need are three patches to fully determine it. 3 patches = 9 readings, 9 equations with 9 unknowns (the coefficients in the matrix), et voilĂ , the one and only precise solution to our problem for the given observer and illuminant.
However, in reality SSFs are not a single matrix multiplication away from Cone Fundamentals, and this introduces non linearities in the system. All we can do with a matrix is then to find a best approximate compromise, and that's why the matrix in this context is properly referred to Compromise Color Matrix. We use more patches in order to build an overdetermined system and 'regress to the mean'. Of course in such situations we have to be mindful of bias and overfitting, as you pointed out. Since the result is just a best compromise, non linear corrections (often LUTs) may become necessary to push down some of the outliers, always keeping bias and overfitting in mind. Your NN performs both functions, linear regression and non-linear corrections, at the same time.
So the result from our overdetermined system is not perfect and may not go through every patch/value: zero and one are just two values like all others. Using 3x3 matrices implies that all three channels will have zero output with zero input. If one were to relax that assumption we could search for 12 coefficients (a 3x4 matrix that allows for an offset) and obtain the best fit that way. I sometimes do it, but typically I want all outputs to go to zero when the input is zero, so I use 3x3 matrices as a matter of course, as I think do most raw converters (and the DNG spec).
In theory the same applies at the clipping end of the color cube. As Oscar will tell you, we could reduce the number of unknowns to 6 by assuming that the 1 end of the cube (white balanced saturation/clipping) occurs at the coordinates of the illuminant (say the XYZ white point or L = 100, a = b = 0), but I generally don't do that because I use it as a rough check on the linearity of the fit (CCT of matrix with [1 1 1] input should be close to that of the actual illuminant) - besides, most landscapes do not need to have more 'perfect' whites at the expense of other tones. In fact, following this line of thinking perhaps I should be using 3x4 matrices...
It would be interesting to see that graph on properly chosen log scales to see deviations from a straight line.
Jack