Patent attributes
Improved techniques for processing large-scale data and various large-scale data applications (e.g., large-scale Data Mining (DM), large-scale data analysis (LSDA)) in computing systems (e.g., Data Information Systems, Database Systems) are disclosed. Redundancy-reduced data (RRDS) can be provided as data that can be used more efficiently by various applications, especially, large-scale data applications. In doing so, at least one assumption about the distribution of a multi-dimensional data set (MDDS) and its corresponding set of responses (Y) can be made in order to reduce the multi-dimensional data set (MDDS). For example, a normal distribution (e.g., bell-shape, symmetric) can be assumed and Mutual information of the combination of a multi-dimensional set (X) and its corresponding responses (Y) can be optimized, for example, by using linear transformations, iterative numerical procedures, one or more constraints associated with the at least one assumption, and using one or more Lagrange multipliers to provide a constraint optimization function.