microFlare - the local optimized LLM

microFlare is a project that leverages asymmetric quantization to compress large models down to a size that can be ran on consumer PCs while maintaining similar quality to the original uncompressed LLM.