* First attempt at GHA CI instead of Azure Pipelines. * Could I have the matrix config wrong? * Oh, oops, I was just looking in the wrong place for the GHA log. * Disable Azure Pipelines Linux job. * Fix some additional small syntax errors. * Minor updates. * Try to figure out what is happening with this test! * Ha, actually remember the argument. * Try to parse the CTest results. * Don't build CLI programs for the debug-only mlpack_test build. * Try to figure out why we don't have the CTest output file. * Seems like maybe the test just needs more iterations. * Fix directory for CTest output. * Try running mlpack_test manually to see if we get better test output. * Remove version numbers and other non-XML from XML output test file. * Try to fix debug build. * Filtering regex. * Try to fail the build when the tests fail. * Minor fixes and cleanups to run Python binding tests. * Print test failure information in the actual build step. * Update Julia testing strategy. * Try to also get outputs from other test bindings. * Fix Julia and Go testing blocks. * Avoid random collisions that cause test failures. * Try moving splitting into the trial loop. * Try a different test reporter action. * Install pytest for Python tests. * Fix working directory for Julia test. * Add test case that will fail to see what it looks like. * Back to the old action, the new one I tried didn't do what I needed. * Try a better way to print the test output. * I still don't know how to mark the job failed yet. * Maybe this stdout character works better. * Try a forked test-summary action. * Try to use distribution version of the modified action. * Try to see if ccache is doing anything. * Always run ccache stats step. * Turn off PCH for ccache. * Let's just see what it looks like with more error output. * Clean up ccache configuration. * Remove fake test because it seems like now I have the output working right. * Try to reduce file size by not printing debug output into it. * Try and refactor binding environment seutp into a composite local action. * Try to fix some syntactical errors. * Refactor test running into a separate composite action. * Start trying to make the R build part of the main CI build. * Oops, the include directory was in the wrong place. * Try actually running R tests in the workflow too. * Try to fix a few action issues. * Try to work around libicu issue. * Remove unnecessary option. * Try to run OS X builds too. * Try to fix matrix configuration. * Revert "Try to fix matrix configuration." This reverts commit 22a08d6dcd0d39de0a27f2df16e885bd68e18fda. * Oops, maybe it was just an extra comma. * Fix naming, and maybe we have to remove stringi first. * Try to set Julia location correctly. * Some fixes to the pipelines. * Oops, I need root for that. * Some additional fixes for jobs. * Hopefully fix a few more builds. * What, how did that get there? * Don't use non-portable -i option with sed. * Run R tests properly to get junit output. * Oops, I have to redirect the output. * Where is the output file? * Try to fix macOS Python location. * Oops, ROOTDIR is just not set. * Try to store the name of the R package correctly. * Oops, we need to specify pip. * Try to get the right directory for the R package. * Try to fix another round of issues. I'm getting closer, at least. * Okay, fine, let's do it the brutalistic way. * What is actually happening with the Go binding tests? * Try a different strategy for printing. * Make some updates to prepare for Windows builds. * What's the output of go test? * Remove -Wall as per golang/go#6883. * Try to fix duplicate library warning. * Maybe I just specified the variable in an invalid way that didn't stick? * This should fix the Go build. * How did that cd go missing? * Try to remove DEBUG output from the debug tests. * Try to be more specific about the ccache key. * Oops, incorrect variable name. * Try to speed up the testing step. * Try to run the Windows build through GHA. * Try to install locally to deps/. * Try to install and set up Windows dependencies correctly. * Try to debug Windows build. * Use bash for CMake configuration. * Some additional path fixes. * A couple fixes, but I don't have the Armadillo path right. * Fix curl command. * Maybe this is closer to the right path... * Okay, look one directory deeper... * Temporarily print STB compilation failure. * Try to use absolute path for STB inclusion test. * Oops, fix syntax. * Try using REALPATH instead. * Try just skipping the check... * Try specifying BLAS and LAPACK locations directly. * Try to figure out what is going on. * Are we even using the correct FindArmadillo script? * Try specifying the libraries manually. * Try disabling the wrapper library. * What version of CMake is this? * Print the configuration. * What if we specify these library locations? * Double-check: this should fail. * Okay, that actually surprisingly did not fail. * Fix style issues. * Compress into only one GHA file so all the jobs show up in the same place. * Try to use absolute path to OpenBLAS. * Try to fix path mangling on Windows. * Fix STB path and no PS please. * Try to fix output on Windows tests. * Call test from the right place. * Remove some unnecessary output. * Try just running the test. * What is going on, why doesn't the test run? * Try just running the test with PowerShell. * Maybe I have the filename wrong? * PowerShell makes me angry... * Right, .exe is the suffix... * I love PowerShell! * Okay, forget PowerShell. * Can we even run the test? * Is it a DLL search path issue? * Okay, I think this might work for Windows. * Fix porting of new r2u/p3m workflow (hopefully). * Fix possibly incorrect variable name. * Reduce the number of macOS jobs. * Fix name of job. * Remove potentially unnecessary step. * Document CI changes and reorganize jobs for better viewing. * Remove invalid link. (Awesome that the link check build picked this up.)
a fast, header-only machine learning library
Home | Download | Documentation | Help |
Download: current stable version (4.5.1)
mlpack is an intuitive, fast, and flexible header-only C++ machine learning library with bindings to other languages. It is meant to be a machine learning analog to LAPACK, and aims to implement a wide array of machine learning methods and functions as a "swiss army knife" for machine learning researchers.
mlpack's lightweight C++ implementation makes it ideal for deployment, and it can also be used for interactive prototyping via C++ notebooks (these can be seen in action on mlpack's homepage).
In addition to its powerful C++ interface, mlpack also provides command-line programs, Python bindings, Julia bindings, Go bindings and R bindings.
Quick links:
- Quickstart guides: C++, CLI, Python, R, Julia, Go
- mlpack homepage
- mlpack documentation
- Examples repository
- Tutorials
- Development Site (Github)
mlpack uses an open governance model and is fiscally sponsored by NumFOCUS. Consider making a tax-deductible donation to help the project pay for developer time, professional services, travel, workshops, and a variety of other needs.
0. Contents
- Citation details
- Dependencies
- Installation
- Usage from C++
- Building mlpack's test suite
- Further resources
1. Citation details
If you use mlpack in your research or software, please cite mlpack using the citation below (given in BibTeX format):
@article{mlpack2023,
title = {mlpack 4: a fast, header-only C++ machine learning library},
author = {Ryan R. Curtin and Marcus Edel and Omar Shrit and
Shubham Agrawal and Suryoday Basak and James J. Balamuta and
Ryan Birmingham and Kartik Dutt and Dirk Eddelbuettel and
Rishabh Garg and Shikhar Jaiswal and Aakash Kaushik and
Sangyeon Kim and Anjishnu Mukherjee and Nanubala Gnana Sai and
Nippun Sharma and Yashwant Singh Parihar and Roshan Swain and
Conrad Sanderson},
journal = {Journal of Open Source Software},
volume = {8},
number = {82},
pages = {5026},
year = {2023},
doi = {10.21105/joss.05026},
url = {https://doi.org/10.21105/joss.05026}
}
Citations are beneficial for the growth and improvement of mlpack.
2. Dependencies
mlpack requires the following additional dependencies:
If the STB library headers are available, image loading support will be available.
If you are compiling Armadillo by hand, ensure that LAPACK and BLAS are enabled.
3. Installation
Detailed installation instructions can be found on the Installing mlpack page.
4. Usage from C++
Once headers are installed with make install, using mlpack in an application
consists only of including it. So, your program should include mlpack:
#include <mlpack.hpp>
and when you link, be sure to link against Armadillo. If your example program
is my_program.cpp, your compiler is GCC, and you would like to compile with
OpenMP support (recommended) and optimizations, compile like this:
g++ -O3 -std=c++17 -o my_program my_program.cpp -larmadillo -fopenmp
Note that if you want to serialize (save or load) neural networks, you should
add #define MLPACK_ENABLE_ANN_SERIALIZATION before including <mlpack.hpp>.
If you don't define MLPACK_ENABLE_ANN_SERIALIZATION and your code serializes a
neural network, a compilation error will occur.
Warning: older versions of OpenBLAS (0.3.26 and older) compiled to use pthreads may use too many threads for computation, causing significant slowdown. OpenBLAS versions compiled with OpenMP do not suffer from this issue. See the test build guide for more details and simple workarounds.
See also:
- the test program compilation section of the installation documentation,
- the C++ quickstart, and
- the examples repository repository for
some examples of mlpack applications in C++, with corresponding
Makefiles.
4.1. Reducing compile time
mlpack is a template-heavy library, and if care is not used, compilation time of a project can be very high. Fortunately, there are a number of ways to reduce compilation time:
-
Include individual headers, like
<mlpack/methods/decision_tree.hpp>, if you are only using one component, instead of<mlpack.hpp>. This reduces the amount of work the compiler has to do. -
Only use the
MLPACK_ENABLE_ANN_SERIALIZATIONdefinition if you are serializing neural networks in your code. When this define is enabled, compilation time will increase significantly, as the compiler must generate code for every possible type of layer. (The large amount of extra compilation overhead is why this is not enabled by default.) -
If you are using mlpack in multiple .cpp files, consider using
extern templatesso that the compiler only instantiates each template once; add an explicit template instantiation for each mlpack template type you want to use in a .cpp file, and then useexterndefinitions elsewhere to let the compiler know it exists in a different file.
Other strategies exist too, such as precompiled headers, compiler options,
ccache, and others.
5. Building mlpack's test suite
See the installation instruction section.
6. Further Resources
More documentation is available for both users and developers.
To learn about the development goals of mlpack in the short- and medium-term future, see the vision document.
If you have problems, find a bug, or need help, you can try visiting
the mlpack help page, or mlpack on
Github. Alternately, mlpack help can be
found on Matrix at #mlpack; see also the
community page.