From cf0613af7d36130b0c2f77c8692122485a933d96 Mon Sep 17 00:00:00 2001 From: Ryan Curtin Date: Mon, 26 Mar 2018 16:38:10 -0400 Subject: [PATCH 1/4] Remove ambiguity in section names so links work right. --- doc/guide/build.hpp | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/doc/guide/build.hpp b/doc/guide/build.hpp index def9318a53..254c60074d 100644 --- a/doc/guide/build.hpp +++ b/doc/guide/build.hpp @@ -1,6 +1,6 @@ /*! @page build Building mlpack From Source -@section buildintro Introduction +@section build_buildintro Introduction This document discusses how to build mlpack from source. However, mlpack is in the repositories of many Linux distributions and so it may be easier to use the @@ -25,7 +25,7 @@ href="https://keon.io/mlpack/mlpack-on-windows/">Keon's excellent tutorial. Download latest mlpack build from here: mlpack-2.2.5 -@section builddir Creating Build Directory +@section build_builddir Creating Build Directory Once the mlpack source is unpacked, you should create a build directory. @@ -37,7 +37,7 @@ $ mkdir build The directory can have any name, not just 'build', but 'build' is sufficient enough. -@section dep Dependencies of mlpack +@section build_dep Dependencies of mlpack mlpack depends on the following libraries, which need to be installed on the system and have headers present: @@ -60,7 +60,7 @@ On Fedora, Red Hat, or CentOS, these same dependencies can be obtained via dnf: armadillo-devel binutils-devel @endcode -@section config Configuring CMake +@section build_config Configuring CMake Running CMake is the equivalent to running `./configure` with autotools. If you are working with the svn trunk version of mlpack and run CMake with no options, @@ -100,7 +100,7 @@ The full list of options mlpack allows: Each option can be specified to CMake with the '-D' flag. Other tools can also be used to configure CMake, but those are not documented here. -@section build Building mlpack +@section build_build Building mlpack Once CMake is configured, building the library is as simple as typing 'make'. This will build all library components as well as 'mlpack_test'. From c2decaf1cd9e2094ba7c9d95f80f67d25f6b291a Mon Sep 17 00:00:00 2001 From: Ryan Curtin Date: Mon, 26 Mar 2018 16:59:20 -0400 Subject: [PATCH 2/4] Refactor and update build documentation. --- doc/guide/build.hpp | 83 ++++++++++++++++++++++++++++++++------------- 1 file changed, 60 insertions(+), 23 deletions(-) diff --git a/doc/guide/build.hpp b/doc/guide/build.hpp index 254c60074d..299c12f428 100644 --- a/doc/guide/build.hpp +++ b/doc/guide/build.hpp @@ -19,23 +19,43 @@ configuration options. One can consult any of numerous CMake tutorials for further documentation, but this tutorial should be enough to get mlpack built and installed on most Linux and UNIX-like systems (including OS X). If you want to build mlpack on Windows, see Keon's excellent tutorial. +href="https://keon.io/mlpack-on-windows/">Keon's excellent tutorial. -@section Download latest mlpack build -Download latest mlpack build from here: +You can download the latest mlpack release from here: mlpack-2.2.5 +@section build_simple Simple Linux build instructions + +Assuming all dependencies are installed in the system, you can run the commands +below directly to build and install mlpack. + +@code +$ wget http://www.mlpack.org/files/mlpack-2.2.5.tar.gz +$ tar -xvzpf mlpack-2.2.5.tar.gz +$ mkdir mlpack-2.2.5/build && cd mlpack-2.2.5/build +$ cmake ../ +$ make -j4 # The -j is the number of cores you want to use for a build. +$ sudo make install +@endcode + +If the \c cmake \c .. command fails, you are probably missing a dependency, so +check the output and install any necessary libraries. (See \ref build_dep.) + +The instructions above are the simplest way to get, build, and install mlpack. +The sections below discuss each of those steps in further detail and show how to +configure mlpack. + @section build_builddir Creating Build Directory -Once the mlpack source is unpacked, you should create a build directory. +First we should unpack the mlpack source and create a build directory. @code +$ tar -xvzpf mlpack-2.2.5.tar.gz $ cd mlpack-2.2.5 $ mkdir build @endcode -The directory can have any name, not just 'build', but 'build' is sufficient -enough. +The directory can have any name, not just 'build', but 'build' is sufficient. @section build_dep Dependencies of mlpack @@ -46,56 +66,69 @@ system and have headers present: - Boost (math_c99, program_options, serialization, unit_test_framework, heap, spirit) >= 1.49 +For Python bindings, the following packages are required: + + - setuptools + - cython >= 0.24 + - numpy + - pandas >= 0.15.0 + In Ubuntu and Debian, you can get all of these dependencies through apt: @code # apt-get install libboost-math-dev libboost-program-options-dev libboost-test-dev libboost-serialization-dev libarmadillo-dev binutils-dev + python-pandas python-numpy python-cython python-setuptools @endcode On Fedora, Red Hat, or CentOS, these same dependencies can be obtained via dnf: @code # dnf install boost-devel boost-test boost-program-options boost-math - armadillo-devel binutils-devel + armadillo-devel binutils-devel python3-Cython python3-setuptools + python3-numpy python3-pandas @endcode @section build_config Configuring CMake Running CMake is the equivalent to running `./configure` with autotools. If you -are working with the svn trunk version of mlpack and run CMake with no options, -it will configure the project to build with debugging symbols and profiling -information: If you are working with a release of mlpack, running CMake with no -options will configure the project to build without debugging or profiling -information (for speed). +run CMake with no options, it will configure the project to build without +debugging or profiling information (for speed). @code $ cd build $ cmake ../ @endcode -You can manually specify options to compile with or without debugging -information and profiling information (i.e. as fast as possible): +You can manually specify options to compile with debugging information and +profiling information (useful if you are developing mlpack): @code $ cd build -$ cmake -D DEBUG=OFF -D PROFILE=OFF ../ +$ cmake -D DEBUG=ON -D PROFILE=ON ../ @endcode The full list of options mlpack allows: - - DEBUG=(ON/OFF): compile with debugging symbols (default ON in svn trunk, OFF - in releases) - - PROFILE=(ON/OFF): compile with profiling symbols (default ON in svn trunk, - OFF in releases) + - DEBUG=(ON/OFF): compile with debugging symbols (default OFF) + - PROFILE=(ON/OFF): compile with profiling symbols (default OFF) - ARMA_EXTRA_DEBUG=(ON/OFF): compile with extra Armadillo debugging symbols (default OFF) - BUILD_TESTS=(ON/OFF): compile the \c mlpack_test program (default ON) - BUILD_CLI_EXECUTABLES=(ON/OFF): compile the mlpack command-line executables (i.e. \c mlpack_knn, \c mlpack_kfn, \c mlpack_logistic_regression, etc.) (default ON) + - BUILD_PYTHON_BINDINGS=(ON/OFF): compile the bindings for Python, if the + necessary Python libraries are available (default ON except on Windows) + - BUILD_SHARED_LIBRARIES=(ON/OFF): compile shared libraries as opposed to + static libraries (default ON) - TEST_VERBOSE=(ON/OFF): run test cases in \c mlpack_test with verbose output (default OFF) + - MATHJAX=(ON/OFF): use MathJax for generated Doxygen documentation (default + OFF) + - FORCE_CXX11=(ON/OFF): assume that the compiler supports C++11 instead of + checking; be sure to specify any necessary flag to enable C++11 as part + of CXXFLAGS (default OFF) Each option can be specified to CMake with the '-D' flag. Other tools can also be used to configure CMake, but those are not documented here. @@ -113,6 +146,9 @@ src/mlpack/CMakeFiles/mlpack.dir/core/optimizers/aug_lagrangian/aug_lagrangian_t <...> @endcode +It's often useful to specify \c -jN to the \c make command, which will build on +\c N processor cores. That can accelerate the build significantly. + You can specify individual components which you want to build, if you do not want to build everything in the library: @@ -139,21 +175,22 @@ and submit an issue and the mlpack developers will quickly help you figure it out: http://mlpack.org/ + http://github.com/mlpack/mlpack Alternately, mlpack help can be found in IRC at \#mlpack on irc.freenode.net. @section install Installing mlpack -If you wish to install mlpack to /usr/include/mlpack/ and /usr/lib/ and -/usr/bin/, once it has built, make sure you have root privileges (or write -permissions to those two directories), and simply type +If you wish to install mlpack to the system, make sure you have root privileges +(or write permissions to those two directories), and simply type @code # make install @endcode You can now run the executables by name; you can link against mlpack with --lmlpack, and the mlpack headers are found in /usr/include/mlpack/. +\c -lmlpack, and the mlpack headers are found in \c /usr/include or +\c /usr/local/include (depending on the system and CMake configuration). */ From 78a9149ccafd6387482c1bbf6d96d5f01137cd50 Mon Sep 17 00:00:00 2001 From: Ryan Curtin Date: Wed, 28 Mar 2018 21:05:24 -0400 Subject: [PATCH 3/4] Refactor documentation and tutorials to bring things up to date. Also fixes a few doxygen issues. --- doc/guide/bindings.hpp | 15 +-- doc/guide/formats.hpp | 88 ++++++++++----- doc/guide/iodoc.hpp | 82 ++++++++++---- doc/guide/matrices.hpp | 11 +- doc/guide/sample.hpp | 12 +- doc/guide/timer.hpp | 15 +-- doc/policies/metrics.hpp | 2 +- doc/tutorials/amf/amf.txt | 103 +++++++++--------- doc/tutorials/ann/ann.txt | 10 +- doc/tutorials/approx_kfn/approx_kfn.txt | 11 ++ doc/tutorials/cf/cf.txt | 6 +- doc/tutorials/kmeans/kmeans.txt | 4 +- .../linear_regression/linear_regression.txt | 2 +- doc/tutorials/tutorials.txt | 16 ++- 14 files changed, 241 insertions(+), 136 deletions(-) diff --git a/doc/guide/bindings.hpp b/doc/guide/bindings.hpp index a1e3a4a7a4..be4ddd8854 100644 --- a/doc/guide/bindings.hpp +++ b/doc/guide/bindings.hpp @@ -816,8 +816,8 @@ the CLI::GetParam() source to see how this might be used. The CLI singleton expects the following functions to be defined in the function map for each type: - - @c "GetParam" -- return a pointer to the parameter in @c output. - - @c "GetPrintableParam" -- return a pointer to a string description of the + - @c GetParam -- return a pointer to the parameter in @c output. + - @c GetPrintableParam -- return a pointer to a string description of the parameter in @c output. If these functions are properly defined, then the CLI module will work @@ -836,11 +836,11 @@ The code for the command-line bindings is found in @c src/mlpack/bindings/cli. @subsection bindings_cli_mlpack_main mlpackMain() definition -Any command-line program must be compiled with the @c MLPACK_BINDING_TYPE macro -set to the value @c cli. This is handled by the CMake macro +Any command-line program must be compiled with the @c BINDING_TYPE macro +set to the value @c BINDING_TYPE_CLI. This is handled by the CMake macro @c add_cli_executable(). -When @c MLPACK_BINDING_TYPE is set to @c cli, the following is set in +When @c BINDING_TYPE is set to @c BINDING_TYPE_CLI, the following is set in @c src/mlpack/core/util/mlpack_main.hpp, which must be included by every mlpack binding: @@ -1064,9 +1064,10 @@ guidance for how to make new bindings that will be applicable to each language. In general, the first thing to handle will be how matrices are passed back and forth between the target language. Typically this might mean getting the memory address of an input matrix and wrapping an @c arma::mat object around that -memory address. This can be handled in the @c "GetParam" function that is part +memory address. This can be handled in the @c GetParam() function that is part of the CLI singleton function map; see @c get_param.hpp for both the CLI and -Python bindings for an example. +Python bindings for an example (in @c src/mlpack/bindings/cli/ and +@c src/mlpack/bindings/python/). Serialization of models is also a tricky consideration; in some languages you will be able to pass a pointer to the model itself. This is generally diff --git a/doc/guide/formats.hpp b/doc/guide/formats.hpp index 24e9a1b43a..1fb2d7debb 100644 --- a/doc/guide/formats.hpp +++ b/doc/guide/formats.hpp @@ -1,4 +1,4 @@ -/*! @page formatdoc File formats in mlpack +/*! @page formatdoc File formats and loading data in mlpack @section formatintro Introduction @@ -7,6 +7,51 @@ command-line programs and in C++ programs using mlpack via the mlpack::data::Load() function. This tutorial discusses the formats that are supported and how to use them. +@section formatsimple Simple examples to load data in C++ + +The example code snippets below load data from different formats into an +Armadillo matrix object (\c arma::mat) or model when using C++. + +@code +using namespace mlpack; + +arma::mat matrix1; +data::Load("dataset.csv", matrix1); +@endcode + +@code +using namespace mlpack; + +arma::mat matrix2; +data::Load("dataset.bin", matrix2); +@endcode + +@code +using namespace mlpack; + +arma::mat matrix3; +data::Load("dataset.h5", matrix3); +@endcode + +@code +using namespace mlpack; + +// ARFF loading is a little different, since sometimes mapping has to be done +// for string types. +arma::mat matrix4; +data::DatasetInfo datasetInfo; +data::Load("dataset.arff", matrix4, datasetInfo); + +// The datasetInfo object now holds information about each dimension. +@endcode + +@code +using namespace mlpack; + +regression::LogisticRegression lr; +data::Load("model.bin", "logistic_regression_model", lr); +@endcode + @section formattypes Supported dataset types Datasets in mlpack are represented internally as sparse or dense numeric @@ -28,12 +73,12 @@ mlpack supports the following file types: - PGM, denoted by .pgm - PPM, denoted by .ppm - Armadillo binary, denoted by .bin - - Raw binary, denoted by .bin \b "(note: this will be loaded as" - \b "one-dimensional data, which is likely not what is desired.)" - - HDF5, denoted by .hdf, .hdf5, .h5, or .he5 (note: HDF5 must be enabled" + - Raw binary, denoted by .bin (note: this will be loaded as + one-dimensional data, which is likely not what is desired.) + - HDF5, denoted by .hdf, .hdf5, .h5, or .he5 (note: HDF5 must be enabled in the Armadillo configuration) - - ARFF, denoted by .arff (note: this is not supported by all mlpack" - command-line programs ; see \ref formatcat ) + - ARFF, denoted by .arff (note: this is not supported by all mlpack + command-line programs ; see \ref formatcat) Datasets that are loaded by mlpack should be stored with one row for one point and one column for one dimension. Therefore, a dataset @@ -46,8 +91,8 @@ would be stored in a csv file as: 5, -5 \endcode -As noted earlier, the format is automatically detected at load time. Therefore, -a dataset can be loaded in many ways: +As noted earlier, for command-line programs, the format is automatically +detected at load time. Therefore, a dataset can be loaded in many ways: \code $ mlpack_logistic_regression -t dataset.csv -v @@ -75,7 +120,7 @@ functions. Matrices in mlpack are column-major, meaning that each column should correspond to a point in the dataset and each row should correspond to a dimension; for -more information, see \ref matrices . This is at odds with how the data is +more information, see \ref matrices. This is at odds with how the data is stored in files; therefore, a transposition is required during load and save. The mlpack::data::Load() and mlpack::data::Save() functions do this automatically (unless otherwise specified), which is why they are preferred over @@ -274,7 +319,7 @@ through the \c --input_model_file (\c -m) and \c --output_model_file (\c -M) options; for more information, see the documentation for each program (accessible by passing \c --help as a parameter). -@section formatmodels Loading and saving models in C++ +@section formatmodelscpp Loading and saving models in C++ mlpack uses the \c boost::serialization library internally to perform loading and saving of models, and provides convenience overloads of mlpack::data::Load() @@ -284,29 +329,22 @@ To be serializable, a class must implement the method \code template -void Serialize(Archive& ar, const unsigned int version); +void serialize(Archive& ar, const unsigned int version); \endcode \note For more information on this method and how it works, see the -boost::serialization documentation at http://www.boost.org/libs/serialization/doc/ -. Note that mlpack uses a \c Serialize() -method and not a \c serialize() method, and also mlpack uses the -mlpack::data::CreateNVP() method instead of \c BOOST_SERIALIZATION_NVP() ; this -is for coherence with the mlpack style guidelines, and is done via a -particularly complex bit of template metaprogramming in -src/mlpack/core/data/serialization_shim.hpp (read that file if you want your -head to hurt!). +boost::serialization documentation at +http://www.boost.org/libs/serialization/doc/. \note -Examples of Serialize() methods can be found in most classes; one fairly -straightforward example is found \ref mlpack::math::Range::Serialize() -"in the mlpack::math::Range class". A more complex example is found \ref -mlpack::tree::BinarySpaceTree::Serialize() -"in the mlpack::tree::BinarySpaceTree class". +Examples of serialize() methods can be found in most classes; one fairly +straightforward example is found \ref mlpack::math::Range::serialize() +"in the mlpack::math::Range class". A more complex example is found +\ref mlpack::tree::BinarySpaceTree::serialize() "in the mlpack::tree::BinarySpaceTree class". Using the mlpack::data::Load() and mlpack::data::Save() classes is easy if the -type being saved has a \c Serialize() method implemented: simply call either +type being saved has a \c serialize() method implemented: simply call either function with a filename, a name for the object to save, and the object itself. The example below, for instance, creates an mlpack::math::Range object and saves it as range.txt. Then, that range is loaded from file into another diff --git a/doc/guide/iodoc.hpp b/doc/guide/iodoc.hpp index 8cdf417b7e..fce2a65f3b 100644 --- a/doc/guide/iodoc.hpp +++ b/doc/guide/iodoc.hpp @@ -1,15 +1,22 @@ -/*! @page iodoc mlpack Input and Output +/*! @page iodoc Writing an mlpack binding @section iointro Introduction +This tutorial gives some simple examples of how to write an mlpack binding that +can be compiled for multiple languages. These bindings make up the core of how +most users will interact with mlpack. + mlpack provides the following: - mlpack::Log, for debugging / informational / warning / fatal output - - mlpack::CLI, for parsing command line options + - mlpack::CLI, for parsing command line options or other option Each of those classes are well-documented, and that documentation should be consulted for further reference. +First, we'll discuss the logging infrastructure, which is useful for giving +output that users can see. + @section simplelog Simple Logging Example mlpack has four logging levels: @@ -21,22 +28,27 @@ mlpack has four logging levels: Output to Log::Debug does not show (and has no performance penalty) when mlpack is compiled without debugging symbols. Output to Log::Info is only shown when -the program is run with the --verbose (or -v) flag. Log::Warn is always shown, -and Log::Fatal will throw a std::runtime_error exception, when a newline is sent -to it only. If mlpack was compiled with debugging symbols, Log::Fatal will -always throw a std::runtime_error exception and print backtrace. +the program is run with the \c --verbose (or \c -v) flag. Log::Warn is always +shown, and Log::Fatal will throw a std::runtime_error exception, after a newline +is sent to it. If mlpack was compiled with debugging symbols, Log::Fatal will +also print a backtrace, if the necessary libraries are available. -Here is a simple example, and its output: +Here is a simple example binding, and its output. Note that instead of +\c int \c main(), we use \c static \c void \c mlpackMain(). This is because the +automatic binding generator (see \ref bindings) will set up the environment and +once that is done, it will call \c mlpackMain(). @code #include +#include +// This definition below means we will only compile for the CLI. +#define BINDING_TYPE BINDING_TYPE_CLI +#include using namespace mlpack; -int main(int argc, char** argv) +static void mlpackMain() { - CLI::ParseCommandLine(argc, argv); - Log::Debug << "Compiled with debugging symbols." << std::endl; Log::Info << "Some test informational output." << std::endl; @@ -49,9 +61,18 @@ int main(int argc, char** argv) } @endcode -With debugging output--verbose, the following is shown: +Assuming mlpack is installed on the system and the code above is saved in +\c test.cpp, this program can be compiled with the following command: @code +$ g++ -o test test.cpp -DDEBUG -g -rdynamic -lmlpack +@endcode + +Since we compiled with \c -DDEBUG, if we run the program as below, the following +output is shown: + +@code +$ ./test --verbose [DEBUG] Compiled with debugging symbols. [INFO ] Some test informational output. [WARN ] A warning! @@ -62,10 +83,12 @@ terminate called after throwing an instance of 'std::runtime_error' Aborted @endcode -With debugging output, compilation flags -g -rdynamic and --verbose, -the following is shown: +The flags \c -g and \c -rdynamic are only necessary for providing a backtrace. +If those flags are not given during compilation, the following output would be +shown: @code +$ ./test --verbose [DEBUG] Compiled with debugging symbols. [INFO ] Some test informational output. [WARN ] A warning! @@ -79,10 +102,11 @@ Aborted The last warning is not reached, because Log::Fatal terminates the program. -Without debugging symbols and without --verbose, the following is shown: +Without debugging symbols (i.e. without \c -g and \c -DDEBUG) and without +--verbose, the following is shown: @code -$ ./main +$ ./test [WARN ] A warning! [FATAL] Program has crashed. terminate called after throwing an instance of 'std::runtime_error' @@ -100,9 +124,12 @@ with the PROGRAM_INFO, PARAM_INT, PARAM_DOUBLE, PARAM_STRING, and PARAM_FLAG macros. Here is a sample use of those macros, extracted from methods/pca/pca_main.cpp. +(Some details have been omitted from the snippet below.) @code #include +#include +#include // Document program. PROGRAM_INFO("Principal Components Analysis", "This program performs principal " @@ -112,19 +139,25 @@ PROGRAM_INFO("Principal Components Analysis", "This program performs principal " "eigenvalues."); // Parameters for program. -PARAM_STRING_REQ("input_file", "Input dataset to perform PCA on.", ""); -PARAM_STRING_REQ("output_file", "Output dataset to perform PCA on.", ""); -PARAM_INT("new_dimensionality", "Desired dimensionality of output dataset.", - "", 0); +PARAM_MATRIX_IN_REQ("input", "Input dataset to perform PCA on.", "i"); +PARAM_MATRIX_OUT("output", "Matrix to save modified dataset to.", "o"); +PARAM_INT_IN("new_dimensionality", "Desired dimensionality of output dataset.", + "d", 0); using namespace mlpack; -int main(int argc, char** argv) +static void mlpackMain() { - // Parse commandline. - CLI::ParseCommandLine(argc, argv); + // Load input dataset. + arma::mat& dataset = CLI::GetParam("input"); + + size_t newDimension = CLI::GetParam("new_dimensionality"); ... + + // Now save the results. + if (CLI::HasParam("output")) + CLI::GetParam("output") = std::move(dataset); } @endcode @@ -143,7 +176,7 @@ Principal Components Analysis Required options: --input_file [string] Input dataset to perform PCA on. - --output_file [string] Output dataset to perform PCA on. + --output_file [string] Matrix to save modified dataset to. Options: @@ -158,6 +191,7 @@ Options: @endcode The mlpack::CLI documentation can be consulted for further and complete -documentation. +documentation. Also useful is to look at other example bindings, found in +\c src/mlpack/methods/. */ diff --git a/doc/guide/matrices.hpp b/doc/guide/matrices.hpp index 1fe962666a..73ba2d95fb 100644 --- a/doc/guide/matrices.hpp +++ b/doc/guide/matrices.hpp @@ -24,11 +24,15 @@ most standard machine learning texts! Major implications of this are for linear algebra. For instance, the covariance of a matrix is typically -@f$ C = X^T X @f$ +@f[ +C = X^T X +@f] but for a column-wise matrix, it is -@f$ C = X X^T @f$ +@f[ +C = X X^T +@f] and this is very important to keep in mind! If your mlpack code is not working, this may be a factor in why. @@ -63,7 +67,8 @@ $ cat data.csv @endcode is actually loaded with 5 rows and 13 columns, not 13 rows and 5 columns like -the CSV is written. +the CSV is written. More information on mlpack's loading functionality can be +found in \ref formatdoc. This is important to remember! diff --git a/doc/guide/sample.hpp b/doc/guide/sample.hpp index bb6886ae61..7650a2f2d2 100644 --- a/doc/guide/sample.hpp +++ b/doc/guide/sample.hpp @@ -4,15 +4,16 @@ On this page, several simple mlpack examples are contained, in increasing order of complexity. If you compile from the command-line, be sure that your compiler -is in C++11 mode. With gcc and clang, this can be accomplished by adding the -@c -std=c++11 option. +is in C++11 mode. With modern gcc and clang, this should already be the +default. @note The command-line programs like @c knn_main.cpp and @c logistic_regression_main.cpp from the directory @c src/mlpack/methods/ cannot be compiled easily by hand (the same is true for the individual tests in @c -src/mlpack/tests/); instead, those should be compiled with CMake. However, any -program that uses mlpack (and is not a part of the library itself) can be +src/mlpack/tests/); instead, those should be compiled with CMake, by running, +e.g., @c make @c mlpack_knn or @c make @c mlpack_test; see @ref build. However, +any program that uses mlpack (and is not a part of the library itself) can be compiled easily with g++ or clang from the command line. @section covariance Covariance Computation @@ -87,7 +88,8 @@ int main() @section other Other examples -For more complex examples, it is useful to refer to the main executables: +For more complex examples, it is useful to refer to the main executables, found +in @c src/mlpack/methods/. A few are listed below. - methods/neighbor_search/knn_main.cpp - methods/neighbor_search/kfn_main.cpp diff --git a/doc/guide/timer.hpp b/doc/guide/timer.hpp index cff500107b..d6d578ebae 100644 --- a/doc/guide/timer.hpp +++ b/doc/guide/timer.hpp @@ -4,7 +4,7 @@ mlpack provides a simple timer interface for the timing of machine learning methods. The results of any timers used during the program are displayed at -output by the mlpack::CLI object, when --verbose is given: +output by any command-line binding, when --verbose is given: @code $ mlpack_knn -r dataset.csv -n neighbors_out.csv -d distances_out.csv -k 5 -v @@ -33,7 +33,7 @@ and the result will be the sum of the runs of the timer. Note that \c Timer::Stop() must be called before \c Timer::Start() is called again, otherwise a std::runtime_error exception will be thrown. -A "total_time" timer is run by default for each mlpack program. +A \c "total_time" timer is run by default for each mlpack program. @section example Timer Example @@ -41,13 +41,14 @@ Below is a very simple example of timer usage in code. @code #include +#include +#define BINDING_TYPE BINDING_TYPE_CLI +#include using namespace mlpack; -int main(int argc, char** argv) +void mlpackMain() { - CLI::ParseCommandLine(argc, argv); - // Start a timer. Timer::Start("some_timer"); @@ -59,7 +60,7 @@ int main(int argc, char** argv) } @endcode -If the --verbose flag was given to this executable, the resultant time that -"some_timer" ran for would be shown. +If the --verbose flag was given to this executable, the time that +\c "some_timer" ran for would be printed at the end of the program's output. */ diff --git a/doc/policies/metrics.hpp b/doc/policies/metrics.hpp index 6237dd4e64..a2f290d01c 100644 --- a/doc/policies/metrics.hpp +++ b/doc/policies/metrics.hpp @@ -108,6 +108,6 @@ policy: - mlpack::metric::ChebyshevDistance - mlpack::metric::MahalanobisDistance - mlpack::metric::LMetric (for arbitrary L-metrics) - - mlpack::metric::IPMetric (requires a \ref kernels KernelType parameter) + - mlpack::metric::IPMetric (requires a \ref kernels "KernelType" parameter) */ diff --git a/doc/tutorials/amf/amf.txt b/doc/tutorials/amf/amf.txt index bc4f8ed3c2..373c2b804c 100644 --- a/doc/tutorials/amf/amf.txt +++ b/doc/tutorials/amf/amf.txt @@ -2,9 +2,9 @@ @file amf.txt @author Sumedh Ghaisas -@brief Tutorial for how to use the AMF class. +@brief Tutorial for how to use the AMF class -@page amftutorial Alternating Matrix Factorization tutorial. +@page amftutorial Alternating Matrix Factorization tutorial @section intro_amftut Introduction @@ -60,51 +60,57 @@ which returns the status of convergence. bool IsConverged(arma::mat& W, arma::mat& H) @endcode -list of all the termination policies +Below is a list of all the termination policies that mlpack contains. - \ref mlpack::amf::SimpleResidueTermination - \ref mlpack::amf::SimpleToleranceTermination - \ref mlpack::amf::ValidationRMSETermination -In SimpleResidueTermination, termination decision depends on two factors, value +In \c SimpleResidueTermination, termination decision depends on two factors, value of residue and number of iteration. If the current value of residue drops below the threshold or the number of iterations goes beyond the threshold, positive termination signal is passed to AMF. -In SimpleToleranceTermination, termination criterion is met when increase in -residue value drops below the given tolerance. To accommodate spikes, certain +In \c SimpleToleranceTermination, termination criterion is met when the increase +in residue value drops below the given tolerance. To accommodate spikes, certain number of successive residue drops are accepted. Secondary termination criterion terminates algorithm when iteration count goes beyond the threshold. -ValidationRMSETermination divids the data into 2 sets, training set and +\c ValidationRMSETermination divides the data into 2 sets, training set and validation set. Entries of validation set are nullifed in the input matrix. Termination criterion is met when increase in validation set RMSe value drops below the given tolerance. To accommodate spikes certain number of successive validation RMSE drops are accepted. This upper imit on successive drops can be -adjusted with reverseStepCount. Secondary termination criterion terminates -algorithm when iteration count goes above the threshold. Though this termination -policy is better measure of convergence than the above 2 termination policies, -it may cause a overhead in performance. +adjusted with \c reverseStepCount. A secondary termination criterion terminates +the algorithm when the iteration count goes above the threshold. Though this +termination policy is better measure of convergence than the above 2 termination +policies, it may cause a decrease in performance since it is computationally +expensive. -On the other hand \ref mlpack::amf::CompleteIncrementalTermination +On the other hand, \ref mlpack::amf::CompleteIncrementalTermination "CompleteIncrementalTermination" and \ref mlpack::amf::IncompleteIncrementalTermination -are just wrapper classes for other termination policies. These policies are used -when AMF is applied with \ref mlpack::amf::SVDCompleteIncrementalLearning -"SVDCompleteIncrementalLearning" and \ref mlpack::amf::SVDIncompleteIncrementalLearning -"SVDIncompleteIncrementalLearning" respectively. +"IncompleteIncrementalTermination" are just wrapper classes for other +termination policies. These policies are used when AMF is applied with +\ref mlpack::amf::SVDCompleteIncrementalLearning +"SVDCompleteIncrementalLearning" and +\ref mlpack::amf::SVDIncompleteIncrementalLearning +"SVDIncompleteIncrementalLearning", respectively. @subsection init_rule_amftut Using different initialization policies -The AMF class comes with 2 initialization policies +mlpack currently has 2 initialization policies implemented for AMF: + - \ref mlpack::amf::RandomInitialization "RandomInitialization" - \ref mlpack::amf::RandomAcolInitialization "RandomAcolInitialization" -RandomInitialization initializes matrices W and H with random uniform distribution -while RandomAcolInitialization initializes the W matrix by averaging p randomly -chosen columns of V. In case of RandomAcolInitialization, p is a template parameter. +\c RandomInitialization initializes matrices W and H with random uniform +distribution while \c RandomAcolInitialization initializes the W matrix by +averaging p randomly chosen columns of V. In the case of +\c RandomAcolInitialization, p is a template parameter. To implement their own initialization policy, users need to define the following function in their class. + @code template inline static void Initialize(const MatType& V, @@ -115,7 +121,8 @@ inline static void Initialize(const MatType& V, @subsection update_rule_amftut Using different update rules -AMF supports following update rules +mlpack implements the following update rules for the AMF class: + - \ref mlpack::amf::NMFALSUpdate "AMFALSUpdate" - \ref mlpack::amf::NMFMultiplicativeDistanceUpdate "NMFMultiplicativeDistanceUpdate" - \ref mlpack::amf::NMFMultiplicativeDivergenceUpdate "NMFMultiplicativeDivergenceUpdate" @@ -123,26 +130,26 @@ AMF supports following update rules - \ref mlpack::amf::SVDIncompleteIncrementalLearning "SVDIncompleteIncrementalLearning" - \ref mlpack::amf::SVDCompleteIncrementalLearning "SVDCompleteIncrementalLearning" -Non-Negative Matrix factorization can be achieved with NMFALSUpdate, -NMFMultiplicativeDivergenceUpdate or NMFMultiplicativeDivergenceUpdate. -NMFALSUpdate implements simple Alternating Least Square optimization while -the other rules implement algorithms given in paper 'Algorithms for Non-negative -Matrix Factorization'. +Non-Negative Matrix factorization can be achieved with \c NMFALSUpdate, +\c NMFMultiplicativeDivergenceUpdate or \c NMFMultiplicativeDivergenceUpdate. +\c NMFALSUpdate implements a simple Alternating Least Squares optimization while +the other rules implement algorithms given in the paper 'Algorithms for +Non-negative Matrix Factorization'. -The remaining update rules perform Singular Value Decomposition of matrix V. -This SVD factorization is optimized for the use by Collaborative Filtering. This -use of SVD factorizers for Collaborative Filtering is described in the paper -'A Guide to singular Value Decomposition' by Chih-Chao Ma. For further details -about the algorithms refer to the respective class documentation. +The remaining update rules perform the singular value decomposition of the matrix V. +This SVD factorization is optimized for use by mlpack's collaborative filtering +code (\ref cftutorial). This use of SVD factorizers for collaborative filtering +is described in the paper 'A Guide to Singular Value Decomposition for +Collaborative Filtering' by Chih-Chao Ma. For further details about the +algorithms refer to the respective class documentation. @subsection nmf_amftut Using Non-Negative Matrix Factorization with AMF The use of AMF for Non-Negative Matrix factorization is simple. The AMF module defines \ref mlpack::amf::NMFALSFactorizer "NMFALSFactorizer" which can be used -directly without knowing the internal structure of AMF. For example - +directly without knowing the internal structure of AMF. For example: @code -#include #include #include @@ -156,28 +163,26 @@ int main() mat W, H; mat V = randu(100, 100); double residue = nmf.Apply(V, W, H); - return 1; } @endcode -NMFALSFactorizer uses SimpleResidueTermination which is most preferred with -Non-Negative Matrix factorizers. Initialization of W and H in NMFALSFactorizer -is random. The Apply function returns the residue obtained by comparing the -constructed matrix W * H with the original matrix V. +\c NMFALSFactorizer uses \c SimpleResidueTermination, which is most preferred +with Non-Negative Matrix factorizers. The initialization of W and H in +\c NMFALSFactorizer is random. The \c Apply() function returns the residue +obtained by comparing the constructed matrix W * H with the original matrix V. @subsection svd_amftut Using Singular Value Decomposition with AMF -AMF implementation supports following SVD factorizers - - \ref mlpack::amf::SVDBatchFactorizer "SVDBatchFactorizer" - - \ref mlpack::amf::SparseSVDBatchFactorizer "SparseSVDBatchFactorizer" - - \ref mlpack::amf::SVDIncompleteIncrementalFactorizer "SVDIncompleteIncrementalFactorizer" - - \ref mlpack::amf::SparseSVDIncompleteIncrementalFactorizer "SparseSVDIncompleteIncrementalFactorizer" - - \ref mlpack::amf::SVDCompleteIncrementalFactorizer "SVDCompleteIncrementalFactorizer" - - \ref mlpack::amf::SparseSVDCompleteIncrementalFactorizer "SparseSVDCompleteIncrementalFactorizer" +mlpack has the following SVD factorizers implemented for AMF: -The sparse version of factorizers can be used with Armadillo's sparse matrix -support. These specialized implementations boost runtime performance when the -matrix to be factorized is relatively sparse. + - \ref mlpack::amf::SVDBatchFactorizer "SVDBatchFactorizer" + - \ref mlpack::amf::SVDIncompleteIncrementalFactorizer "SVDIncompleteIncrementalFactorizer" + - \ref mlpack::amf::SVDCompleteIncrementalFactorizer "SVDCompleteIncrementalFactorizer" + +Each of these factorizers takes a template parameter \c MatType, which specifies +the type of the matrix V (dense or sparse---these have types \c arma::mat and +\c arma::sp_mat, respectively). When the matrix to be factorized is relatively +sparse, specifying \c MatType \c = \c arma::sp_mat can provide a runtime boost. @code #include @@ -192,7 +197,7 @@ int main() sp_mat V = randu(100,100); mat W, H; - SparseSVDBatchFactorizer svd; + SVDBatchFactorizer svd; double residue = svd.Apply(V, W, H); } @endcode diff --git a/doc/tutorials/ann/ann.txt b/doc/tutorials/ann/ann.txt index 546e39de11..5b2371057e 100644 --- a/doc/tutorials/ann/ann.txt +++ b/doc/tutorials/ann/ann.txt @@ -13,9 +13,9 @@ converging in local minima, choosing the best model structure, choosing the best optimizers, and so forth. mlpack implements many of these building blocks, making it very easy to create different neural networks in a modular way. -mlpack currently implements two easy-to-use forms of neural networks: \c Feed- -Forward \c Networks (this includes convolutional neural networks) and \c -Recurrent \c Neural \c Networks. +mlpack currently implements two easy-to-use forms of neural networks: +\b Feed-Forward \b Networks (this includes convolutional neural networks) and +\b Recurrent \b Neural \b Networks. @section toc_anntut Table of Contents @@ -51,7 +51,7 @@ Below is some basic guidance on what should be used. Note that the question of guidance below is just that---guidance---and may not be right for a particular problem. - - \c Feed-forward Networks allow signals or inputs to travel one way only. + - \b Feed-forward \b Networks allow signals or inputs to travel one way only. There is no feedback within the network; for instance, the output of any layer does only affect the upcoming layer. That makes Feed-Forward Networks straightforward and very effective. They are extensively used in pattern @@ -59,7 +59,7 @@ problem. set of input and one or more output variables. - - \c Recurrent Networks allow signals or inputs to travel in both directions by + - \b Recurrent \b Networks allow signals or inputs to travel in both directions by introducing loops in the network. Computations derived from earlier inputs are fed back into the network, which gives the recurrent network some kind of memory. RNNs are currently being used for all kinds of sequential tasks; for diff --git a/doc/tutorials/approx_kfn/approx_kfn.txt b/doc/tutorials/approx_kfn/approx_kfn.txt index 247eb72ff9..0542b2f6ff 100644 --- a/doc/tutorials/approx_kfn/approx_kfn.txt +++ b/doc/tutorials/approx_kfn/approx_kfn.txt @@ -60,6 +60,17 @@ These methods are described in the following papers: } @endcode +@code +@article{curtin2018exploiting, + title={Exploiting the structure of furthest neighbor search for fast + approximate results}, + author={Curtin, Ryan R., and Echauz, Javier, and Gardner, Andrew B.}, + journal={Information Systems}, + year={2018}, + publisher={Elsevier} +} +@endcode + The problem of furthest neighbor search is simple, and is the opposite of the much-more-studied nearest neighbor search problem. Given a set of reference points \f$R\f$ (the set in which we are searching), and a set of query points diff --git a/doc/tutorials/cf/cf.txt b/doc/tutorials/cf/cf.txt index 5e506b5da2..b300aaff22 100644 --- a/doc/tutorials/cf/cf.txt +++ b/doc/tutorials/cf/cf.txt @@ -409,8 +409,10 @@ existing factorizers that can be used in \b mlpack; these were detailed in the The \c FactorizerType class must implement one of the two following methods: - - \c "Apply(arma::mat& data, const size_t rank, arma::mat& W, arma::mat& H);" - - \c "Apply(arma::sp_mat& data, const size_t rank, arma::mat& W, arma::mat& H);" + - Apply(arma::mat& data, const size_t rank, arma::mat& W, arma::mat& + H); + - Apply(arma::sp_mat& data, const size_t rank, arma::mat& W, arma::mat& + H); The difference between these two methods is whether \c arma::mat or \c arma::sp_mat is used as input. If \c arma::mat is used, then the data matrix is diff --git a/doc/tutorials/kmeans/kmeans.txt b/doc/tutorials/kmeans/kmeans.txt index 365c915450..b4a8d5d7cf 100644 --- a/doc/tutorials/kmeans/kmeans.txt +++ b/doc/tutorials/kmeans/kmeans.txt @@ -657,8 +657,8 @@ and must accept two template parameters of its own: The \c LloydStepType policy also mandates three functions: - - a constructor: \c "LloydStepType(const MatType& dataset, MetricType& - metric);" + - a constructor: LloydStepType(const MatType& dataset, MetricType& + metric); - an \c Iterate() function: @code diff --git a/doc/tutorials/linear_regression/linear_regression.txt b/doc/tutorials/linear_regression/linear_regression.txt index d2cac8c968..eb1bc976c1 100644 --- a/doc/tutorials/linear_regression/linear_regression.txt +++ b/doc/tutorials/linear_regression/linear_regression.txt @@ -345,7 +345,7 @@ you would set the parameters for a LinearRegression instance. @code arma::vec parameters; // Your model. -LinearRegression lr(); // Create a new LinearRegression instance or reuse one. +LinearRegression lr; // Create a new LinearRegression instance or reuse one. lr.Parameters() = parameters; // Set the model. @endcode diff --git a/doc/tutorials/tutorials.txt b/doc/tutorials/tutorials.txt index a8df927130..eac3f29ca3 100644 --- a/doc/tutorials/tutorials.txt +++ b/doc/tutorials/tutorials.txt @@ -18,9 +18,6 @@ start. - \ref iodoc - \ref timer - \ref sample - - \ref cv - - \ref hpt - - \ref bindings @section method_tut Method-specific Tutorials @@ -38,9 +35,19 @@ progress to complex, extensible uses. - \ref amftutorial - \ref cftutorial - \ref akfntutorial - - \ref cnetutorial - \ref anntutorial +@section adv_tut Advanced Tutorials + +These tutorials discuss some of the more advanced functionality contained in +mlpack. + + - \ref optimizertutorial + - \ref cnetutorial + - \ref bindings + - \ref cv + - \ref hpt + @section policy_tut Policy Class Documentation mlpack uses templates to achieve its genericity and flexibility. Some of the @@ -51,6 +58,5 @@ types. - \ref metrics - \ref kernels - \ref trees - - \ref optimizertutorial */ From 7138af283d51603ef1112250a996b2311ad34a41 Mon Sep 17 00:00:00 2001 From: Ryan Curtin Date: Thu, 29 Mar 2018 17:34:54 -0400 Subject: [PATCH 4/4] Explicitly specify pytest-runner. --- doc/guide/build.hpp | 1 + 1 file changed, 1 insertion(+) diff --git a/doc/guide/build.hpp b/doc/guide/build.hpp index 299c12f428..4d8362f564 100644 --- a/doc/guide/build.hpp +++ b/doc/guide/build.hpp @@ -72,6 +72,7 @@ For Python bindings, the following packages are required: - cython >= 0.24 - numpy - pandas >= 0.15.0 + - pytest-runner In Ubuntu and Debian, you can get all of these dependencies through apt: