Conda: Difference between revisions

From Grid5000
Jump to navigation Jump to search
No edit summary
No edit summary
Line 5: Line 5:
{{TutorialHeader}}
{{TutorialHeader}}


{{Note|text='''This document was written by consolidating the following different information resources:'''
* Grid'5000 documentation Environment modules, HPC and HTC tutorial, Deep Learning Frameworks
* An [[User:Ibada/Tuto Deep Learning|in-depth tutorial|Tuto Deep Learning tutorial]], Ismael Bada
* [[User:Bjonglez/Debian11/Deep Learning Frameworks|Deep Learning Frameworks tutorial]], Benjamin Jonglez
}}


= Introduction =
= Introduction =

Revision as of 10:43, 20 April 2023

Note.png Note

This page is actively maintained by the Grid'5000 team. If you encounter problems, please report them (see the Support page). Additionally, as it is a wiki page, you are free to make minor corrections yourself if needed. If you would like to suggest a more fundamental change, please contact the Grid'5000 team.


Introduction

Conda is an open source package management system and environment management system for installing multiple versions of software packages and their dependencies and switching easily between them. It works on Linux, OS X and Windows, and was created for Python programs but can package and distribute any software.

The conda package and environment manager is included in all versions of Anaconda®, Miniconda, and Anaconda Repository. Conda is also available on conda-forge, a community channel.

Anaconda or Miniconda?

Anaconda contains a full distribution of packages while Miniconda is a condensed version that contains the essentials for standard purposes.

References

Conda usage

Conda shell activation

Conda shell activation is the process of defining some shell functions that facilitate activating and deactivating Conda environments, as well as some optional features such as updating PS1 to show the active environment.

The conda shell function is mainly a forwarder function. It will delegate most of the commands to the real conda executable driven by the Python library.

Activate your Conda shell environment as follow:

Terminal.png $:
eval "$(conda shell.bash hook)"

By defaut, you are located in the base Conda environment that correspond to the base installation of Conda.

Conda environments

Conda allows you to create separate environments containing files, packages, and their dependencies that will not interact with other environments.

When you begin using conda, you already have a default environment named "base". You can create separate environments to keep your programs isolated from each other. Specifying the environment name confines conda commands to that environment.

  • List all your environments
Terminal.png $:
conda info --envs

or

Terminal.png $:
conda env list
  • Create a new environment
Terminal.png $:
conda create --name ENVNAME
  • Activate this environment before installing package
Terminal.png $:
conda activate ENVNAME

For further information:

Conda package installation

In its default configuration, Conda can install and manage the over 7,500 packages at https://repo.anaconda.com/pkgs/ that are built, reviewed, and maintained by Anaconda. This is the default Conda channel which may require a paid license, as described in the repository terms of service a commercial license.

Terminal.png $:
conda install <package>
  • Install specific version of package:
Terminal.png $:
conda install <package>=<version>
  • Uninstall a package:
Terminal.png $:
conda uninstall <package>

For more information:

Conda package installation from channels

Channels are the locations of the repositories where Conda looks for packages. Channels may point to a Cloud repository or a private location on a remote or local repository that you or your organization created. Useful channels are:

To install a package from a specific channel:

Terminal.png $:
conda install -c <chanel_name> <package>
  • List all packages installed with their source channels
Terminal.png $:
conda list --show-channel-urls

For more information:

Suggested reading

Load conda on Grid'5000

Conda is already available in Grid'5000 as a module. You don't need to install Anaconda or Miniconda on Grid'5000! To make it available on a node or on a frontend, you need to load the Conda module as follow:

  • For Miniconda:
Terminal.png fsophia:
module load miniconda3
Terminal.png fsophia:
conda --version
conda 22.11.1
  • For Anaconda:
Terminal.png node:
module load anaconda3
Terminal.png node:
conda --version
conda 4.12.0

By default, Conda and all packages are installed locally with a user-specific configuration. In the Grid'5000 context, Conda comes with some pre-existing packages in the base environment.

  • To list the emplacement of Conda and the current environment:
Terminal.png fsophia:
conda info
     active environment : base
    active env location : /grid5000/spack/opt/spack/linux-debian11-x86_64/gcc-10.2.0/miniconda3-4.10.3-x6kxdkqihyhysyjs7i4g77wururhgvfg
...
       base environment : /grid5000/spack/opt/spack/linux-debian11-x86_64/gcc-10.2.0/miniconda3-4.10.3-x6kxdkqihyhysyjs7i4g77wururhgvfg  (read only)
...
       envs directories : /home/lmirtain/.conda/envs
                          /grid5000/spack/opt/spack/linux-debian11-x86_64/gcc-10.2.0/miniconda3-4.10.3-x6kxdkqihyhysyjs7i4g77wururhgvfg/envs
...

You can see here that the active environment is base and its emplacement is on a NFS storage (/grid5000/....).

  • To list installed packages in the current environment:
Terminal.png fsophia:
conda list

Create conda environments on Grid'5000

Basic Conda workflow

Warning.png Warning

Installing Conda packages can be time and resource consuming. Preferably use a node (instead of a frontend) to perform such an operation. Note, using a node is mandatory if you need to access specific hardware resources like GPU.

  • Load conda module and activate bash completion
Terminal.png fgrenoble:
module load miniconda3
Terminal.png fgrenoble:
eval "$(conda shell.bash hook)"
  • Create an environment (specify a Python version; otherwise, it is the module default version)
Terminal.png fgrenoble:
conda create -y -n <name> python=x.y
  • Load this environment
Terminal.png fgrenoble:
conda activate <name>
  • Install a package
Terminal.png fgrenoble:
conda install <package_name>
  • Exit from the loaded environment
Terminal.png fgrenoble:
conda deactivate

Remove unused Conda environments

Warning.png Warning

Conda packages are installed in $HOME/.conda. You could, therefore, rapidly saturate your homedir quota (25GB by default). Do not forget to occasionally remove unused Conda environment to free up space.

  • To delete an environment
Terminal.png fgrenoble:
conda env remove --name <name>
  • To remove unused packages and the cache. Do not be concerned if this appears to try to delete the packages of the system environment (ie. non-local).
Terminal.png fgrenoble:
conda clean -a

Use a Conda environment on Grid'5000

As seen in the previous section, the Conda environment is stored by default in user's homedir (at ~/.conda). Once the environment is created and packages installed, it is usable on all nodes from the given site.

For interactive jobs

Terminal.png fgrenoble:
oarsub -I
Terminal.png node:
module load miniconda3
Terminal.png node:
eval "$(conda shell.bash hook)"
Terminal.png node:
conda activate <name>

For batch jobs

Warning.png Warning

As module command is not a real executable but a shell function, it must be executed in an actual shell to work. A simple oarsub "module load miniconda3" will fail.

Terminal.png fgrenoble:
oarsub 'bash -l -c "module load miniconda3; conda activate <name>; <your script>"'

Advanced Conda environment operations

Synchronize Conda environments between Grid'5000 sites

  • To synchronize a Conda directory from a siteA to a siteB:
Terminal.png fsiteA:
rsync --dry-run --delete -avz ~/.conda siteB.grid5000.fr:~

To really do things, the --dry-run argument has to be removed and siteB has to be replaced by a real site name.

Share Conda environments between multiple users

You can use two different approaches to share Conda environments with other users.

Export an environment as a yaml file

  • Export it as follow:
Terminal.png fgrenoble:
conda env export > environment.yml
  • Share it by putting the yaml file in your public folder
Terminal.png fgrenoble:
cp environment.yml ~/public/
  • Other users can create the environment from the environment.yml file
Terminal.png fgrenoble:
conda env create -f ~/<login>/public/environment.yml
  • Advantage : it prevents other users from damaging the environment if they add packages that could conflict with other packages and/or even delete packages that another user might need.
  • Inconvenient : it's not a true shared environment. The environment is duplicated on other users' home directory. Any modification on one Conda environment will not be automatically replicated on others.

Use a group storage

Group Storage gives you the possibility to share a storage between multiple users. You can take advantage of a group storage to share a single Conda environment among multiple users.

  • Create a shared Conda environment (--prefix allows you to specify the path to store the conda environment)
Terminal.png flyon:
conda create --prefix /srv/storage/storage_name@server_hostname_(fqdn)/ENVNAME
  • Activate the shared environment (share this command with the targeted users)
Terminal.png flyon:
conda activate /srv/storage/storage_name@server_hostname_(fqdn)/ENVNAME
  • Advantage : It avoids storing duplicate packages and makes any modification accessible to all users
  • Inconvenient : Users could potentially harm the environment by installing or removing packages.

Mamba as an alternative to Conda

mamba is a reimplementation of the conda package manager in C++. Mamba is fully compatible with Conda packages and supports most of Conda's commands. It consists of:

  • mamba: a Python-based CLI conceived as a drop-in replacement for conda, offering higher speed and more reliable environment solutions
  • micromamba: a pure C++-based CLI, self-contained in a single-file executable
  • libmamba: a C++ library exposing low-level and high-level APIs on top of which both mamba and micromamba are built

Mamba is relatively new and unpopular compared to Conda. That means there are probably more undiscovered bugs, and that new bugs may take longer to be discovered. mamba has to be considerate when using a devops chain in order to test and deploy an environment (i.e., docker images) with continuous integration pipelines. Conda has a reputation for taking time when dealing with complex sets of dependencies so CI jobs can take longer than they need to.

  • Mamba installation when already have Conda
Terminal.png inside:
conda install mamba -c conda-forge
  • Installing packages is similarly easy, example:
Terminal.png inside:
mamba install python=3.8 jupyter -c conda-forge

To go further:

Build your HPC-IA framework with conda

Here are some pointers to help you set up your software environment for HPC or AI with conda