irmsd#

class irmsd.Molecule(symbols: list[str], positions: ~numpy.ndarray, energy: float | None = None, info: dict[str, ~typing.Any] = <factory>, cell: ~numpy.ndarray | None = None, pbc: tuple[bool, bool, bool] | None = None, ids: ~numpy.ndarray | None = None)[source]#

Bases: object

Dependency-free replacement for ase.Atoms.

energy is in Hartree. ids holds optional per-atom canonical atom identifiers, one int32 per atom. Construction raises ValueError on bad shapes or unknown symbols.

cell: ndarray | None = None#
copy() → Molecule[source]#

Return a deep copy; no array, info or symbols is shared.

energy: float | None = None#
get_atomic_numbers() → ndarray[source]#

Return atomic numbers as int32 array.

get_axis() → tuple[ndarray, ndarray, ndarray][source]#

Return rotation constants, average momentum and rotation matrix.

Returns:

  • rot ((3,) ndarray) – Rotation constants in MHz.

  • avmom ((1,) ndarray) – Average momentum in a.u.

  • evec ((3, 3) ndarray)

get_canonical(wbo: ndarray | None = None, invtype: str = 'apsp+', heavy: bool = False) → ndarray[source]#

Return (N,) int32 canonical ranks from the Fortran core.

Parameters:
  • wbo ((N, N) ndarray, optional) – Wiberg bond orders; required for invtype="cangen".

  • invtype ({"apsp+", "cangen", "apsp+nmr"}) – "apsp+nmr" runs apsp+, then splits any rank shared by exactly two non-hydrogen atoms (propagated to attached hydrogens), a hack for NMR magnetic (in)equivalencies. Rank requests only, not iRMSD.

  • heavy (bool) – Rank heavy atoms only.

get_chemical_formula(mode: str = 'hill') → str[source]#

Return the chemical formula, omitting counts of 1 as ASE does.

mode="hill" puts C and H first, then the rest alphabetically; any other mode sorts all elements alphabetically, matching ASE.

get_chemical_symbols() → list[str][source]#
get_cn() → ndarray[source]#

Return (N,) coordination numbers from the Fortran core.

get_ids(copy: bool = True) → ndarray | None[source]#

Return (N,) canonical atom IDs or None; copy=False returns the internal array.

get_point_group(**settings) → str | None[source]#

Return the Schoenflies symbol (e.g. “C2v”), or None if skipped.

settings are forwarded; see irmsd.get_point_group().

get_positions(copy: bool = True) → ndarray[source]#

Return (N, 3) positions; copy=False returns the internal array.

get_potential_energy() → float[source]#
get_symmetry_operations(**settings) → tuple[str | None, list][source]#

Return the Schoenflies symbol and the symmetry operations.

settings are forwarded; see irmsd.get_point_group().

Returns:

  • symbol (str or None) – None if skipped.

  • operations (list[irmsd.SymmetryOperation]) – Identity first; empty if skipped.

ids: ndarray | None = None#
info: dict[str, Any]#
property natoms: int#
pbc: tuple[bool, bool, bool] | None = None#
positions: ndarray#
set_atomic_numbers(numbers: Sequence[int]) → None[source]#
set_ids(ids: Sequence[int] | ndarray | None) → None[source]#

Set canonical atom IDs, one per atom (ValueError otherwise); None clears them.

set_positions(positions: Sequence[Sequence[float]]) → None[source]#
set_potential_energy(energy: float | None)[source]#
symbols: list[str]#
class irmsd.SymmetryOperation(label: str, kind: str, order: int, power: int, matrix: ndarray, translation: ndarray, axis: ndarray | None, point: ndarray | None, permutation: ndarray, maxdev: float | None = None)[source]#

Bases: object

Symmetry operation x' = matrix @ x + translation.

Atom i maps onto permutation[i], so op.apply(pos)[i] matches pos[op.permutation[i]] within the symmetry tolerance.

label#

e.g. “E”, “C3”, “C3^2”, “S6^5”, “i”, “sigma”, “Cinf”.

Type:

str

kind#

One of “E”, “C”, “S”, “i”, “sigma”.

Type:

str

order#

n of C_n^k / S_n^k; 0 for Cinf, 1 for E, 2 for i and sigma.

Type:

int

power#

k of C_n^k / S_n^k; 1 for E, i, sigma.

Type:

int

matrix#

Orthogonal part.

Type:

(3, 3) ndarray of float64

translation#

In Å; nonzero when the element misses the origin.

Type:

(3,) ndarray of float64

axis#

Rotation axis, or plane normal for sigma; None for E and i.

Type:

(3,) ndarray of float64 or None

point#

A point on the element in Å; None for E.

Type:

(3,) ndarray of float64 or None

permutation#

0-based image atom of each atom.

Type:

(N,) ndarray of int32

maxdev#

Largest atom displacement in Å; None for derived operations.

Type:

float or None

apply(positions: ndarray) → ndarray[source]#

Apply the operation to (N, 3) positions in Å.

axis: ndarray | None#
kind: str#
label: str#
matrix: ndarray#
maxdev: float | None = None#
order: int#
permutation: ndarray#
point: ndarray | None#
power: int#
translation: ndarray#
irmsd.ase_to_molecule(atoms: ase.Atoms) → Molecule[source]#
irmsd.ase_to_molecule(atoms: Sequence['ase.Atoms']) → list[Molecule]

Convert ASE Atoms (single or sequence) to irmsd.core.Molecule.

Triggers no calculator evaluation and does not modify the input. Energies are converted from eV to Hartree unless info declares energy_units=Hartree; the marker is dropped from the result’s info. A per-atom canonical_id array is carried over as Molecule.ids.

Parameters:

atoms (ase.Atoms or Sequence[ase.Atoms])

Returns:

A list, in input order, if atoms is a sequence.

Return type:

Molecule or list[Molecule]

Raises:
  • RuntimeError – If ASE is not installed.

  • TypeError – If the input is neither an ASE Atoms instance nor a sequence of them.

irmsd.compute_axis_and_print(molecule_list: Sequence[Molecule], run_multiple: bool = False) → List[dict][source]#

Compute and print rotational constants and principal axes.

Returns:

Per structure, “Rotational constants (MHz)” (3,) and “Rotation matrix” (3, 3).

Return type:

list[dict]

irmsd.compute_canonical_and_print(molecule_list: Sequence[Molecule], heavy: bool = False, run_multiple: bool = False) → List[ndarray][source]#

Compute and print canonical atom ranks (heavy atoms only if heavy); one integer array per structure.

irmsd.compute_cn_and_print(molecule_list: Sequence[Molecule], run_multiple: bool = False) → List[ndarray][source]#

Compute and print coordination numbers; one integer array per structure.

irmsd.compute_irmsd_and_print(molecule_list: Sequence[Molecule], inversion=None, outfile=None, idx_ref=0, idx_align=1) → None[source]#

Align one pair and print the iRMSD (Angstrom).

Parameters:
  • inversion ({"auto", "on", "off"}) – Inversion handling in the iRMSD routine.

  • outfile (str or None) – Write the aligned pair to <stem>_ref and <stem>_aligned instead of printing it.

  • idx_ref (int) – Reference and probe indices in molecule_list.

  • idx_align (int) – Reference and probe indices in molecule_list.

irmsd.compute_quaternion_rmsd_and_print(molecule_list: Sequence[Molecule], heavy=False, outfile=None, idx_ref=0, idx_align=1) → None[source]#

Align one pair and print the Cartesian RMSD (Angstrom) and U matrix.

Parameters:
  • heavy (bool) – Restrict the RMSD to heavy atoms.

  • outfile (str or None) – Write the aligned probe here instead of printing it.

  • idx_ref (int) – Reference and probe indices in molecule_list.

  • idx_align (int) – Reference and probe indices in molecule_list.

irmsd.compute_symmetry_and_print(molecule_list: Sequence[Molecule], run_multiple: bool = False, **settings) → List[dict][source]#

Determine and print the Schoenflies point group and operations.

Parameters:

**settings – threshold, primary_threshold, max_axis_order, max_opt_cycles, max_atoms; see irmsd.get_point_group().

Returns:

Per structure, the printed “Point group” and “Symmetry operations” strings and the irmsd.SymmetryOperation list under “operations”. Structures skipped for size carry only a “skipped” point group.

Return type:

list[dict]

irmsd.cregen(molecule_list: Sequence[Molecule], rthr: float = 0.125, ethr: float = 7.96800686e-05, bthr: float = 0.01, printlvl: int = 0, ewin: float | None = None) → List[Molecule][source]#

High-level wrapper around the Fortran-backed cregen_raw that operates directly on Molecule objects. Returns a pruned & energy-sorted list of structures.

Parameters:
  • molecule_list (Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.

  • rthr (float) – Distance threshold for the sorter (passed through to the backend).

  • ethr (float) – Inter-conformer energy threshold (in Hartree)

  • bthr (float) – Inter-conformer rotational constant threshold (fractional)

  • iinversion (int, optional) – Inversion symmetry flag, passed through to the backend.

  • printlvl (int, optional) – Verbosity level, passed through to the backend.

  • ewin (float | None) – Optional energy window to limit ensembe size around lowest energy structure. In Hartree.

Returns:

new_molecule_list – New Molecule objects reconstructed from the sorted atomic numbers and positions returned by the backend. The list contains only n_structures defined as unique according to the selected thresholds.

Return type:

list[Molecule]

Raises:
  • TypeError – If molecule_list does not contain Molecule instances.

  • ValueError – If molecule_list is empty or if the Molecules do not all have the same number of atoms.

irmsd.delta_irmsd_list_molecule(molecule_list: Sequence[Molecule], iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, use_ids: bool = True) → Tuple[ndarray, List[Molecule]][source]#

High-level wrapper around the Fortran-backed delta_irmsd_list that operates directly on Molecule objects.

Parameters:
  • molecule_list (Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.

  • iinversion (int, optional) – Inversion symmetry flag, passed through to the backend.

  • allcanon (bool, optional) – Canonicalization flag, passed through to the backend.

  • printlvl (int, optional) – Verbosity level, passed through to the backend.

Returns:

  • delta (np.ndarray) – Float array returned by the backend (see delta_irmsd_list for detailed semantics).

  • new_molecule_list (list[Molecule]) – New Molecule objects reconstructed from the atomic numbers and positions returned by the backend. The list has the same length and ordering as molecule_list.

Raises:
  • TypeError – If molecule_list does not contain Molecule instances.

  • ValueError – If molecule_list is empty or if the Molecules do not all have the same number of atoms.

irmsd.get_axis(atom_numbers: ndarray, positions: ndarray) → Tuple[ndarray, ndarray, ndarray][source]#

Core API: call the Fortran routine to calculate the rotation axis, average moment, and eigenvectors

Parameters:
  • atom_numbers ((N,) int32-like) – Atomic numbers (or types).

  • positions ((N, 3) float64-like) – Cartesian coordinates in Å.

Returns:

  • rot ((3,) float64 ndarray) – Rotation axis.

  • avmom ((1,) float64 ndarray) – Average moment.

  • evec ((3, 3) float64 ndarray) – Eigenvectors.

Raises:

ValueError – If positions does not have shape (N, 3).

irmsd.get_axis_ase(atoms) → tuple[ndarray, ndarray, ndarray][source]#

Principal-axis data via Molecule.get_axis().

Parameters:

atoms (ase.Atoms)

Returns:

rot_constants_MHz, avg_momentum_au, rotation_matrix

Return type:

tuple[np.ndarray, np.ndarray, np.ndarray]

irmsd.get_axis_rdkit(molecule, conf_id: None | int | list[int] = None) → Tuple[ndarray, ndarray, ndarray] | List[Tuple[ndarray, ndarray, ndarray]][source]#

Principal axes for the selected conformers.

Parameters:
  • molecule (rdkit.Chem.Mol)

  • conf_id (int, list of int, or None) – Conformer ID(s); None selects all conformers.

Returns:

(rotational constants, average moments, eigenvectors) per conformer.

Return type:

tuple of np.ndarray, or list of such tuples

Raises:

TypeError – If molecule is not an RDKit Mol.

irmsd.get_canonical_ase(atoms, wbo: ndarray | None = None, invtype: str = 'apsp+', heavy: bool = False) → ndarray[source]#

Canonical ranks / invariants via Molecule.get_canonical().

Parameters:
  • atoms (ase.Atoms)

  • wbo (np.ndarray or None, optional) – Wiberg bond order matrix or similar, forwarded to the Fortran backend.

  • invtype (str, optional) – Invariant type selector forwarded to the backend.

  • heavy (bool, optional) – Restrict invariants to heavy atoms.

Return type:

np.ndarray

Raises:
  • RuntimeError – If ASE is not installed.

  • TypeError – If atoms is not an ASE Atoms instance.

irmsd.get_canonical_fortran(atom_numbers: ndarray, positions: ndarray, wbo: ndarray | None = None, invtype: str = 'apsp+', heavy: bool = False) → ndarray[source]#

Core API: call the Fortran routine to calculate the canonical ranking of atoms

Parameters:
  • atom_numbers ((N,) int32-like) – Atomic numbers (or types).

  • positions ((N, 3) float64-like) – Cartesian coordinates in Å.

  • heavy (bool, optional) – Whether to consider only heavy atoms (default: False).

  • wbo ((natoms, natoms) float64, C-contiguous, optional) – Optional Wiberg bond order matrix, required if invtype is ‘cangen’, ignored in case of ‘apsp+’.

  • invtype (str, optional) – alogrithm type for invariants calculation (default: apsp+), alternativly ‘cangen’. The special value ‘apsp+nmr’ runs ‘apsp+’ and then splits any rank shared by exactly two non-hydrogen atoms into two distinct ranks (NMR equivalency hack); intended for rank requests only, not for iRMSD matching.

Returns:

rank – Rank array.

Return type:

(N,) int32

Raises:

ValueError – If positions does not have shape (N, 3).

irmsd.get_canonical_rdkit(molecule, conf_id: None | int | list[int] = None, wbo: None | ndarray = None, invtype='apsp+', heavy: bool = False) → ndarray[source]#

Canonical atom ranks for the selected conformers.

Parameters:
  • molecule (rdkit.Chem.Mol)

  • conf_id (int, list of int, or None) – Conformer ID(s); None selects all conformers.

  • wbo (np.ndarray, optional) – Wiberg bond orders, required for invtype='cangen'. Shape (n_atoms, n_atoms) to share across conformers, or (n_conf, n_atoms, n_atoms) for one per conformer.

  • invtype (str) – Invariant type.

  • heavy (bool) – Rank heavy atoms only.

Returns:

Shape (n_atoms,) for one conformer, (n_conf, n_atoms) for several.

Return type:

np.ndarray

Raises:

TypeError – If molecule is not an RDKit Mol.

irmsd.get_cn_ase(atoms) → ndarray[source]#

Coordination numbers via Molecule.get_cn(), shape (N,).

Parameters:

atoms (ase.Atoms)

irmsd.get_cn_fortran(atom_numbers: ndarray, positions: ndarray) → ndarray[source]#

Core API: call the Fortran routine to calculate CN

Parameters:
  • atom_numbers ((N,) int32-like) – Atomic numbers (or types).

  • positions ((N, 3) float64-like) – Cartesian coordinates in Å.

Returns:

cn – array with coordination numbers

Return type:

(N) float64 ndarray

Raises:

ValueError – If positions does not have shape (N, 3).

irmsd.get_cn_rdkit(molecule, conf_id: None | int | list[int] = None) → ndarray[source]#

Coordination numbers for the selected conformers.

Parameters:
  • molecule (rdkit.Chem.Mol)

  • conf_id (int, list of int, or None) – Conformer ID(s); None selects all conformers.

Returns:

Shape (n_atoms,) for one conformer, (n_conf, n_atoms) for several.

Return type:

np.ndarray

Raises:

TypeError – If molecule is not an RDKit Mol.

irmsd.get_energies_from_molecule_list(molecule_list: Sequence[Molecule]) → ndarray[source]#

Collect potential energies from a sequence of Molecule objects.

For each Molecule, this function calls get_potential_energy() and stores the result in a 1D NumPy array of dtype float. If the energy is not available (for example, if get_potential_energy() raises AttributeError or returns None), the corresponding entry is set to 0.0.

Parameters:

molecule_list (Sequence[Molecule]) – Sequence of Molecule instances.

Returns:

energies – Array of shape (n_structures,) containing one energy per Molecule.

Return type:

np.ndarray

irmsd.get_irmsd(atom_numbers1: ndarray, positions1: ndarray, atom_numbers2: ndarray, positions2: ndarray, iinversion: int = 0, ranks1: ndarray | None = None, ranks2: ndarray | None = None) → Tuple[float, ndarray, ndarray, ndarray, ndarray][source]#

Core API: call the Fortran routine to calculate the iRMSD between two structures

Parameters:
  • atom_numbers1 ((N1,) int32-like) – Atomic numbers (or types) of structure 1.

  • positions1 ((N1, 3) float64-like) – Cartesian coordinates in Å of structure 1.

  • atom_numbers2 ((N2,) int32-like) – Atomic numbers (or types) of structure 2.

  • positions2 ((N2, 3) float64-like) – Cartesian coordinates in Å of structure 2.

  • iinversion (int, optional) – Whether to consider inversion symmetry. Default is 0 (auto). Set to 1 to use inversion, set to 2 to disable inversion.

  • ranks1 ((N1,) int-like, optional) – Externally supplied per-atom canonical ranks for structure 1 (e.g. read from a file). Used directly only if both ranks1 and ranks2 are given, contain no zeros, and pass the internal consistency check; otherwise the canonical ranks are recomputed. A zero entry is the “not provided” sentinel.

  • ranks2 ((N2,) int-like, optional) – Externally supplied per-atom canonical ranks for structure 2. See ranks1.

Returns:

  • rmsdval (float) – The calculated iRMSD value.

  • Z3 ((N1,) int32 ndarray) – Atomic numbers of the aligned structure 1.

  • P3 ((N1, 3) float64 ndarray) – Aligned and centered coordinates of structure 1.

  • Z4 ((N2,) int32 ndarray) – Atomic numbers of the aligned structure 2.

  • P4 ((N2, 3) float64 ndarray) – Aligned and centered coordinates of structure 2.

Notes

The returned coordinates P3 and P4 are centered at the origin.

Raises:

ValueError – If positions1 or positions2 do not have shape (Ni, 3).

irmsd.get_irmsd_ase(atoms1, atoms2, iinversion: int = 0) → Tuple[float, 'ase.Atoms', 'ase.Atoms'][source]#

ASE wrapper for get_irmsd_molecule.

Parameters:
  • atoms1 (ase.Atoms)

  • atoms2 (ase.Atoms)

  • iinversion (int, optional) – 0 = ‘auto’, 1 = ‘on’, 2 = ‘off’.

Returns:

  • irmsd (float) – In Angstrom.

  • new_atoms1, new_atoms2 (ase.Atoms) – New objects for the transformed structures.

Raises:
  • RuntimeError – If ASE is not installed.

  • TypeError – If inputs are not ASE Atoms.

irmsd.get_irmsd_molecule(molecule1: Molecule, molecule2: Molecule, iinversion: int = 0, use_ids: bool = True) → Tuple[float, Molecule, Molecule][source]#

Compute the iRMSD between two Molecule objects using the iRMSD backend.

The backend may reorder atoms and/or change atomic numbers according to its canonicalization / matching logic. This wrapper returns copies of both input Molecules with the updated atomic numbers and positions.

Parameters:
  • molecule1 (Molecule) – First input structure.

  • molecule2 (Molecule) – Second input structure.

  • iinversion (int, optional) – Inversion flag passed directly to the Fortran backend. See the backend documentation for allowed values and meanings.

  • use_ids (bool, optional) – If True (default) and both molecules carry per-atom canonical IDs (Molecule.ids), those IDs are passed to the backend as externally supplied ranks. The backend uses them only if they are mutually consistent, otherwise it recomputes the canonical ranks.

Returns:

  • irmsd (float) – iRMSD value in Ångström.

  • new_molecule1 (Molecule) – Copy of molecule1 with updated atomic numbers and positions.

  • new_molecule2 (Molecule) – Copy of molecule2 with updated atomic numbers and positions.

Raises:

TypeError – If either input is not a Molecule.

irmsd.get_irmsd_rdkit(molecule_ref, molecule_align, conf_id_ref=-1, conf_id_align=-1, iinversion: int = 0) → Tuple[float, 'Mol', 'Mol'][source]#

iRMSD between two conformers after permutation and alignment.

Parameters:
  • molecule_ref (rdkit.Chem.Mol)

  • molecule_align (rdkit.Chem.Mol)

  • conf_id_ref (int) – Conformer IDs; -1 is the RDKit default conformer.

  • conf_id_align (int) – Conformer IDs; -1 is the RDKit default conformer.

  • iinversion (int) – Inversion handling: 0 auto, 1 on, 2 off.

Returns:

  • irmsd (float) – In Angstrom.

  • ref, aligned (rdkit.Chem.Mol) – Processed reference and aligned molecules.

Raises:

TypeError – If either input is not an RDKit Mol.

irmsd.get_point_group(atom_numbers: ndarray, positions: ndarray, *, threshold: float = 0.1, primary_threshold: float = 0.5, max_axis_order: int = 10, max_opt_cycles: int = 100, max_atoms: int | None = 200) → str | None[source]#

Schoenflies point group from the brute-force analyzer ported from CREST.

Parameters:
  • atom_numbers ((N,) int32-like) – Atomic numbers (or types).

  • positions ((N, 3) float64-like) – Cartesian coordinates in Å.

  • threshold (float, optional) – Final symmetry tolerance in Bohr: the largest atom displacement an element may cause to be accepted (default: 0.1, as in CREST). Increase for noisy or loosely optimized geometries; large systems need this more, since the worst atom deviation grows with size. Too tight a value on such structures is also slow, as every near-miss candidate is fully optimized before being rejected.

  • primary_threshold (float, optional) – Tolerance in Bohr for the initial pairing of symmetry-related atoms and the distance prescreening of candidates (default: 0.5). Too small values miss elements of distorted structures; larger values test many more candidates, which dominates the cost for large systems.

  • max_axis_order (int, optional) – Highest rotation axis order searched, for proper and improper axes alike (default: 10). An S2n axis needs at least 2n, so values below 10 lose e.g. the S10 axes of Ih and D5d and with them the group.

  • max_opt_cycles (int, optional) – Maximum optimization cycles per candidate element (default: 100).

  • max_atoms (int or None, optional) – Skip the analysis for structures with more atoms than this, since the search scales steeply with system size (default: 200, as in CREST). None disables the limit.

Returns:

Schoenflies symbol, e.g. “C2v”, “Dinfh”; the highest proper axis (e.g. “C3”) if no tabulated group matches; None if skipped by max_atoms.

Return type:

str or None

Raises:

ValueError – If positions is not (N, 3) or max_axis_order < 2.

irmsd.get_point_group_ase(atoms, **settings) → str | None[source]#

Schoenflies point group via Molecule.get_point_group().

Parameters:
  • atoms (ase.Atoms)

  • **settings – Analyzer settings (threshold, primary_threshold, max_axis_order, max_opt_cycles, max_atoms), see irmsd.get_point_group().

Returns:

None if the analysis was skipped.

Return type:

str or None

irmsd.get_point_group_rdkit(molecule, conf_id: None | int | list[int] = None, **settings) → str | None | List[str | None][source]#

Schoenflies point group for the selected conformers.

Parameters:
  • molecule (rdkit.Chem.Mol)

  • conf_id (int, list of int, or None) – Conformer ID(s); None selects all conformers.

  • **settings – Analyzer settings, see irmsd.get_point_group().

Returns:

Symbol per conformer; None if the analyzer skipped it.

Return type:

str or None, or list of those

Raises:

TypeError – If molecule is not an RDKit Mol.

irmsd.get_quaternion_rmsd_fortran(atom_numbers1: ndarray, positions1: ndarray, atom_numbers2: ndarray, positions2: ndarray, mask: ndarray | None = None) → Tuple[float, ndarray, ndarray][source]#

Pair API: call the Fortran routine on TWO structures.

Parameters:
  • atom_numbers1 ((N1,) int32-like)

  • positions1 ((N1, 3) float64-like)

  • atom_numbers2 ((N2,) int32-like)

  • positions2 ((N2, 3) float64-like)

  • mask ((N1,) bool-like or None)

Returns:

  • rmsdval (float64)

  • new_positions2 ((N2, 3) float64, (positions2 @ Umat.T))

  • Umat ((3, 3) float64 (Fortran-ordered))

Notes

The returned new_positions2 is aligned onto positions1. 1. If mask is provided, only the atoms where mask==True in structure 1 are used to compute the RMSD. 2. The rotation matrix Umat is Fortran-ordered, i.e., to rotate positions2, do: new_positions2 = positions2 @ Umat.T 3. The returned new_positions2 is also translated to have the same barycenter as positions1.

Raises:

ValueError – If the input arrays do not have the correct shapes or types.

irmsd.get_rmsd_ase(atoms1, atoms2, mask=None) → Tuple[float, 'ase.Atoms', np.ndarray][source]#

ASE wrapper for get_rmsd_molecule.

Parameters:
  • atoms1 (ase.Atoms) – Reference structure.

  • atoms2 (ase.Atoms) – Structure aligned onto atoms1.

  • mask (array-like of bool, optional) – Atoms of the first structure that enter the RMSD.

Returns:

  • rmsd (float) – In Angstrom.

  • new_atoms2 (ase.Atoms) – New object with coordinates aligned to atoms1.

  • rotation_matrix (np.ndarray) – 3x3 rotation used for the alignment.

Raises:
  • RuntimeError – If ASE is not installed.

  • TypeError – If inputs are not ASE Atoms.

irmsd.get_rmsd_molecule(molecule1: Molecule, molecule2: Molecule, mask=None) → Tuple[float, Molecule, ndarray][source]#

Compute the RMSD between two Molecule objects using the quaternion-based RMSD backend, and return an aligned copy of the second Molecule.

Parameters:
  • molecule1 (Molecule) – Reference structure.

  • molecule2 (Molecule) – Structure to be rotated/translated onto molecule1.

  • mask (array-like of bool, optional) – Optional mask selecting which atoms of molecule1 participate in the RMSD. Must be broadcastable / compatible with the Fortran backend’s mask semantics.

Returns:

  • rmsd (float) – Root-mean-square deviation in Ångström.

  • new_molecule2 (Molecule) – Copy of molecule2 with its positions replaced by the aligned coordinates returned by the backend.

  • rotation_matrix (np.ndarray) – 3×3 rotation matrix applied to align molecule2 onto molecule1.

Raises:

TypeError – If either input is not a Molecule.

irmsd.get_rmsd_rdkit(molecule_ref, molecule_align, conf_id_ref=-1, conf_id_align=-1, mask=None) → Tuple[float, 'Mol', np.ndarray][source]#

RMSD between two conformers after alignment, without permutation.

Parameters:
  • molecule_ref (rdkit.Chem.Mol)

  • molecule_align (rdkit.Chem.Mol)

  • conf_id_ref (int) – Conformer IDs; -1 is the RDKit default conformer.

  • conf_id_align (int) – Conformer IDs; -1 is the RDKit default conformer.

  • mask (array-like of bool, optional) – Atoms to include in the RMSD.

Returns:

  • rmsd (float) – In Angstrom.

  • aligned (rdkit.Chem.Mol)

  • rotmat (np.ndarray) – Rotation matrix.

Raises:

TypeError – If either input is not an RDKit Mol.

irmsd.get_symmetry_elements(atom_numbers: ndarray, positions: ndarray, *, threshold: float = 0.1, primary_threshold: float = 0.5, max_axis_order: int = 10, max_opt_cycles: int = 100, max_atoms: int | None = 200) → tuple[str | None, list[SymmetryOperation]][source]#

Point group and located symmetry elements, each as its generating operation.

Parameters:
  • atom_numbers ((N,) int32-like) – Atomic numbers (or types).

  • positions ((N, 3) float64-like) – Cartesian coordinates in Å.

  • threshold (float, optional) – Final symmetry tolerance in Bohr: the largest atom displacement an element may cause to be accepted (default: 0.1, as in CREST). Increase for noisy or loosely optimized geometries; large systems need this more, since the worst atom deviation grows with size. Too tight a value on such structures is also slow, as every near-miss candidate is fully optimized before being rejected.

  • primary_threshold (float, optional) – Tolerance in Bohr for the initial pairing of symmetry-related atoms and the distance prescreening of candidates (default: 0.5). Too small values miss elements of distorted structures; larger values test many more candidates, which dominates the cost for large systems.

  • max_axis_order (int, optional) – Highest rotation axis order searched, for proper and improper axes alike (default: 10). An S2n axis needs at least 2n, so values below 10 lose e.g. the S10 axes of Ih and D5d and with them the group.

  • max_opt_cycles (int, optional) – Maximum optimization cycles per candidate element (default: 100).

  • max_atoms (int or None, optional) – Skip the analysis for structures with more atoms than this, since the search scales steeply with system size (default: 200, as in CREST). None disables the limit.

Returns:

  • symbol (str or None) – Schoenflies symbol; None if skipped by max_atoms.

  • elements (list[SymmetryOperation]) – Empty if skipped. A Cinf axis has order == 0.

Raises:

ValueError – If positions is not (N, 3) or max_axis_order < 2.

irmsd.get_symmetry_operations(atom_numbers: ndarray, positions: ndarray, *, threshold: float = 0.1, primary_threshold: float = 0.5, max_axis_order: int = 10, max_opt_cycles: int = 100, max_atoms: int | None = 200) → tuple[str | None, list[SymmetryOperation]][source]#

Point group and all distinct powers of the located elements.

Order: E, i, mirror planes, proper, then improper rotations. For finite groups the length equals the group order; for Cinfv/Dinfh only the operations generated by i, sigma and C2 are returned (the Cinf axis comes from get_symmetry_elements()).

Parameters:
  • atom_numbers ((N,) int32-like) – Atomic numbers (or types).

  • positions ((N, 3) float64-like) – Cartesian coordinates in Å.

  • threshold (float, optional) – Final symmetry tolerance in Bohr: the largest atom displacement an element may cause to be accepted (default: 0.1, as in CREST). Increase for noisy or loosely optimized geometries; large systems need this more, since the worst atom deviation grows with size. Too tight a value on such structures is also slow, as every near-miss candidate is fully optimized before being rejected.

  • primary_threshold (float, optional) – Tolerance in Bohr for the initial pairing of symmetry-related atoms and the distance prescreening of candidates (default: 0.5). Too small values miss elements of distorted structures; larger values test many more candidates, which dominates the cost for large systems.

  • max_axis_order (int, optional) – Highest rotation axis order searched, for proper and improper axes alike (default: 10). An S2n axis needs at least 2n, so values below 10 lose e.g. the S10 axes of Ih and D5d and with them the group.

  • max_opt_cycles (int, optional) – Maximum optimization cycles per candidate element (default: 100).

  • max_atoms (int or None, optional) – Skip the analysis for structures with more atoms than this, since the search scales steeply with system size (default: 200, as in CREST). None disables the limit.

Returns:

  • symbol (str or None) – Schoenflies symbol; None if skipped by max_atoms.

  • operations (list[SymmetryOperation]) – Empty if skipped.

Raises:

ValueError – If positions is not (N, 3) or max_axis_order < 2.

irmsd.molecule_to_ase(molecules: Molecule) → ase.Atoms[source]#
irmsd.molecule_to_ase(molecules: Sequence[Molecule]) → list['ase.Atoms']

Convert Molecule(s) to ASE Atoms without attaching a calculator.

The Hartree energy is written to info["energy"] in eV unless info already has an energy entry; any energy_units marker is dropped. Molecule.ids becomes a per-atom canonical_id array. The returned objects do not share state with the input.

Parameters:

molecules (Molecule or Sequence[Molecule])

Returns:

A list, in input order, if molecules is a sequence.

Return type:

ase.Atoms or list[ase.Atoms]

Raises:
  • RuntimeError – If ASE is not installed.

  • TypeError – If the input is neither a Molecule nor a sequence of Molecules.

irmsd.molecule_to_rdkit(molecule: Molecule) → Mol[source]#
irmsd.molecule_to_rdkit(molecules: Sequence[Molecule]) → list['Mol']

Convert irmsd Molecule(s) to RDKit Mol(s) carrying atoms and one conformer.

No bonds or properties are transferred. Returns a single Mol if exactly one Molecule results, else a list.

Raises:

TypeError – If the input is not a Molecule or a list of them.

irmsd.prune(molecule_list: Sequence[Molecule], rthr: float, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None, use_ids: bool = True) → List[Molecule][source]#

High-level wrapper around the Fortran-backed sorter_irmsd that operates directly on Molecule objects. Returns a pruned list of structures.

Parameters:
  • molecule_list (Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.

  • rthr (float) – Distance threshold for the sorter (passed through to the backend).

  • iinversion (int, optional) – Inversion symmetry flag, passed through to the backend.

  • allcanon (bool, optional) – Canonicalization flag, passed through to the backend.

  • printlvl (int, optional) – Verbosity level, passed through to the backend.

  • ethr (float | None) – Optional energy threshold to accelerate by pre-sorting

  • ewin (float | None) – Optional energy window to limit ensembe size around lowest energy structure. In Hartree.

Returns:

new_molecule_list – New Molecule objects reconstructed from the sorted atomic numbers and positions returned by the backend. The list contains only n_structures defined as unique according to the selected thresholds.

Return type:

list[Molecule]

Raises:
  • TypeError – If molecule_list does not contain Molecule instances.

  • ValueError – If molecule_list is empty or if the Molecules do not all have the same number of atoms.

irmsd.rdkit_to_molecule(molecules: Mol, conf_id: int | list[int] | None = None) → Molecule | list[Molecule][source]#
irmsd.rdkit_to_molecule(molecules: Sequence['Mol'], conf_id: int | list[int] | None = None) → list[Molecule]

Convert RDKit Mol(s) to irmsd Molecule(s), one per selected conformer.

Mol-level and conformer-level properties are merged into info; conformer properties win on key clashes.

Parameters:
  • molecules (rdkit.Chem.Mol or list of rdkit.Chem.Mol)

  • conf_id (int, list of int, or None, optional) – Conformer ID(s); None selects all conformers.

Returns:

A single Molecule if exactly one conformer was converted, else a list.

Return type:

Molecule or list of Molecule

Raises:

TypeError – If the input is not a Mol or list of Mols, or a conformer is not 3D.

irmsd.read_structures(paths: str | Sequence[str]) → List[Molecule][source]#

Read an arbitrary number of structures and return them as Molecule objects.

For each path, this routine behaves as follows:

  • If the file extension is .xyz or .extxyz, it uses the internal read_extxyz helper to obtain one or more Molecule objects.

  • For all other file types, it attempts to import ASE via require_ase(), uses ase.io.read to read one or more ASE Atoms objects, and converts them into Molecule objects using ase_to_molecule.

Multi-frame files:

If a file contains multiple frames, all frames are read and appended to the output list. A short informational message is printed indicating the number of frames that were found.

Parameters:

paths (Sequence[str]) – File paths to read.

Returns:

structures – One Molecule per frame found across all input paths.

Return type:

list[Molecule]

irmsd.run_cregen_and_print(molecule_list: Sequence[Molecule], rthr: float, ethr: float, bthr: float, ewin: float | None = None, printlvl: int = 0, maxprint: int = 25, outfile: str | None = None) → None[source]#

Run CREGEN per sum-formula group and print a summary per group.

Parameters:
  • rthr (float) – RMSD, energy, and rotational-constant thresholds for conformer identification.

  • ethr (float) – RMSD, energy, and rotational-constant thresholds for conformer identification.

  • bthr (float) – RMSD, energy, and rotational-constant thresholds for conformer identification.

  • maxprint (int) – Max rows per printed result table.

  • outfile (str or None) – Write the representatives here; with several formulas, one <root>_<formula><ext> file each.

irmsd.sort_get_delta_irmsd_and_print(molecule_list: Sequence[Molecule], inversion: str = None, allcanon: bool = True, printlvl: int = 0, maxprint: int = 25, outfile: str | None = None) → None[source]#

Split by sum formula, energy-sort, and print the iRMSD between consecutive structures of each group.

Parameters:
  • inversion ({"auto", "on", "off"})

  • maxprint (int) – Max rows per printed result table.

  • outfile (str or None) – Unused; no structures are written.

irmsd.sort_structures_and_print(molecule_list: Sequence[Molecule], rthr: float, inversion: str = None, allcanon: bool = True, printlvl: int = 0, maxprint: int = 25, ethr: float | None = None, ewin: float | None = None, outfile: str | None = None) → None[source]#

Split by sum formula, energy-sort, prune each group with the iRMSD sorter, and print a summary per group.

Parameters:
  • rthr (float) – Distance threshold for the sorter.

  • inversion ({"auto", "on", "off"})

  • maxprint (int) – Max rows per printed result table.

  • ethr (float or None) – Inter-conformer energy threshold for presorting.

  • ewin (float or None) – Energy window around the lowest-energy structure.

  • outfile (str or None) – Write the representatives here; with several formulas, one <root>_<formula><ext> file each.

irmsd.sorter_irmsd(atom_numbers_list: Sequence[ndarray], positions_list: Sequence[ndarray], nat: int, rthr: float, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, energies_list: Sequence[ndarray] | None = None, ids_list: Sequence[ndarray | None] | None = None) → Tuple[ndarray, List[ndarray], List[ndarray]][source]#

High-level API: call the sorter_exposed_xyz_fortran Fortran routine.

Parameters:
  • atom_numbers_list (sequence of (N,) int32 arrays) – Per-structure atom numbers.

  • positions_list (sequence of (N,3) float64 arrays) – Per-structure coordinates.

  • nat (int) – Number of atoms for which the groups array is defined. Must satisfy 1 <= nat <= N.

  • rthr (float) – Distance threshold for the Fortran sorter. In Angström.

  • iinversion (int) – Inversion symmetry flag.

  • allcanon (bool) – Canonicalization flag.

  • printlvl (int) – Verbosity level.

  • ethr (float | None) – Inter-conformer energy threshold (optional). In Hartree.

  • energies_list (sequence of (Nall,) floats | None) – List of energies for the passed structures (optional). In Hartree.

Returns:

  • groups ((nat,) int32) – Group index for each of the first nat atoms.

  • xyz_structs (list of (N,3) float64 arrays) – Updated coordinates for each structure.

  • Z_structs (list of (N,) int32 arrays) – Updated atom numbers for each structure.

Raises:

ValueError – If input arrays have inconsistent shapes or invalid parameters.

irmsd.sorter_irmsd_ase(atoms_list: Sequence['ase.Atoms'], rthr: float = 0.125, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None) → Tuple[np.ndarray, List['ase.Atoms']][source]#

ASE wrapper for sorter_irmsd_molecule.

Parameters:
  • atoms_list (Sequence[ase.Atoms]) – List or tuple; all structures must have the same atom count.

  • rthr (float) – Distance threshold for the sorter.

  • iinversion (int, optional) – 0 = ‘auto’, 1 = ‘on’, 2 = ‘off’.

  • allcanon (bool, optional) – Canonicalization flag.

  • printlvl (int, optional) – Verbosity level.

  • ethr (float or None) – Energy threshold for pre-sorting, in Hartree.

  • ewin (float or None) – Energy window above the lowest structure, in Hartree.

Returns:

  • groups (np.ndarray) – Integer group index per structure.

  • new_atoms_list (list[ase.Atoms]) – Rebuilt from the sorted Molecules.

irmsd.sorter_irmsd_molecule(molecule_list: Sequence[Molecule], rthr: float = 0.125, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None, use_ids: bool = True) → Tuple[ndarray, List[Molecule]][source]#

High-level wrapper around the Fortran-backed sorter_irmsd that operates directly on Molecule objects.

Parameters:
  • molecule_list (Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.

  • rthr (float) – Distance threshold for the sorter (passed through to the backend).

  • iinversion (int, optional) – Inversion symmetry flag, passed through to the backend.

  • allcanon (bool, optional) – Canonicalization flag, passed through to the backend.

  • printlvl (int, optional) – Verbosity level, passed through to the backend.

  • ethr (float | None) – Optional energy threshold to accelerate by pre-sorting

  • ewin (float | None) – Optional energy window to limit ensembe size around lowest energy structure. In Hartree.

Returns:

  • groups (np.ndarray) – Integer array of shape (nat,) with group indices for the first nat atoms (as defined by the backend).

  • new_molecule_list (list[Molecule]) – New Molecule objects reconstructed from the sorted atomic numbers and positions returned by the backend. The list has the same length and ordering as molecule_list.

Raises:
  • TypeError – If molecule_list does not contain Molecule instances.

  • ValueError – If molecule_list is empty or if the Molecules do not all have the same number of atoms.

irmsd.sorter_irmsd_rdkit(molecules: 'Mol' | Sequence['Mol'], rthr: float = 0.125, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None) → Tuple[np.ndarray, List['Mol']][source]#

Group an ensemble by iRMSD.

Parameters:
  • molecules (rdkit.Chem.Mol or list of rdkit.Chem.Mol) – A single Mol must carry more than one conformer.

  • rthr (float) – iRMSD threshold in Angstrom.

  • iinversion (int) – Inversion handling: 0 auto, 1 on, 2 off.

  • allcanon – Canonicalization flag and verbosity, passed to the backend.

  • printlvl – Canonicalization flag and verbosity, passed to the backend.

  • ethr (float or None) – Energy pre-sorting threshold.

  • ewin (float or None) – Energy window in Hartree above the lowest-energy structure.

Returns:

  • groups (np.ndarray) – Integer group index per structure.

  • new_molecules_list (list of rdkit.Chem.Mol) – Sorted structures.

Raises:

TypeError – If the input is not an RDKit Mol or a list of them.

irmsd.write_structures(filename: str | Path, structures: Molecule, mode: str = 'w') → None[source]#
irmsd.write_structures(filename: str | Path, structures: Sequence[Molecule], mode: str = 'w') → None

High-level structure writer for the irmsd Molecule type.

This routine mirrors the behaviour of read_structures on the output side: it chooses the appropriate backend based on the file extension and accepts either a single Molecule or a sequence of Molecule objects.

Dispatch rules#

  • If the filename has no extension, ‘.xyz’ is appended and the internal extended-XYZ writer is used.

  • If the filename ends with ‘.xyz’, ‘.extxyz’ or ‘.trj’ (case-insensitive), the structures are written using the internal extended-XYZ writer (write_extxyz). The mode argument is passed through and controls whether the file is overwritten (‘w’, default) or appended to (‘a’).

  • For all other filename extensions, ASE is used as a backend. The structures are first converted to ASE Atoms objects using molecule_to_ase, and then written via ase.io.write. In this case the mode argument is currently ignored and ASE’s default behaviour for the chosen format is used.

param filename:

Output filename. Its extension determines the backend.

type filename:

str or pathlib.Path

param structures:

A single Molecule or a sequence of Molecules to be written. For extended XYZ, multiple Molecules are written as consecutive frames in one file.

type structures:

Molecule or Sequence[Molecule]

param mode:

File open mode for extended XYZ output. Ignored for non-XYZ formats handled via ASE.

type mode:

{"w", "a"}, optional

raises RuntimeError:

If ASE is required (non-XYZ formats) but not installed.

raises TypeError:

If structures is not a Molecule or a sequence of Molecules.