irmsd#
- class irmsd.Molecule(symbols: list[str], positions: ~numpy.ndarray, energy: float | None = None, info: dict[str, ~typing.Any] = <factory>, cell: ~numpy.ndarray | None = None, pbc: tuple[bool, bool, bool] | None = None, ids: ~numpy.ndarray | None = None)[source]#
Bases:
objectDependency-free replacement for ase.Atoms.
energyis in Hartree.idsholds optional per-atom canonical atom identifiers, one int32 per atom. Construction raises ValueError on bad shapes or unknown symbols.- cell: ndarray | None = None#
- energy: float | None = None#
- get_axis() tuple[ndarray, ndarray, ndarray][source]#
Return rotation constants, average momentum and rotation matrix.
- Returns:
rot (
(3,) ndarray) – Rotation constants in MHz.avmom (
(1,) ndarray) – Average momentum in a.u.evec (
(3,3) ndarray)
- get_canonical(wbo: ndarray | None = None, invtype: str = 'apsp+', heavy: bool = False) ndarray[source]#
Return (N,) int32 canonical ranks from the Fortran core.
- Parameters:
wbo (
(N,N) ndarray, optional) – Wiberg bond orders; required forinvtype="cangen".invtype (
{"apsp+", "cangen", "apsp+nmr"}) –"apsp+nmr"runs apsp+, then splits any rank shared by exactly two non-hydrogen atoms (propagated to attached hydrogens), a hack for NMR magnetic (in)equivalencies. Rank requests only, not iRMSD.heavy (
bool) – Rank heavy atoms only.
- get_chemical_formula(mode: str = 'hill') str[source]#
Return the chemical formula, omitting counts of 1 as ASE does.
mode="hill"puts C and H first, then the rest alphabetically; any other mode sorts all elements alphabetically, matching ASE.
- get_ids(copy: bool = True) ndarray | None[source]#
Return (N,) canonical atom IDs or None;
copy=Falsereturns the internal array.
- get_point_group(**settings) str | None[source]#
Return the Schoenflies symbol (e.g. “C2v”), or None if skipped.
settingsare forwarded; seeirmsd.get_point_group().
- get_positions(copy: bool = True) ndarray[source]#
Return (N, 3) positions;
copy=Falsereturns the internal array.
- get_symmetry_operations(**settings) tuple[str | None, list][source]#
Return the Schoenflies symbol and the symmetry operations.
settingsare forwarded; seeirmsd.get_point_group().- Returns:
symbol (
strorNone) – None if skipped.operations (
list[irmsd.SymmetryOperation]) – Identity first; empty if skipped.
- ids: ndarray | None = None#
- info: dict[str, Any]#
- property natoms: int#
- pbc: tuple[bool, bool, bool] | None = None#
- positions: ndarray#
- set_ids(ids: Sequence[int] | ndarray | None) None[source]#
Set canonical atom IDs, one per atom (ValueError otherwise); None clears them.
- symbols: list[str]#
- class irmsd.SymmetryOperation(label: str, kind: str, order: int, power: int, matrix: ndarray, translation: ndarray, axis: ndarray | None, point: ndarray | None, permutation: ndarray, maxdev: float | None = None)[source]#
Bases:
objectSymmetry operation
x' = matrix @ x + translation.Atom
imaps ontopermutation[i], soop.apply(pos)[i]matchespos[op.permutation[i]]within the symmetry tolerance.- label#
e.g. “E”, “C3”, “C3^2”, “S6^5”, “i”, “sigma”, “Cinf”.
- Type:
str
- kind#
One of “E”, “C”, “S”, “i”, “sigma”.
- Type:
str
- order#
n of C_n^k / S_n^k; 0 for Cinf, 1 for E, 2 for i and sigma.
- Type:
int
- power#
k of C_n^k / S_n^k; 1 for E, i, sigma.
- Type:
int
- matrix#
Orthogonal part.
- Type:
(3,3) ndarrayoffloat64
- translation#
In Å; nonzero when the element misses the origin.
- Type:
(3,) ndarrayoffloat64
- axis#
Rotation axis, or plane normal for sigma; None for E and i.
- Type:
(3,) ndarrayoffloat64orNone
- point#
A point on the element in Å; None for E.
- Type:
(3,) ndarrayoffloat64orNone
- permutation#
0-based image atom of each atom.
- Type:
(N,) ndarrayofint32
- maxdev#
Largest atom displacement in Å; None for derived operations.
- Type:
floatorNone
- axis: ndarray | None#
- kind: str#
- label: str#
- matrix: ndarray#
- maxdev: float | None = None#
- order: int#
- permutation: ndarray#
- point: ndarray | None#
- power: int#
- translation: ndarray#
- irmsd.ase_to_molecule(atoms: ase.Atoms) Molecule[source]#
- irmsd.ase_to_molecule(atoms: Sequence['ase.Atoms']) list[Molecule]
Convert ASE Atoms (single or sequence) to irmsd.core.Molecule.
Triggers no calculator evaluation and does not modify the input. Energies are converted from eV to Hartree unless
infodeclaresenergy_units=Hartree; the marker is dropped from the result’sinfo. A per-atomcanonical_idarray is carried over asMolecule.ids.- Parameters:
atoms (
ase.AtomsorSequence[ase.Atoms])- Returns:
A list, in input order, if atoms is a sequence.
- Return type:
Moleculeorlist[Molecule]- Raises:
RuntimeError – If ASE is not installed.
TypeError – If the input is neither an ASE Atoms instance nor a sequence of them.
- irmsd.compute_axis_and_print(molecule_list: Sequence[Molecule], run_multiple: bool = False) List[dict][source]#
Compute and print rotational constants and principal axes.
- Returns:
Per structure, “Rotational constants (MHz)” (3,) and “Rotation matrix” (3, 3).
- Return type:
list[dict]
- irmsd.compute_canonical_and_print(molecule_list: Sequence[Molecule], heavy: bool = False, run_multiple: bool = False) List[ndarray][source]#
Compute and print canonical atom ranks (heavy atoms only if
heavy); one integer array per structure.
- irmsd.compute_cn_and_print(molecule_list: Sequence[Molecule], run_multiple: bool = False) List[ndarray][source]#
Compute and print coordination numbers; one integer array per structure.
- irmsd.compute_irmsd_and_print(molecule_list: Sequence[Molecule], inversion=None, outfile=None, idx_ref=0, idx_align=1) None[source]#
Align one pair and print the iRMSD (Angstrom).
- Parameters:
inversion (
{"auto", "on", "off"}) – Inversion handling in the iRMSD routine.outfile (
strorNone) – Write the aligned pair to<stem>_refand<stem>_alignedinstead of printing it.idx_ref (
int) – Reference and probe indices inmolecule_list.idx_align (
int) – Reference and probe indices inmolecule_list.
- irmsd.compute_quaternion_rmsd_and_print(molecule_list: Sequence[Molecule], heavy=False, outfile=None, idx_ref=0, idx_align=1) None[source]#
Align one pair and print the Cartesian RMSD (Angstrom) and U matrix.
- Parameters:
heavy (
bool) – Restrict the RMSD to heavy atoms.outfile (
strorNone) – Write the aligned probe here instead of printing it.idx_ref (
int) – Reference and probe indices inmolecule_list.idx_align (
int) – Reference and probe indices inmolecule_list.
- irmsd.compute_symmetry_and_print(molecule_list: Sequence[Molecule], run_multiple: bool = False, **settings) List[dict][source]#
Determine and print the Schoenflies point group and operations.
- Parameters:
**settings – threshold, primary_threshold, max_axis_order, max_opt_cycles, max_atoms; see
irmsd.get_point_group().- Returns:
Per structure, the printed “Point group” and “Symmetry operations” strings and the
irmsd.SymmetryOperationlist under “operations”. Structures skipped for size carry only a “skipped” point group.- Return type:
list[dict]
- irmsd.cregen(molecule_list: Sequence[Molecule], rthr: float = 0.125, ethr: float = 7.96800686e-05, bthr: float = 0.01, printlvl: int = 0, ewin: float | None = None) List[Molecule][source]#
High-level wrapper around the Fortran-backed
cregen_rawthat operates directly on Molecule objects. Returns a pruned & energy-sorted list of structures.- Parameters:
molecule_list (
Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.rthr (
float) – Distance threshold for the sorter (passed through to the backend).ethr (
float) – Inter-conformer energy threshold (in Hartree)bthr (
float) – Inter-conformer rotational constant threshold (fractional)iinversion (
int, optional) – Inversion symmetry flag, passed through to the backend.printlvl (
int, optional) – Verbosity level, passed through to the backend.ewin (
float | None) – Optional energy window to limit ensembe size around lowest energy structure. In Hartree.
- Returns:
new_molecule_list – New Molecule objects reconstructed from the sorted atomic numbers and positions returned by the backend. The list contains only n_structures defined as
uniqueaccording to the selected thresholds.- Return type:
list[Molecule]- Raises:
TypeError – If
molecule_listdoes not contain Molecule instances.ValueError – If
molecule_listis empty or if the Molecules do not all have the same number of atoms.
- irmsd.delta_irmsd_list_molecule(molecule_list: Sequence[Molecule], iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, use_ids: bool = True) Tuple[ndarray, List[Molecule]][source]#
High-level wrapper around the Fortran-backed
delta_irmsd_listthat operates directly on Molecule objects.- Parameters:
molecule_list (
Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.iinversion (
int, optional) – Inversion symmetry flag, passed through to the backend.allcanon (
bool, optional) – Canonicalization flag, passed through to the backend.printlvl (
int, optional) – Verbosity level, passed through to the backend.
- Returns:
delta (
np.ndarray) – Float array returned by the backend (seedelta_irmsd_listfor detailed semantics).new_molecule_list (
list[Molecule]) – New Molecule objects reconstructed from the atomic numbers and positions returned by the backend. The list has the same length and ordering asmolecule_list.
- Raises:
TypeError – If
molecule_listdoes not contain Molecule instances.ValueError – If
molecule_listis empty or if the Molecules do not all have the same number of atoms.
- irmsd.get_axis(atom_numbers: ndarray, positions: ndarray) Tuple[ndarray, ndarray, ndarray][source]#
Core API: call the Fortran routine to calculate the rotation axis, average moment, and eigenvectors
- Parameters:
atom_numbers (
(N,) int32-like) – Atomic numbers (or types).positions (
(N,3) float64-like) – Cartesian coordinates in Å.
- Returns:
rot (
(3,) float64 ndarray) – Rotation axis.avmom (
(1,) float64 ndarray) – Average moment.evec (
(3,3) float64 ndarray) – Eigenvectors.
- Raises:
ValueError – If positions does not have shape (N, 3).
- irmsd.get_axis_ase(atoms) tuple[ndarray, ndarray, ndarray][source]#
Principal-axis data via Molecule.get_axis().
- Parameters:
atoms (
ase.Atoms)- Returns:
rot_constants_MHz, avg_momentum_au, rotation_matrix
- Return type:
tuple[np.ndarray,np.ndarray,np.ndarray]
- irmsd.get_axis_rdkit(molecule, conf_id: None | int | list[int] = None) Tuple[ndarray, ndarray, ndarray] | List[Tuple[ndarray, ndarray, ndarray]][source]#
Principal axes for the selected conformers.
- Parameters:
molecule (
rdkit.Chem.Mol)conf_id (
int,listofint, orNone) – Conformer ID(s); None selects all conformers.
- Returns:
(rotational constants, average moments, eigenvectors) per conformer.
- Return type:
tupleofnp.ndarray, orlistofsuch tuples- Raises:
TypeError – If molecule is not an RDKit Mol.
- irmsd.get_canonical_ase(atoms, wbo: ndarray | None = None, invtype: str = 'apsp+', heavy: bool = False) ndarray[source]#
Canonical ranks / invariants via Molecule.get_canonical().
- Parameters:
atoms (
ase.Atoms)wbo (
np.ndarrayorNone, optional) – Wiberg bond order matrix or similar, forwarded to the Fortran backend.invtype (
str, optional) – Invariant type selector forwarded to the backend.heavy (
bool, optional) – Restrict invariants to heavy atoms.
- Return type:
np.ndarray- Raises:
RuntimeError – If ASE is not installed.
TypeError – If atoms is not an ASE Atoms instance.
- irmsd.get_canonical_fortran(atom_numbers: ndarray, positions: ndarray, wbo: ndarray | None = None, invtype: str = 'apsp+', heavy: bool = False) ndarray[source]#
Core API: call the Fortran routine to calculate the canonical ranking of atoms
- Parameters:
atom_numbers (
(N,) int32-like) – Atomic numbers (or types).positions (
(N,3) float64-like) – Cartesian coordinates in Å.heavy (
bool, optional) – Whether to consider only heavy atoms (default: False).wbo (
(natoms,natoms) float64,C-contiguous, optional) – Optional Wiberg bond order matrix, required if invtype is ‘cangen’, ignored in case of ‘apsp+’.invtype (
str, optional) – alogrithm type for invariants calculation (default: apsp+), alternativly ‘cangen’. The special value ‘apsp+nmr’ runs ‘apsp+’ and then splits any rank shared by exactly two non-hydrogen atoms into two distinct ranks (NMR equivalency hack); intended for rank requests only, not for iRMSD matching.
- Returns:
rank – Rank array.
- Return type:
(N,) int32- Raises:
ValueError – If positions does not have shape (N, 3).
- irmsd.get_canonical_rdkit(molecule, conf_id: None | int | list[int] = None, wbo: None | ndarray = None, invtype='apsp+', heavy: bool = False) ndarray[source]#
Canonical atom ranks for the selected conformers.
- Parameters:
molecule (
rdkit.Chem.Mol)conf_id (
int,listofint, orNone) – Conformer ID(s); None selects all conformers.wbo (
np.ndarray, optional) – Wiberg bond orders, required forinvtype='cangen'. Shape (n_atoms, n_atoms) to share across conformers, or (n_conf, n_atoms, n_atoms) for one per conformer.invtype (
str) – Invariant type.heavy (
bool) – Rank heavy atoms only.
- Returns:
Shape (n_atoms,) for one conformer, (n_conf, n_atoms) for several.
- Return type:
np.ndarray- Raises:
TypeError – If molecule is not an RDKit Mol.
- irmsd.get_cn_ase(atoms) ndarray[source]#
Coordination numbers via Molecule.get_cn(), shape (N,).
- Parameters:
atoms (
ase.Atoms)
- irmsd.get_cn_fortran(atom_numbers: ndarray, positions: ndarray) ndarray[source]#
Core API: call the Fortran routine to calculate CN
- Parameters:
atom_numbers (
(N,) int32-like) – Atomic numbers (or types).positions (
(N,3) float64-like) – Cartesian coordinates in Å.
- Returns:
cn – array with coordination numbers
- Return type:
(N) float64 ndarray- Raises:
ValueError – If positions does not have shape (N, 3).
- irmsd.get_cn_rdkit(molecule, conf_id: None | int | list[int] = None) ndarray[source]#
Coordination numbers for the selected conformers.
- Parameters:
molecule (
rdkit.Chem.Mol)conf_id (
int,listofint, orNone) – Conformer ID(s); None selects all conformers.
- Returns:
Shape (n_atoms,) for one conformer, (n_conf, n_atoms) for several.
- Return type:
np.ndarray- Raises:
TypeError – If molecule is not an RDKit Mol.
- irmsd.get_energies_from_molecule_list(molecule_list: Sequence[Molecule]) ndarray[source]#
Collect potential energies from a sequence of Molecule objects.
For each Molecule, this function calls
get_potential_energy()and stores the result in a 1D NumPy array of dtype float. If the energy is not available (for example, ifget_potential_energy()raisesAttributeErroror returnsNone), the corresponding entry is set to 0.0.- Parameters:
molecule_list (
Sequence[Molecule]) – Sequence of Molecule instances.- Returns:
energies – Array of shape (n_structures,) containing one energy per Molecule.
- Return type:
np.ndarray
- irmsd.get_irmsd(atom_numbers1: ndarray, positions1: ndarray, atom_numbers2: ndarray, positions2: ndarray, iinversion: int = 0, ranks1: ndarray | None = None, ranks2: ndarray | None = None) Tuple[float, ndarray, ndarray, ndarray, ndarray][source]#
Core API: call the Fortran routine to calculate the iRMSD between two structures
- Parameters:
atom_numbers1 (
(N1,) int32-like) – Atomic numbers (or types) of structure 1.positions1 (
(N1,3) float64-like) – Cartesian coordinates in Å of structure 1.atom_numbers2 (
(N2,) int32-like) – Atomic numbers (or types) of structure 2.positions2 (
(N2,3) float64-like) – Cartesian coordinates in Å of structure 2.iinversion (
int, optional) – Whether to consider inversion symmetry. Default is 0 (auto). Set to 1 to use inversion, set to 2 to disable inversion.ranks1 (
(N1,) int-like, optional) – Externally supplied per-atom canonical ranks for structure 1 (e.g. read from a file). Used directly only if bothranks1andranks2are given, contain no zeros, and pass the internal consistency check; otherwise the canonical ranks are recomputed. A zero entry is the “not provided” sentinel.ranks2 (
(N2,) int-like, optional) – Externally supplied per-atom canonical ranks for structure 2. Seeranks1.
- Returns:
rmsdval (
float) – The calculated iRMSD value.Z3 (
(N1,) int32 ndarray) – Atomic numbers of the aligned structure 1.P3 (
(N1,3) float64 ndarray) – Aligned and centered coordinates of structure 1.Z4 (
(N2,) int32 ndarray) – Atomic numbers of the aligned structure 2.P4 (
(N2,3) float64 ndarray) – Aligned and centered coordinates of structure 2.
Notes
The returned coordinates P3 and P4 are centered at the origin.
- Raises:
ValueError – If positions1 or positions2 do not have shape (Ni, 3).
- irmsd.get_irmsd_ase(atoms1, atoms2, iinversion: int = 0) Tuple[float, 'ase.Atoms', 'ase.Atoms'][source]#
ASE wrapper for
get_irmsd_molecule.- Parameters:
atoms1 (
ase.Atoms)atoms2 (
ase.Atoms)iinversion (
int, optional) – 0 = ‘auto’, 1 = ‘on’, 2 = ‘off’.
- Returns:
irmsd (
float) – In Angstrom.new_atoms1, new_atoms2 (
ase.Atoms) – New objects for the transformed structures.
- Raises:
RuntimeError – If ASE is not installed.
TypeError – If inputs are not ASE Atoms.
- irmsd.get_irmsd_molecule(molecule1: Molecule, molecule2: Molecule, iinversion: int = 0, use_ids: bool = True) Tuple[float, Molecule, Molecule][source]#
Compute the iRMSD between two Molecule objects using the iRMSD backend.
The backend may reorder atoms and/or change atomic numbers according to its canonicalization / matching logic. This wrapper returns copies of both input Molecules with the updated atomic numbers and positions.
- Parameters:
molecule1 (
Molecule) – First input structure.molecule2 (
Molecule) – Second input structure.iinversion (
int, optional) – Inversion flag passed directly to the Fortran backend. See the backend documentation for allowed values and meanings.use_ids (
bool, optional) – If True (default) and both molecules carry per-atom canonical IDs (Molecule.ids), those IDs are passed to the backend as externally supplied ranks. The backend uses them only if they are mutually consistent, otherwise it recomputes the canonical ranks.
- Returns:
- Raises:
TypeError – If either input is not a Molecule.
- irmsd.get_irmsd_rdkit(molecule_ref, molecule_align, conf_id_ref=-1, conf_id_align=-1, iinversion: int = 0) Tuple[float, 'Mol', 'Mol'][source]#
iRMSD between two conformers after permutation and alignment.
- Parameters:
molecule_ref (
rdkit.Chem.Mol)molecule_align (
rdkit.Chem.Mol)conf_id_ref (
int) – Conformer IDs; -1 is the RDKit default conformer.conf_id_align (
int) – Conformer IDs; -1 is the RDKit default conformer.iinversion (
int) – Inversion handling: 0 auto, 1 on, 2 off.
- Returns:
irmsd (
float) – In Angstrom.ref, aligned (
rdkit.Chem.Mol) – Processed reference and aligned molecules.
- Raises:
TypeError – If either input is not an RDKit Mol.
- irmsd.get_point_group(atom_numbers: ndarray, positions: ndarray, *, threshold: float = 0.1, primary_threshold: float = 0.5, max_axis_order: int = 10, max_opt_cycles: int = 100, max_atoms: int | None = 200) str | None[source]#
Schoenflies point group from the brute-force analyzer ported from CREST.
- Parameters:
atom_numbers (
(N,) int32-like) – Atomic numbers (or types).positions (
(N,3) float64-like) – Cartesian coordinates in Å.threshold (
float, optional) – Final symmetry tolerance in Bohr: the largest atom displacement an element may cause to be accepted (default: 0.1, as in CREST). Increase for noisy or loosely optimized geometries; large systems need this more, since the worst atom deviation grows with size. Too tight a value on such structures is also slow, as every near-miss candidate is fully optimized before being rejected.primary_threshold (
float, optional) – Tolerance in Bohr for the initial pairing of symmetry-related atoms and the distance prescreening of candidates (default: 0.5). Too small values miss elements of distorted structures; larger values test many more candidates, which dominates the cost for large systems.max_axis_order (
int, optional) – Highest rotation axis order searched, for proper and improper axes alike (default: 10). An S2n axis needs at least 2n, so values below 10 lose e.g. the S10 axes of Ih and D5d and with them the group.max_opt_cycles (
int, optional) – Maximum optimization cycles per candidate element (default: 100).max_atoms (
intorNone, optional) – Skip the analysis for structures with more atoms than this, since the search scales steeply with system size (default: 200, as in CREST). None disables the limit.
- Returns:
Schoenflies symbol, e.g. “C2v”, “Dinfh”; the highest proper axis (e.g. “C3”) if no tabulated group matches; None if skipped by
max_atoms.- Return type:
strorNone- Raises:
ValueError – If positions is not (N, 3) or max_axis_order < 2.
- irmsd.get_point_group_ase(atoms, **settings) str | None[source]#
Schoenflies point group via Molecule.get_point_group().
- Parameters:
atoms (
ase.Atoms)**settings – Analyzer settings (threshold, primary_threshold, max_axis_order, max_opt_cycles, max_atoms), see
irmsd.get_point_group().
- Returns:
None if the analysis was skipped.
- Return type:
strorNone
- irmsd.get_point_group_rdkit(molecule, conf_id: None | int | list[int] = None, **settings) str | None | List[str | None][source]#
Schoenflies point group for the selected conformers.
- Parameters:
molecule (
rdkit.Chem.Mol)conf_id (
int,listofint, orNone) – Conformer ID(s); None selects all conformers.**settings – Analyzer settings, see
irmsd.get_point_group().
- Returns:
Symbol per conformer; None if the analyzer skipped it.
- Return type:
strorNone, orlistofthose- Raises:
TypeError – If molecule is not an RDKit Mol.
- irmsd.get_quaternion_rmsd_fortran(atom_numbers1: ndarray, positions1: ndarray, atom_numbers2: ndarray, positions2: ndarray, mask: ndarray | None = None) Tuple[float, ndarray, ndarray][source]#
Pair API: call the Fortran routine on TWO structures.
- Parameters:
atom_numbers1 (
(N1,) int32-like)positions1 (
(N1,3) float64-like)atom_numbers2 (
(N2,) int32-like)positions2 (
(N2,3) float64-like)mask (
(N1,) bool-likeorNone)
- Returns:
rmsdval (
float64)new_positions2 (
(N2,3) float64,(positions2 @ Umat.T))Umat (
(3,3) float64 (Fortran-ordered))
Notes
The returned new_positions2 is aligned onto positions1. 1. If mask is provided, only the atoms where mask==True in structure 1 are used to compute the RMSD. 2. The rotation matrix Umat is Fortran-ordered, i.e., to rotate positions2, do: new_positions2 = positions2 @ Umat.T 3. The returned new_positions2 is also translated to have the same barycenter as positions1.
- Raises:
ValueError – If the input arrays do not have the correct shapes or types.
- irmsd.get_rmsd_ase(atoms1, atoms2, mask=None) Tuple[float, 'ase.Atoms', np.ndarray][source]#
ASE wrapper for
get_rmsd_molecule.- Parameters:
atoms1 (
ase.Atoms) – Reference structure.atoms2 (
ase.Atoms) – Structure aligned onto atoms1.mask (
array-likeofbool, optional) – Atoms of the first structure that enter the RMSD.
- Returns:
rmsd (
float) – In Angstrom.new_atoms2 (
ase.Atoms) – New object with coordinates aligned to atoms1.rotation_matrix (
np.ndarray) – 3x3 rotation used for the alignment.
- Raises:
RuntimeError – If ASE is not installed.
TypeError – If inputs are not ASE Atoms.
- irmsd.get_rmsd_molecule(molecule1: Molecule, molecule2: Molecule, mask=None) Tuple[float, Molecule, ndarray][source]#
Compute the RMSD between two Molecule objects using the quaternion-based RMSD backend, and return an aligned copy of the second Molecule.
- Parameters:
molecule1 (
Molecule) – Reference structure.molecule2 (
Molecule) – Structure to be rotated/translated ontomolecule1.mask (
array-likeofbool, optional) – Optional mask selecting which atoms ofmolecule1participate in the RMSD. Must be broadcastable / compatible with the Fortran backend’s mask semantics.
- Returns:
rmsd (
float) – Root-mean-square deviation in Ångström.new_molecule2 (
Molecule) – Copy ofmolecule2with its positions replaced by the aligned coordinates returned by the backend.rotation_matrix (
np.ndarray) – 3×3 rotation matrix applied to alignmolecule2ontomolecule1.
- Raises:
TypeError – If either input is not a Molecule.
- irmsd.get_rmsd_rdkit(molecule_ref, molecule_align, conf_id_ref=-1, conf_id_align=-1, mask=None) Tuple[float, 'Mol', np.ndarray][source]#
RMSD between two conformers after alignment, without permutation.
- Parameters:
molecule_ref (
rdkit.Chem.Mol)molecule_align (
rdkit.Chem.Mol)conf_id_ref (
int) – Conformer IDs; -1 is the RDKit default conformer.conf_id_align (
int) – Conformer IDs; -1 is the RDKit default conformer.mask (
array-likeofbool, optional) – Atoms to include in the RMSD.
- Returns:
rmsd (
float) – In Angstrom.aligned (
rdkit.Chem.Mol)rotmat (
np.ndarray) – Rotation matrix.
- Raises:
TypeError – If either input is not an RDKit Mol.
- irmsd.get_symmetry_elements(atom_numbers: ndarray, positions: ndarray, *, threshold: float = 0.1, primary_threshold: float = 0.5, max_axis_order: int = 10, max_opt_cycles: int = 100, max_atoms: int | None = 200) tuple[str | None, list[SymmetryOperation]][source]#
Point group and located symmetry elements, each as its generating operation.
- Parameters:
atom_numbers (
(N,) int32-like) – Atomic numbers (or types).positions (
(N,3) float64-like) – Cartesian coordinates in Å.threshold (
float, optional) – Final symmetry tolerance in Bohr: the largest atom displacement an element may cause to be accepted (default: 0.1, as in CREST). Increase for noisy or loosely optimized geometries; large systems need this more, since the worst atom deviation grows with size. Too tight a value on such structures is also slow, as every near-miss candidate is fully optimized before being rejected.primary_threshold (
float, optional) – Tolerance in Bohr for the initial pairing of symmetry-related atoms and the distance prescreening of candidates (default: 0.5). Too small values miss elements of distorted structures; larger values test many more candidates, which dominates the cost for large systems.max_axis_order (
int, optional) – Highest rotation axis order searched, for proper and improper axes alike (default: 10). An S2n axis needs at least 2n, so values below 10 lose e.g. the S10 axes of Ih and D5d and with them the group.max_opt_cycles (
int, optional) – Maximum optimization cycles per candidate element (default: 100).max_atoms (
intorNone, optional) – Skip the analysis for structures with more atoms than this, since the search scales steeply with system size (default: 200, as in CREST). None disables the limit.
- Returns:
symbol (
strorNone) – Schoenflies symbol; None if skipped bymax_atoms.elements (
list[SymmetryOperation]) – Empty if skipped. A Cinf axis hasorder == 0.
- Raises:
ValueError – If positions is not (N, 3) or max_axis_order < 2.
- irmsd.get_symmetry_operations(atom_numbers: ndarray, positions: ndarray, *, threshold: float = 0.1, primary_threshold: float = 0.5, max_axis_order: int = 10, max_opt_cycles: int = 100, max_atoms: int | None = 200) tuple[str | None, list[SymmetryOperation]][source]#
Point group and all distinct powers of the located elements.
Order: E, i, mirror planes, proper, then improper rotations. For finite groups the length equals the group order; for Cinfv/Dinfh only the operations generated by i, sigma and C2 are returned (the Cinf axis comes from
get_symmetry_elements()).- Parameters:
atom_numbers (
(N,) int32-like) – Atomic numbers (or types).positions (
(N,3) float64-like) – Cartesian coordinates in Å.threshold (
float, optional) – Final symmetry tolerance in Bohr: the largest atom displacement an element may cause to be accepted (default: 0.1, as in CREST). Increase for noisy or loosely optimized geometries; large systems need this more, since the worst atom deviation grows with size. Too tight a value on such structures is also slow, as every near-miss candidate is fully optimized before being rejected.primary_threshold (
float, optional) – Tolerance in Bohr for the initial pairing of symmetry-related atoms and the distance prescreening of candidates (default: 0.5). Too small values miss elements of distorted structures; larger values test many more candidates, which dominates the cost for large systems.max_axis_order (
int, optional) – Highest rotation axis order searched, for proper and improper axes alike (default: 10). An S2n axis needs at least 2n, so values below 10 lose e.g. the S10 axes of Ih and D5d and with them the group.max_opt_cycles (
int, optional) – Maximum optimization cycles per candidate element (default: 100).max_atoms (
intorNone, optional) – Skip the analysis for structures with more atoms than this, since the search scales steeply with system size (default: 200, as in CREST). None disables the limit.
- Returns:
symbol (
strorNone) – Schoenflies symbol; None if skipped bymax_atoms.operations (
list[SymmetryOperation]) – Empty if skipped.
- Raises:
ValueError – If positions is not (N, 3) or max_axis_order < 2.
- irmsd.molecule_to_ase(molecules: Molecule) ase.Atoms[source]#
- irmsd.molecule_to_ase(molecules: Sequence[Molecule]) list['ase.Atoms']
Convert Molecule(s) to ASE Atoms without attaching a calculator.
The Hartree energy is written to
info["energy"]in eV unlessinfoalready has anenergyentry; anyenergy_unitsmarker is dropped.Molecule.idsbecomes a per-atomcanonical_idarray. The returned objects do not share state with the input.- Parameters:
molecules (
MoleculeorSequence[Molecule])- Returns:
A list, in input order, if molecules is a sequence.
- Return type:
ase.Atomsorlist[ase.Atoms]- Raises:
RuntimeError – If ASE is not installed.
TypeError – If the input is neither a Molecule nor a sequence of Molecules.
- irmsd.molecule_to_rdkit(molecule: Molecule) Mol[source]#
- irmsd.molecule_to_rdkit(molecules: Sequence[Molecule]) list['Mol']
Convert irmsd Molecule(s) to RDKit Mol(s) carrying atoms and one conformer.
No bonds or properties are transferred. Returns a single Mol if exactly one Molecule results, else a list.
- Raises:
TypeError – If the input is not a Molecule or a list of them.
- irmsd.prune(molecule_list: Sequence[Molecule], rthr: float, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None, use_ids: bool = True) List[Molecule][source]#
High-level wrapper around the Fortran-backed
sorter_irmsdthat operates directly on Molecule objects. Returns a pruned list of structures.- Parameters:
molecule_list (
Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.rthr (
float) – Distance threshold for the sorter (passed through to the backend).iinversion (
int, optional) – Inversion symmetry flag, passed through to the backend.allcanon (
bool, optional) – Canonicalization flag, passed through to the backend.printlvl (
int, optional) – Verbosity level, passed through to the backend.ethr (
float | None) – Optional energy threshold to accelerate by pre-sortingewin (
float | None) – Optional energy window to limit ensembe size around lowest energy structure. In Hartree.
- Returns:
new_molecule_list – New Molecule objects reconstructed from the sorted atomic numbers and positions returned by the backend. The list contains only n_structures defined as
uniqueaccording to the selected thresholds.- Return type:
list[Molecule]- Raises:
TypeError – If
molecule_listdoes not contain Molecule instances.ValueError – If
molecule_listis empty or if the Molecules do not all have the same number of atoms.
- irmsd.rdkit_to_molecule(molecules: Mol, conf_id: int | list[int] | None = None) Molecule | list[Molecule][source]#
- irmsd.rdkit_to_molecule(molecules: Sequence['Mol'], conf_id: int | list[int] | None = None) list[Molecule]
Convert RDKit Mol(s) to irmsd Molecule(s), one per selected conformer.
Mol-level and conformer-level properties are merged into
info; conformer properties win on key clashes.- Parameters:
molecules (
rdkit.Chem.Molorlistofrdkit.Chem.Mol)conf_id (
int,listofint, orNone, optional) – Conformer ID(s); None selects all conformers.
- Returns:
A single Molecule if exactly one conformer was converted, else a list.
- Return type:
- Raises:
TypeError – If the input is not a Mol or list of Mols, or a conformer is not 3D.
- irmsd.read_structures(paths: str | Sequence[str]) List[Molecule][source]#
Read an arbitrary number of structures and return them as Molecule objects.
For each path, this routine behaves as follows:
If the file extension is
.xyzor.extxyz, it uses the internalread_extxyzhelper to obtain one or more Molecule objects.For all other file types, it attempts to import ASE via
require_ase(), usesase.io.readto read one or more ASE Atoms objects, and converts them into Molecule objects usingase_to_molecule.
- Multi-frame files:
If a file contains multiple frames, all frames are read and appended to the output list. A short informational message is printed indicating the number of frames that were found.
- Parameters:
paths (
Sequence[str]) – File paths to read.- Returns:
structures – One Molecule per frame found across all input paths.
- Return type:
list[Molecule]
- irmsd.run_cregen_and_print(molecule_list: Sequence[Molecule], rthr: float, ethr: float, bthr: float, ewin: float | None = None, printlvl: int = 0, maxprint: int = 25, outfile: str | None = None) None[source]#
Run CREGEN per sum-formula group and print a summary per group.
- Parameters:
rthr (
float) – RMSD, energy, and rotational-constant thresholds for conformer identification.ethr (
float) – RMSD, energy, and rotational-constant thresholds for conformer identification.bthr (
float) – RMSD, energy, and rotational-constant thresholds for conformer identification.maxprint (
int) – Max rows per printed result table.outfile (
strorNone) – Write the representatives here; with several formulas, one<root>_<formula><ext>file each.
- irmsd.sort_get_delta_irmsd_and_print(molecule_list: Sequence[Molecule], inversion: str = None, allcanon: bool = True, printlvl: int = 0, maxprint: int = 25, outfile: str | None = None) None[source]#
Split by sum formula, energy-sort, and print the iRMSD between consecutive structures of each group.
- Parameters:
inversion (
{"auto", "on", "off"})maxprint (
int) – Max rows per printed result table.outfile (
strorNone) – Unused; no structures are written.
- irmsd.sort_structures_and_print(molecule_list: Sequence[Molecule], rthr: float, inversion: str = None, allcanon: bool = True, printlvl: int = 0, maxprint: int = 25, ethr: float | None = None, ewin: float | None = None, outfile: str | None = None) None[source]#
Split by sum formula, energy-sort, prune each group with the iRMSD sorter, and print a summary per group.
- Parameters:
rthr (
float) – Distance threshold for the sorter.inversion (
{"auto", "on", "off"})maxprint (
int) – Max rows per printed result table.ethr (
floatorNone) – Inter-conformer energy threshold for presorting.ewin (
floatorNone) – Energy window around the lowest-energy structure.outfile (
strorNone) – Write the representatives here; with several formulas, one<root>_<formula><ext>file each.
- irmsd.sorter_irmsd(atom_numbers_list: Sequence[ndarray], positions_list: Sequence[ndarray], nat: int, rthr: float, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, energies_list: Sequence[ndarray] | None = None, ids_list: Sequence[ndarray | None] | None = None) Tuple[ndarray, List[ndarray], List[ndarray]][source]#
High-level API: call the sorter_exposed_xyz_fortran Fortran routine.
- Parameters:
atom_numbers_list (
sequenceof(N,) int32 arrays) – Per-structure atom numbers.positions_list (
sequenceof(N,3) float64 arrays) – Per-structure coordinates.nat (
int) – Number of atoms for which the groups array is defined. Must satisfy 1 <= nat <= N.rthr (
float) – Distance threshold for the Fortran sorter. In Angström.iinversion (
int) – Inversion symmetry flag.allcanon (
bool) – Canonicalization flag.printlvl (
int) – Verbosity level.ethr (
float | None) – Inter-conformer energy threshold (optional). In Hartree.energies_list (
sequenceof(Nall,) floats | None) – List of energies for the passed structures (optional). In Hartree.
- Returns:
groups (
(nat,) int32) – Group index for each of the first nat atoms.xyz_structs (
listof(N,3) float64 arrays) – Updated coordinates for each structure.Z_structs (
listof(N,) int32 arrays) – Updated atom numbers for each structure.
- Raises:
ValueError – If input arrays have inconsistent shapes or invalid parameters.
- irmsd.sorter_irmsd_ase(atoms_list: Sequence['ase.Atoms'], rthr: float = 0.125, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None) Tuple[np.ndarray, List['ase.Atoms']][source]#
ASE wrapper for
sorter_irmsd_molecule.- Parameters:
atoms_list (
Sequence[ase.Atoms]) – List or tuple; all structures must have the same atom count.rthr (
float) – Distance threshold for the sorter.iinversion (
int, optional) – 0 = ‘auto’, 1 = ‘on’, 2 = ‘off’.allcanon (
bool, optional) – Canonicalization flag.printlvl (
int, optional) – Verbosity level.ethr (
floatorNone) – Energy threshold for pre-sorting, in Hartree.ewin (
floatorNone) – Energy window above the lowest structure, in Hartree.
- Returns:
groups (
np.ndarray) – Integer group index per structure.new_atoms_list (
list[ase.Atoms]) – Rebuilt from the sorted Molecules.
- irmsd.sorter_irmsd_molecule(molecule_list: Sequence[Molecule], rthr: float = 0.125, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None, use_ids: bool = True) Tuple[ndarray, List[Molecule]][source]#
High-level wrapper around the Fortran-backed
sorter_irmsdthat operates directly on Molecule objects.- Parameters:
molecule_list (
Sequence[Molecule]) – Sequence of Molecule objects. All molecules must have the same number of atoms.rthr (
float) – Distance threshold for the sorter (passed through to the backend).iinversion (
int, optional) – Inversion symmetry flag, passed through to the backend.allcanon (
bool, optional) – Canonicalization flag, passed through to the backend.printlvl (
int, optional) – Verbosity level, passed through to the backend.ethr (
float | None) – Optional energy threshold to accelerate by pre-sortingewin (
float | None) – Optional energy window to limit ensembe size around lowest energy structure. In Hartree.
- Returns:
groups (
np.ndarray) – Integer array of shape (nat,) with group indices for the firstnatatoms (as defined by the backend).new_molecule_list (
list[Molecule]) – New Molecule objects reconstructed from the sorted atomic numbers and positions returned by the backend. The list has the same length and ordering asmolecule_list.
- Raises:
TypeError – If
molecule_listdoes not contain Molecule instances.ValueError – If
molecule_listis empty or if the Molecules do not all have the same number of atoms.
- irmsd.sorter_irmsd_rdkit(molecules: 'Mol' | Sequence['Mol'], rthr: float = 0.125, iinversion: int = 0, allcanon: bool = True, printlvl: int = 0, ethr: float | None = None, ewin: float | None = None) Tuple[np.ndarray, List['Mol']][source]#
Group an ensemble by iRMSD.
- Parameters:
molecules (
rdkit.Chem.Molorlistofrdkit.Chem.Mol) – A single Mol must carry more than one conformer.rthr (
float) – iRMSD threshold in Angstrom.iinversion (
int) – Inversion handling: 0 auto, 1 on, 2 off.allcanon – Canonicalization flag and verbosity, passed to the backend.
printlvl – Canonicalization flag and verbosity, passed to the backend.
ethr (
floatorNone) – Energy pre-sorting threshold.ewin (
floatorNone) – Energy window in Hartree above the lowest-energy structure.
- Returns:
groups (
np.ndarray) – Integer group index per structure.new_molecules_list (
listofrdkit.Chem.Mol) – Sorted structures.
- Raises:
TypeError – If the input is not an RDKit Mol or a list of them.
- irmsd.write_structures(filename: str | Path, structures: Molecule, mode: str = 'w') None[source]#
- irmsd.write_structures(filename: str | Path, structures: Sequence[Molecule], mode: str = 'w') None
High-level structure writer for the irmsd Molecule type.
This routine mirrors the behaviour of read_structures on the output side: it chooses the appropriate backend based on the file extension and accepts either a single Molecule or a sequence of Molecule objects.
Dispatch rules#
If the filename has no extension, ‘.xyz’ is appended and the internal extended-XYZ writer is used.
If the filename ends with ‘.xyz’, ‘.extxyz’ or ‘.trj’ (case-insensitive), the structures are written using the internal extended-XYZ writer (write_extxyz). The mode argument is passed through and controls whether the file is overwritten (‘w’, default) or appended to (‘a’).
For all other filename extensions, ASE is used as a backend. The structures are first converted to ASE Atoms objects using molecule_to_ase, and then written via ase.io.write. In this case the mode argument is currently ignored and ASE’s default behaviour for the chosen format is used.
- param filename:
Output filename. Its extension determines the backend.
- type filename:
strorpathlib.Path- param structures:
A single Molecule or a sequence of Molecules to be written. For extended XYZ, multiple Molecules are written as consecutive frames in one file.
- type structures:
MoleculeorSequence[Molecule]- param mode:
File open mode for extended XYZ output. Ignored for non-XYZ formats handled via ASE.
- type mode:
{"w", "a"}, optional- raises RuntimeError:
If ASE is required (non-XYZ formats) but not installed.
- raises TypeError:
If structures is not a Molecule or a sequence of Molecules.