Compare commits

..
238 changed files with 3622 additions and 17129 deletions
+3 -2
View File
@@ -23,8 +23,9 @@ install:
- set MSMPI_LIB64=C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x64
- set MSMPI_INC=C:\Program Files (x86)\Microsoft SDKs\MPI\Include
# Install METIS, use MFEM's mirror because the original source server is often
# down and we don't support yet the new repo https://github.com/KarypisLab/METIS
# Install METIS, use a mirror because the original source server is not always
# up. Original url:
# http://glaros.dtc.umn.edu/gkhome/fetch/sw/metis/metis-5.1.0.tar.gz
- ps: Start-FileDownload 'https://mfem.github.io/tpls/metis-5.1.0.tar.gz'
- 7z x metis-5.1.0.tar.gz -so | 7z x -si -ttar > nul
- cd metis-5.1.0
+1
View File
@@ -49,6 +49,7 @@ jobs:
# Details on CodeQL's query packs refer to : https://docs.github.com/en/code-security/code-scanning/automatically-scanning-your-code-for-vulnerabilities-and-errors/configuring-code-scanning#using-queries-in-ql-packs
# queries: security-extended,security-and-quality
queries: lgtm
# Autobuild attempts to build any compiled languages (C/C++, C#, or Java).
+1 -1
View File
@@ -107,7 +107,7 @@ jobs:
run: |
sudo apt-get install doxygen graphviz
cd doc
doxygen -u CodeDocumentation.conf.in
doxygen -u CodeDocumentation.conf.in 2>/dev/null
- name: build documentation
run: |
-8
View File
@@ -18,11 +18,6 @@ CMakeFiles/
# Backup files
*~
*.sqlite
*.nsys-rep
*.qdstrm
*.csv
# Default install location
/mfem/
@@ -280,7 +275,6 @@ miniapps/tools/load-dc
miniapps/tools/convert-dc
miniapps/tools/lor-transfer
miniapps/tools/get-values
miniapps/tools/check-tmop-metric
miniapps/toys/automata
miniapps/toys/life
@@ -313,7 +307,6 @@ miniapps/solvers/sol.*
miniapps/parelag/MultilevelHcurlHdivSolver
miniapps/parelag/*.mesh
miniapps/multidomain/multidomain
miniapps/hooke/hooke
# Unit test binary and outputs
@@ -334,7 +327,6 @@ tests/benchmarks/bench_ceed
tests/benchmarks/bench_tmop
tests/benchmarks/bench_vector
tests/benchmarks/bench_virtuals
tests/benchmarks/bench_lor
# Test script output
tests/scripts/*.err
+39 -85
View File
@@ -8,48 +8,31 @@
https://mfem.org
Version 4.5, released on October 22, 2022
=========================================
Version 4.4.1 (development)
===========================
Meshing improvements
--------------------
- Added new SubMesh and ParSubMesh classes that can be used to extract a subset
of a given Mesh. These classes have the same functionality as Mesh and ParMesh
and work with all existing MFEM interfaces like finite element spaces etc.
- Added a method, ParMesh::GetSerialMesh(), that reconstructs a partitioned
parallel mesh on a given single rank. Also, added ParMesh::PrintAsSerial(),
which saves the reconstructed serial mesh to a C++ stream on rank 0.
- Added more 3D TMOP metrics, as well as specialized metrics for mesh
untangling and worst-case quality improvement.
- Added a new method, Mesh::NodesUpdated, which should be called after the mesh
node coordinates have changed, e.g. after the mesh has moved. This is
necessary, for example, with device assembly of linear and bilinear forms.
- Added support for mixed meshes and pyramids in GSLIB-FindPoints.
Discretization improvements
---------------------------
- Added support for assembling low-order-refined matrices using a GPU-enabled
"batched" algorithm. The lor_solvers and plor_solvers now fully support GPU
acceleration.
- Added support for partial assembly and fully matrix-free operators on mixed
meshes (different element types and p-adaptivity) through libCEED, including
device acceleration, e.g. with NVIDIA and AMD GPUs. The p-adaptivity is
currently limited by MFEM capabilities, i.e. 2D serial meshes. All mixed
element topologies are supported in serial and parallel: segment, triangle,
square, tetrahedron, cube, prism, and pyramid.
- Added full assembly and device support for several LinearForm integrators:
* DomainLF: (f, v)
* VectorDomainLF: ((f1,...,fn), (v1,...,vn))
* DomainLFGrad: (f, grad(v))
* VectorDomainLFGrad: ((f1x,f1y,f1z,...,fnx,fny,fnz), grad(v1,...,vn))
The device assembly of linear forms has to be explicitly enabled by calling
LinearForm::UseFastAssembly(true), otherwise the legacy linear form assembly
is used by default.
- Added support for assembling low-order-refined matrices using a GPU-enabled
"batched" algorithm. The lor_solvers and plor_solvers now fully support GPU
acceleration with arbitrary user-supplied coefficients.
- Added a new class FaceQuadratureSpace that allows for the construction of
QuadratureFunctions on the interior or boundary faces of a mesh.
- Added a class CoefficientVector for efficient access of variable coefficient
values at quadrature points (in particular for GPU/device kernels).
- Added WhiteGaussianNoiseDomainLFIntegrator: a LinearFormIntegrator class for
spatial Gaussian white noise.
@@ -57,25 +40,8 @@ Discretization improvements
- Added a new Zienkiewicz-Zhu patch recovery-based a posteriori error estimator.
See fem/estimators.hpp.
- Various fixes and improvements in LinearFormExtension.
Linear and nonlinear solvers
----------------------------
- Added a new class DGMassInverse that performs a local element-wise CG
iteration to solve systems involving the discontinuous Galerkin mass matrix,
including support for device/GPU acceleration.
- Added more flexibility to the constrained solver classes:
* PenaltyConstrainedSolver now allows for a vector of penalty parameters
(necessary for penalty contact)
* PenaltyConstrainedSolver and EliminationSolver can use GMRES or PCG
* All constraint solver classes can take a user-defined preconditioner
- Added functions to toggle additional options for the SuperLU_Dist and Hypre
preconditioners (ParaSails, Euclid, ILU).
- Added boundary elimination with device support for `SparseMatrix` and
`HypreParMatrix`.
New and updated examples and miniapps
-------------------------------------
@@ -85,28 +51,12 @@ New and updated examples and miniapps
automatic differentiation tools like a native dual number implementation or a
third party library such as Enzyme. See miniapps/elasticity for more details.
- Added example for body-fitted volumetric and shape integration using the
Algoim library in miniapps/shifted.
- Add a new example code, Example 33/33p, to demonstrate the solution of
spectral fractional PDEs with MFEM.
Integrations, testing and documentation
---------------------------------------
- Added a Dockerfile for a simple MFEM container, see config/docker/README.md.
More sophisticated developer containers are available in the new repo
https://github.com/mfem/containers.
- Added support for the LLVM-based automatic differentiation tool Enzyme, see
https://github.com/EnzymeAD/Enzyme. Build system flags and a convenience
header are provided. The functionality and interaction are demonstrated in a
new miniapp in miniapps/elasticity.
- Added support for partial assembly and fully matrix-free operators on mixed
meshes (different element types and p-adaptivity) through libCEED, including
device acceleration, e.g. with NVIDIA and AMD GPUs. The p-adaptivity is
currently limited to 2D serial meshes. All mixed element topologies are
supported in both serial and parallel.
- Added support for ParMoonolith, https://bitbucket.org/zulianp/par_moonolith,
which provides parallel non-conforming, non-matching, variational, volumetric
@@ -114,40 +64,29 @@ Integrations, testing and documentation
between arbitrarily distributed and unrelated finite element meshes in a
variationally consistent way.
- Fully encapsulated SUNDIALS `N_Vector` object within the `SundialsNVector`
class by removing deprecated (e.g. `HypreParVector::ToNVector`) and
non-deprecated (e.g. `Vector::ToNVector`) functions in other classes.
- Added support for the LLVM-based automatic differentiation tool Enzyme, see
https://github.com/EnzymeAD/Enzyme. Build system flags and a convenience
header are provided. The functionality and interaction are demonstrated in a
new miniapp in miniapps/elasticity.
- New benchmark for the different assembly levels inspired by the CEED
Bake-Off Problems, see tests/benchmarks/bench_assembly_levels.cpp.
- Added example for body-fitted volumetric and shape integration using the
Algoim library.
- Added Windows 2022 CI testing with GitHub actions.
Miscellaneous
-------------
- The method SparseMatrix::EnsureMultTranspose() is now automatically called
by the methods AddMultTranspose(), MultTranspose(), and AbsMultTranspose().
Added a method with the same name to class HypreParMatrix which is also called
automatically by the HypreParMatrix::MultTranspose() methods.
- Various other simplifications, extensions, and bugfixes in the code.
- Updated various MemoryUsage methods to return 'std::size_t' instead of 'long'
since the latter is 32-bit in Win64 builds.
- Added boundary elimination with device support for `SparseMatrix` and
`HypreParMatrix`.
- When using `AssemblyLevel::FULL`, `FABilinearFormExtension::FormSystemMatrix`
outputs an `OperatorHandle` containing a `SparseMatrix` in serial, and an
`HypreParMatrix` in parallel (instead of a `ConstrainedOperator`).
- In various places in the library, replace the use of 'long' with 'long long'
to better support Win64 builds where 'long' is 32-bit and 'long long' is
64-bit. On Linux and MacOS, both types are typically 64-bit.
- The behavior of GridFunction::GetTrueVector() has been changed to not return
an empty true vector.
- Added support for ordering search points byVDIM in FindPointsGSLIB.
- Various other simplifications, extensions, and bugfixes in the code.
- Added TMOP metrics for mesh untangling and worst-case quality improvement.
Version 4.4, released on March 21, 2022
=======================================
@@ -180,6 +119,11 @@ Meshing improvements
- Added a simpler interface to access mesh face information, see FaceInformation
and GetFaceInformation in the Mesh class.
- Added the method ParMesh::GetSerialMesh() that reconstructs a partitioned
parallel mesh on a given single rank. Also, added the method
ParMesh::PrintAsSerial() that saves the reconstructed serial mesh to a C++
stream on rank 0.
- Gmsh meshes where all elements have zero physical tag (the default Gmsh output
format if no physical groups are defined) are now successfully loaded, and
elements are reassigned attribute number 1.
@@ -269,6 +213,9 @@ Integrations, testing and documentation
- Switched from Artistic Style (astyle) version 2.05.1 to version 3.1 for code
formatting. See the "make style" target.
- New benchmark for the different assembly levels inspired by the CEED
Bake-Off Problems, see tests/benchmarks/bench_assembly_levels.cpp.
Miscellaneous
-------------
- Added a simple singleton class, Mpi, as a replacement for MPI_Session. New
@@ -279,6 +226,13 @@ Miscellaneous
- Fixed several MinGW build issues on Windows.
- In various places in the library, replace the use of 'long' with 'long long'
to better support Win64 builds where 'long' is 32-bit and 'long long' is
64-bit. On Linux and MacOS, both types are typically 64-bit.
- Update various "MemoryUsage" methods to return 'std::size_t' instead of 'long'
since the latter is 32-bit in Win64 builds.
- Added 'double' atomicAdd implementation for previous versions of CUDA.
- HypreParVector and Vector now support C++ move semantics, and the copy
+5 -18
View File
@@ -51,7 +51,7 @@ project(mfem NONE)
# Current version of MFEM, see also `makefile`.
# mfem_VERSION = (string)
# MFEM_VERSION = (int) [automatically derived from mfem_VERSION]
set(${PROJECT_NAME}_VERSION 4.5.0)
set(${PROJECT_NAME}_VERSION 4.4.1)
# Prohibit in-source build
if (${PROJECT_SOURCE_DIR} STREQUAL ${PROJECT_BINARY_DIR})
@@ -81,10 +81,6 @@ if (MFEM_USE_STRUMPACK)
# Just needed to find the MPI_Fortran libraries to link with
set(XSDK_ENABLE_Fortran ON)
endif()
# SUNDIALS >= 6.4.0 requires C++14:
if (MFEM_USE_SUNDIALS AND ("${CMAKE_CXX_STANDARD}" LESS "14"))
set(CMAKE_CXX_STANDARD 14)
endif()
if (MFEM_USE_GINKGO AND ("${CMAKE_CXX_STANDARD}" LESS "14"))
set(CMAKE_CXX_STANDARD 14)
endif()
@@ -141,7 +137,7 @@ if (MFEM_USE_CUDA)
set(CUSPARSE_FOUND TRUE)
set(CUSPARSE_LIBRARIES "cusparse")
set(CUBLAS_FOUND TRUE)
set(CUBLAS_LIBRARIES "cublas")
set(CUSBLAS_LIBRARIES "cublas")
endif()
if (XSDK_ENABLE_C)
@@ -204,10 +200,10 @@ if (MFEM_USE_MPI)
find_package(MPI REQUIRED)
set(MPI_CXX_INCLUDE_DIRS ${MPI_CXX_INCLUDE_PATH})
if (MFEM_MPIEXEC)
string(REPLACE " " ";" MPIEXEC ${MFEM_MPIEXEC})
set(MPIEXEC ${MFEM_MPIEXEC})
endif()
if (MFEM_MPIEXEC_NP)
string(REPLACE " " ";" MPIEXEC_NUMPROC_FLAG ${MFEM_MPIEXEC_NP})
set(MPIEXEC_NUMPROC_FLAG ${MFEM_MPIEXEC_NP})
endif()
# Parallel MFEM depends on hypre
find_package(HYPRE REQUIRED)
@@ -481,21 +477,12 @@ if (NOT DEFINED MFEM_TIMER_TYPE)
endif()
endif()
# Without this, CMake 3.21.1 (and 3.20.2) run into CMake Errors like the following:
# CMake Error at config/cmake/modules/MfemCmakeUtilities.cmake:60 (add_library):
# Target "mfem" links to target "Threads::Threads" but the target was not
# found. Perhaps a find_package() call is missing for an IMPORTED target, or
# an ALIAS target is missing?
# Call Stack (most recent call first):
# CMakeLists.txt:474 (mfem_add_library)
find_package(Threads REQUIRED)
# List all possible libraries in order of dependencies.
# [METIS < SuiteSparse]:
# With newer versions of SuiteSparse which include METIS header using 64-bit
# integers, the METIS header (with 32-bit indices, as used by mfem) needs to
# be before SuiteSparse.
set(MFEM_TPLS OPENMP HYPRE LAPACK BLAS SuperLUDist METIS SuiteSparse SUNDIALS
set(MFEM_TPLS OPENMP HYPRE BLAS LAPACK SuperLUDist METIS SuiteSparse SUNDIALS
PETSC SLEPC MESQUITE MUMPS STRUMPACK AXOM FMS CONDUIT Ginkgo GNUTLS GSLIB
NETCDF MPFR PUMI HIOP POSIXCLOCKS MFEMBacktrace ZLIB OCCA CEED RAJA UMPIRE
ADIOS2 CUBLAS CUSPARSE MKL_CPARDISO AMGX CALIPER CODIPACK BENCHMARK PARELAG
+2 -8
View File
@@ -102,9 +102,7 @@ The MFEM source code has the following structure:
.
├── config
│ ├── cmake
── docker
│ ├── githooks
│ └── vcpkg
── githooks
├── data
├── doc
├── examples
@@ -113,7 +111,6 @@ The MFEM source code has the following structure:
│ ├── ginkgo
│ ├── hiop
│ ├── jupyter
│ ├── moonolith
│ ├── petsc
│ ├── pumi
│ ├── sundials
@@ -121,15 +118,13 @@ The MFEM source code has the following structure:
├── fem
│ ├── ceed
│ ├── fe
│ ├── lor
│ ├── moonolith
│ ├── qinterp
│ ├── moonolith
│ └── tmop
├── general
├── linalg
│ └── simd
├── mesh
│ └── submesh
├── miniapps
│ ├── adjoint
│ ├── autodiff
@@ -139,7 +134,6 @@ The MFEM source code has the following structure:
│ ├── hooke
│ ├── meshing
│ ├── mtop
│ ├── multidomain
│ ├── navier
│ ├── nurbs
│ ├── parelag
+9 -17
View File
@@ -7,10 +7,6 @@
https://mfem.org
This file provides a detailed description of how to build and install the MFEM
library. For a simple build, see the step-by-step instructions on the website
at https://mfem.org/building.
The MFEM library has a serial and an MPI-based parallel version, which largely
share the same code base. The only prerequisite for building the serial version
of MFEM is a (modern) C++ compiler, such as g++. The parallel version of MFEM
@@ -20,11 +16,7 @@ requires an MPI C++ compiler, as well as the following external libraries:
https://github.com/hypre-space/hypre
- METIS (a family of multilevel partitioning algorithms)
https://github.com/mfem/tpls
Note: We recommend our mirror of metis-4.0.3/5.1.0 above because the METIS
webpage, http://glaros.dtc.umn.edu/gkhome/metis/metis/overview, is often down
and we don't support yet the new repo https://github.com/KarypisLab/METIS.
http://glaros.dtc.umn.edu/gkhome/metis/metis/overview
The hypre dependency can be downloaded as a tarball from GitHub or from the
project webpage https://www.llnl.gov/casc/hypre. For example, the 2.24.0 release
@@ -480,10 +472,10 @@ MFEM_USE_CODIPACK = YES/NO
www.scicomp.uni-kl.de/codi/
MFEM_USE_ALGOIM = YES/NO
Enable the usage of Algoim - a collection of high-order accurate numerical
methods and C++ algorithms for working with implicitly-defined geometry and
level set methods. The Algoim library requires the Blitz++ library. The MFEM
provides interface to Algoim v1. Thus, to check out the specific state use:
Enable the usage of Algoim - a collection of high-order accurate numerical
methods and C++ algorithms for working with implicitly-defined geometry and
level set methods. The Algoim library requires the Blitz++ library. The MFEM
provides interface to Algoim v1. Thus, to check out the specific state use:
git checkout 9c9ca0ef094d8ab0390ed36367a1151b459bbe0a
https://algoim.github.io
@@ -558,7 +550,7 @@ MFEM_USE_FMS = YES/NO
Enables support for the FMS library which consists of the DataCollection
sub-class mfem::FMSDataCollection for I/O in FMS formats, see the header file
fem/fmsdatacollection.hpp. In addition, this option enables in-memory
conversion routines between FMS's FmsDataCollection structure and MFEM's
convetion routines between FMS's FmsDataCollection structure and MFEM's
DataCollection class, see the header file fem/fmsconvert.hpp.
MFEM_USE_PARELAG = YES/NO
@@ -605,7 +597,7 @@ The specific libraries and their options are:
- METIS, used when MFEM_USE_METIS = YES. If using METIS 5, set
MFEM_USE_METIS_5 = YES (default is to use METIS 4).
URL: https://github.com/mfem/tpls (MFEM mirror, see above)
URL: http://glaros.dtc.umn.edu/gkhome/metis/metis/overview
Options: METIS_OPT, METIS_LIB.
Versions: METIS 4.0.3 or 5.1.0.
@@ -762,12 +754,12 @@ The specific libraries and their options are:
Options: GSLIB_OPT, GSLIB_LIB.
Versions: GSLIB >= 1.0.7.
- ALGOIM (optional), used when MFEM_USE_ALGOIM=YES. The library provides only
- ALGOIM (optional), used when MFE_USE_ALGOIM=YES. The library provides only
headers so it just needs to be downloaded at the same level as MFEM. Download
the specific version we use as:
"git clone https://github.com/algoim/algoim.git;
git checkout 9c9ca0ef094d8ab0390ed36367a1151b459bbe0a"
ALGOIM depends on BLITZ and the library must be built prior to the MFEM build.
ALGOIM depends on BLITZ and rhe library must be built prior to the MFEM build.
Download v1.0.2, untar it at the same level as MFEM and create a symbolic link:
"ln -s blitz-1.0.2 blitz".
Build Blitz using CMake as:
+1 -1
View File
@@ -19,7 +19,7 @@ if(EXISTS "${ENZYME_DIR}/ClangEnzyme-${ENZYME_VERSION}.so")
# Set ENZYME_FOUND
set(ENZYME_FOUND TRUE CACHE BOOL "ENZYME was found." FORCE)
# Set CXX flags to accommodate the Enzyme Clang plugin
# Set CXX flags to accomodate the Enzyme Clang plugin
set(CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -Xclang -load -Xclang ${ENZYME_DIR}/ClangEnzyme-${ENZYME_VERSION}.so -mllvm -enzyme-loose-types=1")
set(MFEM_USE_ENZYME YES)
else()
-57
View File
@@ -1,57 +0,0 @@
# Copyright (c) 2010-2022, Lawrence Livermore National Security, LLC. Produced
# at the Lawrence Livermore National Laboratory. All Rights reserved. See files
# LICENSE and NOTICE for details. LLNL-CODE-806117.
#
# This file is part of the MFEM library. For more information and source code
# availability visit https://mfem.org.
#
# MFEM is free software; you can redistribute it and/or modify it under the
# terms of the BSD-3 license. We welcome feedback and contributions, see file
# CONTRIBUTING.md for details.
# Defines the following variables:
# - HDF5_FOUND - If HDF5 was found
# - HDF5_LIBRARIES - The HDF5 libraries
# - HDF5_INCLUDE_DIRS - The HDF5 include directories
# First Check for HDF5_DIR
if(NOT HDF5_DIR)
MESSAGE(FATAL_ERROR "Could not find HDF5. HDF5 support needs explicit HDF5_DIR")
endif()
# Find includes
find_path( HDF5_INCLUDE_DIRS hdf5.h
PATHS ${HDF5_DIR}/include/
NO_DEFAULT_PATH
NO_CMAKE_ENVIRONMENT_PATH
NO_CMAKE_PATH
NO_SYSTEM_ENVIRONMENT_PATH
NO_CMAKE_SYSTEM_PATH)
find_library( __HDF5_LIBRARY NAMES hdf5 libhdf5 libhdf5_D libhdf5_debug
PATHS ${HDF5_DIR}/lib
NO_DEFAULT_PATH
NO_CMAKE_ENVIRONMENT_PATH
NO_CMAKE_PATH
NO_SYSTEM_ENVIRONMENT_PATH
NO_CMAKE_SYSTEM_PATH)
find_library( __HDF5_HL_LIBRARY NAMES hdf5_hl libhdf5_hl libhdf5_hl_D libhdf5_hl_debug
PATHS ${HDF5_DIR}/lib
NO_DEFAULT_PATH
NO_CMAKE_ENVIRONMENT_PATH
NO_CMAKE_PATH
NO_SYSTEM_ENVIRONMENT_PATH
NO_CMAKE_SYSTEM_PATH)
set(HDF5_LIBRARIES ${__HDF5_HL_LIBRARY} ${__HDF5_LIBRARY})
include(FindPackageHandleStandardArgs)
# Handle the QUIETLY and REQUIRED arguments and set HDF5_FOUND to TRUE if all
# listed variables are TRUE
find_package_handle_standard_args(HDF5 DEFAULT_MSG
HDF5_INCLUDE_DIRS
__HDF5_LIBRARY
__HDF5_HL_LIBRARY
HDF5_LIBRARIES )
+3 -3
View File
@@ -14,6 +14,6 @@
# - UMPIRE_LIBRARIES
# - UMPIRE_INCLUDE_DIRS
find_package(umpire REQUIRED CONFIG)
set(UMPIRE_FOUND ${umpire_FOUND})
set(UMPIRE_LIBRARIES "umpire")
include(MfemCmakeUtilities)
mfem_find_package(UMPIRE UMPIRE UMPIRE_DIR "include" "umpire/Umpire.hpp" "lib" "umpire"
"Paths to headers required by UMPIRE." "Libraries required by UMPIRE.")
+12 -4
View File
@@ -43,14 +43,22 @@ function(convert_filenames_to_full_paths NAMES)
set(${NAMES} ${tmp_names} PARENT_SCOPE)
endfunction()
# Wrapper for add_executable
# Wrapper for add_executable that calls the HIP wrapper if applicable
macro(mfem_add_executable NAME)
add_executable(${NAME} ${ARGN})
if (MFEM_USE_HIP)
add_executable(${NAME} ${ARGN})
else()
add_executable(${NAME} ${ARGN})
endif()
endmacro()
# Wrapper for add_library
# Wrapper for add_library that calls the HIP wrapper if applicable
macro(mfem_add_library NAME)
add_library(${NAME} ${ARGN})
if (MFEM_USE_HIP)
add_library(${NAME} ${ARGN})
else()
add_library(${NAME} ${ARGN})
endif()
endmacro()
# Simple shortcut to add_custom_target() with option to add the target to the
-2
View File
@@ -31,11 +31,9 @@
// Windows specific options
#ifdef _WIN32
#ifndef _USE_MATH_DEFINES
// Macro needed to get defines like M_PI from <cmath>. (Visual Studio C++ only?)
#define _USE_MATH_DEFINES
#endif
#endif
// On Cygwin the option -std=c++11 prevents the definition of M_PI. Defining
// the following macro allows us to get M_PI and some needed functions, e.g.
// posix_memalign(), strdup(), strerror_r().
+7 -11
View File
@@ -179,7 +179,7 @@ ifeq ($(MFEM_USE_MPI)$(MFEM_USE_HIP),YESYES)
endif
# ROCM/HIP directory such that ROCM/HIP libraries like rocsparse and rocrand are
# found in $(HIP_DIR)/lib, usually as links. Typically, this directory is of
# found in $(HIP_DIR)/lib, usually as links. Typically, this directoory is of
# the form /opt/rocm-X.Y.Z which is called ROCM_PATH by hipconfig.
ifeq ($(MFEM_USE_HIP),YES)
HIP_DIR := $(patsubst %/,%,$(dir $(shell which $(HIP_CXX))))
@@ -251,16 +251,12 @@ POSIX_CLOCKS_LIB = -lrt
# SUNDIALS library configuration
# For sundials_nvecmpiplusx and nvecparallel remember to build with MPI_ENABLE=ON
# and modify cmake variables for hypre for sundials
SUNDIALS_DIR = @MFEM_DIR@/../sundials-5.0.0/instdir
# SUNDIALS >= 6.4.0 requires C++14:
ifeq ($(MFEM_USE_SUNDIALS),YES)
BASE_FLAGS = -std=c++14
endif
SUNDIALS_OPT = -I$(SUNDIALS_DIR)/include
SUNDIALS_LIB = $(XLINKER)-rpath,$(SUNDIALS_DIR)/lib64\
$(XLINKER)-rpath,$(SUNDIALS_DIR)/lib\
-L$(SUNDIALS_DIR)/lib64 -L$(SUNDIALS_DIR)/lib\
SUNDIALS_DIR = @MFEM_DIR@/../sundials-5.0.0/instdir
SUNDIALS_OPT = -I$(SUNDIALS_DIR)/include
SUNDIALS_LIBDIR = $(wildcard $(SUNDIALS_DIR)/lib*)
SUNDIALS_LIB = $(XLINKER)-rpath,$(SUNDIALS_LIBDIR) -L$(SUNDIALS_LIBDIR)\
-lsundials_arkode -lsundials_cvodes -lsundials_nvecserial -lsundials_kinsol
ifeq ($(MFEM_USE_MPI),YES)
SUNDIALS_LIB += -lsundials_nvecparallel -lsundials_nvecmpiplusx
endif
@@ -313,7 +309,7 @@ SCALAPACK_LIB = -L$(SCALAPACK_DIR)/lib -lscalapack $(LAPACK_LIB)
MPI_FORTRAN_LIB = -lmpifort
# OpenMPI:
# MPI_FORTRAN_LIB = -lmpi_mpifh
# Additional Fortran library:
# Additional Fortan library:
# MPI_FORTRAN_LIB += -lgfortran
# MUMPS library configuration
+3 -2
View File
@@ -554,14 +554,15 @@ function go()
local cmd_line="${1##+( )}"
cmd_line="${cmd_line%%+( )}"
shopt -u extglob
eval local cmd=(${cmd_line})
local res=""
echo $sep
echo "<${group}>" "${cmd_line}"
echo $sep
if [ "${timing}" == "yes" ]; then
timed_run eval "${cmd_line}"
timed_run "${cmd[@]}"
else
eval "${cmd_line}"
"${cmd[@]}"
fi
if [ "$?" -eq 0 ]; then
res="${green} OK ${none}"
+1 -1
View File
@@ -3,5 +3,5 @@
"version-string": "5.1.0",
"port-version": 0,
"description": "Serial Graph Partitioning and Fill-reducing Matrix Ordering",
"homepage": "http://glaros.dtc.umn.edu/gkhome/metis/metis/overview"
"homepage": "https://glaros.dtc.umn.edu/gkhome/metis/metis/overview"
}
+1 -1
View File
@@ -1,7 +1,7 @@
MFEM mesh v1.0
#
# MFEM Geometry Types (see mesh/geom.hpp):
# MFEM Geomety Types (see mesh/geom.hpp):
#
# POINT = 0
# SEGMENT = 1
+1 -1
View File
@@ -38,7 +38,7 @@ PROJECT_NAME = "MFEM"
# could be handy for archiving the generated documentation or if some version
# control system is used.
PROJECT_NUMBER = v4.5.0
PROJECT_NUMBER = v4.4.1
# Using the PROJECT_BRIEF tag one can provide an optional one line description
# for a project that appears at the top of each page and should give viewer a
+1 -1
View File
@@ -195,7 +195,7 @@ int main(int argc, char *argv[])
Array<int> ess_tdof_list(0);
if (h1 && pmesh.bdr_attributes.Size())
{
// For a continuous basis the linear system must be modified to enforce an
// For a continuous basis the linear system must be modifed to enforce an
// essential (Dirichlet) boundary condition. In the DG case this is not
// necessary as the boundary condition will only be enforced weakly.
fespace.GetEssentialTrueDofs(dbc_bdr, ess_tdof_list);
+1
View File
@@ -197,6 +197,7 @@ int main(int argc, char *argv[])
SparseMatrix &M(mVarf->SpMat());
SparseMatrix &B(bVarf->SpMat());
B *= -1.;
B.EnsureMultTranspose();
Bt = new TransposeOperator(&B);
darcyOp.SetBlock(0,0, &M);
+1
View File
@@ -187,6 +187,7 @@ int main(int argc, char *argv[])
{
// 1. Initialize MPI and HYPRE.
Mpi::Init(argc, argv);
int num_procs = Mpi::WorldSize();
int myid = Mpi::WorldRank();
Hypre::Init();
+6 -10
View File
@@ -248,10 +248,7 @@ int main(int argc, char *argv[])
// constraints for non-conforming AMR, static condensation, etc.
if (myid == 0) { cout << "matrix ... " << flush; }
if (static_cond) { a->EnableStaticCondensation(); }
// Here we want to try out block-size aware AMG solver in PETSc.
// For that to work properly, we need a fully-compliant block-size
// structure and we do not skip zeros when assembling.
a->Assemble(use_petsc ? 0 : 1);
a->Assemble();
Vector B, X;
if (!use_petsc)
@@ -297,14 +294,13 @@ int main(int argc, char *argv[])
cout << "done." << endl;
cout << "Size of linear system: " << A.M() << endl;
}
// Tell PETSc the matrix has a block structure
A.SetBlockSize(dim);
// The preconditioner for the PCG solver can be specified in the
// PETSc config file
PetscPCGSolver *pcg = new PetscPCGSolver(A);
// The preconditioner for the PCG solver defined below is specified in the
// PETSc config file, rc_ex2p, since a Krylov solver in PETSc can also
// customize its preconditioner.
PetscPreconditioner *prec = NULL;
if (use_nonoverlapping) // Specialized BDDC construction
if (use_nonoverlapping)
{
// Compute dofs belonging to the natural boundary
Array<int> nat_tdof_list, nat_bdr(pmesh->bdr_attributes.Max());
+1 -1
View File
@@ -450,7 +450,7 @@ int main(int argc, char *argv[])
for (int ti = 0; !done; )
{
// We cannot match exactly the time history of the Run method
// since we are explicitly telling PETSc to use a time step
// since we are explictly telling PETSc to use a time step
double dt_real = min(dt, t_final - t);
ode_solver->Step(*U, t, dt_real);
ti++;
-2
View File
@@ -78,7 +78,6 @@ EX1_ARGS_CUDA := -m ../../data/star.mesh --usepetsc --partial-assembly -
EX1_ARGS_CUDAAMG := -m ../../data/star.mesh --usepetsc --device cuda --petscopts rc_ex1p_cudaamg
EX2_ARGS := -m ../../data/beam-quad.mesh --usepetsc --petscopts rc_ex2p
EX2_ARGS_BDDC := -m ../../data/beam-tri.mesh --usepetsc --nonoverlapping --petscopts rc_ex2p_bddc
EX2_ARGS_ASM := -m ../../data/beam-quad.mesh --usepetsc --petscopts rc_ex2p_asm
EX3_ARGS := -m ../../data/klein-bottle.mesh -o 2 -f 0.1 --usepetsc --petscopts rc_ex3p_bddc --nonoverlapping
EX4_ARGS := -m ../../data/klein-bottle.mesh -o 2 --usepetsc --petscopts rc_ex4p_bddc --nonoverlapping
EX4_HYB_ARGS := -m ../../data/klein-bottle.mesh -o 2 --usepetsc --petscopts rc_ex4p_bddc --nonoverlapping --hybridization
@@ -110,7 +109,6 @@ endif
ex2p-test-par: ex2p
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX2_ARGS))
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX2_ARGS_BDDC))
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX2_ARGS_ASM))
ex3p-test-par: ex3p
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX3_ARGS))
ex4p-test-par: ex4p
+2 -1
View File
@@ -1,7 +1,8 @@
-ksp_converged_reason
# GAMG is still not used at its best,
# since we are not exploiting the RBMs
# since we are not exploiting the
# block size (Ordering::byVDIM) and the RBMs
-ksp_view
-pc_type gamg
-10
View File
@@ -1,10 +0,0 @@
# Additive Schwarz with Overlap
# This is not a good solver for elasticity
# These options are here only to describe
# the setup of the solver
-ksp_converged_reason
-ksp_view
-ksp_max_it 10
-pc_type asm
-pc_asm_overlap 1
-sub_pc_type icc
-3
View File
@@ -210,9 +210,6 @@ void visualize(ostream &os, Mesh *mesh, GridFunction *deformed_nodes,
int main(int argc, char *argv[])
{
// 0. Initialize SUNDIALS.
Sundials::Init();
// 1. Parse command-line options.
const char *mesh_file = "../../data/beam-quad.mesh";
int ref_levels = 2;
+1 -2
View File
@@ -215,11 +215,10 @@ void visualize(ostream &os, ParMesh *mesh, ParGridFunction *deformed_nodes,
int main(int argc, char *argv[])
{
// 1. Initialize MPI, HYPRE, and SUNDIALS.
// 1. Initialize MPI and HYPRE.
Mpi::Init(argc, argv);
int myid = Mpi::WorldRank();
Hypre::Init();
Sundials::Init();
// 2. Parse command-line options.
const char *mesh_file = "../../data/beam-quad.mesh";
+1 -7
View File
@@ -109,9 +109,6 @@ double InitialTemperature(const Vector &x);
int main(int argc, char *argv[])
{
// 0. Initialize SUNDIALS.
Sundials::Init();
// 1. Parse command-line options.
const char *mesh_file = "../../data/star.mesh";
int ref_levels = 2;
@@ -293,10 +290,7 @@ int main(int argc, char *argv[])
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 11)
{
arkode->SetERKTableNum(ARKODE_FEHLBERG_13_7_8);
}
if (ode_solver_type == 11) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
case 12:
arkode = new ARKStepSolver(ARKStepSolver::IMPLICIT);
+2 -6
View File
@@ -101,12 +101,11 @@ double InitialTemperature(const Vector &x);
int main(int argc, char *argv[])
{
// 1. Initialize MPI, HYPRE, and SUNDIALS.
// 1. Initialize MPI and HYPRE.
Mpi::Init(argc, argv);
int num_procs = Mpi::WorldSize();
int myid = Mpi::WorldRank();
Hypre::Init();
Sundials::Init();
// 2. Parse command-line options.
const char *mesh_file = "../../data/star.mesh";
@@ -328,10 +327,7 @@ int main(int argc, char *argv[])
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 11)
{
arkode->SetERKTableNum(ARKODE_FEHLBERG_13_7_8);
}
if (ode_solver_type == 11) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
case 12:
arkode = new ARKStepSolver(MPI_COMM_WORLD, ARKStepSolver::IMPLICIT);
+1 -4
View File
@@ -140,9 +140,6 @@ public:
int main(int argc, char *argv[])
{
// 0. Initialize SUNDIALS.
Sundials::Init();
// 1. Parse command-line options.
problem = 0;
const char *mesh_file = "../../data/periodic-hexagon.mesh";
@@ -411,7 +408,7 @@ int main(int argc, char *argv[])
arkode->Init(adv);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
arkode->SetERKTableNum(ARKODE_FEHLBERG_13_7_8);
arkode->SetERKTableNum(FEHLBERG_13_7_8);
ode_solver = arkode; break;
}
+2 -6
View File
@@ -152,12 +152,11 @@ public:
int main(int argc, char *argv[])
{
// 1. Initialize MPI, HYPRE, and SUNDIALS.
// 1. Initialize MPI and HYPRE.
Mpi::Init(argc, argv);
int num_procs = Mpi::WorldSize();
int myid = Mpi::WorldRank();
Hypre::Init();
Sundials::Init();
// 2. Parse command-line options.
problem = 0;
@@ -488,10 +487,7 @@ int main(int argc, char *argv[])
arkode->Init(adv);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 9)
{
arkode->SetERKTableNum(ARKODE_FEHLBERG_13_7_8);
}
if (ode_solver_type == 9) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
}
+1 -1
View File
@@ -35,7 +35,7 @@ add_mfem_examples(SUPERLU_EXAMPLES_SRCS ${PFX} "" test_superlu)
if (MFEM_ENABLE_TESTING)
# Command line options for the tests.
# Example 1: Test SuperLU on the simple Poisson problem
set(EX1_COMMON_OPTS -m ../../data/star.mesh)
set(EX1_COMMON_OPTS -m ../../data/star.mesh -p 2)
set(EX1P_TEST_OPTS ${EX1_COMMON_OPTS})
# Add the tests: one test per source file.
-9
View File
@@ -39,7 +39,6 @@ set(SRCS
complex_fem.cpp
convergence.cpp
datacollection.cpp
dgmassinv.cpp
doftrans.cpp
eltrans.cpp
estimators.cpp
@@ -73,7 +72,6 @@ set(SRCS
linearform.cpp
linearform_ext.cpp
lininteg.cpp
lininteg_boundary.cpp
lininteg_domain.cpp
lininteg_domain_grad.cpp
lor/lor.cpp
@@ -90,7 +88,6 @@ set(SRCS
fespacehierarchy.cpp
nonlininteg_vectorconvection.cpp
nonlininteg_vectorconvection_mf.cpp
qfunction.cpp
qinterp/det.cpp
qinterp/eval_by_nodes.cpp
qinterp/eval_by_vdim.cpp
@@ -98,7 +95,6 @@ set(SRCS
qinterp/grad_by_vdim.cpp
qinterp/grad_phys_by_nodes.cpp
qinterp/grad_phys_by_vdim.cpp
qspace.cpp
quadinterpolator.cpp
quadinterpolator_face.cpp
restriction.cpp
@@ -140,13 +136,10 @@ set(HDRS
bilinearform.hpp
bilinearform_ext.hpp
bilininteg.hpp
bilininteg_mass_pa.hpp
coefficient.hpp
complex_fem.hpp
convergence.hpp
datacollection.hpp
dgmassinv.hpp
dgmassinv_kernels.hpp
doftrans.hpp
eltrans.hpp
estimators.hpp
@@ -196,11 +189,9 @@ set(HDRS
nonlinearform.hpp
nonlinearform_ext.hpp
nonlininteg.hpp
qfunction.hpp
qinterp/dispatch.hpp
qinterp/eval.hpp
qinterp/grad.hpp
qspace.hpp
quadinterpolator.hpp
quadinterpolator_face.hpp
restriction.hpp
+1 -2
View File
@@ -136,7 +136,7 @@ void BilinearForm::SetAssemblyLevel(AssemblyLevel assembly_level)
ext = new MFBilinearFormExtension(this);
break;
default:
MFEM_ABORT("BilinearForm: unknown assembly level");
mfem_error("Unknown assembly level");
}
}
@@ -992,7 +992,6 @@ void BilinearForm::EliminateVDofs(const Array<int> &vdofs_,
mat_e = new SparseMatrix(height);
}
vdofs_.HostRead();
for (int i = 0; i < vdofs_.Size(); i++)
{
int vdof = vdofs_[i];
+4 -5
View File
@@ -26,8 +26,7 @@ namespace mfem
{
/** @brief Enumeration defining the assembly level for bilinear and nonlinear
form classes derived from Operator. For more details, see
https://mfem.org/howto/assembly_levels */
form classes derived from Operator. */
enum class AssemblyLevel
{
/// In the case of a BilinearForm LEGACY corresponds to a fully assembled
@@ -178,7 +177,7 @@ public:
- AssemblyLevel::ELEMENT
- AssemblyLevel::NONE
If used, this method must be called before assembly. */
This method must be called before assembly. */
void SetAssemblyLevel(AssemblyLevel assembly_level);
/// Returns the assembly level
@@ -334,7 +333,7 @@ public:
/** @brief Nullifies the internal matrix \f$ M \f$ and returns a pointer
to it. Used for transferring ownership. */
to it. Used for transfering ownership. */
SparseMatrix *LoseMat() { SparseMatrix *tmp = mat; mat = NULL; return tmp; }
/** @brief Returns a const reference to the sparse matrix of eliminated b.c.:
@@ -775,7 +774,7 @@ public:
SparseMatrix &SpMat() { return *mat; }
/** @brief Nullifies the internal matrix \f$ M \f$ and returns a pointer
to it. Used for transferring ownership. */
to it. Used for transfering ownership. */
SparseMatrix *LoseMat() { SparseMatrix *tmp = mat; mat = NULL; return tmp; }
/// Adds a domain integrator. Assumes ownership of @a bfi.
+12 -25
View File
@@ -18,8 +18,6 @@
#include "pgridfunc.hpp"
#include "ceed/interface/util.hpp"
#include "../general/nvtx.hpp"
namespace mfem
{
@@ -162,7 +160,7 @@ void MFBilinearFormExtension::Mult(const Vector &x, Vector &y) const
{
intFaceIntegrators[i]->AddMultMF(int_face_X, int_face_Y);
}
int_face_restrict_lex->AddMultTransposeInPlace(int_face_Y, y);
int_face_restrict_lex->AddMultTranspose(int_face_Y, y);
}
}
@@ -178,7 +176,7 @@ void MFBilinearFormExtension::Mult(const Vector &x, Vector &y) const
{
bdrFaceIntegrators[i]->AddMultMF(bdr_face_X, bdr_face_Y);
}
bdr_face_restrict_lex->AddMultTransposeInPlace(bdr_face_Y, y);
bdr_face_restrict_lex->AddMultTranspose(bdr_face_Y, y);
}
}
}
@@ -219,7 +217,7 @@ void MFBilinearFormExtension::MultTranspose(const Vector &x, Vector &y) const
{
intFaceIntegrators[i]->AddMultTransposeMF(int_face_X, int_face_Y);
}
int_face_restrict_lex->AddMultTransposeInPlace(int_face_Y, y);
int_face_restrict_lex->AddMultTranspose(int_face_Y, y);
}
}
@@ -235,7 +233,7 @@ void MFBilinearFormExtension::MultTranspose(const Vector &x, Vector &y) const
{
bdrFaceIntegrators[i]->AddMultTransposeMF(bdr_face_X, bdr_face_Y);
}
bdr_face_restrict_lex->AddMultTransposeInPlace(bdr_face_Y, y);
bdr_face_restrict_lex->AddMultTranspose(bdr_face_Y, y);
}
}
}
@@ -291,10 +289,6 @@ void PABilinearFormExtension::SetupRestrictionOperators(const L2FaceValues m)
void PABilinearFormExtension::Assemble()
{
#undef MFEM_NVTX_COLOR
#define MFEM_NVTX_COLOR NavyBlue
NVTX("HO Assemble");
SetupRestrictionOperators(L2FaceValues::DoubleValued);
Array<BilinearFormIntegrator*> &integrators = *a->GetDBFI();
@@ -389,10 +383,6 @@ void PABilinearFormExtension::FormLinearSystem(const Array<int> &ess_tdof_list,
void PABilinearFormExtension::Mult(const Vector &x, Vector &y) const
{
#undef MFEM_NVTX_COLOR
#define MFEM_NVTX_COLOR MediumSpringGreen
NVTX("HO Apply");
Array<BilinearFormIntegrator*> &integrators = *a->GetDBFI();
const int iSz = integrators.Size();
@@ -428,7 +418,7 @@ void PABilinearFormExtension::Mult(const Vector &x, Vector &y) const
{
intFaceIntegrators[i]->AddMultPA(int_face_X, int_face_Y);
}
int_face_restrict_lex->AddMultTransposeInPlace(int_face_Y, y);
int_face_restrict_lex->AddMultTranspose(int_face_Y, y);
}
}
@@ -444,7 +434,7 @@ void PABilinearFormExtension::Mult(const Vector &x, Vector &y) const
{
bdrFaceIntegrators[i]->AddMultPA(bdr_face_X, bdr_face_Y);
}
bdr_face_restrict_lex->AddMultTransposeInPlace(bdr_face_Y, y);
bdr_face_restrict_lex->AddMultTranspose(bdr_face_Y, y);
}
}
}
@@ -485,7 +475,7 @@ void PABilinearFormExtension::MultTranspose(const Vector &x, Vector &y) const
{
intFaceIntegrators[i]->AddMultTransposePA(int_face_X, int_face_Y);
}
int_face_restrict_lex->AddMultTransposeInPlace(int_face_Y, y);
int_face_restrict_lex->AddMultTranspose(int_face_Y, y);
}
}
@@ -501,7 +491,7 @@ void PABilinearFormExtension::MultTranspose(const Vector &x, Vector &y) const
{
bdrFaceIntegrators[i]->AddMultTransposePA(bdr_face_X, bdr_face_Y);
}
bdr_face_restrict_lex->AddMultTransposeInPlace(bdr_face_Y, y);
bdr_face_restrict_lex->AddMultTranspose(bdr_face_Y, y);
}
}
}
@@ -678,7 +668,7 @@ void EABilinearFormExtension::Mult(const Vector &x, Vector &y) const
Y(j, 0, f) += res;
});
// Apply the Interior Face Restriction transposed
int_face_restrict_lex->AddMultTransposeInPlace(int_face_Y, y);
int_face_restrict_lex->AddMultTranspose(int_face_Y, y);
}
}
@@ -709,7 +699,7 @@ void EABilinearFormExtension::Mult(const Vector &x, Vector &y) const
Y(j, f) += res;
});
// Apply the Boundary Face Restriction transposed
bdr_face_restrict_lex->AddMultTransposeInPlace(bdr_face_Y, y);
bdr_face_restrict_lex->AddMultTranspose(bdr_face_Y, y);
}
}
}
@@ -806,7 +796,7 @@ void EABilinearFormExtension::MultTranspose(const Vector &x, Vector &y) const
Y(j, 0, f) += res;
});
// Apply the Interior Face Restriction transposed
int_face_restrict_lex->AddMultTransposeInPlace(int_face_Y, y);
int_face_restrict_lex->AddMultTranspose(int_face_Y, y);
}
}
@@ -837,7 +827,7 @@ void EABilinearFormExtension::MultTranspose(const Vector &x, Vector &y) const
Y(j, f) += res;
});
// Apply the Boundary Face Restriction transposed
bdr_face_restrict_lex->AddMultTransposeInPlace(bdr_face_Y, y);
bdr_face_restrict_lex->AddMultTranspose(bdr_face_Y, y);
}
}
}
@@ -985,9 +975,6 @@ void FABilinearFormExtension::RAP(OperatorHandle &A)
void FABilinearFormExtension::EliminateBC(const Array<int> &ess_dofs,
OperatorHandle &A)
{
MFEM_VERIFY(a->diag_policy == DiagonalPolicy::DIAG_ONE,
"Only DiagonalPolicy::DIAG_ONE supported with"
" FABilinearFormExtension.");
#ifdef MFEM_USE_MPI
if ( dynamic_cast<ParBilinearForm*>(a) )
{
+2 -205
View File
@@ -2003,83 +2003,6 @@ void CurlCurlIntegrator::AssembleElementMatrix
}
}
void CurlCurlIntegrator::AssembleElementMatrix2(const FiniteElement &trial_fe,
const FiniteElement &test_fe,
ElementTransformation &Trans,
DenseMatrix &elmat)
{
int tr_nd = trial_fe.GetDof();
int te_nd = test_fe.GetDof();
dim = trial_fe.GetDim();
int dimc = trial_fe.GetCurlDim();
double w;
#ifdef MFEM_THREAD_SAFE
Vector D;
DenseMatrix curlshape(tr_nd,dimc), curlshape_dFt(tr_nd,dimc), M;
DenseMatrix te_curlshape(te_nd,dimc), te_curlshape_dFt(te_nd,dimc);
#else
curlshape.SetSize(tr_nd,dimc);
curlshape_dFt.SetSize(tr_nd,dimc);
te_curlshape.SetSize(te_nd,dimc);
te_curlshape_dFt.SetSize(te_nd,dimc);
#endif
elmat.SetSize(te_nd, tr_nd);
if (MQ) { M.SetSize(dimc); }
if (DQ) { D.SetSize(dimc); }
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int order;
if (trial_fe.Space() == FunctionSpace::Pk)
{
order = test_fe.GetOrder() + trial_fe.GetOrder() - 2;
}
else
{
order = test_fe.GetOrder() + trial_fe.GetOrder() + trial_fe.GetDim() - 1;
}
ir = &IntRules.Get(trial_fe.GetGeomType(), order);
}
elmat = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
{
const IntegrationPoint &ip = ir->IntPoint(i);
Trans.SetIntPoint(&ip);
w = ip.weight * Trans.Weight();
trial_fe.CalcPhysCurlShape(Trans, curlshape_dFt);
test_fe.CalcPhysCurlShape(Trans, te_curlshape_dFt);
if (MQ)
{
MQ->Eval(M, Trans, ip);
M *= w;
Mult(te_curlshape_dFt, M, te_curlshape);
AddMultABt(te_curlshape, curlshape_dFt, elmat);
}
else if (DQ)
{
DQ->Eval(D, Trans, ip);
D *= w;
AddMultADBt(te_curlshape_dFt,D,curlshape_dFt,elmat);
}
else
{
if (Q)
{
w *= Q->Eval(Trans, ip);
}
curlshape_dFt *= w;
AddMultABt(te_curlshape_dFt, curlshape_dFt, elmat);
}
}
}
void CurlCurlIntegrator
::ComputeElementFlux(const FiniteElement &el, ElementTransformation &Trans,
Vector &u, const FiniteElement &fluxelem, Vector &flux,
@@ -2317,84 +2240,6 @@ double VectorCurlCurlIntegrator::GetElementEnergy(
return 0.5 * energy;
}
void MixedCurlIntegrator::AssembleElementMatrix2(
const FiniteElement &trial_fe, const FiniteElement &test_fe,
ElementTransformation &Trans, DenseMatrix &elmat)
{
int dim = trial_fe.GetDim();
int trial_dof = trial_fe.GetDof();
int test_dof = test_fe.GetDof();
int dimc = (dim == 3) ? 3 : 1;
MFEM_VERIFY(trial_fe.GetMapType() == mfem::FiniteElement::H_CURL ||
(dim == 2 && trial_fe.GetMapType() == mfem::FiniteElement::VALUE),
"Trial finite element must be either 2D/3D H(Curl) or 2D H1");
MFEM_VERIFY(test_fe.GetMapType() == mfem::FiniteElement::VALUE ||
test_fe.GetMapType() == mfem::FiniteElement::INTEGRAL,
"Test finite element must be in H1/L2");
bool spaceH1 = (trial_fe.GetMapType() == mfem::FiniteElement::VALUE);
if (spaceH1)
{
dshape.SetSize(trial_dof,dim);
curlshape.SetSize(dim*trial_dof,1);
dimc = dim;
}
else
{
curlshape.SetSize(trial_dof,dimc);
elmat_comp.SetSize(test_dof, trial_dof);
}
elmat.SetSize(dimc * test_dof, trial_dof);
shape.SetSize(test_dof);
elmat = 0.0;
double c;
Vector d_col;
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int order = trial_fe.GetOrder() + test_fe.GetOrder() + Trans.OrderJ();
ir = &IntRules.Get(trial_fe.GetGeomType(), order);
}
for (int i = 0; i < ir->GetNPoints(); i++)
{
const IntegrationPoint &ip = ir->IntPoint(i);
Trans.SetIntPoint(&ip);
if (spaceH1)
{
trial_fe.CalcPhysDShape(Trans, dshape);
dshape.GradToCurl(curlshape);
}
else
{
trial_fe.CalcPhysCurlShape(Trans, curlshape);
}
test_fe.CalcPhysShape(Trans, shape);
c = ip.weight*Trans.Weight();
if (Q)
{
c *= Q->Eval(Trans, ip);
}
shape *= c;
for (int d = 0; d < dimc; ++d)
{
double * curldata = &(curlshape.GetData())[d*trial_dof];
for (int jj = 0; jj < trial_dof; ++jj)
{
for (int ii = 0; ii < test_dof; ++ii)
{
elmat(d * test_dof + ii, jj) += shape(ii) * curldata[jj];
}
}
}
}
}
void VectorFEMassIntegrator::AssembleElementMatrix(
const FiniteElement &el,
@@ -2741,54 +2586,6 @@ void DivDivIntegrator::AssembleElementMatrix(
}
}
void DivDivIntegrator::AssembleElementMatrix2(
const FiniteElement &trial_fe,
const FiniteElement &test_fe,
ElementTransformation &Trans,
DenseMatrix &elmat)
{
int tr_nd = trial_fe.GetDof();
int te_nd = test_fe.GetDof();
double c;
#ifdef MFEM_THREAD_SAFE
Vector divshape(tr_nd);
Vector te_divshape(te_nd);
#else
divshape.SetSize(tr_nd);
te_divshape.SetSize(te_nd);
#endif
elmat.SetSize(te_nd,tr_nd);
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int order = 2 * max(test_fe.GetOrder(),
trial_fe.GetOrder()) - 2; // <--- OK for RTk
ir = &IntRules.Get(test_fe.GetGeomType(), order);
}
elmat = 0.0;
for (int i = 0; i < ir -> GetNPoints(); i++)
{
const IntegrationPoint &ip = ir->IntPoint(i);
trial_fe.CalcDivShape(ip,divshape);
test_fe.CalcDivShape(ip,te_divshape);
Trans.SetIntPoint (&ip);
c = ip.weight / Trans.Weight();
if (Q)
{
c *= Q -> Eval (Trans, ip);
}
te_divshape *= c;
AddMultVWt(te_divshape, divshape, elmat);
}
}
void VectorDiffusionIntegrator::AssembleElementMatrix(
const FiniteElement &el,
@@ -3983,7 +3780,7 @@ void NormalTraceJumpIntegrator::AssembleFaceMatrix(
for (i = 0; i < ndof1; i++)
for (j = 0; j < face_ndof; j++)
{
elmat(i, j) += shape1_n(i) * face_shape(j);
elmat(i, j) -= shape1_n(i) * face_shape(j);
}
if (ndof2)
{
@@ -3991,7 +3788,7 @@ void NormalTraceJumpIntegrator::AssembleFaceMatrix(
for (i = 0; i < ndof2; i++)
for (j = 0; j < face_ndof; j++)
{
elmat(ndof1+i, j) -= shape2_n(i) * face_shape(j);
elmat(ndof1+i, j) += shape2_n(i) * face_shape(j);
}
}
}
+5 -47
View File
@@ -215,10 +215,10 @@ public:
function by any coefficients describing the
integrator.
@param[in] ir If passed (the default value is NULL), the implementation
of the method will ignore the integration rule provided
by the @a fluxelem parameter and, instead, compute the
discrete flux at the points specified by the integration
rule @a ir.
of the method will ignore the integration rule provided
by the @a fluxelem parameter and, instead, compute the
discrete flux at the points specified by the integration
rule @a ir.
*/
virtual void ComputeElementFlux(const FiniteElement &el,
ElementTransformation &Trans,
@@ -2174,7 +2174,6 @@ public:
/** Class for local mass matrix assembling a(u,v) := (Q u, v) */
class MassIntegrator: public BilinearFormIntegrator
{
friend class DGMassInverse;
protected:
#ifndef MFEM_THREAD_SAFE
Vector shape, te_shape;
@@ -2525,7 +2524,6 @@ private:
#ifndef MFEM_THREAD_SAFE
Vector D;
DenseMatrix curlshape, curlshape_dFt, M;
DenseMatrix te_curlshape, te_curlshape_dFt;
DenseMatrix vshape, projcurl;
#endif
@@ -2559,11 +2557,6 @@ public:
ElementTransformation &Trans,
DenseMatrix &elmat);
virtual void AssembleElementMatrix2(const FiniteElement &trial_fe,
const FiniteElement &test_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
virtual void ComputeElementFlux(const FiniteElement &el,
ElementTransformation &Trans,
Vector &u, const FiniteElement &fluxelem,
@@ -2609,35 +2602,6 @@ public:
const Vector &elfun);
};
/** Class for integrating the bilinear form a(u,v) := (Q curl u, v) where Q is
an optional scalar coefficient, and v is a vector with components v_i in
the L2 or H1 space. This integrator handles 3 cases:
(a) u H(curl) in 3D, v is a 3D vector with components v_i in L^2 or H^1
(b) u H(curl) in 2D, v is a scalar field in L^2 or H^1
(c) u is a scalar field in H^1, i.e, curl u := [0 1;-1 0]grad u and v is a
2D vector field with components v_i in L^2 or H^1 space.
Note: Case (b) can also be handled by MixedScalarCurlIntegrator */
class MixedCurlIntegrator : public BilinearFormIntegrator
{
protected:
Coefficient *Q;
private:
Vector shape;
DenseMatrix dshape;
DenseMatrix curlshape;
DenseMatrix elmat_comp;
public:
MixedCurlIntegrator() : Q{NULL} { }
MixedCurlIntegrator(Coefficient *q_) : Q{q_} { }
MixedCurlIntegrator(Coefficient &q) : Q{&q} { }
virtual void AssembleElementMatrix2(const FiniteElement &trial_fe,
const FiniteElement &test_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
};
/** Integrator for (Q u, v), where Q is an optional coefficient (of type scalar,
vector (diagonal matrix), or matrix), trial function u is in H(Curl) or
H(Div), and test function v is in H(Curl), H(Div), or v=(v1,...,vn), where
@@ -2761,7 +2725,7 @@ protected:
private:
#ifndef MFEM_THREAD_SAFE
Vector divshape, te_divshape;
Vector divshape;
#endif
// PA extension
@@ -2779,12 +2743,6 @@ public:
virtual void AssembleElementMatrix(const FiniteElement &el,
ElementTransformation &Trans,
DenseMatrix &elmat);
virtual void AssembleElementMatrix2(const FiniteElement &trial_fe,
const FiniteElement &test_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
const Coefficient *GetCoefficient() const { return Q; }
};
+58 -3
View File
@@ -12,7 +12,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qfunction.hpp"
#include "ceed/integrators/convection/convection.hpp"
#include "quadinterpolator.hpp"
@@ -1409,10 +1408,66 @@ void ConvectionIntegrator::AssemblePA(const FiniteElementSpace &fes)
dofs1D = maps->ndof;
quad1D = maps->nqpt;
pa_data.SetSize(symmDims * nq * ne, mt);
Vector vel;
if (VectorConstantCoefficient *cQ =
dynamic_cast<VectorConstantCoefficient*>(Q))
{
vel = cQ->GetVec();
}
else if (VectorGridFunctionCoefficient *vgfQ =
dynamic_cast<VectorGridFunctionCoefficient*>(Q))
{
vel.SetSize(dim * nq * ne, mt);
QuadratureSpace qs(*mesh, *ir);
CoefficientVector vel(*Q, qs, CoefficientStorage::COMPRESSED);
const GridFunction *gf = vgfQ->GetGridFunction();
const FiniteElementSpace &gf_fes = *gf->FESpace();
const QuadratureInterpolator *qi(gf_fes.GetQuadratureInterpolator(*ir));
const bool use_tensor_products = UsesTensorBasis(gf_fes);
const ElementDofOrdering ordering = use_tensor_products ?
ElementDofOrdering::LEXICOGRAPHIC :
ElementDofOrdering::NATIVE;
const Operator *R = gf_fes.GetElementRestriction(ordering);
Vector xe(R->Height(), mt);
xe.UseDevice(true);
R->Mult(*gf, xe);
qi->SetOutputLayout(QVectorLayout::byVDIM);
qi->DisableTensorProducts(!use_tensor_products);
qi->Values(xe,vel);
}
else if (VectorQuadratureFunctionCoefficient* vqfQ =
dynamic_cast<VectorQuadratureFunctionCoefficient*>(Q))
{
const QuadratureFunction &qFun = vqfQ->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == dim * nq * ne,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
vel.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
vel.SetSize(dim * nq * ne);
auto C = Reshape(vel.HostWrite(), dim, nq, ne);
DenseMatrix MQ_ir;
for (int e = 0; e < ne; ++e)
{
ElementTransformation& T = *fes.GetElementTransformation(e);
Q->Eval(MQ_ir, T, *ir);
for (int q = 0; q < nq; ++q)
{
for (int i = 0; i < dim; ++i)
{
C(i,q,e) = MQ_ir(i,q);
}
}
}
}
PAConvectionSetup(dim, nq, ne, ir->GetWeights(), geom->J,
vel, alpha, pa_data);
}
+103 -37
View File
@@ -12,7 +12,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qfunction.hpp"
#include "restriction.hpp"
using namespace std;
@@ -162,24 +161,88 @@ void DGTraceIntegrator::SetupPA(const FiniteElementSpace &fes, FaceType type)
dofs1D = maps->ndof;
quad1D = maps->nqpt;
pa_data.SetSize(symmDims * nq * nf, Device::GetMemoryType());
FaceQuadratureSpace qs(*mesh, *ir, type);
CoefficientVector vel(*u, qs, CoefficientStorage::COMPRESSED);
CoefficientVector r(qs, CoefficientStorage::COMPRESSED);
if (rho == nullptr)
Vector vel;
if (VectorConstantCoefficient *c_u = dynamic_cast<VectorConstantCoefficient*>
(u))
{
r.SetConstant(1.0);
vel = c_u->GetVec();
}
else if (ConstantCoefficient *const_rho = dynamic_cast<ConstantCoefficient*>
(rho))
else if (VectorQuadratureFunctionCoefficient* qf_u =
dynamic_cast<VectorQuadratureFunctionCoefficient*>(u))
{
r.SetConstant(const_rho->constant);
// Assumed to be in lexicographical ordering
const QuadratureFunction &qFun = qf_u->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == dim * nq * nf,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
vel.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
vel.SetSize(dim * nq * nf);
auto C = Reshape(vel.HostWrite(), dim, nq, nf);
Vector Vq(dim);
int f_ind = 0;
for (int f = 0; f < mesh->GetNumFacesWithGhost(); ++f)
{
Mesh::FaceInformation face = mesh->GetFaceInformation(f);
if (face.IsNonconformingCoarse())
{
// We skip nonconforming coarse faces as they are treated
// by the corresponding nonconforming fine faces.
continue;
}
else if ( face.IsOfFaceType(type) )
{
const int mask = FaceElementTransformations::HAVE_ELEM1 |
FaceElementTransformations::HAVE_LOC1;
FaceElementTransformations &T =
*fes.GetMesh()->GetFaceElementTransformations(f, mask);
for (int q = 0; q < nq; ++q)
{
// Convert to lexicographic ordering
int iq = ToLexOrdering(dim, face.element[0].local_face_id,
quad1D, q);
T.SetAllIntPoints(&ir->IntPoint(q));
const IntegrationPoint &eip1 = T.GetElement1IntPoint();
u->Eval(Vq, *T.Elem1, eip1);
for (int i = 0; i < dim; ++i)
{
C(i,iq,f_ind) = Vq(i);
}
}
f_ind++;
}
}
MFEM_VERIFY(f_ind==nf, "Incorrect number of faces.");
}
Vector r;
if (rho==nullptr)
{
r.SetSize(1);
r(0) = 1.0;
}
else if (ConstantCoefficient *c_rho = dynamic_cast<ConstantCoefficient*>(rho))
{
r.SetSize(1);
r(0) = c_rho->constant;
}
else if (QuadratureFunctionCoefficient* qf_rho =
dynamic_cast<QuadratureFunctionCoefficient*>(rho))
{
r.MakeRef(qf_rho->GetQuadFunction());
const QuadratureFunction &qFun = qf_rho->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == nq * nf,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
r.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
@@ -191,42 +254,45 @@ void DGTraceIntegrator::SetupPA(const FiniteElementSpace &fes, FaceType type)
for (int f = 0; f < mesh->GetNumFacesWithGhost(); ++f)
{
Mesh::FaceInformation face = mesh->GetFaceInformation(f);
if (face.IsNonconformingCoarse() || !face.IsOfFaceType(type))
if (face.IsNonconformingCoarse())
{
// We skip nonconforming coarse faces as they are treated
// by the corresponding nonconforming fine faces.
continue;
}
FaceElementTransformations &T =
*fes.GetMesh()->GetFaceElementTransformations(f);
for (int q = 0; q < nq; ++q)
else if ( face.IsOfFaceType(type) )
{
// Convert to lexicographic ordering
int iq = ToLexOrdering(dim, face.element[0].local_face_id,
quad1D, q);
T.SetAllIntPoints(&ir->IntPoint(q));
const IntegrationPoint &eip1 = T.GetElement1IntPoint();
const IntegrationPoint &eip2 = T.GetElement2IntPoint();
double rq;
if (face.IsBoundary())
FaceElementTransformations &T =
*fes.GetMesh()->GetFaceElementTransformations(f);
for (int q = 0; q < nq; ++q)
{
rq = rho->Eval(*T.Elem1, eip1);
}
else
{
double udotn = 0.0;
for (int d=0; d<dim; ++d)
// Convert to lexicographic ordering
int iq = ToLexOrdering(dim, face.element[0].local_face_id,
quad1D, q);
T.SetAllIntPoints(&ir->IntPoint(q));
const IntegrationPoint &eip1 = T.GetElement1IntPoint();
const IntegrationPoint &eip2 = T.GetElement2IntPoint();
double rq;
if ( face.IsBoundary() )
{
udotn += C_vel(d,iq,f_ind)*n(iq,d,f_ind);
rq = rho->Eval(*T.Elem1, eip1);
}
if (udotn >= 0.0) { rq = rho->Eval(*T.Elem2, eip2); }
else { rq = rho->Eval(*T.Elem1, eip1); }
else
{
double udotn = 0.0;
for (int d=0; d<dim; ++d)
{
udotn += C_vel(d,iq,f_ind)*n(iq,d,f_ind);
}
if (udotn >= 0.0) { rq = rho->Eval(*T.Elem2, eip2); }
else { rq = rho->Eval(*T.Elem1, eip1); }
}
C(iq,f_ind) = rq;
}
C(iq,f_ind) = rq;
f_ind++;
}
f_ind++;
}
MFEM_VERIFY(f_ind==nf, "Incorrect number of faces.");
}
+112 -21
View File
@@ -12,7 +12,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qfunction.hpp"
#include "ceed/integrators/diffusion/diffusion.hpp"
using namespace std;
@@ -391,21 +390,120 @@ void DiffusionIntegrator::AssemblePA(const FiniteElementSpace &fes)
maps = &el.GetDofToQuad(*ir, DofToQuad::TENSOR);
dofs1D = maps->ndof;
quad1D = maps->nqpt;
int coeffDim = 1;
Vector coeff;
const int MQfullDim = MQ ? MQ->GetHeight() * MQ->GetWidth() : 0;
if (auto *SMQ = dynamic_cast<SymmetricMatrixCoefficient *>(MQ))
{
MFEM_VERIFY(SMQ->GetSize() == dim, "");
coeffDim = symmDims;
coeff.SetSize(symmDims * nq * ne);
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(qs, CoefficientStorage::COMPRESSED);
DenseSymmetricMatrix sym_mat;
sym_mat.SetSize(dim);
if (MQ) { coeff.ProjectTranspose(*MQ); }
else if (VQ) { coeff.Project(*VQ); }
else if (Q) { coeff.Project(*Q); }
else { coeff.SetConstant(1.0); }
auto C = Reshape(coeff.HostWrite(), symmDims, nq, ne);
const int coeff_dim = coeff.GetVDim();
symmetric = (coeff_dim != dims*dims);
const int pa_size = symmetric ? symmDims : dims*dims;
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
SMQ->Eval(sym_mat, *tr, ir->IntPoint(p));
int cnt = 0;
for (int i=0; i<dim; ++i)
for (int j=i; j<dim; ++j, ++cnt)
{
C(cnt, p, e) = sym_mat(i,j);
}
}
}
}
else if (MQ)
{
symmetric = false;
MFEM_VERIFY(MQ->GetHeight() == dim && MQ->GetWidth() == dim, "");
pa_data.SetSize(pa_size * nq * ne, mt);
PADiffusionSetup(dim, sdim, dofs1D, quad1D, coeff_dim, ne, ir->GetWeights(),
coeffDim = MQfullDim;
coeff.SetSize(MQfullDim * nq * ne);
DenseMatrix mat;
mat.SetSize(dim);
auto C = Reshape(coeff.HostWrite(), MQfullDim, nq, ne);
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
MQ->Eval(mat, *tr, ir->IntPoint(p));
for (int i=0; i<dim; ++i)
for (int j=0; j<dim; ++j)
{
C(j+(i*dim), p, e) = mat(i,j);
}
}
}
}
else if (VQ)
{
MFEM_VERIFY(VQ->GetVDim() == dim, "");
coeffDim = VQ->GetVDim();
coeff.SetSize(coeffDim * nq * ne);
auto C = Reshape(coeff.HostWrite(), coeffDim, nq, ne);
Vector DM(coeffDim);
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
VQ->Eval(DM, *tr, ir->IntPoint(p));
for (int i=0; i<coeffDim; ++i)
{
C(i, p, e) = DM[i];
}
}
}
}
else if (Q == nullptr)
{
coeff.SetSize(1);
coeff(0) = 1.0;
}
else if (ConstantCoefficient* cQ = dynamic_cast<ConstantCoefficient*>(Q))
{
coeff.SetSize(1);
coeff(0) = cQ->constant;
}
else if (QuadratureFunctionCoefficient* qfQ =
dynamic_cast<QuadratureFunctionCoefficient*>(Q))
{
const QuadratureFunction &qFun = qfQ->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == ne*nq,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
coeff.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
coeff.SetSize(nq * ne);
auto C = Reshape(coeff.HostWrite(), nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation& T = *fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
C(q,e) = Q->Eval(T, ir->IntPoint(q));
}
}
}
pa_data.SetSize((symmetric ? symmDims : MQfullDim) * nq * ne, mt);
PADiffusionSetup(dim, sdim, dofs1D, quad1D, coeffDim, ne, ir->GetWeights(),
geom->J, coeff, pa_data);
}
@@ -1686,7 +1784,7 @@ static void PADiffusionApply(const int dim,
case 0x77: return SmemPADiffusionApply2D<7,7,4>(NE,symm,B,G,D,X,Y);
case 0x88: return SmemPADiffusionApply2D<8,8,2>(NE,symm,B,G,D,X,Y);
case 0x99: return SmemPADiffusionApply2D<9,9,2>(NE,symm,B,G,D,X,Y);
// default: return PADiffusionApply2D(NE,symm,B,G,Bt,Gt,D,X,Y,D1D,Q1D);
default: return PADiffusionApply2D(NE,symm,B,G,Bt,Gt,D,X,Y,D1D,Q1D);
}
}
@@ -1704,14 +1802,7 @@ static void PADiffusionApply(const int dim,
case 0x67: return SmemPADiffusionApply3D<6,7>(NE,symm,B,G,D,X,Y);
case 0x78: return SmemPADiffusionApply3D<7,8>(NE,symm,B,G,D,X,Y);
case 0x89: return SmemPADiffusionApply3D<8,9>(NE,symm,B,G,D,X,Y);
case 0x33: return SmemPADiffusionApply3D<3,3>(NE,symm,B,G,D,X,Y);
case 0x44: return SmemPADiffusionApply3D<4,4>(NE,symm,B,G,D,X,Y);
case 0x55: return SmemPADiffusionApply3D<5,5>(NE,symm,B,G,D,X,Y);
case 0x66: return SmemPADiffusionApply3D<6,6>(NE,symm,B,G,D,X,Y);
case 0x77: return SmemPADiffusionApply3D<7,7>(NE,symm,B,G,D,X,Y);
case 0x88: return SmemPADiffusionApply3D<8,8>(NE,symm,B,G,D,X,Y);
case 0x99: return SmemPADiffusionApply3D<9,9>(NE,symm,B,G,D,X,Y);
// default: return PADiffusionApply3D(NE,symm,B,G,Bt,Gt,D,X,Y,D1D,Q1D);
default: return PADiffusionApply3D(NE,symm,B,G,Bt,Gt,D,X,Y,D1D,Q1D);
}
}
MFEM_ABORT("Unknown kernel: 0x"<<std::hex << id << std::dec);
+39 -3
View File
@@ -12,7 +12,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qfunction.hpp"
using namespace std;
@@ -210,8 +209,44 @@ void GradientIntegrator::AssemblePA(const FiniteElementSpace &trial_fes,
"PA requires test and trial space to have same number of quadrature points!");
pa_data.SetSize(nq * dimsToStore * ne, Device::GetMemoryType());
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
Vector coeff;
if (Q == nullptr)
{
coeff.SetSize(1);
coeff(0) = 1.0;
}
else if (ConstantCoefficient* cQ = dynamic_cast<ConstantCoefficient*>(Q))
{
coeff.SetSize(1);
coeff(0) = cQ->constant;
}
else if (QuadratureFunctionCoefficient* qfQ =
dynamic_cast<QuadratureFunctionCoefficient*>(Q))
{
const QuadratureFunction &qFun = qfQ->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == ne*nq,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
coeff.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
coeff.SetSize(nq * ne);
auto C = Reshape(coeff.HostWrite(), nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation& T = *trial_fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
C(q,e) = Q->Eval(T, ir->IntPoint(q));
}
}
}
PAGradientSetup(dim, trial_dofs1D, test_dofs1D, quad1D,
ne, ir->GetWeights(), geom->J, coeff, pa_data);
@@ -830,3 +865,4 @@ void GradientIntegrator::AddMultTransposePA(const Vector &x, Vector &y) const
}
} // namespace mfem
+169 -33
View File
@@ -12,7 +12,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qspace.hpp"
using namespace std;
@@ -968,6 +967,8 @@ void CurlCurlIntegrator::AssemblePA(const FiniteElementSpace &fes)
dim = mesh->Dimension();
MFEM_VERIFY(dim == 2 || dim == 3, "");
const int dimc = (dim == 3) ? 3 : 1;
ne = fes.GetNE();
geom = mesh->GetGeometricFactors(*ir, GeometricFactors::JACOBIANS);
mapsC = &el->GetDofToQuad(*ir, DofToQuad::TENSOR);
@@ -977,19 +978,88 @@ void CurlCurlIntegrator::AssemblePA(const FiniteElementSpace &fes)
MFEM_VERIFY(dofs1D == mapsO->ndof + 1 && quad1D == mapsO->nqpt, "");
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(qs, CoefficientStorage::SYMMETRIC);
if (Q) { coeff.Project(*Q); }
else if (MQ) { coeff.ProjectTranspose(*MQ); }
else if (DQ) { coeff.Project(*DQ); }
else { coeff.SetConstant(1.0); }
auto SMQ = dynamic_cast<SymmetricMatrixCoefficient *>(MQ);
const int coeff_dim = coeff.GetVDim();
symmetric = (coeff_dim != dim*dim);
const int sym_dims = (dims * (dims + 1)) / 2; // 1x1: 1, 2x2: 3, 3x3: 6
const int ndata = (dim == 2) ? 1 : (symmetric ? sym_dims : dim*dim);
const int MQsymmDim = SMQ ? (SMQ->GetSize() * (SMQ->GetSize() + 1)) / 2 : 0;
const int MQfullDim = MQ ? (MQ->GetHeight() * MQ->GetWidth()) : 0;
const int MQdim = SMQ ? MQsymmDim : MQfullDim;
const int coeffDim = MQ ? MQdim : (DQ ? DQ->GetVDim() : 1);
symmetric = (SMQ || MQ == NULL);
const int symmDims = (dims * (dims + 1)) / 2; // 1x1: 1, 2x2: 3, 3x3: 6
const int ndata = (dim == 2) ? 1 : (symmetric ? symmDims : MQfullDim);
pa_data.SetSize(ndata * nq * ne, Device::GetMemoryType());
Vector coeff(coeffDim * ne * nq);
coeff = 1.0;
auto coeffh = Reshape(coeff.HostWrite(), coeffDim, nq, ne);
if (Q || DQ || MQ)
{
Vector DM(DQ ? coeffDim : 0);
DenseMatrix GM;
DenseSymmetricMatrix SM;
if (DQ)
{
MFEM_VERIFY(coeffDim == dimc, "");
}
if (SMQ)
{
SM.SetSize(dimc);
MFEM_VERIFY(SMQ->GetSize() == dimc, "");
}
else if (MQ)
{
GM.SetSize(dimc);
MFEM_VERIFY(coeffDim == MQdim, "");
MFEM_VERIFY(MQ->GetHeight() == dimc && MQ->GetWidth() == dimc, "");
}
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
if (SMQ)
{
SMQ->Eval(SM, *tr, ir->IntPoint(p));
int cnt = 0;
for (int i=0; i<dimc; ++i)
for (int j=i; j<dimc; ++j, ++cnt)
{
coeffh(cnt, p, e) = SM(i,j);
}
}
else if (MQ)
{
MQ->Eval(GM, *tr, ir->IntPoint(p));
for (int i=0; i<dimc; ++i)
for (int j=0; j<dimc; ++j)
{
coeffh(j+(i*dimc), p, e) = GM(i,j);
}
}
else if (DQ)
{
DQ->Eval(DM, *tr, ir->IntPoint(p));
for (int i=0; i<coeffDim; ++i)
{
coeffh(i, p, e) = DM[i];
}
}
else
{
coeffh(0, p, e) = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
}
if (el->GetDerivType() != mfem::FiniteElement::CURL)
{
MFEM_ABORT("Unknown kernel.");
@@ -997,7 +1067,7 @@ void CurlCurlIntegrator::AssemblePA(const FiniteElementSpace &fes)
if (dim == 3)
{
PACurlCurlSetup3D(quad1D, coeff_dim, ne, ir->GetWeights(), geom->J, coeff,
PACurlCurlSetup3D(quad1D, coeffDim, ne, ir->GetWeights(), geom->J, coeff,
pa_data);
}
else
@@ -2710,7 +2780,7 @@ void CurlCurlIntegrator::AssembleDiagonalPA(Vector& diag)
}
}
// Apply to x corresponding to DOFs in H^1 (trial), whose gradients are
// Apply to x corresponding to DOF's in H^1 (trial), whose gradients are
// integrated against H(curl) test functions corresponding to y.
void PAHcurlH1Apply3D(const int D1D,
const int Q1D,
@@ -2900,7 +2970,7 @@ void PAHcurlH1Apply3D(const int D1D,
}); // end of element loop
}
// Apply to x corresponding to DOFs in H(curl), integrated
// Apply to x corresponding to DOF's in H(curl), integrated
// against gradients of H^1 functions corresponding to y.
void PAHcurlH1ApplyTranspose3D(const int D1D,
const int Q1D,
@@ -3099,7 +3169,7 @@ void PAHcurlH1ApplyTranspose3D(const int D1D,
}); // end of element loop
}
// Apply to x corresponding to DOFs in H^1 (trial), whose gradients are
// Apply to x corresponding to DOF's in H^1 (trial), whose gradients are
// integrated against H(curl) test functions corresponding to y.
void PAHcurlH1Apply2D(const int D1D,
const int Q1D,
@@ -3223,7 +3293,7 @@ void PAHcurlH1Apply2D(const int D1D,
}); // end of element loop
}
// Apply to x corresponding to DOFs in H(curl), integrated
// Apply to x corresponding to DOF's in H(curl), integrated
// against gradients of H^1 functions corresponding to y.
void PAHcurlH1ApplyTranspose2D(const int D1D,
const int Q1D,
@@ -3419,8 +3489,20 @@ void MixedScalarCurlIntegrator::AssemblePA(const FiniteElementSpace &trial_fes,
pa_data.SetSize(nq * ne, Device::GetMemoryType());
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::FULL);
Vector coeff(ne * nq);
coeff = 1.0;
auto coeffh = Reshape(coeff.HostWrite(), nq, ne);
if (Q)
{
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
coeffh(p, e) = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
if (dim == 2)
{
@@ -3511,11 +3593,38 @@ void MixedVectorCurlIntegrator::AssemblePA(const FiniteElementSpace &trial_fes,
const int ndata = curlSpaces ? (coeffDim == 1 ? 1 : 9) : symmDims;
pa_data.SetSize(ndata * nq * ne, Device::GetMemoryType());
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(qs, CoefficientStorage::FULL);
if (Q) { coeff.Project(*Q); }
else if (DQ) { coeff.Project(*DQ); }
else { coeff.SetConstant(1.0); }
Vector coeff(coeffDim * nq * ne);
coeff = 1.0;
auto coeffh = Reshape(coeff.HostWrite(), coeffDim, nq, ne);
if (Q || DQ)
{
Vector V(coeffDim);
if (DQ)
{
MFEM_VERIFY(DQ->GetVDim() == coeffDim, "");
}
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
if (DQ)
{
DQ->Eval(V, *tr, ir->IntPoint(p));
for (int i=0; i<coeffDim; ++i)
{
coeffh(i, p, e) = V[i];
}
}
else
{
coeffh(0, p, e) = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
}
if (testType == mfem::FiniteElement::CURL &&
trialType == mfem::FiniteElement::CURL && dim == 3)
@@ -3543,7 +3652,7 @@ void MixedVectorCurlIntegrator::AssemblePA(const FiniteElementSpace &trial_fes,
}
}
// Apply to x corresponding to DOFs in H(curl) (trial), whose curl is
// Apply to x corresponding to DOF's in H(curl) (trial), whose curl is
// integrated against H(curl) test functions corresponding to y.
template<int MAX_D1D = HCURL_MAX_D1D, int MAX_Q1D = HCURL_MAX_Q1D>
static void PAHcurlL2Apply3D(const int D1D,
@@ -3906,7 +4015,7 @@ static void PAHcurlL2Apply3D(const int D1D,
}); // end of element loop
}
// Apply to x corresponding to DOFs in H(curl) (trial), whose curl is
// Apply to x corresponding to DOF's in H(curl) (trial), whose curl is
// integrated against H(curl) test functions corresponding to y.
template<int MAX_D1D = HCURL_MAX_D1D, int MAX_Q1D = HCURL_MAX_Q1D>
static void SmemPAHcurlL2Apply3D(const int D1D,
@@ -4216,7 +4325,7 @@ static void SmemPAHcurlL2Apply3D(const int D1D,
ForallWrap<3>(true, NE, device_kernel, host_kernel, Q1D, Q1D, Q1D);
}
// Apply to x corresponding to DOFs in H(curl) (trial), whose curl is
// Apply to x corresponding to DOF's in H(curl) (trial), whose curl is
// integrated against H(div) test functions corresponding to y.
template<int MAX_D1D = HCURL_MAX_D1D, int MAX_Q1D = HCURL_MAX_Q1D>
static void PAHcurlHdivApply3D(const int D1D,
@@ -4572,7 +4681,7 @@ static void PAHcurlHdivApply3D(const int D1D,
}); // end of element loop
}
// Apply to x corresponding to DOFs in H(div) (test), integrated against the
// Apply to x corresponding to DOF's in H(div) (test), integrated against the
// curl of H(curl) trial functions corresponding to y.
template<int MAX_D1D = HCURL_MAX_D1D, int MAX_Q1D = HCURL_MAX_Q1D>
static void PAHcurlHdivApply3DTranspose(const int D1D,
@@ -5037,11 +5146,38 @@ void MixedVectorWeakCurlIntegrator::AssemblePA(const FiniteElementSpace
pa_data.SetSize(ndata * nq * ne, Device::GetMemoryType());
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(qs, CoefficientStorage::FULL);
if (Q) { coeff.Project(*Q); }
else if (DQ) { coeff.Project(*DQ); }
else { coeff.SetConstant(1.0); }
Vector coeff(coeffDim * nq * ne);
coeff = 1.0;
auto coeffh = Reshape(coeff.HostWrite(), coeffDim, nq, ne);
if (Q || DQ)
{
Vector V(coeffDim);
if (DQ)
{
MFEM_VERIFY(DQ->GetVDim() == coeffDim, "");
}
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
if (DQ)
{
DQ->Eval(V, *tr, ir->IntPoint(p));
for (int i=0; i<coeffDim; ++i)
{
coeffh(i, p, e) = V[i];
}
}
else
{
coeffh(0, p, e) = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
}
if (trialType == mfem::FiniteElement::CURL && dim == 3)
{
@@ -5067,7 +5203,7 @@ void MixedVectorWeakCurlIntegrator::AssemblePA(const FiniteElementSpace
}
}
// Apply to x corresponding to DOFs in H(curl) (trial), integrated against curl
// Apply to x corresponding to DOF's in H(curl) (trial), integrated against curl
// of H(curl) test functions corresponding to y.
template<int MAX_D1D = HCURL_MAX_D1D, int MAX_Q1D = HCURL_MAX_Q1D>
static void PAHcurlL2Apply3DTranspose(const int D1D,
+28 -7
View File
@@ -12,7 +12,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qspace.hpp"
using namespace std;
@@ -1514,8 +1513,19 @@ void DivDivIntegrator::AssemblePA(const FiniteElementSpace &fes)
pa_data.SetSize(nq * ne, Device::GetMemoryType());
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::FULL);
Vector coeff(ne * nq);
coeff = 1.0;
if (Q)
{
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
coeff[p + (e * nq)] = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
if (el->GetDerivType() == mfem::FiniteElement::DIV && dim == 3)
{
@@ -1773,8 +1783,19 @@ VectorFEDivergenceIntegrator::AssemblePA(const FiniteElementSpace &trial_fes,
pa_data.SetSize(nq * ne, Device::GetMemoryType());
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::FULL);
Vector coeff(ne * nq);
coeff = 1.0;
if (Q)
{
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
coeff[p + (e * nq)] = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
if (test_el->GetMapType() == FiniteElement::INTEGRAL)
{
@@ -1797,7 +1818,7 @@ VectorFEDivergenceIntegrator::AssemblePA(const FiniteElementSpace &trial_fes,
}
}
// Apply to x corresponding to DOFs in H(div) (trial), whose divergence is
// Apply to x corresponding to DOF's in H(div) (trial), whose divergence is
// integrated against L_2 test functions corresponding to y.
static void PAHdivL2Apply3D(const int D1D,
const int Q1D,
@@ -1960,7 +1981,7 @@ static void PAHdivL2Apply3D(const int D1D,
}); // end of element loop
}
// Apply to x corresponding to DOFs in H(div) (trial), whose divergence is
// Apply to x corresponding to DOF's in H(div) (trial), whose divergence is
// integrated against L_2 test functions corresponding to y.
static void PAHdivL2Apply2D(const int D1D,
const int Q1D,
+544 -35
View File
@@ -12,9 +12,7 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qfunction.hpp"
#include "ceed/integrators/mass/mass.hpp"
#include "bilininteg_mass_pa.hpp"
using namespace std;
@@ -62,10 +60,43 @@ void MassIntegrator::AssemblePA(const FiniteElementSpace &fes)
dofs1D = maps->ndof;
quad1D = maps->nqpt;
pa_data.SetSize(ne*nq, mt);
Vector coeff;
if (Q == nullptr)
{
coeff.SetSize(1);
coeff(0) = 1.0;
}
else if (ConstantCoefficient* cQ = dynamic_cast<ConstantCoefficient*>(Q))
{
coeff.SetSize(1);
coeff(0) = cQ->constant;
}
else if (QuadratureFunctionCoefficient* qfQ =
dynamic_cast<QuadratureFunctionCoefficient*>(Q))
{
const QuadratureFunction &qFun = qfQ->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == nq * ne,
"Incompatible QuadratureFunction dimension \n");
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
coeff.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
coeff.SetSize(nq * ne);
auto C = Reshape(coeff.HostWrite(), nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation& T = *fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
C(q,e) = Q->Eval(T, ir->IntPoint(q));
}
}
}
if (dim==1) { MFEM_ABORT("Not supported yet... stay tuned!"); }
if (dim==2)
{
@@ -559,18 +590,85 @@ static void PAMassApply2D(const int NE,
const int d1d = 0,
const int q1d = 0)
{
MFEM_VERIFY(T_D1D ? T_D1D : d1d <= MAX_D1D, "");
MFEM_VERIFY(T_Q1D ? T_Q1D : q1d <= MAX_Q1D, "");
const auto B = b_.Read();
const auto Bt = bt_.Read();
const auto D = d_.Read();
const auto X = x_.Read();
auto Y = y_.ReadWrite();
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
MFEM_VERIFY(D1D <= MAX_D1D, "");
MFEM_VERIFY(Q1D <= MAX_Q1D, "");
auto B = Reshape(b_.Read(), Q1D, D1D);
auto Bt = Reshape(bt_.Read(), D1D, Q1D);
auto D = Reshape(d_.Read(), Q1D, Q1D, NE);
auto X = Reshape(x_.Read(), D1D, D1D, NE);
auto Y = Reshape(y_.ReadWrite(), D1D, D1D, NE);
MFEM_FORALL(e, NE,
{
internal::PAMassApply2D_Element(e, NE, B, Bt, D, X, Y, d1d, q1d);
const int D1D = T_D1D ? T_D1D : d1d; // nvcc workaround
const int Q1D = T_Q1D ? T_Q1D : q1d;
// the following variables are evaluated at compile time
constexpr int max_D1D = T_D1D ? T_D1D : MAX_D1D;
constexpr int max_Q1D = T_Q1D ? T_Q1D : MAX_Q1D;
double sol_xy[max_Q1D][max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] = 0.0;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
double sol_x[max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
sol_x[qy] = 0.0;
}
for (int dx = 0; dx < D1D; ++dx)
{
const double s = X(dx,dy,e);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] += B(qx,dx)* s;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
const double d2q = B(qy,dy);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] += d2q * sol_x[qx];
}
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] *= D(qx,qy,e);
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
double sol_x[max_D1D];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] = 0.0;
}
for (int qx = 0; qx < Q1D; ++qx)
{
const double s = sol_xy[qy][qx];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] += Bt(dx,qx) * s;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
const double q2d = Bt(dy,qy);
for (int dx = 0; dx < D1D; ++dx)
{
Y(dx,dy,e) += q2d * sol_x[dx];
}
}
}
});
}
@@ -592,13 +690,108 @@ static void SmemPAMassApply2D(const int NE,
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
MFEM_VERIFY(D1D <= MD1, "");
MFEM_VERIFY(Q1D <= MQ1, "");
const auto b = b_.Read();
const auto D = d_.Read();
const auto x = x_.Read();
auto Y = y_.ReadWrite();
auto b = Reshape(b_.Read(), Q1D, D1D);
auto D = Reshape(d_.Read(), Q1D, Q1D, NE);
auto x = Reshape(x_.Read(), D1D, D1D, NE);
auto Y = Reshape(y_.ReadWrite(), D1D, D1D, NE);
MFEM_FORALL_2D(e, NE, Q1D, Q1D, NBZ,
{
internal::SmemPAMassApply2D_Element<T_D1D,T_Q1D,T_NBZ>(e, NE, b, D, x, Y, d1d, q1d);
const int tidz = MFEM_THREAD_ID(z);
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int NBZ = T_NBZ ? T_NBZ : 1;
constexpr int MQ1 = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
constexpr int MDQ = (MQ1 > MD1) ? MQ1 : MD1;
MFEM_SHARED double BBt[MQ1*MD1];
double (*B)[MD1] = (double (*)[MD1]) BBt;
double (*Bt)[MQ1] = (double (*)[MQ1]) BBt;
MFEM_SHARED double sm0[NBZ][MDQ*MDQ];
MFEM_SHARED double sm1[NBZ][MDQ*MDQ];
double (*X)[MD1] = (double (*)[MD1]) (sm0 + tidz);
double (*DQ)[MQ1] = (double (*)[MQ1]) (sm1 + tidz);
double (*QQ)[MQ1] = (double (*)[MQ1]) (sm0 + tidz);
double (*QD)[MD1] = (double (*)[MD1]) (sm1 + tidz);
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
X[dy][dx] = x(dx,dy,e);
}
}
if (tidz == 0)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
B[q][dy] = b(q,dy);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double dq = 0.0;
for (int dx = 0; dx < D1D; ++dx)
{
dq += X[dy][dx] * B[qx][dx];
}
DQ[dy][qx] = dq;
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double qq = 0.0;
for (int dy = 0; dy < D1D; ++dy)
{
qq += DQ[dy][qx] * B[qy][dy];
}
QQ[qy][qx] = qq * D(qx, qy, e);
}
}
MFEM_SYNC_THREAD;
if (tidz == 0)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
Bt[dy][q] = b(q,dy);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double dq = 0.0;
for (int qx = 0; qx < Q1D; ++qx)
{
dq += QQ[qy][qx] * Bt[dx][qx];
}
QD[qy][dx] = dq;
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double dd = 0.0;
for (int qy = 0; qy < Q1D; ++qy)
{
dd += (QD[qy][dx] * Bt[dy][qy]);
}
Y(dx, dy, e) += dd;
}
}
});
}
@@ -612,18 +805,134 @@ static void PAMassApply3D(const int NE,
const int d1d = 0,
const int q1d = 0)
{
MFEM_VERIFY(T_D1D ? T_D1D : d1d <= MAX_D1D, "");
MFEM_VERIFY(T_Q1D ? T_Q1D : q1d <= MAX_Q1D, "");
const auto B = b_.Read();
const auto Bt = bt_.Read();
const auto D = d_.Read();
const auto X = x_.Read();
auto Y = y_.ReadWrite();
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
MFEM_VERIFY(D1D <= MAX_D1D, "");
MFEM_VERIFY(Q1D <= MAX_Q1D, "");
auto B = Reshape(b_.Read(), Q1D, D1D);
auto Bt = Reshape(bt_.Read(), D1D, Q1D);
auto D = Reshape(d_.Read(), Q1D, Q1D, Q1D, NE);
auto X = Reshape(x_.Read(), D1D, D1D, D1D, NE);
auto Y = Reshape(y_.ReadWrite(), D1D, D1D, D1D, NE);
MFEM_FORALL(e, NE,
{
internal::PAMassApply3D_Element(e, NE, B, Bt, D, X, Y, d1d, q1d);
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int max_D1D = T_D1D ? T_D1D : MAX_D1D;
constexpr int max_Q1D = T_Q1D ? T_Q1D : MAX_Q1D;
double sol_xyz[max_Q1D][max_Q1D][max_Q1D];
for (int qz = 0; qz < Q1D; ++qz)
{
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] = 0.0;
}
}
}
for (int dz = 0; dz < D1D; ++dz)
{
double sol_xy[max_Q1D][max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] = 0.0;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
double sol_x[max_Q1D];
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] = 0;
}
for (int dx = 0; dx < D1D; ++dx)
{
const double s = X(dx,dy,dz,e);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] += B(qx,dx) * s;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
const double wy = B(qy,dy);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] += wy * sol_x[qx];
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
const double wz = B(qz,dz);
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] += wz * sol_xy[qy][qx];
}
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] *= D(qx,qy,qz,e);
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
double sol_xy[max_D1D][max_D1D];
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
sol_xy[dy][dx] = 0;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
double sol_x[max_D1D];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] = 0;
}
for (int qx = 0; qx < Q1D; ++qx)
{
const double s = sol_xyz[qz][qy][qx];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] += Bt(dx,qx) * s;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
const double wy = Bt(dy,qy);
for (int dx = 0; dx < D1D; ++dx)
{
sol_xy[dy][dx] += wy * sol_x[dx];
}
}
}
for (int dz = 0; dz < D1D; ++dz)
{
const double wz = Bt(dz,qz);
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
Y(dx,dy,dz,e) += wz * sol_xy[dy][dx];
}
}
}
}
});
}
@@ -644,13 +953,213 @@ static void SmemPAMassApply3D(const int NE,
constexpr int M1D = T_D1D ? T_D1D : MAX_D1D;
MFEM_VERIFY(D1D <= M1D, "");
MFEM_VERIFY(Q1D <= M1Q, "");
auto b = b_.Read();
auto d = d_.Read();
auto x = x_.Read();
auto y = y_.ReadWrite();
auto b = Reshape(b_.Read(), Q1D, D1D);
auto d = Reshape(d_.Read(), Q1D, Q1D, Q1D, NE);
auto x = Reshape(x_.Read(), D1D, D1D, D1D, NE);
auto y = Reshape(y_.ReadWrite(), D1D, D1D, D1D, NE);
MFEM_FORALL_3D(e, NE, Q1D, Q1D, 1,
{
internal::SmemPAMassApply3D_Element<T_D1D,T_Q1D>(e, NE, b, d, x, y, d1d, q1d);
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int MQ1 = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
constexpr int MDQ = (MQ1 > MD1) ? MQ1 : MD1;
MFEM_SHARED double sDQ[MQ1*MD1];
double (*B)[MD1] = (double (*)[MD1]) sDQ;
double (*Bt)[MQ1] = (double (*)[MQ1]) sDQ;
MFEM_SHARED double sm0[MDQ*MDQ*MDQ];
MFEM_SHARED double sm1[MDQ*MDQ*MDQ];
double (*X)[MD1][MD1] = (double (*)[MD1][MD1]) sm0;
double (*DDQ)[MD1][MQ1] = (double (*)[MD1][MQ1]) sm1;
double (*DQQ)[MQ1][MQ1] = (double (*)[MQ1][MQ1]) sm0;
double (*QQQ)[MQ1][MQ1] = (double (*)[MQ1][MQ1]) sm1;
double (*QQD)[MQ1][MD1] = (double (*)[MQ1][MD1]) sm0;
double (*QDD)[MD1][MD1] = (double (*)[MD1][MD1]) sm1;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
X[dz][dy][dx] = x(dx,dy,dz,e);
}
}
MFEM_FOREACH_THREAD(dx,x,Q1D)
{
B[dx][dy] = b(dx,dy);
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u[D1D];
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
u[dz] = 0;
}
MFEM_UNROLL(MD1)
for (int dx = 0; dx < D1D; ++dx)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
u[dz] += X[dz][dy][dx] * B[qx][dx];
}
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
DDQ[dz][dy][qx] = u[dz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u[D1D];
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
u[dz] = 0;
}
MFEM_UNROLL(MD1)
for (int dy = 0; dy < D1D; ++dy)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
u[dz] += DDQ[dz][dy][qx] * B[qy][dy];
}
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
DQQ[dz][qy][qx] = u[dz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u[Q1D];
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; qz++)
{
u[qz] = 0;
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; qz++)
{
u[qz] += DQQ[dz][qy][qx] * B[qz][dz];
}
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; qz++)
{
QQQ[qz][qy][qx] = u[qz] * d(qx,qy,qz,e);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(d,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
Bt[d][q] = b(q,d);
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u[Q1D];
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] = 0;
}
MFEM_UNROLL(MQ1)
for (int qx = 0; qx < Q1D; ++qx)
{
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] += QQQ[qz][qy][qx] * Bt[dx][qx];
}
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
QQD[qz][qy][dx] = u[qz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u[Q1D];
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] = 0;
}
MFEM_UNROLL(MQ1)
for (int qy = 0; qy < Q1D; ++qy)
{
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] += QQD[qz][qy][dx] * Bt[dy][qy];
}
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
QDD[qz][dy][dx] = u[qz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u[D1D];
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
u[dz] = 0;
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
u[dz] += QDD[qz][dy][dx] * Bt[dz][qz];
}
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
y(dx,dy,dz,e) += u[dz];
}
}
}
});
}
-632
View File
@@ -1,632 +0,0 @@
// Copyright (c) 2010-2022, Lawrence Livermore National Security, LLC. Produced
// at the Lawrence Livermore National Laboratory. All Rights reserved. See files
// LICENSE and NOTICE for details. LLNL-CODE-806117.
//
// This file is part of the MFEM library. For more information and source code
// availability visit https://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the BSD-3 license. We welcome feedback and contributions, see file
// CONTRIBUTING.md for details.
#ifndef MFEM_BILININTEG_MASS_PA_HPP
#define MFEM_BILININTEG_MASS_PA_HPP
#include "../config/config.hpp"
#include "../general/forall.hpp"
#include "../linalg/dtensor.hpp"
namespace mfem
{
namespace internal
{
template <bool ACCUMULATE = true>
MFEM_HOST_DEVICE inline
void PAMassApply2D_Element(const int e,
const int NE,
const double *b_,
const double *bt_,
const double *d_,
const double *x_,
double *y_,
const int d1d = 0,
const int q1d = 0)
{
const int D1D = d1d;
const int Q1D = q1d;
auto B = ConstDeviceMatrix(b_, Q1D, D1D);
auto Bt = ConstDeviceMatrix(bt_, D1D, Q1D);
auto D = ConstDeviceCube(d_, Q1D, Q1D, NE);
auto X = ConstDeviceCube(x_, D1D, D1D, NE);
auto Y = DeviceCube(y_, D1D, D1D, NE);
if (!ACCUMULATE)
{
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
Y(dx, dy, e) = 0.0;
}
}
}
constexpr int max_D1D = MAX_D1D;
constexpr int max_Q1D = MAX_Q1D;
double sol_xy[max_Q1D][max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] = 0.0;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
double sol_x[max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
sol_x[qy] = 0.0;
}
for (int dx = 0; dx < D1D; ++dx)
{
const double s = X(dx,dy,e);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] += B(qx,dx)* s;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
const double d2q = B(qy,dy);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] += d2q * sol_x[qx];
}
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] *= D(qx,qy,e);
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
double sol_x[max_D1D];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] = 0.0;
}
for (int qx = 0; qx < Q1D; ++qx)
{
const double s = sol_xy[qy][qx];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] += Bt(dx,qx) * s;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
const double q2d = Bt(dy,qy);
for (int dx = 0; dx < D1D; ++dx)
{
Y(dx,dy,e) += q2d * sol_x[dx];
}
}
}
}
template<int T_D1D, int T_Q1D, int T_NBZ, bool ACCUMULATE = true>
MFEM_HOST_DEVICE inline
void SmemPAMassApply2D_Element(const int e,
const int NE,
const double *b_,
const double *d_,
const double *x_,
double *y_,
int d1d = 0,
int q1d = 0)
{
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int NBZ = T_NBZ ? T_NBZ : 1;
constexpr int MQ1 = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
constexpr int MDQ = (MQ1 > MD1) ? MQ1 : MD1;
auto b = ConstDeviceMatrix(b_, Q1D, D1D);
auto D = ConstDeviceCube(d_, Q1D, Q1D, NE);
auto x = ConstDeviceCube(x_, D1D, D1D, NE);
auto Y = DeviceCube(y_, D1D, D1D, NE);
const int tidz = MFEM_THREAD_ID(z);
MFEM_SHARED double BBt[MQ1*MD1];
double (*B)[MD1] = (double (*)[MD1]) BBt;
double (*Bt)[MQ1] = (double (*)[MQ1]) BBt;
MFEM_SHARED double sm0[NBZ][MDQ*MDQ];
MFEM_SHARED double sm1[NBZ][MDQ*MDQ];
double (*X)[MD1] = (double (*)[MD1]) (sm0 + tidz);
double (*DQ)[MQ1] = (double (*)[MQ1]) (sm1 + tidz);
double (*QQ)[MQ1] = (double (*)[MQ1]) (sm0 + tidz);
double (*QD)[MD1] = (double (*)[MD1]) (sm1 + tidz);
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
X[dy][dx] = x(dx,dy,e);
}
}
if (tidz == 0)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
B[q][dy] = b(q,dy);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double dq = 0.0;
for (int dx = 0; dx < D1D; ++dx)
{
dq += X[dy][dx] * B[qx][dx];
}
DQ[dy][qx] = dq;
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double qq = 0.0;
for (int dy = 0; dy < D1D; ++dy)
{
qq += DQ[dy][qx] * B[qy][dy];
}
QQ[qy][qx] = qq * D(qx, qy, e);
}
}
MFEM_SYNC_THREAD;
if (tidz == 0)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
Bt[dy][q] = b(q,dy);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double dq = 0.0;
for (int qx = 0; qx < Q1D; ++qx)
{
dq += QQ[qy][qx] * Bt[dx][qx];
}
QD[qy][dx] = dq;
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double dd = 0.0;
for (int qy = 0; qy < Q1D; ++qy)
{
dd += (QD[qy][dx] * Bt[dy][qy]);
}
if (ACCUMULATE)
{
Y(dx, dy, e) += dd;
}
else
{
Y(dx, dy, e) = dd;
}
}
}
}
template <bool ACCUMULATE = true>
MFEM_HOST_DEVICE inline
void PAMassApply3D_Element(const int e,
const int NE,
const double *b_,
const double *bt_,
const double *d_,
const double *x_,
double *y_,
const int d1d,
const int q1d)
{
const int D1D = d1d;
const int Q1D = q1d;
auto B = ConstDeviceMatrix(b_, Q1D, D1D);
auto Bt = ConstDeviceMatrix(bt_, D1D, Q1D);
auto D = DeviceTensor<4,const double>(d_, Q1D, Q1D, Q1D, NE);
auto X = DeviceTensor<4,const double>(x_, D1D, D1D, D1D, NE);
auto Y = DeviceTensor<4,double>(y_, D1D, D1D, D1D, NE);
if (!ACCUMULATE)
{
for (int dz = 0; dz < D1D; ++dz)
{
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
Y(dx, dy, dz, e) = 0.0;
}
}
}
}
constexpr int max_D1D = MAX_D1D;
constexpr int max_Q1D = MAX_Q1D;
double sol_xyz[max_Q1D][max_Q1D][max_Q1D];
for (int qz = 0; qz < Q1D; ++qz)
{
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] = 0.0;
}
}
}
for (int dz = 0; dz < D1D; ++dz)
{
double sol_xy[max_Q1D][max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] = 0.0;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
double sol_x[max_Q1D];
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] = 0;
}
for (int dx = 0; dx < D1D; ++dx)
{
const double s = X(dx,dy,dz,e);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] += B(qx,dx) * s;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
const double wy = B(qy,dy);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] += wy * sol_x[qx];
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
const double wz = B(qz,dz);
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] += wz * sol_xy[qy][qx];
}
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] *= D(qx,qy,qz,e);
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
double sol_xy[max_D1D][max_D1D];
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
sol_xy[dy][dx] = 0;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
double sol_x[max_D1D];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] = 0;
}
for (int qx = 0; qx < Q1D; ++qx)
{
const double s = sol_xyz[qz][qy][qx];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] += Bt(dx,qx) * s;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
const double wy = Bt(dy,qy);
for (int dx = 0; dx < D1D; ++dx)
{
sol_xy[dy][dx] += wy * sol_x[dx];
}
}
}
for (int dz = 0; dz < D1D; ++dz)
{
const double wz = Bt(dz,qz);
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
Y(dx,dy,dz,e) += wz * sol_xy[dy][dx];
}
}
}
}
}
template<int T_D1D, int T_Q1D, bool ACCUMULATE = true>
MFEM_HOST_DEVICE inline
void SmemPAMassApply3D_Element(const int e,
const int NE,
const double *b_,
const double *d_,
const double *x_,
double *y_,
const int d1d = 0,
const int q1d = 0)
{
constexpr int D1D = T_D1D ? T_D1D : d1d;
constexpr int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int MQ1 = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
constexpr int MDQ = (MQ1 > MD1) ? MQ1 : MD1;
auto b = ConstDeviceMatrix(b_, Q1D, D1D);
auto d = DeviceTensor<4,const double>(d_, Q1D, Q1D, Q1D, NE);
auto x = DeviceTensor<4,const double>(x_, D1D, D1D, D1D, NE);
auto y = DeviceTensor<4,double>(y_, D1D, D1D, D1D, NE);
MFEM_SHARED double sDQ[MQ1*MD1];
double (*B)[MD1] = (double (*)[MD1]) sDQ;
double (*Bt)[MQ1] = (double (*)[MQ1]) sDQ;
MFEM_SHARED double sm0[MDQ*MDQ*MDQ];
MFEM_SHARED double sm1[MDQ*MDQ*MDQ];
double (*X)[MD1][MD1] = (double (*)[MD1][MD1]) sm0;
double (*DDQ)[MD1][MQ1] = (double (*)[MD1][MQ1]) sm1;
double (*DQQ)[MQ1][MQ1] = (double (*)[MQ1][MQ1]) sm0;
double (*QQQ)[MQ1][MQ1] = (double (*)[MQ1][MQ1]) sm1;
double (*QQD)[MQ1][MD1] = (double (*)[MQ1][MD1]) sm0;
double (*QDD)[MD1][MD1] = (double (*)[MD1][MD1]) sm1;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
X[dz][dy][dx] = x(dx,dy,dz,e);
}
}
MFEM_FOREACH_THREAD(dx,x,Q1D)
{
B[dx][dy] = b(dx,dy);
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u[D1D];
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
u[dz] = 0;
}
MFEM_UNROLL(MD1)
for (int dx = 0; dx < D1D; ++dx)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
u[dz] += X[dz][dy][dx] * B[qx][dx];
}
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
DDQ[dz][dy][qx] = u[dz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u[D1D];
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
u[dz] = 0;
}
MFEM_UNROLL(MD1)
for (int dy = 0; dy < D1D; ++dy)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
u[dz] += DDQ[dz][dy][qx] * B[qy][dy];
}
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; dz++)
{
DQQ[dz][qy][qx] = u[dz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u[Q1D];
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; qz++)
{
u[qz] = 0;
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; qz++)
{
u[qz] += DQQ[dz][qy][qx] * B[qz][dz];
}
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; qz++)
{
QQQ[qz][qy][qx] = u[qz] * d(qx,qy,qz,e);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(di,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
Bt[di][q] = b(q,di);
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u[Q1D];
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] = 0;
}
MFEM_UNROLL(MQ1)
for (int qx = 0; qx < Q1D; ++qx)
{
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] += QQQ[qz][qy][qx] * Bt[dx][qx];
}
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
QQD[qz][qy][dx] = u[qz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u[Q1D];
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] = 0;
}
MFEM_UNROLL(MQ1)
for (int qy = 0; qy < Q1D; ++qy)
{
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
u[qz] += QQD[qz][qy][dx] * Bt[dy][qy];
}
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
QDD[qz][dy][dx] = u[qz];
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u[D1D];
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
u[dz] = 0;
}
MFEM_UNROLL(MQ1)
for (int qz = 0; qz < Q1D; ++qz)
{
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
u[dz] += QDD[qz][dy][dx] * Bt[dz][qz];
}
}
MFEM_UNROLL(MD1)
for (int dz = 0; dz < D1D; ++dz)
{
if (ACCUMULATE)
{
y(dx,dy,dz,e) += u[dz];
}
else
{
y(dx,dy,dz,e) = u[dz];
}
}
}
}
MFEM_SYNC_THREAD;
}
} // namespace internal
} // namespace mfem
#endif
+36 -3
View File
@@ -12,7 +12,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
#include "qfunction.hpp"
#include "ceed/integrators/diffusion/diffusion.hpp"
using namespace std;
@@ -176,9 +175,43 @@ void VectorDiffusionIntegrator::AssemblePA(const FiniteElementSpace &fes)
MFEM_VERIFY(!VQ && !MQ,
"Only scalar coefficient supported for partial assembly for VectorDiffusionIntegrator");
Vector coeff;
if (Q == nullptr)
{
coeff.SetSize(1);
coeff(0) = 1.0;
}
else if (ConstantCoefficient* cQ = dynamic_cast<ConstantCoefficient*>(Q))
{
coeff.SetSize(1);
coeff(0) = cQ->constant;
}
else if (QuadratureFunctionCoefficient* qfQ =
dynamic_cast<QuadratureFunctionCoefficient*>(Q))
{
const QuadratureFunction &qFun = qfQ->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == ne*nq,
"Incompatible QuadratureFunction dimension \n");
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
coeff.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
coeff.SetSize(nq * ne);
auto Co = Reshape(coeff.HostWrite(), nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation& T = *fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
Co(q,e) = Q->Eval(T, ir->IntPoint(q));
}
}
}
const Array<double> &w = ir->GetWeights();
const Vector &j = geom->J;
+110 -23
View File
@@ -11,7 +11,6 @@
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "qspace.hpp"
#include "gridfunc.hpp"
namespace mfem
@@ -794,63 +793,140 @@ void VectorFEMassIntegrator::AssemblePA(const FiniteElementSpace &trial_fes,
trial_fetype = trial_el->GetDerivType();
test_fetype = test_el->GetDerivType();
auto SMQ = dynamic_cast<SymmetricMatrixCoefficient *>(MQ);
const int MQsymmDim = SMQ ? (SMQ->GetSize() * (SMQ->GetSize() + 1)) / 2 : 0;
const int MQfullDim = MQ ? (MQ->GetHeight() * MQ->GetWidth()) : 0;
const int MQdim = SMQ ? MQsymmDim : MQfullDim;
const int coeffDim = MQ ? MQdim : (DQ ? DQ->GetVDim() : 1);
symmetric = (SMQ || MQ == NULL);
const bool trial_curl = (trial_fetype == mfem::FiniteElement::CURL);
const bool trial_div = (trial_fetype == mfem::FiniteElement::DIV);
const bool test_curl = (test_fetype == mfem::FiniteElement::CURL);
const bool test_div = (test_fetype == mfem::FiniteElement::DIV);
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(qs, CoefficientStorage::SYMMETRIC);
if (Q) { coeff.Project(*Q); }
else if (MQ) { coeff.ProjectTranspose(*MQ); }
else if (DQ) { coeff.Project(*DQ); }
else { coeff.SetConstant(1.0); }
const int coeff_dim = coeff.GetVDim();
symmetric = (coeff_dim != dim*dim);
if ((trial_curl && test_div) || (trial_div && test_curl))
pa_data.SetSize((coeff_dim == 1 ? 1 : dim*dim) * nq * ne,
pa_data.SetSize((coeffDim == 1 ? 1 : dim*dim) * nq * ne,
Device::GetMemoryType());
else
pa_data.SetSize((symmetric ? symmDims : dims*dims) * nq * ne,
pa_data.SetSize((symmetric ? symmDims : MQfullDim) * nq * ne,
Device::GetMemoryType());
Vector coeff;
auto *qf_c = dynamic_cast<QuadratureFunctionCoefficient*>(Q);
if (qf_c)
{
const QuadratureFunction &qf = qf_c->GetQuadFunction();
qf.Read();
coeff.MakeRef(const_cast<QuadratureFunction&>(qf), 0);
}
else
{
coeff.SetSize(coeffDim * ne * nq);
coeff = 1.0;
auto coeffh = Reshape(coeff.HostWrite(), coeffDim, nq, ne);
if (Q || DQ || MQ)
{
Vector DM(DQ ? coeffDim : 0);
DenseMatrix M;
DenseSymmetricMatrix SM;
if (DQ)
{
MFEM_VERIFY(coeffDim == dim, "");
}
if (SMQ)
{
MFEM_VERIFY(SMQ->GetSize() == dim, "");
SM.SetSize(dim);
}
else if (MQ)
{
MFEM_VERIFY(coeffDim == MQdim, "");
MFEM_VERIFY(MQ->GetHeight() == dim && MQ->GetWidth() == dim, "");
M.SetSize(dim);
}
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
if (SMQ)
{
SMQ->Eval(SM, *tr, ir->IntPoint(p));
int cnt = 0;
for (int i=0; i<dim; ++i)
for (int j=i; j<dim; ++j, ++cnt)
{
coeffh(cnt, p, e) = SM(i,j);
}
}
else if (MQ)
{
MQ->Eval(M, *tr, ir->IntPoint(p));
for (int i=0; i<dim; ++i)
for (int j=0; j<dim; ++j)
{
coeffh(j+(i*dim), p, e) = M(i,j);
}
}
else if (DQ)
{
DQ->Eval(DM, *tr, ir->IntPoint(p));
for (int i=0; i<coeffDim; ++i)
{
coeffh(i, p, e) = DM[i];
}
}
else
{
coeffh(0, p, e) = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
}
}
if (trial_curl && test_curl && dim == 3)
{
PADiffusionSetup3D(quad1D, coeff_dim, ne, ir->GetWeights(), geom->J,
PADiffusionSetup3D(quad1D, coeffDim, ne, ir->GetWeights(), geom->J,
coeff, pa_data);
}
else if (trial_curl && test_curl && dim == 2)
{
PADiffusionSetup2D<2>(quad1D, coeff_dim, ne, ir->GetWeights(), geom->J,
PADiffusionSetup2D<2>(quad1D, coeffDim, ne, ir->GetWeights(), geom->J,
coeff, pa_data);
}
else if (trial_div && test_div && dim == 3)
{
PAHdivSetup3D(quad1D, coeff_dim, ne, ir->GetWeights(), geom->J,
PAHdivSetup3D(quad1D, coeffDim, ne, ir->GetWeights(), geom->J,
coeff, pa_data);
}
else if (trial_div && test_div && dim == 2)
{
PAHdivSetup2D(quad1D, coeff_dim, ne, ir->GetWeights(), geom->J,
PAHdivSetup2D(quad1D, coeffDim, ne, ir->GetWeights(), geom->J,
coeff, pa_data);
}
else if (((trial_curl && test_div) || (trial_div && test_curl)) &&
test_fel->GetOrder() == trial_fel->GetOrder())
{
if (coeff_dim == 1)
if (coeffDim == 1)
{
PAHcurlL2Setup(nq, coeff_dim, ne, ir->GetWeights(), coeff, pa_data);
PAHcurlL2Setup(nq, coeffDim, ne, ir->GetWeights(), coeff, pa_data);
}
else
{
const bool tr = (trial_div && test_curl);
if (dim == 3)
PAHcurlHdivSetup3D(quad1D, coeff_dim, ne, tr, ir->GetWeights(),
PAHcurlHdivSetup3D(quad1D, coeffDim, ne, tr, ir->GetWeights(),
geom->J, coeff, pa_data);
else
PAHcurlHdivSetup2D(quad1D, coeff_dim, ne, tr, ir->GetWeights(),
PAHcurlHdivSetup2D(quad1D, coeffDim, ne, tr, ir->GetWeights(),
geom->J, coeff, pa_data);
}
}
@@ -1092,8 +1168,19 @@ void MixedVectorGradientIntegrator::AssemblePA(const FiniteElementSpace
pa_data.SetSize(symmDims * nq * ne, Device::GetMemoryType());
QuadratureSpace qs(*mesh, *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::FULL);
Vector coeff(ne * nq);
coeff = 1.0;
if (Q)
{
for (int e=0; e<ne; ++e)
{
ElementTransformation *tr = mesh->GetElementTransformation(e);
for (int p=0; p<nq; ++p)
{
coeff[p + (e * nq)] = Q->Eval(*tr, ir->IntPoint(p));
}
}
}
// Use the same setup functions as VectorFEMassIntegrator.
if (test_el->GetDerivType() == mfem::FiniteElement::CURL && dim == 3)
+1 -1
View File
@@ -112,7 +112,7 @@ static void InitBasisImpl(const FiniteElementSpace &fes,
const bool tensor = dynamic_cast<const mfem::TensorBasisElement *>
(&fe) != nullptr;
// Init or retrieve key values
// Init or retreive key values
if (basis_itr == mfem::internal::ceed_basis_map.end())
{
if ( tensor )
+4 -5
View File
@@ -20,7 +20,6 @@
#include "../../../linalg/dtensor.hpp"
#include "../../../mesh/mesh.hpp"
#include "../../gridfunc.hpp"
#include "../../qfunction.hpp"
#include "util.hpp"
#include "ceed.hpp"
@@ -122,7 +121,7 @@ void InitCoefficient(mfem::Coefficient *Q, mfem::Mesh &mesh,
MFEM_VERIFY(qFun.Size() == nq * ne,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetIntRule(0),
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
@@ -196,7 +195,7 @@ void InitCoefficient(mfem::VectorCoefficient *VQ, mfem::Mesh &mesh,
MFEM_VERIFY(qFun.Size() == dim * nq * ne,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetIntRule(0),
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
@@ -280,7 +279,7 @@ void InitCoefficientWithIndices(mfem::Coefficient *Q, mfem::Mesh &mesh,
MFEM_VERIFY(qFun.Size() == nq * ne,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetIntRule(0),
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
ceedCoeff->coeff.SetSize(nq * nelem);
@@ -370,7 +369,7 @@ void InitCoefficientWithIndices(mfem::VectorCoefficient *VQ, mfem::Mesh &mesh,
MFEM_VERIFY(qFun.Size() == dim * nq * ne,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetIntRule(0),
MFEM_VERIFY(&ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
ceedCoeff->coeff.SetSize(dim * nq * nelem);
+3 -3
View File
@@ -232,7 +232,7 @@ void InitRestriction(const FiniteElementSpace &fes,
RestrKey restr_key(&fes, nelem, P, ncomp, restr_type::Standard);
auto restr_itr = mfem::internal::ceed_restr_map.find(restr_key);
// Init or retrieve key values
// Init or retreive key values
if (restr_itr == mfem::internal::ceed_restr_map.end())
{
InitRestrictionImpl(fes, ceed, restr);
@@ -257,7 +257,7 @@ void InitRestrictionWithIndices(const FiniteElementSpace &fes,
RestrKey restr_key(&fes, nelem, P, ncomp, restr_type::Standard);
auto restr_itr = mfem::internal::ceed_restr_map.find(restr_key);
// Init or retrieve key values
// Init or retreive key values
if (restr_itr == mfem::internal::ceed_restr_map.end())
{
InitRestrictionWithIndicesImpl(fes, nelem, indices, ceed, restr);
@@ -281,7 +281,7 @@ void InitCoeffRestrictionWithIndices(const FiniteElementSpace &fes,
RestrKey restr_key(&fes, nelem, nquads, ncomp, restr_type::Coeff);
auto restr_itr = mfem::internal::ceed_restr_map.find(restr_key);
// Init or retrieve key values
// Init or retreive key values
if (restr_itr == mfem::internal::ceed_restr_map.end())
{
InitCoeffRestrictionWithIndicesImpl(fes, nelem, indices, nquads, ncomp,
+10 -9
View File
@@ -676,6 +676,7 @@ AlgebraicSpaceHierarchy::AlgebraicSpaceHierarchy(FiniteElementSpace &fes)
const SparseMatrix *R = fespaces[ilevel+1]->GetRestrictionMatrix();
if (R)
{
R->EnsureMultTranspose();
R_tr[ilevel] = new TransposeOperator(*R);
}
else
@@ -744,7 +745,7 @@ ParAlgebraicCoarseSpace::ParAlgebraicCoarseSpace(
ldof_group.SetSize(lsize);
ldof_group = 0;
const GroupTopology &group_topo = gc_fine->GetGroupTopology();
GroupTopology &group_topo = gc_fine->GetGroupTopology();
gc = new GroupCommunicator(group_topo);
Table &group_ldof = gc->GroupLDofTable();
group_ldof.MakeI(group_ldof_fine.Size());
@@ -821,11 +822,11 @@ HypreParMatrix *ParAlgebraicCoarseSpace::GetProlongationHypreParMatrix()
ParMesh *pmesh = dynamic_cast<ParMesh*>(mesh);
MFEM_VERIFY(pmesh != NULL, "");
Array<HYPRE_BigInt> dof_offsets, tdof_offsets, tdof_nb_offsets;
Array<HYPRE_BigInt> *offsets[2] = {&dof_offsets, &tdof_offsets};
Array<HYPRE_Int> dof_offsets, tdof_offsets, tdof_nb_offsets;
Array<HYPRE_Int> *offsets[2] = {&dof_offsets, &tdof_offsets};
int lsize = P->Height();
int ltsize = P->Width();
HYPRE_BigInt loc_sizes[2] = {lsize, ltsize};
HYPRE_Int loc_sizes[2] = {lsize, ltsize};
pmesh->GenerateOffsets(2, loc_sizes, offsets);
MPI_Comm comm = pmesh->GetComm();
@@ -869,12 +870,12 @@ HypreParMatrix *ParAlgebraicCoarseSpace::GetProlongationHypreParMatrix()
HYPRE_Int *j_offd = Memory<HYPRE_Int>(lsize-ltsize);
int offd_counter;
HYPRE_BigInt *cmap = Memory<HYPRE_BigInt>(lsize-ltsize);
HYPRE_Int *cmap = Memory<HYPRE_Int>(lsize-ltsize);
HYPRE_BigInt *col_starts = tdof_offsets;
HYPRE_BigInt *row_starts = dof_offsets;
HYPRE_Int *col_starts = tdof_offsets;
HYPRE_Int *row_starts = dof_offsets;
Array<Pair<HYPRE_BigInt, int> > cmap_j_offd(lsize-ltsize);
Array<Pair<HYPRE_Int, int> > cmap_j_offd(lsize-ltsize);
i_diag[0] = i_offd[0] = 0;
diag_counter = offd_counter = 0;
@@ -908,7 +909,7 @@ HypreParMatrix *ParAlgebraicCoarseSpace::GetProlongationHypreParMatrix()
i_offd[i_ldof+1] = offd_counter;
}
SortPairs<HYPRE_BigInt, int>(cmap_j_offd, offd_counter);
SortPairs<HYPRE_Int, int>(cmap_j_offd, offd_counter);
for (int i = 0; i < offd_counter; i++)
{
+3 -316
View File
@@ -48,31 +48,6 @@ ElementTransformation *RefinedToCoarse(
return coarse_T;
}
void Coefficient::Project(QuadratureFunction &qf)
{
QuadratureSpaceBase &qspace = *qf.GetSpace();
const int ne = qspace.GetNE();
Vector values;
for (int iel = 0; iel < ne; ++iel)
{
qf.GetValues(iel, values);
const IntegrationRule &ir = qspace.GetIntRule(iel);
ElementTransformation& T = *qspace.GetTransformation(iel);
for (int iq = 0; iq < ir.Size(); ++iq)
{
const IntegrationPoint &ip = ir[iq];
T.SetIntPoint(&ip);
const int iq_p = qspace.GetPermutedIndex(iel, iq);
values[iq_p] = Eval(T, ip);
}
}
}
void ConstantCoefficient::Project(QuadratureFunction &qf)
{
qf = constant;
}
double PWConstCoefficient::Eval(ElementTransformation & T,
const IntegrationPoint & ip)
{
@@ -160,11 +135,6 @@ double GridFunctionCoefficient::Eval (ElementTransformation &T,
}
}
void GridFunctionCoefficient::Project(QuadratureFunction &qf)
{
qf.ProjectGridFunction(*GridF);
}
void TransformedCoefficient::SetTime(double t)
{
if (Q1) { Q1->SetTime(t); }
@@ -233,29 +203,6 @@ void VectorCoefficient::Eval(DenseMatrix &M, ElementTransformation &T,
}
}
void VectorCoefficient::Project(QuadratureFunction &qf)
{
MFEM_VERIFY(vdim == qf.GetVDim(), "Wrong sizes.");
QuadratureSpaceBase &qspace = *qf.GetSpace();
const int ne = qspace.GetNE();
DenseMatrix values;
Vector col;
for (int iel = 0; iel < ne; ++iel)
{
qf.GetValues(iel, values);
const IntegrationRule &ir = qspace.GetIntRule(iel);
ElementTransformation& T = *qspace.GetTransformation(iel);
for (int iq = 0; iq < ir.Size(); ++iq)
{
const IntegrationPoint &ip = ir[iq];
T.SetIntPoint(&ip);
const int iq_p = qspace.GetPermutedIndex(iel, iq);
values.GetColumnReference(iq_p, col);
Eval(col, T, ip);
}
}
}
void PWVectorCoefficient::InitMap(const Array<int> & attr,
const Array<VectorCoefficient*> & coefs)
{
@@ -421,11 +368,6 @@ void VectorGridFunctionCoefficient::Eval(
}
}
void VectorGridFunctionCoefficient::Project(QuadratureFunction &qf)
{
qf.ProjectGridFunction(*GridFunc);
}
GradientGridFunctionCoefficient::GradientGridFunctionCoefficient (
const GridFunction *gf)
: VectorCoefficient((gf) ?
@@ -575,29 +517,6 @@ void VectorRestrictedCoefficient::Eval(
}
}
void MatrixCoefficient::Project(QuadratureFunction &qf, bool transpose)
{
MFEM_VERIFY(qf.GetVDim() == height*width, "Wrong sizes.");
QuadratureSpaceBase &qspace = *qf.GetSpace();
const int ne = qspace.GetNE();
DenseMatrix values, matrix;
for (int iel = 0; iel < ne; ++iel)
{
qf.GetValues(iel, values);
const IntegrationRule &ir = qspace.GetIntRule(iel);
ElementTransformation& T = *qspace.GetTransformation(iel);
for (int iq = 0; iq < ir.Size(); ++iq)
{
const IntegrationPoint &ip = ir[iq];
T.SetIntPoint(&ip);
const int iq_p = qspace.GetPermutedIndex(iel, iq);
matrix.UseExternalData(&values(0, iq_p), height, width);
Eval(matrix, T, ip);
if (transpose) { matrix.Transpose(); }
}
}
}
void PWMatrixCoefficient::InitMap(const Array<int> & attr,
const Array<MatrixCoefficient*> & coefs)
{
@@ -750,31 +669,6 @@ void MatrixFunctionCoefficient::EvalSymmetric(Vector &K,
}
}
void SymmetricMatrixCoefficient::ProjectSymmetric(QuadratureFunction &qf)
{
const int vdim = qf.GetVDim();
MFEM_VERIFY(vdim == height*(height+1)/2, "Wrong sizes.");
QuadratureSpaceBase &qspace = *qf.GetSpace();
const int ne = qspace.GetNE();
DenseMatrix values;
DenseSymmetricMatrix matrix;
for (int iel = 0; iel < ne; ++iel)
{
qf.GetValues(iel, values);
const IntegrationRule &ir = qspace.GetIntRule(iel);
ElementTransformation& T = *qspace.GetTransformation(iel);
for (int iq = 0; iq < ir.Size(); ++iq)
{
const IntegrationPoint &ip = ir[iq];
T.SetIntPoint(&ip);
matrix.UseExternalData(&values(0, iq), vdim);
Eval(matrix, T, ip);
}
}
}
void SymmetricMatrixCoefficient::Eval(DenseMatrix &K, ElementTransformation &T,
const IntegrationPoint &ip)
{
@@ -1543,12 +1437,12 @@ void VectorQuadratureFunctionCoefficient::Eval(Vector &V,
if (index == 0 && vdim == QuadF.GetVDim())
{
QuadF.GetValues(T.ElementNo, ip.index, V);
QuadF.GetElementValues(T.ElementNo, ip.index, V);
}
else
{
Vector temp;
QuadF.GetValues(T.ElementNo, ip.index, temp);
QuadF.GetElementValues(T.ElementNo, ip.index, temp);
V.SetSize(vdim);
for (int i = 0; i < vdim; i++)
{
@@ -1559,11 +1453,6 @@ void VectorQuadratureFunctionCoefficient::Eval(Vector &V,
return;
}
void VectorQuadratureFunctionCoefficient::Project(QuadratureFunction &qf)
{
qf = QuadF;
}
QuadratureFunctionCoefficient::QuadratureFunctionCoefficient(
QuadratureFunction &qf) : QuadF(qf)
{
@@ -1575,210 +1464,8 @@ double QuadratureFunctionCoefficient::Eval(ElementTransformation &T,
{
QuadF.HostRead();
Vector temp(1);
QuadF.GetValues(T.ElementNo, ip.index, temp);
QuadF.GetElementValues(T.ElementNo, ip.index, temp);
return temp[0];
}
void QuadratureFunctionCoefficient::Project(QuadratureFunction &qf)
{
qf = QuadF;
}
CoefficientVector::CoefficientVector(
QuadratureSpaceBase &qs_, CoefficientStorage storage_)
: Vector(), storage(storage_), vdim(0), qs(qs_), qf(NULL)
{
UseDevice(true);
}
CoefficientVector::CoefficientVector(Coefficient *coeff,
QuadratureSpaceBase &qs_,
CoefficientStorage storage_)
: CoefficientVector(qs_, storage_)
{
if (coeff == NULL)
{
SetConstant(1.0);
}
else
{
Project(*coeff);
}
}
CoefficientVector::CoefficientVector(Coefficient &coeff,
QuadratureSpaceBase &qs_,
CoefficientStorage storage_)
: CoefficientVector(qs_, storage_)
{
Project(coeff);
}
CoefficientVector::CoefficientVector(VectorCoefficient &coeff,
QuadratureSpaceBase &qs_,
CoefficientStorage storage_)
: CoefficientVector(qs_, storage_)
{
Project(coeff);
}
CoefficientVector::CoefficientVector(MatrixCoefficient &coeff,
QuadratureSpaceBase &qs_,
CoefficientStorage storage_)
: CoefficientVector(qs_, storage_)
{
Project(coeff);
}
void CoefficientVector::Project(Coefficient &coeff)
{
vdim = 1;
if (auto *const_coeff = dynamic_cast<ConstantCoefficient*>(&coeff))
{
SetConstant(const_coeff->constant);
}
else if (auto *qf_coeff = dynamic_cast<QuadratureFunctionCoefficient*>(&coeff))
{
MakeRef(qf_coeff->GetQuadFunction());
}
else
{
if (qf == nullptr) { qf = new QuadratureFunction(qs); }
qf->SetVDim(1);
coeff.Project(*qf);
Vector::MakeRef(*qf, 0, qf->Size());
}
}
void CoefficientVector::Project(VectorCoefficient &coeff)
{
vdim = coeff.GetVDim();
if (auto *const_coeff = dynamic_cast<VectorConstantCoefficient*>(&coeff))
{
SetConstant(const_coeff->GetVec());
}
else if (auto *qf_coeff =
dynamic_cast<VectorQuadratureFunctionCoefficient*>(&coeff))
{
MakeRef(qf_coeff->GetQuadFunction());
}
else
{
if (qf == nullptr) { qf = new QuadratureFunction(qs, vdim); }
qf->SetVDim(vdim);
coeff.Project(*qf);
Vector::MakeRef(*qf, 0, qf->Size());
}
}
void CoefficientVector::Project(MatrixCoefficient &coeff, bool transpose)
{
if (auto *const_coeff = dynamic_cast<MatrixConstantCoefficient*>(&coeff))
{
SetConstant(const_coeff->GetMatrix());
}
else if (auto *const_sym_coeff =
dynamic_cast<SymmetricMatrixConstantCoefficient*>(&coeff))
{
SetConstant(const_sym_coeff->GetMatrix());
}
else
{
auto *sym_coeff = dynamic_cast<SymmetricMatrixCoefficient*>(&coeff);
const bool sym = sym_coeff && (storage & CoefficientStorage::SYMMETRIC);
const int height = coeff.GetHeight();
const int width = coeff.GetWidth();
vdim = sym ? height*(height + 1)/2 : width*height;
if (qf == nullptr) { qf = new QuadratureFunction(qs, vdim); }
qf->SetVDim(vdim);
if (sym) { sym_coeff->ProjectSymmetric(*qf); }
else { coeff.Project(*qf, transpose); }
Vector::MakeRef(*qf, 0, qf->Size());
}
}
void CoefficientVector::ProjectTranspose(MatrixCoefficient &coeff)
{
Project(coeff, true);
}
void CoefficientVector::MakeRef(const QuadratureFunction &qf_)
{
vdim = qf_.GetVDim();
const QuadratureSpaceBase *qs2 = qf_.GetSpace();
MFEM_CONTRACT_VAR(qs2); // qs2 used only for asserts
MFEM_VERIFY(qs2 != NULL, "Invalid QuadratureSpace.")
MFEM_VERIFY(qs2->GetMesh() == qs.GetMesh(), "Meshes differ.");
MFEM_VERIFY(qs2->GetOrder() == qs.GetOrder(), "Orders differ.");
Vector::MakeRef(const_cast<QuadratureFunction&>(qf_), 0, qf_.Size());
}
void CoefficientVector::SetConstant(double constant)
{
const int nq = (storage & CoefficientStorage::CONSTANTS) ? 1 : qs.GetSize();
vdim = 1;
SetSize(nq);
Vector::operator=(constant);
}
void CoefficientVector::SetConstant(const Vector &constant)
{
const int nq = (storage & CoefficientStorage::CONSTANTS) ? 1 : qs.GetSize();
vdim = constant.Size();
SetSize(nq*vdim);
for (int iq = 0; iq < nq; ++iq)
{
for (int vd = 0; vd<vdim; ++vd)
{
(*this)[vd + iq*vdim] = constant[vd];
}
}
}
void CoefficientVector::SetConstant(const DenseMatrix &constant)
{
const int nq = (storage & CoefficientStorage::CONSTANTS) ? 1 : qs.GetSize();
const int width = constant.Width();
const int height = constant.Height();
vdim = width*height;
SetSize(nq*vdim);
for (int iq = 0; iq < nq; ++iq)
{
for (int j = 0; j < width; ++j)
{
for (int i = 0; i < height; ++i)
{
(*this)[i + j*height + iq*vdim] = constant(i, j);
}
}
}
}
void CoefficientVector::SetConstant(const DenseSymmetricMatrix &constant)
{
const int nq = (storage & CoefficientStorage::CONSTANTS) ? 1 : qs.GetSize();
const int height = constant.Height();
const bool sym = storage & CoefficientStorage::SYMMETRIC;
vdim = sym ? height*(height + 1)/2 : height*height;
SetSize(nq*vdim);
for (int iq = 0; iq < nq; ++iq)
{
for (int vd = 0; vd < vdim; ++vd)
{
const double value = sym ? constant.GetData()[vd] : constant(vd % height,
vd / height);
(*this)[vd + iq*vdim] = value;
}
}
}
int CoefficientVector::GetVDim() const { return vdim; }
CoefficientVector::~CoefficientVector()
{
delete qf;
}
}
+3 -171
View File
@@ -23,8 +23,6 @@ namespace mfem
{
class Mesh;
class QuadratureSpaceBase;
class QuadratureFunction;
#ifdef MFEM_USE_MPI
class ParMesh;
@@ -72,10 +70,6 @@ public:
return Eval(T, ip);
}
/// @brief Fill the QuadratureFunction @a qf by evaluating the coefficient at
/// the quadrature points.
virtual void Project(QuadratureFunction &qf);
virtual ~Coefficient() { }
};
@@ -93,9 +87,6 @@ public:
virtual double Eval(ElementTransformation &T,
const IntegrationPoint &ip)
{ return (constant); }
/// Fill the QuadratureFunction @a qf with the constant value.
void Project(QuadratureFunction &qf);
};
/** @brief A piecewise constant coefficient with the constants keyed
@@ -283,13 +274,6 @@ public:
/// Evaluate the coefficient at @a ip.
virtual double Eval(ElementTransformation &T,
const IntegrationPoint &ip);
/// @brief Fill the QuadratureFunction @a qf by evaluating the coefficient at
/// the quadrature points.
///
/// This function uses the efficient QuadratureFunction::ProjectGridFunction
/// to fill the QuadratureFunction.
virtual void Project(QuadratureFunction &qf);
};
@@ -487,13 +471,6 @@ public:
virtual void Eval(DenseMatrix &M, ElementTransformation &T,
const IntegrationRule &ir);
/// @brief Fill the QuadratureFunction @a qf by evaluating the coefficient at
/// the quadrature points.
///
/// The @a vdim of the VectorCoefficient should be equal to the @a vdim of
/// the QuadratureFunction.
virtual void Project(QuadratureFunction &qf);
virtual ~VectorCoefficient() { }
};
@@ -514,7 +491,7 @@ public:
const IntegrationPoint &ip) { V = vec; }
/// Return a reference to the constant vector in this class.
const Vector& GetVec() const { return vec; }
const Vector& GetVec() { return vec; }
};
/** @brief A piecewise vector-valued coefficient with the pieces keyed off the
@@ -711,13 +688,6 @@ public:
virtual void Eval(DenseMatrix &M, ElementTransformation &T,
const IntegrationRule &ir);
/// @brief Fill the QuadratureFunction @a qf by evaluating the coefficient at
/// the quadrature points.
///
/// This function uses the efficient QuadratureFunction::ProjectGridFunction
/// to fill the QuadratureFunction.
virtual void Project(QuadratureFunction &qf);
virtual ~VectorGridFunctionCoefficient() { }
};
@@ -945,14 +915,6 @@ public:
virtual void Eval(DenseMatrix &K, ElementTransformation &T,
const IntegrationPoint &ip) = 0;
/// @brief Fill the QuadratureFunction @a qf by evaluating the coefficient at
/// the quadrature points. The matrix will be transposed or not according to
/// the boolean argument @a transpose.
///
/// The @a vdim of the QuadratureFunction should be equal to the height times
/// the width of the matrix.
virtual void Project(QuadratureFunction &qf, bool transpose=false);
/// (DEPRECATED) Evaluate a symmetric matrix coefficient.
/** @brief Evaluate the upper triangular entries of the matrix coefficient
in the symmetric case, similarly to Eval. Matrix entry (i,j) is stored
@@ -981,8 +943,6 @@ public:
/// Evaluate the matrix coefficient at @a ip.
virtual void Eval(DenseMatrix &M, ElementTransformation &T,
const IntegrationPoint &ip) { M = mat; }
/// Return a reference to the constant matrix.
const DenseMatrix& GetMatrix() { return mat; }
};
@@ -1186,8 +1146,6 @@ public:
can be overridden with the @a own parameter. */
void Set(int i, int j, Coefficient * c, bool own=true);
using MatrixCoefficient::Eval;
/// Evaluate coefficient located at (i,j) in the matrix using integration
/// point @a ip.
double Eval(int i, int j, ElementTransformation &T, const IntegrationPoint &ip)
@@ -1302,15 +1260,6 @@ public:
/// Get the size of the matrix.
int GetSize() const { return height; }
/// @brief Fill the QuadratureFunction @a qf by evaluating the coefficient at
/// the quadrature points.
///
/// @note As opposed to MatrixCoefficient::Project, this function stores only
/// the @a symmetric part of the matrix at each quadrature point.
///
/// The @a vdim of the coefficient should be equal to height*(height+1)/2.
virtual void ProjectSymmetric(QuadratureFunction &qf);
/** @brief Evaluate the matrix coefficient in the element described by @a T
at the point @a ip, storing the result as a symmetric matrix @a K. */
/** @note When this method is called, the caller must make sure that the
@@ -1331,9 +1280,6 @@ public:
virtual void Eval(DenseMatrix &K, ElementTransformation &T,
const IntegrationPoint &ip);
/// Return a reference to the constant matrix.
const DenseSymmetricMatrix& GetMatrix() { return mat; }
virtual ~SymmetricMatrixCoefficient() { }
};
@@ -2103,6 +2049,8 @@ public:
};
///@}
class QuadratureFunction;
/** @brief Vector quadrature function coefficient which requires that the
quadrature rules used for this vector coefficient be the same as those that
live within the supplied QuadratureFunction. */
@@ -2127,8 +2075,6 @@ public:
virtual void Eval(Vector &V, ElementTransformation &T,
const IntegrationPoint &ip);
virtual void Project(QuadratureFunction &qf);
virtual ~VectorQuadratureFunctionCoefficient() { }
};
@@ -2148,123 +2094,9 @@ public:
virtual double Eval(ElementTransformation &T, const IntegrationPoint &ip);
virtual void Project(QuadratureFunction &qf);
virtual ~QuadratureFunctionCoefficient() { }
};
/// Flags that determine what storage optimizations to use in CoefficientVector
enum class CoefficientStorage : int
{
FULL = 0, ///< Store the coefficient as a full QuadratureFunction.
CONSTANTS = 1 << 0, ///< Store constants using only @a vdim entries.
SYMMETRIC = 1 << 1, ///< Store the triangular part of symmetric matrices.
COMPRESSED = CONSTANTS | SYMMETRIC ///< Enable all above compressions.
};
inline CoefficientStorage operator|(CoefficientStorage a, CoefficientStorage b)
{
return CoefficientStorage(int(a) | int(b));
}
inline int operator&(CoefficientStorage a, CoefficientStorage b)
{
return int(a) & int(b);
}
/// @brief Class to represent a coefficient evaluated at quadrature points.
///
/// In the general case, a CoefficientVector is the same as a QuadratureFunction
/// with a coefficient projected onto it.
///
/// This class allows for some "compression" of the coefficient data, according
/// to the storage flags given by CoefficientStorage. For example, constant
/// coefficients can be stored using only @a vdim values, and symmetric matrices
/// can be stored using e.g. the upper triangular part of the matrix.
class CoefficientVector : public Vector
{
protected:
CoefficientStorage storage; ///< Storage optimizations (see CoefficientStorage).
int vdim; ///< Number of values per quadrature point.
QuadratureSpaceBase &qs; ///< Associated QuadratureSpaceBase.
QuadratureFunction *qf; ///< Internal QuadratureFunction (owned, may be NULL).
public:
/// Create an empty CoefficientVector.
CoefficientVector(QuadratureSpaceBase &qs_,
CoefficientStorage storage_ = CoefficientStorage::FULL);
/// @brief Create a CoefficientVector from the given Coefficient and
/// QuadratureSpaceBase.
///
/// If @a coeff is NULL, it will be interpreted as a constant with value one.
/// @sa CoefficientStorage for a description of @a storage_.
CoefficientVector(Coefficient *coeff, QuadratureSpaceBase &qs,
CoefficientStorage storage_ = CoefficientStorage::FULL);
/// @brief Create a CoefficientVector from the given Coefficient and
/// QuadratureSpaceBase.
///
/// @sa CoefficientStorage for a description of @a storage_.
CoefficientVector(Coefficient &coeff, QuadratureSpaceBase &qs,
CoefficientStorage storage_ = CoefficientStorage::FULL);
/// @brief Create a CoefficientVector from the given VectorCoefficient and
/// QuadratureSpaceBase.
///
/// @sa CoefficientStorage for a description of @a storage_.
CoefficientVector(VectorCoefficient &coeff, QuadratureSpaceBase &qs,
CoefficientStorage storage_ = CoefficientStorage::FULL);
/// @brief Create a CoefficientVector from the given MatrixCoefficient and
/// QuadratureSpaceBase.
///
/// @sa CoefficientStorage for a description of @a storage_.
CoefficientVector(MatrixCoefficient &coeff, QuadratureSpaceBase &qs,
CoefficientStorage storage_ = CoefficientStorage::FULL);
/// @brief Evaluate the given Coefficient at the quadrature points defined by
/// @ref qs.
void Project(Coefficient &coeff);
/// @brief Evaluate the given VectorCoefficient at the quadrature points
/// defined by @ref qs.
///
/// @sa CoefficientVector for a description of the @a compress argument.
void Project(VectorCoefficient &coeff);
/// @brief Evaluate the given MatrixCoefficient at the quadrature points
/// defined by @ref qs.
///
/// @sa CoefficientVector for a description of the @a compress argument.
void Project(MatrixCoefficient &coeff, bool transpose=false);
/// @brief Project the transpose of @a coeff.
///
/// @sa Project(MatrixCoefficient&, QuadratureSpace&, bool, bool)
void ProjectTranspose(MatrixCoefficient &coeff);
/// Make this vector a reference to the given QuadratureFunction.
void MakeRef(const QuadratureFunction &qf_);
/// Set this vector to the given constant.
void SetConstant(double constant);
/// Set this vector to the given constant vector.
void SetConstant(const Vector &constant);
/// Set this vector to the given constant matrix.
void SetConstant(const DenseMatrix &constant);
/// Set this vector to the given constant symmetric matrix.
void SetConstant(const DenseSymmetricMatrix &constant);
/// Return the number of values per quadrature point.
int GetVDim() const;
~CoefficientVector();
};
/** @brief Compute the Lp norm of a function f.
\f$ \| f \|_{Lp} = ( \int_\Omega | f |^p d\Omega)^{1/p} \f$ */
double ComputeLpNorm(double p, Coefficient &coeff, Mesh &mesh,
+1 -1
View File
@@ -442,7 +442,7 @@ void VisItDataCollection::RegisterQField(const std::string& name,
{
int locLOD = GlobGeometryRefiner.GetRefinementLevelFromElems(
mesh->GetElementBaseGeometry(e),
qf->GetIntRule(e).GetNPoints());
qf->GetElementIntRule(e).GetNPoints());
LOD = std::max(LOD,locLOD);
}
-1
View File
@@ -14,7 +14,6 @@
#include "../config/config.hpp"
#include "gridfunc.hpp"
#include "qfunction.hpp"
#ifdef MFEM_USE_MPI
#include "pgridfunc.hpp"
#endif
-315
View File
@@ -1,315 +0,0 @@
// Copyright (c) 2010-2022, Lawrence Livermore National Security, LLC. Produced
// at the Lawrence Livermore National Laboratory. All Rights reserved. See files
// LICENSE and NOTICE for details. LLNL-CODE-806117.
//
// This file is part of the MFEM library. For more information and source code
// availability visit https://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the BSD-3 license. We welcome feedback and contributions, see file
// CONTRIBUTING.md for details.
#include "dgmassinv.hpp"
#include "bilinearform.hpp"
#include "dgmassinv_kernels.hpp"
#include "../general/forall.hpp"
namespace mfem
{
DGMassInverse::DGMassInverse(FiniteElementSpace &fes_orig, Coefficient *coeff,
const IntegrationRule *ir,
int btype)
: Solver(fes_orig.GetTrueVSize()),
fec(fes_orig.GetMaxElementOrder(),
fes_orig.GetMesh()->Dimension(),
btype,
fes_orig.GetFE(0)->GetMapType()),
fes(fes_orig.GetMesh(), &fec)
{
MFEM_VERIFY(fes.IsDGSpace(), "Space must be DG.");
MFEM_VERIFY(!fes.IsVariableOrder(), "Variable orders not supported.");
const int btype_orig =
static_cast<const L2_FECollection*>(fes_orig.FEColl())->GetBasisType();
if (btype_orig == btype)
{
// No change of basis required
d2q = nullptr;
}
else
{
// original basis to solver basis
const auto mode = DofToQuad::TENSOR;
d2q = &fes_orig.GetFE(0)->GetDofToQuad(fes.GetFE(0)->GetNodes(), mode);
int n = d2q->ndof;
Array<double> B_inv = d2q->B; // deep copy
Array<int> ipiv(n);
// solver basis to original
LUFactors lu(B_inv.HostReadWrite(), ipiv.HostWrite());
lu.Factor(n);
B_.SetSize(n*n);
lu.GetInverseMatrix(n, B_.HostWrite());
Bt_.SetSize(n*n);
DenseMatrix B_matrix(B_.HostReadWrite(), n, n);
DenseMatrix Bt_matrix(Bt_.HostWrite(), n, n);
Bt_matrix.Transpose(B_matrix);
}
if (coeff) { m = new MassIntegrator(*coeff, ir); }
else { m = new MassIntegrator(ir); }
diag_inv.SetSize(height);
// Workspace vectors used for CG
r_.SetSize(height);
d_.SetSize(height);
z_.SetSize(height);
// Only need transformed RHS if basis is different
if (btype_orig != btype) { b2_.SetSize(height); }
M = new BilinearForm(&fes);
M->AddDomainIntegrator(m); // M assumes ownership of m
M->SetAssemblyLevel(AssemblyLevel::PARTIAL);
// Assemble the bilinear form and its diagonal (for preconditioning).
Update();
}
DGMassInverse::DGMassInverse(FiniteElementSpace &fes_, Coefficient &coeff,
int btype)
: DGMassInverse(fes_, &coeff, nullptr, btype) { }
DGMassInverse::DGMassInverse(FiniteElementSpace &fes_, Coefficient &coeff,
const IntegrationRule &ir, int btype)
: DGMassInverse(fes_, &coeff, &ir, btype) { }
DGMassInverse::DGMassInverse(FiniteElementSpace &fes_,
const IntegrationRule &ir, int btype)
: DGMassInverse(fes_, nullptr, &ir, btype) { }
DGMassInverse::DGMassInverse(FiniteElementSpace &fes_, int btype)
: DGMassInverse(fes_, nullptr, nullptr, btype) { }
void DGMassInverse::SetOperator(const Operator &op)
{
MFEM_ABORT("SetOperator not supported with DGMassInverse.")
}
void DGMassInverse::SetRelTol(const double rel_tol_) { rel_tol = rel_tol_; }
void DGMassInverse::SetAbsTol(const double abs_tol_) { abs_tol = abs_tol_; }
void DGMassInverse::SetMaxIter(const double max_iter_) { max_iter = max_iter_; }
void DGMassInverse::Update()
{
M->Assemble();
M->AssembleDiagonal(diag_inv);
internal::MakeReciprocal(diag_inv.Size(), diag_inv.ReadWrite());
}
DGMassInverse::~DGMassInverse()
{
delete M;
}
template<int DIM, int D1D, int Q1D>
void DGMassInverse::DGMassCGIteration(const Vector &b_, Vector &u_) const
{
using namespace internal; // host/device kernel functions
const int NE = fes.GetNE();
const int d1d = m->dofs1D;
const int q1d = m->quad1D;
const int ND = static_cast<int>(pow(d1d, DIM));
const auto B = m->maps->B.Read();
const auto Bt = m->maps->Bt.Read();
const auto pa_data = m->pa_data.Read();
const auto dinv = diag_inv.Read();
auto r = r_.Write();
auto d = d_.Write();
auto z = z_.Write();
auto u = u_.ReadWrite();
const double RELTOL = rel_tol;
const double ABSTOL = abs_tol;
const double MAXIT = max_iter;
const bool IT_MODE = iterative_mode;
const bool CHANGE_BASIS = (d2q != nullptr);
// b is the right-hand side (if no change of basis, this just points to the
// incoming RHS vector, if we have to change basis, this points to the
// internal b2 vector where we put the transformed RHS)
const double *b;
// the following are non-null if we have to change basis
double *b2 = nullptr; // non-const access to b2
const double *b_orig = nullptr; // RHS vector in "original" basis
const double *d2q_B = nullptr; // matrix to transform initial guess
const double *q2d_B = nullptr; // matrix to transform solution
const double *q2d_Bt = nullptr; // matrix to transform RHS
if (CHANGE_BASIS)
{
d2q_B = d2q->B.Read();
q2d_B = B_.Read();
q2d_Bt = Bt_.Read();
b2 = b2_.Write();
b_orig = b_.Read();
b = b2;
}
else
{
b = b_.Read();
}
constexpr int NB = Q1D ? Q1D : 1; // block size
MFEM_FORALL_2D(e, NE, NB, NB, 1,
{
constexpr int NB = Q1D ? Q1D : 1; // redefine here for some compilers
// Perform change of basis if needed
if (CHANGE_BASIS)
{
// Transform RHS
DGMassBasis<DIM,D1D,MAX_D1D>(e, NE, q2d_Bt, b_orig, b2, d1d);
if (IT_MODE)
{
// Transform initial guess
DGMassBasis<DIM,D1D,MAX_D1D>(e, NE, d2q_B, u, u, d1d);
}
}
const int tid = MFEM_THREAD_ID(x) + NB*MFEM_THREAD_ID(y);
// Compute first residual
if (IT_MODE)
{
DGMassApply<DIM,D1D,Q1D>(e, NE, B, Bt, pa_data, u, r, d1d, q1d);
DGMassAxpy(e, NE, ND, 1.0, b, -1.0, r, r); // r = b - r
}
else
{
// if not in iterative mode, use zero initial guess
const int BX = MFEM_THREAD_SIZE(x);
const int BY = MFEM_THREAD_SIZE(y);
const int bxy = BX*BY;
const auto B = ConstDeviceMatrix(b, ND, NE);
auto U = DeviceMatrix(u, ND, NE);
auto R = DeviceMatrix(r, ND, NE);
for (int i = tid; i < ND; i += bxy)
{
U(i, e) = 0.0;
R(i, e) = B(i, e);
}
MFEM_SYNC_THREAD;
}
DGMassPreconditioner(e, NE, ND, dinv, r, z);
DGMassAxpy(e, NE, ND, 1.0, z, 0.0, z, d); // d = z
double nom = DGMassDot<NB>(e, NE, ND, d, r);
if (nom < 0.0) { return; /* Not positive definite */ }
double r0 = fmax(nom*RELTOL*RELTOL, ABSTOL*ABSTOL);
if (nom <= r0) { return; /* Converged */ }
DGMassApply<DIM,D1D,Q1D>(e, NE, B, Bt, pa_data, d, z, d1d, q1d);
double den = DGMassDot<NB>(e, NE, ND, z, d);
if (den <= 0.0)
{
DGMassDot<NB>(e, NE, ND, d, d);
// d2 > 0 => not positive definite
if (den == 0.0) { return; }
}
// start iteration
int i = 1;
while (true)
{
const double alpha = nom/den;
DGMassAxpy(e, NE, ND, 1.0, u, alpha, d, u); // u = u + alpha*d
DGMassAxpy(e, NE, ND, 1.0, r, -alpha, z, r); // r = r - alpha*A*d
DGMassPreconditioner(e, NE, ND, dinv, r, z);
double betanom = DGMassDot<NB>(e, NE, ND, r, z);
if (betanom < 0.0) { return; /* Not positive definite */ }
if (betanom <= r0) { break; /* Converged */ }
if (++i > MAXIT) { break; }
const double beta = betanom/nom;
DGMassAxpy(e, NE, ND, 1.0, z, beta, d, d); // d = z + beta*d
DGMassApply<DIM,D1D,Q1D>(e, NE, B, Bt, pa_data, d, z, d1d, q1d); // z = A d
den = DGMassDot<NB>(e, NE, ND, d, z);
if (den <= 0.0)
{
DGMassDot<NB>(e, NE, ND, d, d);
// d2 > 0 => not positive definite
if (den == 0.0) { break; }
}
nom = betanom;
}
if (CHANGE_BASIS)
{
DGMassBasis<DIM,D1D,MAX_D1D>(e, NE, q2d_B, u, u, d1d);
}
});
}
void DGMassInverse::Mult(const Vector &Mu, Vector &u) const
{
// Dispatch to templated version based on dim, d1d, and q1d.
const int dim = fes.GetMesh()->Dimension();
const int d1d = m->dofs1D;
const int q1d = m->quad1D;
const int id = (d1d << 4) | q1d;
if (dim == 2)
{
switch (id)
{
case 0x11: return DGMassCGIteration<2,1,1>(Mu, u);
case 0x22: return DGMassCGIteration<2,2,2>(Mu, u);
case 0x33: return DGMassCGIteration<2,3,3>(Mu, u);
case 0x35: return DGMassCGIteration<2,3,5>(Mu, u);
case 0x44: return DGMassCGIteration<2,4,4>(Mu, u);
case 0x46: return DGMassCGIteration<2,4,6>(Mu, u);
case 0x55: return DGMassCGIteration<2,5,5>(Mu, u);
case 0x57: return DGMassCGIteration<2,5,7>(Mu, u);
case 0x66: return DGMassCGIteration<2,6,6>(Mu, u);
case 0x68: return DGMassCGIteration<2,6,8>(Mu, u);
default: return DGMassCGIteration<2>(Mu, u); // Fallback
}
}
else if (dim == 3)
{
switch (id)
{
case 0x22: return DGMassCGIteration<3,2,2>(Mu, u);
case 0x23: return DGMassCGIteration<3,2,3>(Mu, u);
case 0x33: return DGMassCGIteration<3,3,3>(Mu, u);
case 0x34: return DGMassCGIteration<3,3,4>(Mu, u);
case 0x35: return DGMassCGIteration<3,3,5>(Mu, u);
case 0x44: return DGMassCGIteration<3,4,4>(Mu, u);
case 0x45: return DGMassCGIteration<3,4,5>(Mu, u);
case 0x46: return DGMassCGIteration<3,4,6>(Mu, u);
case 0x48: return DGMassCGIteration<3,4,8>(Mu, u);
case 0x55: return DGMassCGIteration<3,5,5>(Mu, u);
case 0x56: return DGMassCGIteration<3,5,6>(Mu, u);
case 0x57: return DGMassCGIteration<3,5,7>(Mu, u);
case 0x58: return DGMassCGIteration<3,5,8>(Mu, u);
case 0x66: return DGMassCGIteration<3,6,6>(Mu, u);
case 0x67: return DGMassCGIteration<3,6,7>(Mu, u);
default: return DGMassCGIteration<3>(Mu, u); // Fallback
}
}
}
} // namespace mfem
-112
View File
@@ -1,112 +0,0 @@
// Copyright (c) 2010-2022, Lawrence Livermore National Security, LLC. Produced
// at the Lawrence Livermore National Laboratory. All Rights reserved. See files
// LICENSE and NOTICE for details. LLNL-CODE-806117.
//
// This file is part of the MFEM library. For more information and source code
// availability visit https://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the BSD-3 license. We welcome feedback and contributions, see file
// CONTRIBUTING.md for details.
#ifndef MFEM_DGMASSINV_HPP
#define MFEM_DGMASSINV_HPP
#include "../linalg/operator.hpp"
#include "fespace.hpp"
namespace mfem
{
/// @brief Solver for the discontinuous Galerkin mass matrix.
///
/// This class performs a @a local (diagonally preconditioned) conjugate
/// gradient iteration for each element. Optionally, a change of basis is
/// performed to iterate on a better-conditioned system. This class fully
/// supports execution on device (GPU).
class DGMassInverse : public Solver
{
protected:
DG_FECollection fec; ///< FE collection in requested basis.
FiniteElementSpace fes; ///< FE space in requested basis.
const DofToQuad *d2q; ///< Change of basis. Not owned.
Array<double> B_; ///< Inverse of change of basis.
Array<double> Bt_; ///< Inverse of change of basis, transposed.
class BilinearForm *M; ///< Mass bilinear form, owned.
class MassIntegrator *m; ///< Mass integrator, owned by the form @ref M.
Vector diag_inv; ///< Jacobi preconditioner.
double rel_tol = 1e-12; ///< Relative CG tolerance.
double abs_tol = 1e-12; ///< Absolute CG tolerance.
int max_iter = 100; ///< Maximum number of CG iterations;
/// @name Intermediate vectors needed for CG three-term recurrence.
///@{
mutable Vector r_, d_, z_, b2_;
///@}
/// @brief Protected constructor, used internally.
///
/// Custom coefficient and integration rule are used if @a coeff and @a ir
/// are non-NULL.
DGMassInverse(FiniteElementSpace &fes_, Coefficient *coeff,
const IntegrationRule *ir, int btype);
public:
/// @brief Construct the DG inverse mass operator for @a fes_.
///
/// The basis type @a btype determines which basis should be used internally
/// in the solver. This <b>does not</b> have to be the same basis as @a fes_.
/// The best choice is typically BasisType::GaussLegendre because it is
/// well-preconditioned by its diagonal.
///
/// The solution and right-hand side used for the solver are not affected by
/// this basis (they correspond to the basis of @a fes_). @a btype is only
/// used internally, and only has an effect on the convergence rate.
DGMassInverse(FiniteElementSpace &fes_, int btype=BasisType::GaussLegendre);
/// @brief Construct the DG inverse mass operator for @a fes_ with
/// Coefficient @a coeff.
///
/// @sa DGMassInverse(FiniteElementSpace&, int) for information about @a
/// btype.
DGMassInverse(FiniteElementSpace &fes_, Coefficient &coeff,
int btype=BasisType::GaussLegendre);
/// @brief Construct the DG inverse mass operator for @a fes_ with
/// Coefficient @a coeff and IntegrationRule @a ir.
///
/// @sa DGMassInverse(FiniteElementSpace&, int) for information about @a
/// btype.
DGMassInverse(FiniteElementSpace &fes_, Coefficient &coeff,
const IntegrationRule &ir, int btype=BasisType::GaussLegendre);
/// @brief Construct the DG inverse mass operator for @a fes_ with
/// IntegrationRule @a ir.
///
/// @sa DGMassInverse(FiniteElementSpace&, int) for information about @a
/// btype.
DGMassInverse(FiniteElementSpace &fes_, const IntegrationRule &ir,
int btype=BasisType::GaussLegendre);
/// @brief Solve the system M b = u.
///
/// If @ref iterative_mode is @a true, @a u is used as an initial guess.
void Mult(const Vector &b, Vector &u) const;
/// Not implemented. Aborts.
void SetOperator(const Operator &op);
/// Set the relative tolerance.
void SetRelTol(const double rel_tol_);
/// Set the absolute tolerance.
void SetAbsTol(const double abs_tol_);
/// Set the maximum number of iterations.
void SetMaxIter(const double max_iter_);
/// Recompute operator and preconditioner (when coefficient or mesh changes).
void Update();
~DGMassInverse();
/// @brief Solve the system M b = u. <b>Not part of the public interface.</b>
/// @note This member function must be public because it contains an
/// MFEM_FORALL kernel (nvcc limitation)
template<int DIM, int D1D = 0, int Q1D = 0>
void DGMassCGIteration(const Vector &b_, Vector &u_) const;
};
} // namespace mfem
#endif
-295
View File
@@ -1,295 +0,0 @@
// Copyright (c) 2010-2022, Lawrence Livermore National Security, LLC. Produced
// at the Lawrence Livermore National Laboratory. All Rights reserved. See files
// LICENSE and NOTICE for details. LLNL-CODE-806117.
//
// This file is part of the MFEM library. For more information and source code
// availability visit https://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the BSD-3 license. We welcome feedback and contributions, see file
// CONTRIBUTING.md for details.
#ifndef MFEM_DGMASSINV_KERNELS_HPP
#define MFEM_DGMASSINV_KERNELS_HPP
#include "bilininteg_mass_pa.hpp"
#include "../linalg/kernels.hpp"
#include "kernels.hpp"
namespace mfem
{
namespace internal
{
void MakeReciprocal(int n, double *x)
{
MFEM_FORALL(i, n, x[i] = 1.0/x[i]; );
}
template <int DIM, int D1D, int Q1D>
MFEM_HOST_DEVICE inline
void DGMassApply(const int e,
const int NE,
const double *B,
const double *Bt,
const double *pa_data,
const double *x,
double *y,
const int d1d = 0,
const int q1d = 0)
{
constexpr bool use_smem = (D1D > 0 && Q1D > 0);
constexpr bool ACCUM = false;
constexpr int NBZ = 1;
if (use_smem)
{
// cannot specialize functions below with D1D or Q1D equal to zero
// (this branch only runs with D1D and Q1D are both positive)
constexpr int TD1D = D1D ? D1D : 1;
constexpr int TQ1D = Q1D ? Q1D : 1;
if (DIM == 2)
{
SmemPAMassApply2D_Element<TD1D,TQ1D,NBZ,ACCUM>(e, NE, B, pa_data, x, y);
}
else if (DIM == 3)
{
SmemPAMassApply3D_Element<TD1D,TQ1D,ACCUM>(e, NE, B, pa_data, x, y);
}
else
{
MFEM_ABORT_KERNEL("Unsupported dimension.");
}
}
else
{
if (DIM == 2)
{
PAMassApply2D_Element<ACCUM>(e, NE, B, Bt, pa_data, x, y, d1d, q1d);
}
else if (DIM == 3)
{
PAMassApply3D_Element<ACCUM>(e, NE, B, Bt, pa_data, x, y, d1d, q1d);
}
else
{
MFEM_ABORT_KERNEL("Unsupported dimension.");
}
}
}
MFEM_HOST_DEVICE inline
void DGMassPreconditioner(const int e,
const int NE,
const int ND,
const double *dinv,
const double *x,
double *y)
{
const auto X = ConstDeviceMatrix(x, ND, NE);
const auto D = ConstDeviceMatrix(dinv, ND, NE);
auto Y = DeviceMatrix(y, ND, NE);
const int tid = MFEM_THREAD_ID(x) + MFEM_THREAD_SIZE(x)*MFEM_THREAD_ID(y);
const int bxy = MFEM_THREAD_SIZE(x)*MFEM_THREAD_SIZE(y);
for (int i = tid; i < ND; i += bxy)
{
Y(i, e) = D(i, e)*X(i, e);
}
MFEM_SYNC_THREAD;
}
MFEM_HOST_DEVICE inline
void DGMassAxpy(const int e,
const int NE,
const int ND,
const double a,
const double *x,
const double b,
const double *y,
double *z)
{
const auto X = ConstDeviceMatrix(x, ND, NE);
const auto Y = ConstDeviceMatrix(y, ND, NE);
auto Z = DeviceMatrix(z, ND, NE);
const int tid = MFEM_THREAD_ID(x) + MFEM_THREAD_SIZE(x)*MFEM_THREAD_ID(y);
const int bxy = MFEM_THREAD_SIZE(x)*MFEM_THREAD_SIZE(y);
for (int i = tid; i < ND; i += bxy)
{
Z(i, e) = a*X(i, e) + b*Y(i, e);
}
MFEM_SYNC_THREAD;
}
template <int NB>
MFEM_HOST_DEVICE inline
double DGMassDot(const int e,
const int NE,
const int ND,
const double *x,
const double *y)
{
const auto X = ConstDeviceMatrix(x, ND, NE);
const auto Y = ConstDeviceMatrix(y, ND, NE);
const int tid = MFEM_THREAD_ID(x) + MFEM_THREAD_SIZE(x)*MFEM_THREAD_ID(y);
const int bxy = MFEM_THREAD_SIZE(x)*MFEM_THREAD_SIZE(y);
MFEM_SHARED double s_dot[NB*NB];
s_dot[tid] = 0.0;
for (int i = tid; i < ND; i += bxy) { s_dot[tid] += X(i,e)*Y(i,e); }
MFEM_SYNC_THREAD;
if (bxy > 512 && tid + 512 < bxy) { s_dot[tid] += s_dot[tid + 512]; }
MFEM_SYNC_THREAD;
if (bxy > 256 && tid < 256 && tid + 256 < bxy) { s_dot[tid] += s_dot[tid + 256]; }
MFEM_SYNC_THREAD;
if (bxy > 128 && tid < 128 && tid + 128 < bxy) { s_dot[tid] += s_dot[tid + 128]; }
MFEM_SYNC_THREAD;
if (bxy > 64 && tid < 64 && tid + 64 < bxy) { s_dot[tid] += s_dot[tid + 64]; }
MFEM_SYNC_THREAD;
if (bxy > 32 && tid < 32 && tid + 32 < bxy) { s_dot[tid] += s_dot[tid + 32]; }
MFEM_SYNC_THREAD;
if (bxy > 16 && tid < 16 && tid + 16 < bxy) { s_dot[tid] += s_dot[tid + 16]; }
MFEM_SYNC_THREAD;
if (bxy > 8 && tid < 8 && tid + 8 < bxy) { s_dot[tid] += s_dot[tid + 8]; }
MFEM_SYNC_THREAD;
if (bxy > 4 && tid < 4 && tid + 4 < bxy) { s_dot[tid] += s_dot[tid + 4]; }
MFEM_SYNC_THREAD;
if (bxy > 2 && tid < 2 && tid + 2 < bxy) { s_dot[tid] += s_dot[tid + 2]; }
MFEM_SYNC_THREAD;
if (bxy > 1 && tid < 1 && tid + 1 < bxy) { s_dot[tid] += s_dot[tid + 1]; }
MFEM_SYNC_THREAD;
return s_dot[0];
}
template<int T_D1D = 0, int MAX_D1D = 0>
MFEM_HOST_DEVICE inline
void DGMassBasis2D(const int e,
const int NE,
const double *b_,
const double *x_,
double *y_,
const int d1d = 0)
{
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
const int D1D = T_D1D ? T_D1D : d1d;
const auto b = Reshape(b_, D1D, D1D);
const auto x = Reshape(x_, D1D, D1D, NE);
auto y = Reshape(y_, D1D, D1D, NE);
MFEM_SHARED double sB[MD1*MD1];
MFEM_SHARED double sm0[MD1*MD1];
MFEM_SHARED double sm1[MD1*MD1];
kernels::internal::LoadB<MD1,MD1>(D1D,D1D,b,sB);
ConstDeviceMatrix B(sB, D1D,D1D);
DeviceMatrix DD(sm0, MD1, MD1);
DeviceMatrix DQ(sm1, MD1, MD1);
DeviceMatrix QQ(sm0, MD1, MD1);
kernels::internal::LoadX(e,D1D,x,DD);
kernels::internal::EvalX(D1D,D1D,B,DD,DQ);
kernels::internal::EvalY(D1D,D1D,B,DQ,QQ);
MFEM_SYNC_THREAD; // sync here to allow in-place evaluations
MFEM_FOREACH_THREAD(qy,y,D1D)
{
MFEM_FOREACH_THREAD(qx,x,D1D)
{
y(qx,qy,e) = QQ(qx,qy);
}
}
MFEM_SYNC_THREAD;
}
template<int T_D1D = 0, int MAX_D1D = 0>
MFEM_HOST_DEVICE inline
void DGMassBasis3D(const int e,
const int NE,
const double *b_,
const double *x_,
double *y_,
const int d1d = 0)
{
const int D1D = T_D1D ? T_D1D : d1d;
const auto b = Reshape(b_, D1D, D1D);
const auto x = Reshape(x_, D1D, D1D, D1D, NE);
auto y = Reshape(y_, D1D, D1D, D1D, NE);
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
MFEM_SHARED double sB[MD1*MD1];
MFEM_SHARED double sm0[MD1*MD1*MD1];
MFEM_SHARED double sm1[MD1*MD1*MD1];
kernels::internal::LoadB<MD1,MD1>(D1D,D1D,b,sB);
ConstDeviceMatrix B(sB, D1D,D1D);
DeviceCube DDD(sm0, MD1,MD1,MD1);
DeviceCube DDQ(sm1, MD1,MD1,MD1);
DeviceCube DQQ(sm0, MD1,MD1,MD1);
DeviceCube QQQ(sm1, MD1,MD1,MD1);
kernels::internal::LoadX(e,D1D,x,DDD);
kernels::internal::EvalX(D1D,D1D,B,DDD,DDQ);
kernels::internal::EvalY(D1D,D1D,B,DDQ,DQQ);
kernels::internal::EvalZ(D1D,D1D,B,DQQ,QQQ);
MFEM_SYNC_THREAD; // sync here to allow in-place evaluation
MFEM_FOREACH_THREAD(qz,z,D1D)
{
MFEM_FOREACH_THREAD(qy,y,D1D)
{
for (int qx = 0; qx < D1D; ++qx)
{
y(qx,qy,qz,e) = QQQ(qz,qy,qx);
}
}
}
MFEM_SYNC_THREAD;
}
template<int DIM, int T_D1D = 0, int MAX_D1D = 0>
MFEM_HOST_DEVICE inline
void DGMassBasis(const int e,
const int NE,
const double *b_,
const double *x_,
double *y_,
const int d1d = 0)
{
if (DIM == 2)
{
DGMassBasis2D<T_D1D, MAX_D1D>(e, NE, b_, x_, y_, d1d);
}
else if (DIM == 3)
{
DGMassBasis3D<T_D1D, MAX_D1D>(e, NE, b_, x_, y_, d1d);
}
else
{
MFEM_ABORT_KERNEL("Dimension not supported.");
}
}
} // namespace internal
} // namespace mfem
#endif
+2 -2
View File
@@ -159,7 +159,7 @@ void KellyErrorEstimator::ComputeEstimates()
// the FaceInfo class [1]. Also, the FaceElementTransformations
// documentation [2] may be helpful to grasp what is going on. Note
// that the FaceElementTransformations also works in the non-
// conforming case to transfer the Gauss points from the slave to
// conforming case to transfer the gauss points from the slave to
// the master element.
// [1]
// https://github.com/mfem/mfem/blob/02d0bfe9c18ce049c3c93a6a4208080fcfc96991/mesh/mesh.hpp#L94
@@ -417,7 +417,7 @@ void KellyErrorEstimator::ComputeEstimates()
Vector val(flux_space->GetVDim());
flux->GetVectorValue(FT->Elem2No, ip, val);
// Evaluate Gauss point
// Evaluate gauss point
Vector normal(mesh->SpaceDimension());
FT->Face->SetIntPoint(&fip);
if (mesh->Dimension() == mesh->SpaceDimension())
+2 -2
View File
@@ -314,7 +314,7 @@ void ND_HexahedronElement::CalcVShape(const IntegrationPoint &ip,
#ifdef MFEM_THREAD_SAFE
Vector shape_cx(p + 1), shape_ox(p), shape_cy(p + 1), shape_oy(p);
Vector shape_cz(p + 1), shape_oz(p);
Vector dshape_cx(p + 1), dshape_cy(p + 1), dshape_cz(p + 1);
Vector dshape_cx, dshape_cy, dshape_cz;
#endif
if (obasis1d.IsIntegratedType())
@@ -656,7 +656,7 @@ void ND_QuadrilateralElement::CalcVShape(const IntegrationPoint &ip,
#ifdef MFEM_THREAD_SAFE
Vector shape_cx(p + 1), shape_ox(p), shape_cy(p + 1), shape_oy(p);
Vector dshape_cx(p + 1), dshape_cy(p + 1);
Vector dshape_cx, dshape_cy;
#endif
if (obasis1d.IsIntegratedType())
+2 -2
View File
@@ -145,7 +145,7 @@ void RT_QuadrilateralElement::CalcVShape(const IntegrationPoint &ip,
#ifdef MFEM_THREAD_SAFE
Vector shape_cx(pp1 + 1), shape_ox(pp1), shape_cy(pp1 + 1), shape_oy(pp1);
Vector dshape_cx(pp1 + 1), dshape_cy(pp1 + 1);
Vector dshape_cx, dshape_cy;
#endif
if (obasis1d.IsIntegratedType())
@@ -473,7 +473,7 @@ void RT_HexahedronElement::CalcVShape(const IntegrationPoint &ip,
#ifdef MFEM_THREAD_SAFE
Vector shape_cx(pp1 + 1), shape_ox(pp1), shape_cy(pp1 + 1), shape_oy(pp1);
Vector shape_cz(pp1 + 1), shape_oz(pp1);
Vector dshape_cx(pp1 + 1), dshape_cy(pp1 + 1), dshape_cz(pp1 + 1);
Vector dshape_cx, dshape_cy, dshape_cz;
#endif
if (obasis1d.IsIntegratedType())
+6 -10
View File
@@ -53,12 +53,8 @@ public:
virtual int DofForGeometry(Geometry::Type GeomType) const = 0;
/** @brief Returns an array, say p, that maps a local permuted index i to a
local base index: base_i = p[i].
@note Only provides information about interior dofs. See
FiniteElementCollection::SubDofOrder if interior \a and boundary dof
order is needed. */
/** @brief Returns an array, say p, that maps a local permuted index i to
a local base index: base_i = p[i]. */
virtual const int *DofOrderForOrientation(Geometry::Type GeomType,
int Or) const = 0;
@@ -99,10 +95,10 @@ public:
| RT_ValTrace_[DIM]_[ORDER] | H^{1/2} | * | 1 / 0 | VALUE | H^{1/2}-conforming trace elements for H(div) defined on the interface between mesh elements (faces) |
| RT_Trace@[BTYPE]_[DIM]_[ORDER] | H^{1/2} | * | 1 / 0 | INTEGRAL | H^{1/2}-conforming trace elements for H(div) defined on the interface between mesh elements (faces) |
| RT_ValTrace@[BTYPE]_[DIM]_[ORDER] | H^{1/2} | * | 1 / 0 | VALUE | H^{1/2}-conforming trace elements for H(div) defined on the interface between mesh elements (faces) |
| L2_[DIM]_[ORDER] | L2 | * | 0 | VALUE | Discontinuous L2 elements |
| L2_T[BTYPE]_[DIM]_[ORDER] | L2 | * | 0 | VALUE | Discontinuous L2 elements |
| L2Int_[DIM]_[ORDER] | L2 | * | 0 | INTEGRAL | Discontinuous L2 elements |
| L2Int_T[BTYPE]_[DIM]_[ORDER] | L2 | * | 0 | INTEGRAL | Discontinuous L2 elements |
| L2_[DIM]_[ORDER] | L2 | * | 0 | VALUE | Discontinous L2 elements |
| L2_T[BTYPE]_[DIM]_[ORDER] | L2 | * | 0 | VALUE | Discontinous L2 elements |
| L2Int_[DIM]_[ORDER] | L2 | * | 0 | INTEGRAL | Discontinous L2 elements |
| L2Int_T[BTYPE]_[DIM]_[ORDER] | L2 | * | 0 | INTEGRAL | Discontinous L2 elements |
| DG_Iface_[DIM]_[ORDER] | - | * | 0 | VALUE | Discontinuous elements on the interface between mesh elements (faces) |
| DG_Iface@[BTYPE]_[DIM]_[ORDER] | - | * | 0 | VALUE | Discontinuous elements on the interface between mesh elements (faces) |
| DG_IntIface_[DIM]_[ORDER] | - | * | 0 | INTEGRAL | Discontinuous elements on the interface between mesh elements (faces) |
-1
View File
@@ -45,7 +45,6 @@
#include "multigrid.hpp"
#include "ceed/solvers/algebraic.hpp"
#include "lor/lor.hpp"
#include "dgmassinv.hpp"
#ifdef MFEM_USE_MPI
#include "pfespace.hpp"
+60 -4
View File
@@ -1199,6 +1199,8 @@ void FiniteElementSpace::BuildConformingInterpolation() const
MakeVDimMatrix(*cR);
if (cR_hp) { MakeVDimMatrix(*cR_hp); }
}
cP->EnsureMultTranspose();
}
void FiniteElementSpace::MakeVDimMatrix(SparseMatrix &mat) const
@@ -1256,7 +1258,7 @@ int FiniteElementSpace::GetNConformingDofs() const
return P ? (P->Width() / vdim) : ndofs;
}
const ElementRestrictionOperator *FiniteElementSpace::GetElementRestriction(
const Operator *FiniteElementSpace::GetElementRestriction(
ElementDofOrdering e_ordering) const
{
// Check if we have a discontinuous space using the FE collection:
@@ -1271,7 +1273,7 @@ const ElementRestrictionOperator *FiniteElementSpace::GetElementRestriction(
// The output E-vector layout is: ND x VDIM x NE.
L2E_nat.Reset(new L2ElementRestriction(*this));
}
return L2E_nat.Is<ElementRestrictionOperator>();
return L2E_nat.Ptr();
}
if (e_ordering == ElementDofOrdering::LEXICOGRAPHIC)
{
@@ -1279,14 +1281,14 @@ const ElementRestrictionOperator *FiniteElementSpace::GetElementRestriction(
{
L2E_lex.Reset(new ElementRestriction(*this, e_ordering));
}
return L2E_lex.Is<ElementRestrictionOperator>();
return L2E_lex.Ptr();
}
// e_ordering == ElementDofOrdering::NATIVE
if (L2E_nat.Ptr() == NULL)
{
L2E_nat.Reset(new ElementRestriction(*this, e_ordering));
}
return L2E_nat.Is<ElementRestrictionOperator>();
return L2E_nat.Ptr();
}
const FaceRestriction *FiniteElementSpace::GetFaceRestriction(
@@ -3611,4 +3613,58 @@ FiniteElementCollection *FiniteElementSpace::Load(Mesh *m, std::istream &input)
return r_fec;
}
void QuadratureSpace::Construct()
{
// protected method
int offset = 0;
const int num_elem = mesh->GetNE();
element_offsets = new int[num_elem + 1];
for (int g = 0; g < Geometry::NumGeom; g++)
{
int_rule[g] = NULL;
}
for (int i = 0; i < num_elem; i++)
{
element_offsets[i] = offset;
int geom = mesh->GetElementBaseGeometry(i);
if (int_rule[geom] == NULL)
{
int_rule[geom] = &IntRules.Get(geom, order);
}
offset += int_rule[geom]->GetNPoints();
}
element_offsets[num_elem] = size = offset;
}
QuadratureSpace::QuadratureSpace(Mesh *mesh_, std::istream &in)
: mesh(mesh_)
{
const char *msg = "invalid input stream";
string ident;
in >> ident; MFEM_VERIFY(ident == "QuadratureSpace", msg);
in >> ident; MFEM_VERIFY(ident == "Type:", msg);
in >> ident;
if (ident == "default_quadrature")
{
in >> ident; MFEM_VERIFY(ident == "Order:", msg);
in >> order;
}
else
{
MFEM_ABORT("unknown QuadratureSpace type: " << ident);
return;
}
Construct();
}
void QuadratureSpace::Save(std::ostream &os) const
{
os << "QuadratureSpace\n"
<< "Type: default_quadrature\n"
<< "Order: " << order << '\n';
}
} // namespace mfem
+56 -29
View File
@@ -47,14 +47,6 @@ public:
static void DofsToVDofs(int ndofs, int vdim, Array<int> &dofs);
};
/// @brief Type describing possible layouts for Q-vectors.
/// @sa QuadratureInterpolator and FaceQuadratureInterpolator.
enum class QVectorLayout
{
byNODES, ///< NQPT x VDIM x NE (values) / NQPT x VDIM x DIM x NE (grads)
byVDIM ///< VDIM x NQPT x NE (values) / VDIM x DIM x NQPT x NE (grads)
};
template <> inline int
Ordering::Map<Ordering::byNODES>(int ndofs, int vdim, int dof, int vd)
{
@@ -404,7 +396,7 @@ public:
FiniteElementSpace();
/** @brief Copy constructor: deep copy all data from @a orig except the Mesh,
the FiniteElementCollection, and some derived data. */
the FiniteElementCollection, ans some derived data. */
/** If the @a mesh or @a fec pointers are NULL (default), then the new
FiniteElementSpace will reuse the respective pointers from @a orig. If
any of these pointers is not NULL, the given pointer will be used instead
@@ -516,8 +508,7 @@ public:
L2ElementRestriction class.
The returned Operator is owned by the FiniteElementSpace. */
const ElementRestrictionOperator *GetElementRestriction(
ElementDofOrdering e_ordering) const;
const Operator *GetElementRestriction(ElementDofOrdering e_ordering) const;
/// Return an Operator that converts L-vectors to E-vectors on each face.
virtual const FaceRestriction *GetFaceRestriction(
@@ -531,12 +522,7 @@ public:
Operator returned by GetElementRestriction().
All elements will use the same IntegrationRule, @a ir as the target
quadrature points.
@note The returned pointer is shared. A good practice, before using it,
is to set all its properties to their expected values, as other parts of
the code may also change them. That is, it's good to call
SetOutputLayout() and DisableTensorProducts() before interpolating. */
quadrature points. */
const QuadratureInterpolator *GetQuadratureInterpolator(
const IntegrationRule &ir) const;
@@ -547,22 +533,12 @@ public:
Operator returned by GetElementRestriction().
The target quadrature points in the elements are described by the given
QuadratureSpace, @a qs.
@note The returned pointer is shared. A good practice, before using it,
is to set all its properties to their expected values, as other parts of
the code may also change them. That is, it's good to call
SetOutputLayout() and DisableTensorProducts() before interpolating. */
QuadratureSpace, @a qs. */
const QuadratureInterpolator *GetQuadratureInterpolator(
const QuadratureSpace &qs) const;
/** @brief Return a FaceQuadratureInterpolator that interpolates E-vectors to
quadrature point values and/or derivatives (Q-vectors).
@note The returned pointer is shared. A good practice, before using it,
is to set all its properties to their expected values, as other parts of
the code may also change them. That is, it's good to call
SetOutputLayout() and DisableTensorProducts() before interpolating. */
quadrature point values and/or derivatives (Q-vectors). */
const FaceQuadratureInterpolator *GetFaceQuadratureInterpolator(
const IntegrationRule &ir, FaceType type) const;
@@ -953,6 +929,57 @@ public:
virtual ~FiniteElementSpace();
};
/// Class representing the storage layout of a QuadratureFunction.
/** Multiple QuadratureFunction%s can share the same QuadratureSpace. */
class QuadratureSpace
{
protected:
friend class QuadratureFunction; // Uses the element_offsets.
Mesh *mesh;
int order;
int size;
const IntegrationRule *int_rule[Geometry::NumGeom];
int *element_offsets; // scalar offsets; size = number of elements + 1
// protected functions
// Assuming mesh and order are set, construct the members: int_rule,
// element_offsets, and size.
void Construct();
public:
/// Create a QuadratureSpace based on the global rules from #IntRules.
QuadratureSpace(Mesh *mesh_, int order_)
: mesh(mesh_), order(order_) { Construct(); }
/// Read a QuadratureSpace from the stream @a in.
QuadratureSpace(Mesh *mesh_, std::istream &in);
virtual ~QuadratureSpace() { delete [] element_offsets; }
/// Return the total number of quadrature points.
int GetSize() const { return size; }
/// Return the order of the quadrature rule(s) used by all elements.
int GetOrder() const { return order; }
/// Returns the mesh
inline Mesh *GetMesh() const { return mesh; }
/// Returns number of elements in the mesh.
inline int GetNE() const { return mesh->GetNE(); }
/// Get the IntegrationRule associated with mesh element @a idx.
const IntegrationRule &GetElementIntRule(int idx) const
{ return *int_rule[mesh->GetElementBaseGeometry(idx)]; }
/// Write the QuadratureSpace to the stream @a out.
void Save(std::ostream &out) const;
};
/// @brief Return true if the mesh contains only one topology and the elements are tensor elements.
inline bool UsesTensorBasis(const FiniteElementSpace& fes)
{
+201 -36
View File
@@ -12,7 +12,6 @@
// Implementation of GridFunction
#include "gridfunc.hpp"
#include "quadinterpolator.hpp"
#include "../mesh/nurbs.hpp"
#include "../general/text.hpp"
@@ -189,8 +188,6 @@ void GridFunction::Update()
{
SetSize(fes->GetVSize());
}
if (t_vec.Size() > 0) { SetTrueVector(); }
}
void GridFunction::SetSpace(FiniteElementSpace *f)
@@ -2764,8 +2761,7 @@ void GridFunction::ProjectBdrCoefficientTangent(
}
double GridFunction::ComputeL2Error(
Coefficient *exsol[], const IntegrationRule *irs[],
const Array<int> *elems) const
Coefficient *exsol[], const IntegrationRule *irs[]) const
{
double error = 0.0, a;
const FiniteElement *fe;
@@ -2776,7 +2772,6 @@ double GridFunction::ComputeL2Error(
for (i = 0; i < fes->GetNE(); i++)
{
if (elems != NULL && (*elems)[i] == 0) { continue; }
fe = fes->GetFE(i);
fdof = fe->GetDof();
transf = fes->GetElementTransformation(i);
@@ -2820,7 +2815,7 @@ double GridFunction::ComputeL2Error(
double GridFunction::ComputeL2Error(
VectorCoefficient &exsol, const IntegrationRule *irs[],
const Array<int> *elems) const
Array<int> *elems) const
{
double error = 0.0;
const FiniteElement *fe;
@@ -3239,7 +3234,7 @@ double GridFunction::ComputeMaxError(
double GridFunction::ComputeW11Error(
Coefficient *exsol, VectorCoefficient *exgrad, int norm_type,
const Array<int> *elems, const IntegrationRule *irs[]) const
Array<int> *elems, const IntegrationRule *irs[]) const
{
// assuming vdim is 1
int i, fdof, dim, intorder, j, k;
@@ -3345,8 +3340,7 @@ double GridFunction::ComputeW11Error(
double GridFunction::ComputeLpError(const double p, Coefficient &exsol,
Coefficient *weight,
const IntegrationRule *irs[],
const Array<int> *elems) const
const IntegrationRule *irs[]) const
{
double error = 0.0;
const FiniteElement *fe;
@@ -3355,7 +3349,6 @@ double GridFunction::ComputeLpError(const double p, Coefficient &exsol,
for (int i = 0; i < fes->GetNE(); i++)
{
if (elems != NULL && (*elems)[i] == 0) { continue; }
fe = fes->GetFE(i);
const IntegrationRule *ir;
if (irs)
@@ -3975,6 +3968,178 @@ void GridFunction::LegacyNCReorder()
Vector::Swap(tmp);
}
QuadratureFunction::QuadratureFunction(Mesh *mesh, std::istream &in)
{
const char *msg = "invalid input stream";
string ident;
qspace = new QuadratureSpace(mesh, in);
own_qspace = true;
in >> ident; MFEM_VERIFY(ident == "VDim:", msg);
in >> vdim;
Load(in, vdim*qspace->GetSize());
}
QuadratureFunction & QuadratureFunction::operator=(double value)
{
Vector::operator=(value);
return *this;
}
QuadratureFunction & QuadratureFunction::operator=(const Vector &v)
{
MFEM_ASSERT(qspace && v.Size() == this->Size(), "");
Vector::operator=(v);
return *this;
}
QuadratureFunction & QuadratureFunction::operator=(const QuadratureFunction &v)
{
return this->operator=((const Vector &)v);
}
void QuadratureFunction::Save(std::ostream &os) const
{
qspace->Save(os);
os << "VDim: " << vdim << '\n'
<< '\n';
Vector::Print(os, vdim);
os.flush();
}
std::ostream &operator<<(std::ostream &os, const QuadratureFunction &qf)
{
qf.Save(os);
return os;
}
void QuadratureFunction::SaveVTU(std::ostream &os, VTKFormat format,
int compression_level) const
{
os << R"(<VTKFile type="UnstructuredGrid" version="0.1")";
if (compression_level != 0)
{
os << R"( compressor="vtkZLibDataCompressor")";
}
os << " byte_order=\"" << VTKByteOrder() << "\">\n";
os << "<UnstructuredGrid>\n";
const char *fmt_str = (format == VTKFormat::ASCII) ? "ascii" : "binary";
const char *type_str = (format != VTKFormat::BINARY32) ? "Float64" : "Float32";
std::vector<char> buf;
int np = qspace->GetSize();
int ne = qspace->GetNE();
int sdim = qspace->GetMesh()->SpaceDimension();
// For quadrature functions, each point is a vertex cell, so number of cells
// is equal to number of points
os << "<Piece NumberOfPoints=\"" << np
<< "\" NumberOfCells=\"" << np << "\">\n";
// print out the points
os << "<Points>\n";
os << "<DataArray type=\"" << type_str
<< "\" NumberOfComponents=\"3\" format=\"" << fmt_str << "\">\n";
Vector pt(sdim);
for (int i = 0; i < ne; i++)
{
ElementTransformation &T = *qspace->GetMesh()->GetElementTransformation(i);
const IntegrationRule &ir = GetElementIntRule(i);
for (int j = 0; j < ir.Size(); j++)
{
T.Transform(ir[j], pt);
WriteBinaryOrASCII(os, buf, pt[0], " ", format);
if (sdim > 1) { WriteBinaryOrASCII(os, buf, pt[1], " ", format); }
else { WriteBinaryOrASCII(os, buf, 0.0, " ", format); }
if (sdim > 2) { WriteBinaryOrASCII(os, buf, pt[2], "", format); }
else { WriteBinaryOrASCII(os, buf, 0.0, "", format); }
if (format == VTKFormat::ASCII) { os << '\n'; }
}
}
if (format != VTKFormat::ASCII)
{
WriteBase64WithSizeAndClear(os, buf, compression_level);
}
os << "</DataArray>\n";
os << "</Points>\n";
// Write cells (each cell is just a vertex)
os << "<Cells>\n";
// Connectivity
os << R"(<DataArray type="Int32" Name="connectivity" format=")"
<< fmt_str << "\">\n";
for (int i=0; i<np; ++i) { WriteBinaryOrASCII(os, buf, i, "\n", format); }
if (format != VTKFormat::ASCII)
{
WriteBase64WithSizeAndClear(os, buf, compression_level);
}
os << "</DataArray>\n";
// Offsets
os << R"(<DataArray type="Int32" Name="offsets" format=")"
<< fmt_str << "\">\n";
for (int i=0; i<np; ++i) { WriteBinaryOrASCII(os, buf, i, "\n", format); }
if (format != VTKFormat::ASCII)
{
WriteBase64WithSizeAndClear(os, buf, compression_level);
}
os << "</DataArray>\n";
// Types
os << R"(<DataArray type="UInt8" Name="types" format=")"
<< fmt_str << "\">\n";
for (int i = 0; i < np; i++)
{
uint8_t vtk_cell_type = VTKGeometry::POINT;
WriteBinaryOrASCII(os, buf, vtk_cell_type, "\n", format);
}
if (format != VTKFormat::ASCII)
{
WriteBase64WithSizeAndClear(os, buf, compression_level);
}
os << "</DataArray>\n";
os << "</Cells>\n";
os << "<PointData>\n";
os << "<DataArray type=\"" << type_str << "\" Name=\"u\" format=\""
<< fmt_str << "\" NumberOfComponents=\"" << vdim << "\">\n";
for (int i = 0; i < ne; i++)
{
DenseMatrix vals;
GetElementValues(i, vals);
for (int j = 0; j < vals.Size(); ++j)
{
for (int vd = 0; vd < vdim; ++vd)
{
WriteBinaryOrASCII(os, buf, vals(vd, j), " ", format);
}
if (format == VTKFormat::ASCII) { os << '\n'; }
}
}
if (format != VTKFormat::ASCII)
{
WriteBase64WithSizeAndClear(os, buf, compression_level);
}
os << "</DataArray>\n";
os << "</PointData>\n";
os << "</Piece>\n";
os << "</UnstructuredGrid>\n";
os << "</VTKFile>" << std::endl;
}
void QuadratureFunction::SaveVTU(const std::string &filename, VTKFormat format,
int compression_level) const
{
std::ofstream f(filename + ".vtu");
SaveVTU(f, format, compression_level);
}
double ZZErrorEstimator(BilinearFormIntegrator &blfi,
GridFunction &u,
GridFunction &flux, Vector &error_estimates,
@@ -4092,7 +4257,7 @@ void TensorProductLegendre(int dim, // input
}
else
{
// Bounding box is not reoriented no need to change orientation
// Bounding box is not reorientated no need to change orientation
x = x_in;
}
@@ -4116,44 +4281,44 @@ void TensorProductLegendre(int dim, // input
switch (dim)
{
case 1:
{
for (int i = 0; i <= order; i++)
{
poly(i) = poly_x(i);
}
}
break;
case 2:
{
for (int j = 0; j <= order; j++)
{
for (int i = 0; i <= order; i++)
{
int cnt = i + (order+1) * j;
poly(cnt) = poly_x(i) * poly_y(j);
poly(i) = poly_x(i);
}
}
}
break;
case 3:
{
for (int k = 0; k <= order; k++)
break;
case 2:
{
for (int j = 0; j <= order; j++)
{
for (int i = 0; i <= order; i++)
{
int cnt = i + (order+1) * j + (order+1) * (order+1) * k;
poly(cnt) = poly_x(i) * poly_y(j) * poly_z(k);
int cnt = i + (order+1) * j;
poly(cnt) = poly_x(i) * poly_y(j);
}
}
}
}
break;
break;
case 3:
{
for (int k = 0; k <= order; k++)
{
for (int j = 0; j <= order; j++)
{
for (int i = 0; i <= order; i++)
{
int cnt = i + (order+1) * j + (order+1) * (order+1) * k;
poly(cnt) = poly_x(i) * poly_y(j) * poly_z(k);
}
}
}
}
break;
default:
{
MFEM_ABORT("TensorProductLegendre: invalid value of dim");
}
{
MFEM_ABORT("TensorProductLegendre: invalid value of dim");
}
}
}
+276 -29
View File
@@ -129,21 +129,19 @@ public:
int CurlDim() const;
/// Read only access to the (optional) internal true-dof Vector.
const Vector &GetTrueVector() const
{
MFEM_VERIFY(t_vec.Size() > 0, "SetTrueVector() before GetTrueVector()");
return t_vec;
}
/** Note that the returned Vector may be empty, if not previously allocated
or set. */
const Vector &GetTrueVector() const { return t_vec; }
/// Read and write access to the (optional) internal true-dof Vector.
/** Note that @a t_vec is set if it is not allocated or set already.*/
Vector &GetTrueVector()
{ if (t_vec.Size() == 0) { SetTrueVector(); } return t_vec; }
/** Note that the returned Vector may be empty, if not previously allocated
or set. */
Vector &GetTrueVector() { return t_vec; }
/// Extract the true-dofs from the GridFunction.
void GetTrueDofs(Vector &tv) const;
/// Shortcut for calling GetTrueDofs() with GetTrueVector() as argument.
void SetTrueVector() { GetTrueDofs(t_vec); }
void SetTrueVector() { GetTrueDofs(GetTrueVector()); }
/// Set the GridFunction from the given true-dof vector.
virtual void SetFromTrueDofs(const Vector &tv);
@@ -478,27 +476,21 @@ public:
Array<int> &bdr_attr);
virtual double ComputeL2Error(Coefficient &exsol,
const IntegrationRule *irs[] = NULL) const
{ return ComputeLpError(2.0, exsol, NULL, irs); }
virtual double ComputeL2Error(Coefficient *exsol[],
const IntegrationRule *irs[] = NULL) const;
virtual double ComputeL2Error(VectorCoefficient &exsol,
const IntegrationRule *irs[] = NULL,
const Array<int> *elems = NULL) const;
Array<int> *elems = NULL) const;
/// Returns ||grad u_ex - grad u_h||_L2 in element ielem for H1 or L2 elements
virtual double ComputeElementGradError(int ielem, VectorCoefficient *exgrad,
const IntegrationRule *irs[] = NULL) const;
/// Returns ||u_ex - u_h||_L2 for H1 or L2 elements
/* The @a elems input variable expects a list of markers:
an elem marker equal to 1 will compute the L2 error on that element
an elem marker equal to 0 will not compute the L2 error on that element */
virtual double ComputeL2Error(Coefficient &exsol,
const IntegrationRule *irs[] = NULL,
const Array<int> *elems = NULL) const
{ return GridFunction::ComputeLpError(2.0, exsol, NULL, irs, elems); }
virtual double ComputeL2Error(VectorCoefficient &exsol,
const IntegrationRule *irs[] = NULL,
const Array<int> *elems = NULL) const;
/// Returns ||grad u_ex - grad u_h||_L2 for H1 or L2 elements
virtual double ComputeGradError(VectorCoefficient *exgrad,
const IntegrationRule *irs[] = NULL) const;
@@ -572,20 +564,16 @@ public:
{ return ComputeLpError(1.0, exsol, NULL, irs); }
virtual double ComputeW11Error(Coefficient *exsol, VectorCoefficient *exgrad,
int norm_type, const Array<int> *elems = NULL,
int norm_type, Array<int> *elems = NULL,
const IntegrationRule *irs[] = NULL) const;
virtual double ComputeL1Error(VectorCoefficient &exsol,
const IntegrationRule *irs[] = NULL) const
{ return ComputeLpError(1.0, exsol, NULL, NULL, irs); }
/* The @a elems input variable expects a list of markers:
an elem marker equal to 1 will compute the L2 error on that element
an elem marker equal to 0 will not compute the L2 error on that element */
virtual double ComputeLpError(const double p, Coefficient &exsol,
Coefficient *weight = NULL,
const IntegrationRule *irs[] = NULL,
const Array<int> *elems = NULL) const;
const IntegrationRule *irs[] = NULL) const;
/** Compute the Lp error in each element of the mesh and store the results in
the Vector @a error. The result should be of length number of elements,
@@ -767,6 +755,176 @@ public:
}
};
/** @brief Class representing a function through its values (scalar or vector)
at quadrature points. */
class QuadratureFunction : public Vector
{
protected:
QuadratureSpace *qspace; ///< Associated QuadratureSpace
int vdim; ///< Vector dimension
bool own_qspace; ///< QuadratureSpace ownership flag
public:
/// Create an empty QuadratureFunction.
/** The object can be initialized later using the SetSpace() methods. */
QuadratureFunction()
: qspace(NULL), vdim(0), own_qspace(false) { }
/** @brief Copy constructor. The QuadratureSpace ownership flag, #own_qspace,
in the new object is set to false. */
QuadratureFunction(const QuadratureFunction &orig)
: Vector(orig),
qspace(orig.qspace), vdim(orig.vdim), own_qspace(false) { }
/// Create a QuadratureFunction based on the given QuadratureSpace.
/** The QuadratureFunction does not assume ownership of the QuadratureSpace.
@note The Vector data is not initialized. */
QuadratureFunction(QuadratureSpace *qspace_, int vdim_ = 1)
: Vector(vdim_*qspace_->GetSize()),
qspace(qspace_), vdim(vdim_), own_qspace(false) { }
/** @brief Create a QuadratureFunction based on the given QuadratureSpace,
using the external data, @a qf_data. */
/** The QuadratureFunction does not assume ownership of neither the
QuadratureSpace nor the external data. */
QuadratureFunction(QuadratureSpace *qspace_, double *qf_data, int vdim_ = 1)
: Vector(qf_data, vdim_*qspace_->GetSize()),
qspace(qspace_), vdim(vdim_), own_qspace(false) { }
/// Read a QuadratureFunction from the stream @a in.
/** The QuadratureFunction assumes ownership of the read QuadratureSpace. */
QuadratureFunction(Mesh *mesh, std::istream &in);
virtual ~QuadratureFunction() { if (own_qspace) { delete qspace; } }
/// Get the associated QuadratureSpace.
QuadratureSpace *GetSpace() const { return qspace; }
/// Change the QuadratureSpace and optionally the vector dimension.
/** If the new QuadratureSpace is different from the current one, the
QuadratureFunction will not assume ownership of the new space; otherwise,
the ownership flag remains the same.
If the new vector dimension @a vdim_ < 0, the vector dimension remains
the same.
The data size is updated by calling Vector::SetSize(). */
inline void SetSpace(QuadratureSpace *qspace_, int vdim_ = -1);
/** @brief Change the QuadratureSpace, the data array, and optionally the
vector dimension. */
/** If the new QuadratureSpace is different from the current one, the
QuadratureFunction will not assume ownership of the new space; otherwise,
the ownership flag remains the same.
If the new vector dimension @a vdim_ < 0, the vector dimension remains
the same.
The data array is replaced by calling Vector::NewDataAndSize(). */
inline void SetSpace(QuadratureSpace *qspace_, double *qf_data,
int vdim_ = -1);
/// Get the vector dimension.
int GetVDim() const { return vdim; }
/// Set the vector dimension, updating the size by calling Vector::SetSize().
void SetVDim(int vdim_)
{ vdim = vdim_; SetSize(vdim*qspace->GetSize()); }
/// Get the QuadratureSpace ownership flag.
bool OwnsSpace() { return own_qspace; }
/// Set the QuadratureSpace ownership flag.
void SetOwnsSpace(bool own) { own_qspace = own; }
/// Redefine '=' for QuadratureFunction = constant.
QuadratureFunction &operator=(double value);
/// Copy the data from @a v.
/** The size of @a v must be equal to the size of the associated
QuadratureSpace #qspace times the QuadratureFunction dimension
i.e. QuadratureFunction::Size(). */
QuadratureFunction &operator=(const Vector &v);
/// Copy assignment. Only the data of the base class Vector is copied.
/** The QuadratureFunctions @a v and @a *this must have QuadratureSpaces with
the same size.
@note Defining this method overwrites the implicitly defined copy
assignment operator. */
QuadratureFunction &operator=(const QuadratureFunction &v);
/// Get the IntegrationRule associated with mesh element @a idx.
const IntegrationRule &GetElementIntRule(int idx) const
{ return qspace->GetElementIntRule(idx); }
/// Return all values associated with mesh element @a idx in a Vector.
/** The result is stored in the Vector @a values as a reference to the
global values.
Inside the Vector @a values, the index `i+vdim*j` corresponds to the
`i`-th vector component at the `j`-th quadrature point.
*/
inline void GetElementValues(int idx, Vector &values);
/// Return all values associated with mesh element @a idx in a Vector.
/** The result is stored in the Vector @a values as a copy of the
global values.
Inside the Vector @a values, the index `i+vdim*j` corresponds to the
`i`-th vector component at the `j`-th quadrature point.
*/
inline void GetElementValues(int idx, Vector &values) const;
/// Return the quadrature function values at an integration point.
/** The result is stored in the Vector @a values as a reference to the
global values. */
inline void GetElementValues(int idx, const int ip_num, Vector &values);
/// Return the quadrature function values at an integration point.
/** The result is stored in the Vector @a values as a copy to the
global values. */
inline void GetElementValues(int idx, const int ip_num, Vector &values) const;
/// Return all values associated with mesh element @a idx in a DenseMatrix.
/** The result is stored in the DenseMatrix @a values as a reference to the
global values.
Inside the DenseMatrix @a values, the `(i,j)` entry corresponds to the
`i`-th vector component at the `j`-th quadrature point.
*/
inline void GetElementValues(int idx, DenseMatrix &values);
/// Return all values associated with mesh element @a idx in a const DenseMatrix.
/** The result is stored in the DenseMatrix @a values as a copy of the
global values.
Inside the DenseMatrix @a values, the `(i,j)` entry corresponds to the
`i`-th vector component at the `j`-th quadrature point.
*/
inline void GetElementValues(int idx, DenseMatrix &values) const;
/// Write the QuadratureFunction to the stream @a out.
void Save(std::ostream &out) const;
/// @brief Write the QuadratureFunction to @a out in VTU (ParaView) format.
///
/// The data will be uncompressed if @a compression_level is zero, or if the
/// format is VTKFormat::ASCII. Otherwise, zlib compression will be used for
/// binary data.
void SaveVTU(std::ostream &out, VTKFormat format=VTKFormat::ASCII,
int compression_level=0) const;
/// @brief Save the QuadratureFunction to a VTU (ParaView) file.
///
/// The extension ".vtu" will be appended to @a filename.
/// @sa SaveVTU(std::ostream &out, VTKFormat format=VTKFormat::ASCII,
/// int compression_level=0)
void SaveVTU(const std::string &filename, VTKFormat format=VTKFormat::ASCII,
int compression_level=0) const;
};
/// Overload operator<< for std::ostream and QuadratureFunction.
std::ostream &operator<<(std::ostream &out, const QuadratureFunction &qf);
@@ -854,6 +1012,95 @@ public:
GridFunction *Extrude1DGridFunction(Mesh *mesh, Mesh *mesh2d,
GridFunction *sol, const int ny);
// Inline methods
inline void QuadratureFunction::SetSpace(QuadratureSpace *qspace_, int vdim_)
{
if (qspace_ != qspace)
{
if (own_qspace) { delete qspace; }
qspace = qspace_;
own_qspace = false;
}
vdim = (vdim_ < 0) ? vdim : vdim_;
SetSize(vdim*qspace->GetSize());
}
inline void QuadratureFunction::SetSpace(QuadratureSpace *qspace_,
double *qf_data, int vdim_)
{
if (qspace_ != qspace)
{
if (own_qspace) { delete qspace; }
qspace = qspace_;
own_qspace = false;
}
vdim = (vdim_ < 0) ? vdim : vdim_;
NewDataAndSize(qf_data, vdim*qspace->GetSize());
}
inline void QuadratureFunction::GetElementValues(int idx, Vector &values)
{
const int s_offset = qspace->element_offsets[idx];
const int sl_size = qspace->element_offsets[idx+1] - s_offset;
values.NewDataAndSize(data + vdim*s_offset, vdim*sl_size);
}
inline void QuadratureFunction::GetElementValues(int idx, Vector &values) const
{
const int s_offset = qspace->element_offsets[idx];
const int sl_size = qspace->element_offsets[idx+1] - s_offset;
values.SetSize(vdim*sl_size);
const double *q = data + vdim*s_offset;
for (int i = 0; i<values.Size(); i++)
{
values(i) = *(q++);
}
}
inline void QuadratureFunction::GetElementValues(int idx, const int ip_num,
Vector &values)
{
const int s_offset = qspace->element_offsets[idx] * vdim + ip_num * vdim;
values.NewDataAndSize(data + s_offset, vdim);
}
inline void QuadratureFunction::GetElementValues(int idx, const int ip_num,
Vector &values) const
{
const int s_offset = qspace->element_offsets[idx] * vdim + ip_num * vdim;
values.SetSize(vdim);
const double *q = data + s_offset;
for (int i = 0; i < values.Size(); i++)
{
values(i) = *(q++);
}
}
inline void QuadratureFunction::GetElementValues(int idx, DenseMatrix &values)
{
const int s_offset = qspace->element_offsets[idx];
const int sl_size = qspace->element_offsets[idx+1] - s_offset;
values.Reset(data + vdim*s_offset, vdim, sl_size);
}
inline void QuadratureFunction::GetElementValues(int idx,
DenseMatrix &values) const
{
const int s_offset = qspace->element_offsets[idx];
const int sl_size = qspace->element_offsets[idx+1] - s_offset;
values.SetSize(vdim, sl_size);
const double *q = data + vdim*s_offset;
for (int j = 0; j<sl_size; j++)
{
for (int i = 0; i<vdim; i++)
{
values(i,j) = *(q++);
}
}
}
} // namespace mfem
#endif
+44 -88
View File
@@ -168,8 +168,7 @@ void FindPointsGSLIB::Setup(Mesh &m, const double bb_t, const double newt_tol,
setupflag = true;
}
void FindPointsGSLIB::FindPoints(const Vector &point_pos,
int point_pos_ordering)
void FindPointsGSLIB::FindPoints(const Vector &point_pos)
{
MFEM_VERIFY(setupflag, "Use FindPointsGSLIB::Setup before finding points.");
points_cnt = point_pos.Size() / dim;
@@ -179,24 +178,14 @@ void FindPointsGSLIB::FindPoints(const Vector &point_pos,
gsl_ref.SetSize(points_cnt * dim);
gsl_dist.SetSize(points_cnt);
const double *xv_base[dim];
unsigned xv_stride[dim];
for (int d = 0; d < dim; d++)
{
if (point_pos_ordering == Ordering::byNODES)
{
xv_base[d] = point_pos.GetData() + d*points_cnt;
xv_stride[d] = sizeof(double);
}
else
{
xv_base[d] = point_pos.GetData() + d;
xv_stride[d] = dim*sizeof(double);
}
}
if (dim == 2)
{
const double *xv_base[2];
xv_base[0] = point_pos.GetData();
xv_base[1] = point_pos.GetData() + points_cnt;
unsigned xv_stride[2];
xv_stride[0] = sizeof(double);
xv_stride[1] = sizeof(double);
findpts_2(gsl_code.GetData(), sizeof(unsigned int),
gsl_proc.GetData(), sizeof(unsigned int),
gsl_elem.GetData(), sizeof(unsigned int),
@@ -206,6 +195,14 @@ void FindPointsGSLIB::FindPoints(const Vector &point_pos,
}
else
{
const double *xv_base[3];
xv_base[0] = point_pos.GetData();
xv_base[1] = point_pos.GetData() + points_cnt;
xv_base[2] = point_pos.GetData() + 2*points_cnt;
unsigned xv_stride[3];
xv_stride[0] = sizeof(double);
xv_stride[1] = sizeof(double);
xv_stride[2] = sizeof(double);
findpts_3(gsl_code.GetData(), sizeof(unsigned int),
gsl_proc.GetData(), sizeof(unsigned int),
gsl_elem.GetData(), sizeof(unsigned int),
@@ -230,29 +227,27 @@ void FindPointsGSLIB::FindPoints(const Vector &point_pos,
}
void FindPointsGSLIB::FindPoints(Mesh &m, const Vector &point_pos,
int point_pos_ordering, const double bb_t,
const double newt_tol, const int npt_max)
const double bb_t, const double newt_tol,
const int npt_max)
{
if (!setupflag || (mesh != &m) )
{
Setup(m, bb_t, newt_tol, npt_max);
}
FindPoints(point_pos, point_pos_ordering);
FindPoints(point_pos);
}
void FindPointsGSLIB::Interpolate(const Vector &point_pos,
const GridFunction &field_in, Vector &field_out,
int point_pos_ordering)
const GridFunction &field_in, Vector &field_out)
{
FindPoints(point_pos, point_pos_ordering);
FindPoints(point_pos);
Interpolate(field_in, field_out);
}
void FindPointsGSLIB::Interpolate(Mesh &m, const Vector &point_pos,
const GridFunction &field_in, Vector &field_out,
int point_pos_ordering)
const GridFunction &field_in, Vector &field_out)
{
FindPoints(m, point_pos, point_pos_ordering);
FindPoints(m, point_pos);
Interpolate(field_in, field_out);
}
@@ -865,9 +860,7 @@ void FindPointsGSLIB::Interpolate(const GridFunction &field_in,
{
for (int i = 0; i < indl2.Size(); i++)
{
int idx = field_in.FESpace()->GetOrdering() == Ordering::byNODES ?
indl2[i] + j*points_cnt:
indl2[i]*ncomp + j;
int idx = indl2[i] + j*points_cnt;
field_out(idx) = field_out_l2(idx);
}
}
@@ -892,17 +885,7 @@ void FindPointsGSLIB::InterpolateH1(const GridFunction &field_in,
{
const int dataptrin = i*points_fld,
dataptrout = i*points_cnt;
if (field_in.FESpace()->GetOrdering() == Ordering::byNODES)
{
field_in_scalar.NewDataAndSize(field_in.GetData()+dataptrin, points_fld);
}
else
{
for (int j = 0; j < points_fld; j++)
{
field_in_scalar(j) = field_in(i + j*ncomp);
}
}
field_in_scalar.NewDataAndSize(field_in.GetData()+dataptrin, points_fld);
GetNodalValues(&field_in_scalar, node_vals);
if (dim==2)
@@ -924,17 +907,6 @@ void FindPointsGSLIB::InterpolateH1(const GridFunction &field_in,
points_cnt, node_vals.GetData(), fdata3D);
}
}
if (field_in.FESpace()->GetOrdering() == Ordering::byVDIM)
{
Vector field_out_temp = field_out;
for (int i = 0; i < ncomp; i++)
{
for (int j = 0; j < points_cnt; j++)
{
field_out(i + j*ncomp) = field_out_temp(j + i*points_cnt);
}
}
}
}
void FindPointsGSLIB::InterpolateGeneral(const GridFunction &field_in,
@@ -957,19 +929,9 @@ void FindPointsGSLIB::InterpolateGeneral(const GridFunction &field_in,
if (dim == 3) { ip.z = gsl_mfem_ref(index*dim + 2); }
Vector localval(ncomp);
field_in.GetVectorValue(gsl_mfem_elem[index], ip, localval);
if (field_in.FESpace()->GetOrdering() == Ordering::byNODES)
for (int i = 0; i < ncomp; i++)
{
for (int i = 0; i < ncomp; i++)
{
field_out(index + i*npt) = localval(i);
}
}
else //byVDIM
{
for (int i = 0; i < ncomp; i++)
{
field_out(index*ncomp + i) = localval(i);
}
field_out(index + i*npt) = localval(i);
}
}
}
@@ -1082,9 +1044,7 @@ void FindPointsGSLIB::InterpolateGeneral(const GridFunction &field_in,
sdpt = (struct send_pt *)sendpt->ptr;
for (int index = 0; index < sendpt->n; index++)
{
int idx = field_in.FESpace()->GetOrdering() == Ordering::byNODES ?
sdpt->index + j*nptorig :
sdpt->index*ncomp + j;
int idx = sdpt->index + j*nptorig;
field_out(idx) = sdpt->ival;
++sdpt;
}
@@ -1179,8 +1139,7 @@ void OversetFindPointsGSLIB::Setup(Mesh &m, const int meshid,
}
void OversetFindPointsGSLIB::FindPoints(const Vector &point_pos,
Array<unsigned int> &point_id,
int point_pos_ordering)
Array<unsigned int> &point_id)
{
MFEM_VERIFY(setupflag, "Use OversetFindPointsGSLIB::Setup before "
"finding points.");
@@ -1194,24 +1153,14 @@ void OversetFindPointsGSLIB::FindPoints(const Vector &point_pos,
gsl_ref.SetSize(points_cnt * dim);
gsl_dist.SetSize(points_cnt);
const double *xv_base[dim];
unsigned xv_stride[dim];
for (int d = 0; d < dim; d++)
{
if (point_pos_ordering == Ordering::byNODES)
{
xv_base[d] = point_pos.GetData() + d*points_cnt;
xv_stride[d] = sizeof(double);
}
else
{
xv_base[d] = point_pos.GetData() + d;
xv_stride[d] = dim*sizeof(double);
}
}
if (dim == 2)
{
const double *xv_base[2];
xv_base[0] = point_pos.GetData();
xv_base[1] = point_pos.GetData() + points_cnt;
unsigned xv_stride[2];
xv_stride[0] = sizeof(double);
xv_stride[1] = sizeof(double);
findptsms_2(gsl_code.GetData(), sizeof(unsigned int),
gsl_proc.GetData(), sizeof(unsigned int),
gsl_elem.GetData(), sizeof(unsigned int),
@@ -1223,6 +1172,14 @@ void OversetFindPointsGSLIB::FindPoints(const Vector &point_pos,
}
else
{
const double *xv_base[3];
xv_base[0] = point_pos.GetData();
xv_base[1] = point_pos.GetData() + points_cnt;
xv_base[2] = point_pos.GetData() + 2*points_cnt;
unsigned xv_stride[3];
xv_stride[0] = sizeof(double);
xv_stride[1] = sizeof(double);
xv_stride[2] = sizeof(double);
findptsms_3(gsl_code.GetData(), sizeof(unsigned int),
gsl_proc.GetData(), sizeof(unsigned int),
gsl_elem.GetData(), sizeof(unsigned int),
@@ -1251,10 +1208,9 @@ void OversetFindPointsGSLIB::FindPoints(const Vector &point_pos,
void OversetFindPointsGSLIB::Interpolate(const Vector &point_pos,
Array<unsigned int> &point_id,
const GridFunction &field_in,
Vector &field_out,
int point_pos_ordering)
Vector &field_out)
{
FindPoints(point_pos, point_id, point_pos_ordering);
FindPoints(point_pos, point_id);
Interpolate(field_in, field_out);
}
+13 -27
View File
@@ -121,9 +121,8 @@ public:
void Setup(Mesh &m, const double bb_t = 0.1,
const double newt_tol = 1.0e-12,
const int npt_max = 256);
/** Searches positions given in physical space by @a point_pos.
These positions can be ordered byNodes: (XXX...,YYY...,ZZZ) or
byVDim: (XYZ,XYZ,....XYZ) specified by @a point_pos_ordering.
/** Searches positions given in physical space by @a point_pos. These positions
must by ordered by nodes: (XXX...,YYY...,ZZZ).
This function populates the following member variables:
#gsl_code Return codes for each point: inside element (0),
element boundary (1), not found (2).
@@ -141,11 +140,9 @@ public:
Defaults to 0 for points that were not found.
#gsl_dist Distance between the sought and the found point
in physical space. */
void FindPoints(const Vector &point_pos,
int point_pos_ordering = Ordering::byNODES);
void FindPoints(const Vector &point_pos);
/// Setup FindPoints and search positions
void FindPoints(Mesh &m, const Vector &point_pos,
int point_pos_ordering = Ordering::byNODES,
const double bb_t = 0.1,
const double newt_tol = 1.0e-12, const int npt_max = 256);
@@ -157,18 +154,12 @@ public:
@param[out] field_out Interpolated values. For points that are not found
the value is set to #default_interp_value. */
virtual void Interpolate(const GridFunction &field_in, Vector &field_out);
/** Search positions and interpolate. The ordering (byNODES or byVDIM) of
the output values in @a field_out corresponds to the ordering used
in the input GridFunction @a field_in. */
/** Search positions and interpolate */
void Interpolate(const Vector &point_pos, const GridFunction &field_in,
Vector &field_out,
int point_pos_ordering = Ordering::byNODES);
/** Setup FindPoints, search positions and interpolate. The ordering (byNODES
or byVDIM) of the output values in @a field_out corresponds to the
ordering used in the input GridFunction @a field_in. */
Vector &field_out);
/** Setup FindPoints, search positions and interpolate */
void Interpolate(Mesh &m, const Vector &point_pos,
const GridFunction &field_in, Vector &field_out,
int point_pos_ordering = Ordering::byNODES);
const GridFunction &field_in, Vector &field_out);
/// Average type to be used for L2 functions in-case a point is located at
/// an element boundary where the function might be multi-valued.
@@ -256,20 +247,15 @@ public:
/** Searches positions given in physical space by @a point_pos. All output
Arrays and Vectors are expected to have the correct size.
@param[in] point_pos Positions to be found.
@param[in] point_id Index of the mesh that the point belongs
to (corresponding to @a meshid in Setup).
@param[in] point_pos_ordering Ordering of the points:
byNodes: (XXX...,YYY...,ZZZ) or
byVDim: (XYZ,XYZ,....XYZ) */
void FindPoints(const Vector &point_pos,
Array<unsigned int> &point_id,
int point_pos_ordering = Ordering::byNODES);
@param[in] point_pos Positions to be found. Must by ordered by nodes
(XXX...,YYY...,ZZZ).
@param[in] point_id Index of the mesh that the point belongs to
(corresponding to @a meshid in Setup). */
void FindPoints(const Vector &point_pos, Array<unsigned int> &point_id);
/** Search positions and interpolate */
void Interpolate(const Vector &point_pos, Array<unsigned int> &point_id,
const GridFunction &field_in, Vector &field_out,
int point_pos_ordering = Ordering::byNODES);
const GridFunction &field_in, Vector &field_out);
using FindPointsGSLIB::Interpolate;
};
+1
View File
@@ -827,6 +827,7 @@ void Hybridization::ReduceRHS(const Vector &b, Vector &b_r) const
}
else
{
Ct->EnsureMultTranspose();
Ct->MultTranspose(bf, bl);
}
b_r.SetSize(pH.Ptr()->Height());
+9 -17
View File
@@ -120,22 +120,6 @@ MFEM_HOST_DEVICE inline void LoadBGt(const int D1D, const int Q1D,
MFEM_SYNC_THREAD;
}
/// Load 2D input scalar into given DeviceMatrix
MFEM_HOST_DEVICE inline void LoadX(const int e, const int D1D,
const DeviceTensor<3, const double> &x,
DeviceMatrix &DD)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
DD(dx,dy) = x(dx,dy,e);
}
}
MFEM_SYNC_THREAD;
}
/// Load 2D input scalar into shared memory
template<int MD1, int NBZ>
MFEM_HOST_DEVICE inline void LoadX(const int e, const int D1D,
@@ -144,7 +128,15 @@ MFEM_HOST_DEVICE inline void LoadX(const int e, const int D1D,
{
const int tidz = MFEM_THREAD_ID(z);
DeviceMatrix X(sX[tidz], D1D, D1D);
LoadX(e, D1D, x, X);
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
X(dx,dy) = x(dx,dy,e);
}
}
MFEM_SYNC_THREAD;
}
/// Load 2D input scalar into shared memory, with comp
+25 -51
View File
@@ -19,14 +19,13 @@ namespace mfem
LinearForm::LinearForm(FiniteElementSpace *f, LinearForm *lf)
: Vector(f->GetVSize())
{
ext = nullptr;
extern_lfs = 1;
fast_assembly = false;
fes = f;
// Linear forms are stored on the device
UseDevice(true);
fes = f;
ext = nullptr;
extern_lfs = 1;
// Copy the pointers to the integrators
domain_integs = lf->domain_integs;
@@ -103,45 +102,25 @@ void LinearForm::AddInteriorFaceIntegrator(LinearFormIntegrator *lfi)
bool LinearForm::SupportsDevice()
{
// return false for NURBS meshes, so we dont convert it to non-NURBS
// return false for NURBS meshs, so we dont convert it to non-NURBS
// through Assemble, AssembleDevice, GetGeometricFactors and EnsureNodes
const Mesh &mesh = *fes->GetMesh();
if (mesh.NURBSext != nullptr) { return false; }
if (fes->GetMesh()->NURBSext != nullptr) { return false; }
// scan integrators to verify that all can use device assembly
auto IntegratorsSupportDevice = [](const Array<LinearFormIntegrator*> &integ)
// scan domain integrator to verify that all can use device assembly
if (domain_integs.Size() > 0)
{
for (int k = 0; k < integ.Size(); k++)
for (int k = 0; k < domain_integs.Size(); k++)
{
if (!integ[k]->SupportsDevice()) { return false; }
}
return true;
};
if (!IntegratorsSupportDevice(domain_integs)) { return false; }
if (!IntegratorsSupportDevice(boundary_integs)) { return false; }
if (boundary_face_integs.Size() > 0 || interior_face_integs.Size() > 0 ||
domain_delta_integs.Size() > 0) { return false; }
if (boundary_integs.Size() > 0)
{
// Make sure there are no boundary faces that are not boundary elements
if (fes->GetNFbyType(FaceType::Boundary) != fes->GetNBE())
{
return false;
}
// Make sure every boundary element corresponds to a boundary face
for (int be = 0; be < fes->GetNBE(); ++be)
{
const int f = mesh.GetBdrElementEdgeIndex(be);
const auto face_info = mesh.GetFaceInformation(f);
if (!face_info.IsBoundary())
{
return false;
}
if (!domain_integs[k]->SupportsDevice()) { return false; }
}
}
// boundary, delta and face integrators are not supported yet
if (GetBLFI()->Size() > 0 || GetFLFI()->Size() > 0 ||
GetDLFI_Delta()->Size() > 0 || GetIFLFI()->Size() > 0) { return false; }
const Mesh &mesh = *fes->GetMesh();
// no support for elements with varying polynomial orders
if (fes->IsVariableOrder()) { return false; }
@@ -155,30 +134,25 @@ bool LinearForm::SupportsDevice()
return true;
}
void LinearForm::UseFastAssembly(bool use_fa)
{
fast_assembly = use_fa;
if (fast_assembly && SupportsDevice() && !ext)
{
ext = new LinearFormExtension(this);
}
}
void LinearForm::Assemble()
void LinearForm::Assemble(bool use_device)
{
Array<int> vdofs;
ElementTransformation *eltrans;
DofTransformation *doftrans;
Vector elemvect;
if (!ext && use_device && SupportsDevice())
{
ext = new LinearFormExtension(this);
}
Vector::operator=(0.0);
// The above operation is executed on device because of UseDevice().
// The first use of AddElementVector() below will move it back to host
// because both 'vdofs' and 'elemvect' are on host.
if (fast_assembly && ext) { return ext->Assemble(); }
if (ext) { return ext->Assemble(); }
if (domain_integs.Size())
{
@@ -199,8 +173,8 @@ void LinearForm::Assemble()
int elem_attr = fes->GetMesh()->GetAttribute(i);
for (int k = 0; k < domain_integs.Size(); k++)
{
const Array<int> * const markers = domain_integs_marker[k];
if ( markers == NULL || (*markers)[elem_attr-1] == 1 )
if ( domain_integs_marker[k] == NULL ||
(*(domain_integs_marker[k]))[elem_attr-1] == 1 )
{
doftrans = fes -> GetElementVDofs (i, vdofs);
eltrans = fes -> GetElementTransformation (i);
+9 -21
View File
@@ -30,16 +30,12 @@ protected:
FiniteElementSpace *fes;
/** @brief Extension for supporting different assembly levels. */
LinearFormExtension *ext = nullptr;
/// @brief Should we use the device-compatible fast assembly algorithm (false
/// by default)
bool fast_assembly = false;
LinearFormExtension *ext;
/** @brief Indicates the LinearFormIntegrator%s stored in #domain_integs,
#domain_delta_integs, #boundary_integs, and #boundary_face_integs are
owned by another LinearForm. */
int extern_lfs = 0;
int extern_lfs;
/// Set of Domain Integrators to be applied.
Array<LinearFormIntegrator*> domain_integs;
@@ -86,7 +82,7 @@ public:
/// Creates linear form associated with FE space @a *f.
/** The pointer @a f is not owned by the newly constructed object. */
LinearForm(FiniteElementSpace *f) : Vector(f->GetVSize())
{ fes = f; UseDevice(true); }
{ fes = f; ext = nullptr; extern_lfs = 0; UseDevice(true); }
/** @brief Create a LinearForm on the FiniteElementSpace @a f, using the
same integrators as the LinearForm @a lf.
@@ -101,8 +97,7 @@ public:
/** The associated FiniteElementSpace can be set later using one of the
methods: Update(FiniteElementSpace *) or
Update(FiniteElementSpace *, Vector &, int). */
LinearForm()
{ fes = NULL; UseDevice(true); }
LinearForm() { fes = NULL; ext = nullptr; extern_lfs = 0; UseDevice(true); }
/// Construct a LinearForm using previously allocated array @a data.
/** The LinearForm does not assume ownership of @a data which is assumed to
@@ -110,7 +105,7 @@ public:
for externally allocated array, the pointer @a data can be NULL. The data
array can be replaced later using the method SetData(). */
LinearForm(FiniteElementSpace *f, double *data) : Vector(data, f->GetVSize())
{ fes = f; }
{ fes = f; ext = nullptr; extern_lfs = 0; }
/// Copy assignment. Only the data of the base class Vector is copied.
/** It is assumed that this object and @a rhs use FiniteElementSpace%s that
@@ -190,20 +185,13 @@ public:
corresponding pointer (to Array<int>) will be NULL. */
Array<Array<int>*> *GetFLFI_Marker() { return &boundary_face_integs_marker; }
/// @brief Which assembly algorithm to use: the new device-compatible fast
/// assembly (true), or the legacy CPU-only algorithm (false).
/** If not set, the default value is false. If used, this method must be
called before assembly. */
void UseFastAssembly(bool use_fa);
/// Assembles the linear form i.e. sums over all domain/bdr integrators.
/** When @ref UseFastAssembly "UseFastAssembly(true)" has been called and the
linearform assembly is compatible with device execution, it will be
executed on the device. */
void Assemble();
/// When @a use_device is set to true and the linearform assembly is
/// compatible with device execution, it will be executed on the device.
void Assemble(bool use_device = true);
/// Return true if assembly on device is supported, false otherwise.
virtual bool SupportsDevice();
bool SupportsDevice();
/// Assembles delta functions of the linear form
void AssembleDelta();
+14 -97
View File
@@ -25,8 +25,7 @@ void LinearFormExtension::Assemble()
"match the number of vector dofs!");
const Array<Array<int>*> &domain_integs_marker = *lf->GetDLFI_Marker();
const int mesh_attributes_max = fes.GetMesh()->attributes.Size() ?
fes.GetMesh()->attributes.Max() : 0;
const int mesh_attributes_size = fes.GetMesh()->attributes.Size();
const Array<LinearFormIntegrator*> &domain_integs = *lf->GetDLFI();
for (int k = 0; k < domain_integs.Size(); ++k)
@@ -40,7 +39,7 @@ void LinearFormExtension::Assemble()
if (has_markers_k)
{
// Element attribute marker should be of length mesh->attributes
MFEM_VERIFY(mesh_attributes_max == domain_integs_marker_k->Size(),
MFEM_VERIFY(mesh_attributes_size == domain_integs_marker_k->Size(),
"invalid element marker for domain linear form "
"integrator #" << k << ", counting from zero");
}
@@ -60,47 +59,7 @@ void LinearFormExtension::Assemble()
// Assemble the linear form
b = 0.0;
domain_integs[k]->AssembleDevice(fes, markers, b);
if (k == 0) { elem_restrict_lex->MultTranspose(b, *lf); }
else { elem_restrict_lex->AddMultTranspose(b, *lf); }
}
const Array<Array<int>*> &boundary_integs_marker = lf->boundary_integs_marker;
const int bdr_attributes_max = fes.GetMesh()->bdr_attributes.Size() ?
fes.GetMesh()->bdr_attributes.Max() : 0;
const Array<LinearFormIntegrator*> &boundary_integs = lf->boundary_integs;
for (int k = 0; k < boundary_integs.Size(); ++k)
{
// Get the markers for this integrator
const Array<int> *boundary_integs_marker_k = boundary_integs_marker[k];
// check if there are markers for this integrator
const bool has_markers_k = boundary_integs_marker_k != nullptr;
if (has_markers_k)
{
// Element attribute marker should be of length mesh->attributes
MFEM_VERIFY(bdr_attributes_max == boundary_integs_marker_k->Size(),
"invalid boundary marker for boundary linear form "
"integrator #" << k << ", counting from zero");
}
// if there are no markers, just use the whole linear form (1)
if (!has_markers_k) { bdr_markers.HostReadWrite(); bdr_markers = 1; }
else
{
// scan the attributes to set the markers to 0 or 1
const int NBE = bdr_attributes.Size();
const auto attr = bdr_attributes.Read();
const auto attr_markers = boundary_integs_marker_k->Read();
auto markers_w = bdr_markers.Write();
MFEM_FORALL(e, NBE, markers_w[e] = attr_markers[attr[e]-1] == 1;);
}
// Assemble the linear form
bdr_b = 0.0;
boundary_integs[k]->AssembleDevice(fes, bdr_markers, bdr_b);
bdr_restrict_lex->AddMultTranspose(bdr_b, *lf);
elem_restrict_lex->MultTranspose(b, *lf);
}
}
@@ -108,64 +67,22 @@ void LinearFormExtension::Update()
{
const FiniteElementSpace &fes = *lf->FESpace();
const Mesh &mesh = *fes.GetMesh();
constexpr ElementDofOrdering ordering = ElementDofOrdering::LEXICOGRAPHIC;
const int NE = fes.GetNE();
MFEM_VERIFY(lf->Size() == fes.GetVSize(), "");
if (lf->domain_integs.Size() > 0)
{
const int NE = fes.GetNE();
markers.SetSize(NE);
//markers.UseDevice(true);
markers.SetSize(NE);
//markers.UseDevice(true);
// Gather the attributes on the host from all the elements
attributes.SetSize(NE);
for (int i = 0; i < NE; ++i) { attributes[i] = mesh.GetAttribute(i); }
// Gather the attributes on the host from all the elements
attributes.SetSize(NE);
for (int i = 0; i < NE; ++i) { attributes[i] = mesh.GetAttribute(i); }
elem_restrict_lex = fes.GetElementRestriction(ordering);
MFEM_VERIFY(elem_restrict_lex, "Element restriction not available");
b.SetSize(elem_restrict_lex->Height(), Device::GetMemoryType());
b.UseDevice(true);
}
if (lf->boundary_integs.Size() > 0)
{
const int nf_bdr = fes.GetNFbyType(FaceType::Boundary);
bdr_markers.SetSize(nf_bdr);
// bdr_markers.UseDevice(true);
// The face restriction will give us "face E-vectors" on the boundary that
// are numbered in the order of the faces of mesh. This numbering will be
// different than the numbering of the boundary elements. We compute
// mappings so that the array `bdr_attributes[i]` gives the boundary
// attribute of the `i`th boundary face in the mesh face order.
std::unordered_map<int,int> f_to_be;
for (int i = 0; i < mesh.GetNBE(); ++i)
{
const int f = mesh.GetBdrElementEdgeIndex(i);
f_to_be[f] = i;
}
MFEM_VERIFY(size_t(nf_bdr) == f_to_be.size(), "Incompatible sizes");
bdr_attributes.SetSize(nf_bdr);
int f_ind = 0;
for (int f = 0; f < mesh.GetNumFaces(); ++f)
{
if (f_to_be.find(f) != f_to_be.end())
{
const int be = f_to_be[f];
bdr_attributes[f_ind] = mesh.GetBdrAttribute(be);
++f_ind;
}
}
bdr_restrict_lex =
dynamic_cast<const FaceRestriction*>(
fes.GetFaceRestriction(ordering, FaceType::Boundary,
L2FaceValues::SingleValued));
MFEM_VERIFY(bdr_restrict_lex, "Face restriction not available");
bdr_b.SetSize(bdr_restrict_lex->Height(), Device::GetMemoryType());
bdr_b.UseDevice(true);
}
constexpr ElementDofOrdering ordering = ElementDofOrdering::LEXICOGRAPHIC;
elem_restrict_lex = fes.GetElementRestriction(ordering);
MFEM_VERIFY(elem_restrict_lex, "Element restriction not available");
b.SetSize(elem_restrict_lex->Height(), Device::GetMemoryType());
b.UseDevice(true);
}
} // namespace mfem
+4 -7
View File
@@ -25,22 +25,19 @@ class LinearForm;
class LinearFormExtension
{
/// Attributes of all mesh elements.
Array<int> attributes, bdr_attributes;
Array<int> attributes;
/// Temporary markers for device kernels.
Array<int> markers, bdr_markers;
Array<int> markers;
/// Linear form from which this extension depends. Not owned.
LinearForm *lf;
/// Operator that converts FiniteElementSpace L-vectors to E-vectors.
const ElementRestrictionOperator *elem_restrict_lex; // Not owned
/// Operator that converts L-vectors to boundary E-vectors.
const FaceRestriction *bdr_restrict_lex; // Not owned
const Operator *elem_restrict_lex; // Not owned
/// Internal E-vectors.
mutable Vector b, bdr_b;
mutable Vector b;
public:
+2 -2
View File
@@ -1046,7 +1046,7 @@ void VectorQuadratureLFIntegrator::AssembleRHSElementVect(
const FiniteElement &fe, ElementTransformation &Tr, Vector &elvect)
{
const IntegrationRule *ir =
&vqfc.GetQuadFunction().GetSpace()->GetIntRule(Tr.ElementNo);
&vqfc.GetQuadFunction().GetSpace()->GetElementIntRule(Tr.ElementNo);
const int nqp = ir->GetNPoints();
const int vdim = vqfc.GetVDim();
@@ -1078,7 +1078,7 @@ void QuadratureLFIntegrator::AssembleRHSElementVect(const FiniteElement &fe,
Vector &elvect)
{
const IntegrationRule *ir =
&qfc.GetQuadFunction().GetSpace()->GetIntRule(Tr.ElementNo);
&qfc.GetQuadFunction().GetSpace()->GetElementIntRule(Tr.ElementNo);
const int nqp = ir->GetNPoints();
const int ndofs = fe.GetDof();
-14
View File
@@ -187,13 +187,6 @@ public:
BoundaryLFIntegrator(Coefficient &QG, int a = 1, int b = 1)
: Q(QG), oa(a), ob(b) { }
virtual bool SupportsDevice() { return true; }
/// Method defining assembly on device
virtual void AssembleDevice(const FiniteElementSpace &fes,
const Array<int> &markers,
Vector &b);
/** Given a particular boundary Finite Element and a transformation (Tr)
computes the element boundary vector, elvect. */
virtual void AssembleRHSElementVect(const FiniteElement &el,
@@ -217,13 +210,6 @@ public:
BoundaryNormalLFIntegrator(VectorCoefficient &QG, int a = 1, int b = 1)
: Q(QG), oa(a), ob(b) { }
virtual bool SupportsDevice() { return true; }
/// Method defining assembly on device
virtual void AssembleDevice(const FiniteElementSpace &fes,
const Array<int> &markers,
Vector &b);
virtual void AssembleRHSElementVect(const FiniteElement &el,
ElementTransformation &Tr,
Vector &elvect);
-241
View File
@@ -1,241 +0,0 @@
// Copyright (c) 2010-2022, Lawrence Livermore National Security, LLC. Produced
// at the Lawrence Livermore National Laboratory. All Rights reserved. See files
// LICENSE and NOTICE for details. LLNL-CODE-806117.
//
// This file is part of the MFEM library. For more information and source code
// availability visit https://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the BSD-3 license. We welcome feedback and contributions, see file
// CONTRIBUTING.md for details.
#include "fem.hpp"
#include "../fem/kernels.hpp"
#include "../general/forall.hpp"
namespace mfem
{
template<int T_D1D = 0, int T_Q1D = 0> static
void BLFEvalAssemble2D(const int vdim, const int nbe, const int d, const int q,
const bool normals, const int *markers, const double *b,
const double *detj, const double *n, const double *weights,
const Vector &coeff, double *y)
{
const auto F = coeff.Read();
const auto M = Reshape(markers, nbe);
const auto B = Reshape(b, q, d);
const auto detJ = Reshape(detj, q, nbe);
const auto N = Reshape(n, q, 2, nbe);
const auto W = Reshape(weights, q);
const int cvdim = normals ? 2 : 1;
const bool cst = coeff.Size() == cvdim;
const auto C = cst ? Reshape(F,cvdim,1,1) : Reshape(F,cvdim,q,nbe);
auto Y = Reshape(y, d, vdim, nbe);
MFEM_FORALL(e, nbe,
{
if (M(e) == 0) { return; } // ignore
constexpr int Q = T_Q1D ? T_Q1D : MAX_Q1D;
double QQ[Q];
for (int c = 0; c < vdim; ++c)
{
for (int qx = 0; qx < q; ++qx)
{
double coeff_val = 0.0;
if (normals)
{
for (int cd = 0; cd < 2; ++cd)
{
const double cval = cst ? C(cd,0,0) : C(cd,qx,e);
coeff_val += cval * N(qx, cd, e);
}
}
else
{
coeff_val = cst ? C(0,0,0) : C(0,qx,e);
}
QQ[qx] = W(qx) * coeff_val * detJ(qx,e);
}
for (int dx = 0; dx < d; ++dx)
{
double u = 0;
for (int qx = 0; qx < q; ++qx) { u += QQ[qx] * B(qx,dx); }
Y(dx,c,e) += u;
}
}
});
}
template<int T_D1D = 0, int T_Q1D = 0> static
void BLFEvalAssemble3D(const int vdim, const int nbe, const int d, const int q,
const bool normals, const int *markers, const double *b,
const double *detj, const double *n, const double *weights,
const Vector &coeff, double *y)
{
const auto F = coeff.Read();
const auto M = Reshape(markers, nbe);
const auto B = Reshape(b, q, d);
const auto detJ = Reshape(detj, q, q, nbe);
const auto N = Reshape(n, q, q, 3, nbe);
const auto W = Reshape(weights, q, q);
const int cvdim = normals ? 3 : 1;
const bool cst = coeff.Size() == cvdim;
const auto C = cst ? Reshape(F,cvdim,1,1,1) : Reshape(F,cvdim,q,q,nbe);
auto Y = Reshape(y, d, d, vdim, nbe);
MFEM_FORALL_2D(e, nbe, q, q, 1,
{
if (M(e) == 0) { return; } // ignore
constexpr int Q = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int D = T_D1D ? T_D1D : MAX_D1D;
MFEM_SHARED double sBt[Q*D];
MFEM_SHARED double sQQ[Q*Q];
MFEM_SHARED double sQD[Q*D];
const DeviceMatrix Bt(sBt, d, q);
kernels::internal::LoadB<D,Q>(d, q, B, sBt);
const DeviceMatrix QQ(sQQ, q, q);
const DeviceMatrix QD(sQD, q, d);
for (int c = 0; c < vdim; ++c)
{
MFEM_FOREACH_THREAD(x,x,q)
{
MFEM_FOREACH_THREAD(y,y,q)
{
double coeff_val = 0.0;
if (normals)
{
for (int cd = 0; cd < 3; ++cd)
{
double cval = cst ? C(cd,0,0,0) : C(cd,x,y,e);
coeff_val += cval * N(x,y,cd,e);
}
}
else
{
coeff_val = cst ? C(0,0,0,0) : C(0,x,y,e);
}
QQ(y,x) = W(x,y) * coeff_val * detJ(x,y,e);
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,q)
{
MFEM_FOREACH_THREAD(dx,x,d)
{
double u = 0.0;
for (int qx = 0; qx < q; ++qx) { u += QQ(qy,qx) * Bt(dx,qx); }
QD(qy,dx) = u;
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,d)
{
MFEM_FOREACH_THREAD(dx,x,d)
{
double u = 0.0;
for (int qy = 0; qy < q; ++qy) { u += QD(qy,dx) * Bt(dy,qy); }
Y(dx,dy,c,e) += u;
}
}
MFEM_SYNC_THREAD;
}
});
}
static void BLFEvalAssemble(const FiniteElementSpace &fes,
const IntegrationRule &ir,
const Array<int> &markers,
const Vector &coeff,
const bool normals,
Vector &y)
{
Mesh &mesh = *fes.GetMesh();
const int dim = mesh.Dimension();
const FiniteElement &el = *fes.GetBE(0);
const MemoryType mt = Device::GetDeviceMemoryType();
const DofToQuad &maps = el.GetDofToQuad(ir, DofToQuad::TENSOR);
const int d = maps.ndof, q = maps.nqpt;
int flags = FaceGeometricFactors::DETERMINANTS;
if (normals) { flags |= FaceGeometricFactors::NORMALS; }
const FaceGeometricFactors *geom = mesh.GetFaceGeometricFactors(
ir, flags, FaceType::Boundary, mt);
auto ker = (dim == 2) ? BLFEvalAssemble2D<> : BLFEvalAssemble3D<>;
if (dim==2)
{
if (d==1 && q==1) { ker=BLFEvalAssemble2D<1,1>; }
if (d==2 && q==2) { ker=BLFEvalAssemble2D<2,2>; }
if (d==3 && q==3) { ker=BLFEvalAssemble2D<3,3>; }
if (d==4 && q==4) { ker=BLFEvalAssemble2D<4,4>; }
if (d==5 && q==5) { ker=BLFEvalAssemble2D<5,5>; }
if (d==2 && q==3) { ker=BLFEvalAssemble2D<2,3>; }
if (d==3 && q==4) { ker=BLFEvalAssemble2D<3,4>; }
if (d==4 && q==5) { ker=BLFEvalAssemble2D<4,5>; }
if (d==5 && q==6) { ker=BLFEvalAssemble2D<5,6>; }
}
if (dim==3)
{
if (d==1 && q==1) { ker=BLFEvalAssemble3D<1,1>; }
if (d==2 && q==2) { ker=BLFEvalAssemble3D<2,2>; }
if (d==3 && q==3) { ker=BLFEvalAssemble3D<3,3>; }
if (d==4 && q==4) { ker=BLFEvalAssemble3D<4,4>; }
if (d==5 && q==5) { ker=BLFEvalAssemble3D<5,5>; }
if (d==2 && q==3) { ker=BLFEvalAssemble3D<2,3>; }
if (d==3 && q==4) { ker=BLFEvalAssemble3D<3,4>; }
if (d==4 && q==5) { ker=BLFEvalAssemble3D<4,5>; }
if (d==5 && q==6) { ker=BLFEvalAssemble3D<5,6>; }
}
MFEM_VERIFY(ker, "No kernel ndof " << d << " nqpt " << q);
const int vdim = fes.GetVDim();
const int nbe = fes.GetMesh()->GetNFbyType(FaceType::Boundary);
const int *M = markers.Read();
const double *B = maps.B.Read();
const double *detJ = geom->detJ.Read();
const double *n = geom->normal.Read();
const double *W = ir.GetWeights().Read();
double *Y = y.ReadWrite();
ker(vdim, nbe, d, q, normals, M, B, detJ, n, W, coeff, Y);
}
void BoundaryLFIntegrator::AssembleDevice(const FiniteElementSpace &fes,
const Array<int> &markers,
Vector &b)
{
const FiniteElement &fe = *fes.GetBE(0);
const int qorder = oa * fe.GetOrder() + ob;
const Geometry::Type gtype = fe.GetGeomType();
const IntegrationRule &ir = IntRule ? *IntRule : IntRules.Get(gtype, qorder);
Mesh &mesh = *fes.GetMesh();
FaceQuadratureSpace qs(mesh, ir, FaceType::Boundary);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
BLFEvalAssemble(fes, ir, markers, coeff, false, b);
}
void BoundaryNormalLFIntegrator::AssembleDevice(const FiniteElementSpace &fes,
const Array<int> &markers,
Vector &b)
{
const FiniteElement &fe = *fes.GetBE(0);
const int qorder = oa * fe.GetOrder() + ob;
const Geometry::Type gtype = fe.GetGeomType();
const IntegrationRule &ir = IntRule ? *IntRule : IntRules.Get(gtype, qorder);
Mesh &mesh = *fes.GetMesh();
FaceQuadratureSpace qs(mesh, ir, FaceType::Boundary);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
BLFEvalAssemble(fes, ir, markers, coeff, true, b);
}
} // namespace mfem
+112 -14
View File
@@ -19,13 +19,13 @@ namespace mfem
template<int T_D1D = 0, int T_Q1D = 0> static
void DLFEvalAssemble2D(const int vdim, const int ne, const int d, const int q,
const int map_type, const int *markers, const double *b,
const double *detj, const double *weights,
const double *j, const double *weights,
const Vector &coeff, double *y)
{
const auto F = coeff.Read();
const auto M = Reshape(markers, ne);
const auto B = Reshape(b, q, d);
const auto DETJ = Reshape(detj, q, q, ne);
const auto J = Reshape(j, q, q, 2,2, ne);
const auto W = Reshape(weights, q, q);
const bool cst = coeff.Size() == vdim;
const auto C = cst ? Reshape(F,vdim,1,1,1) : Reshape(F,vdim,q,q,ne);
@@ -55,7 +55,19 @@ void DLFEvalAssemble2D(const int vdim, const int ne, const int d, const int q,
{
MFEM_FOREACH_THREAD(y,y,q)
{
const double detJ = (map_type == FiniteElement::VALUE) ? DETJ(x,y,e) : 1.0;
double detJ;
if (map_type == FiniteElement::VALUE)
{
const double J11 = J(x,y,0,0,e);
const double J21 = J(x,y,1,0,e);
const double J12 = J(x,y,0,1,e);
const double J22 = J(x,y,1,1,e);
detJ = J11 * J22 - J21 * J12;
}
else
{
detJ = 1.0;
}
const double coeff_val = cst ? cst_val : C(c,x,y,e);
QQ(y,x) = W(x,y) * coeff_val * detJ;
}
@@ -88,13 +100,13 @@ void DLFEvalAssemble2D(const int vdim, const int ne, const int d, const int q,
template<int T_D1D = 0, int T_Q1D = 0> static
void DLFEvalAssemble3D(const int vdim, const int ne, const int d, const int q,
const int map_type, const int *markers, const double *b,
const double *detj, const double *weights,
const double *j, const double *weights,
const Vector &coeff, double *y)
{
const auto F = coeff.Read();
const auto M = Reshape(markers, ne);
const auto B = Reshape(b, q,d);
const auto DETJ = Reshape(detj, q, q, q, ne);
const auto J = Reshape(j, q,q,q, 3,3, ne);
const auto W = Reshape(weights, q,q,q);
const bool cst_coeff = coeff.Size() == vdim;
const auto C = cst_coeff ? Reshape(F,vdim,1,1,1,1):Reshape(F,vdim,q,q,q,ne);
@@ -126,7 +138,26 @@ void DLFEvalAssemble3D(const int vdim, const int ne, const int d, const int q,
{
for (int z = 0; z < q; ++z)
{
const double detJ = (map_type == FiniteElement::VALUE) ? DETJ(x,y,z,e) : 1.0;
double detJ;
if (map_type == FiniteElement::VALUE)
{
const double J11 = J(x,y,z,0,0,e);
const double J21 = J(x,y,z,1,0,e);
const double J31 = J(x,y,z,2,0,e);
const double J12 = J(x,y,z,0,1,e);
const double J22 = J(x,y,z,1,1,e);
const double J32 = J(x,y,z,2,1,e);
const double J13 = J(x,y,z,0,2,e);
const double J23 = J(x,y,z,1,2,e);
const double J33 = J(x,y,z,2,2,e);
detJ = J11 * (J22 * J33 - J32 * J23) -
/* */ J21 * (J12 * J33 - J32 * J13) +
/* */ J31 * (J12 * J23 - J22 * J13);
}
else
{
detJ = 1.0;
}
const double coeff_val = cst_coeff ? cst_val : C(c,x,y,z,e);
QQQ(z,y,x) = W(x,y,z) * coeff_val * detJ;
}
@@ -191,7 +222,7 @@ static void DLFEvalAssemble(const FiniteElementSpace &fes,
const MemoryType mt = Device::GetDeviceMemoryType();
const DofToQuad &maps = el.GetDofToQuad(*ir, DofToQuad::TENSOR);
const int d = maps.ndof, q = maps.nqpt;
constexpr int flags = GeometricFactors::DETERMINANTS;
constexpr int flags = GeometricFactors::JACOBIANS;
const GeometricFactors *geom = mesh->GetGeometricFactors(*ir, flags, mt);
const int map_type = fes.GetFE(0)->GetMapType();
decltype(&DLFEvalAssemble2D<>) ker =
@@ -229,10 +260,10 @@ static void DLFEvalAssemble(const FiniteElementSpace &fes,
const int ne = fes.GetMesh()->GetNE();
const int *M = markers.Read();
const double *B = maps.B.Read();
const double *detJ = geom->detJ.Read();
const double *J = geom->J.Read();
const double *W = ir->GetWeights().Read();
double *Y = y.ReadWrite();
ker(vdim, ne, d, q, map_type, M, B, detJ, W, coeff, Y);
ker(vdim, ne, d, q, map_type, M, B, J, W, coeff, Y);
}
void DomainLFIntegrator::AssembleDevice(const FiniteElementSpace &fes,
@@ -243,9 +274,42 @@ void DomainLFIntegrator::AssembleDevice(const FiniteElementSpace &fes,
const int qorder = oa * fe.GetOrder() + ob;
const Geometry::Type gtype = fe.GetGeomType();
const IntegrationRule *ir = IntRule ? IntRule : &IntRules.Get(gtype, qorder);
const int nq = ir->GetNPoints(), ne = fes.GetMesh()->GetNE();
QuadratureSpace qs(*fes.GetMesh(), *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
Vector coeff;
if (ConstantCoefficient *cQ =
dynamic_cast<ConstantCoefficient*>(&Q))
{
coeff.SetSize(1);
coeff(0) = cQ->constant;
}
else if (QuadratureFunctionCoefficient *qfQ =
dynamic_cast<QuadratureFunctionCoefficient*>(&Q))
{
const QuadratureFunction &qfun = qfQ->GetQuadFunction();
MFEM_VERIFY(qfun.Size() == fes.GetVDim()*ne*nq,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qfun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different.\n");
qfun.Read();
coeff.MakeRef(const_cast<QuadratureFunction&>(qfun),0);
}
else
{
coeff.SetSize(nq * ne);
auto C = Reshape(coeff.HostWrite(), nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation& Tr = *fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
const IntegrationPoint &ip = ir->IntPoint(q);
Tr.SetIntPoint(&ip);
C(q,e) = Q.Eval(Tr, ip);
}
}
}
DLFEvalAssemble(fes, ir, markers, coeff, b);
}
@@ -253,14 +317,48 @@ void VectorDomainLFIntegrator::AssembleDevice(const FiniteElementSpace &fes,
const Array<int> &markers,
Vector &b)
{
const int vdim = fes.GetVDim();
const FiniteElement &fe = *fes.GetFE(0);
const int qorder = 2 * fe.GetOrder();
const Geometry::Type gtype = fe.GetGeomType();
const IntegrationRule *ir = IntRule ? IntRule : &IntRules.Get(gtype, qorder);
const int nq = ir->GetNPoints(), ne = fes.GetMesh()->GetNE();
QuadratureSpace qs(*fes.GetMesh(), *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
DLFEvalAssemble(fes, ir, markers, coeff, b);
if (VectorConstantCoefficient *vcQ =
dynamic_cast<VectorConstantCoefficient*>(&Q))
{
Qvec = vcQ->GetVec();
}
else if (VectorQuadratureFunctionCoefficient *vQ =
dynamic_cast<VectorQuadratureFunctionCoefficient*>(&Q))
{
const QuadratureFunction &qfun = vQ->GetQuadFunction();
MFEM_VERIFY(qfun.Size() == vdim*ne*nq,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qfun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different.\n");
qfun.Read();
Qvec.MakeRef(const_cast<QuadratureFunction&>(qfun),0);
}
else
{
Vector qv(vdim);
Qvec.SetSize(vdim * nq * ne);
auto C = Reshape(Qvec.HostWrite(), vdim, nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation& Tr = *fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
const IntegrationPoint &ip = ir->IntPoint(q);
Tr.SetIntPoint(&ip);
Q.Eval(qv, Tr, ip);
for (int c=0; c<vdim; ++c) { C(c,q,e) = qv[c]; }
}
}
}
DLFEvalAssemble(fes, ir, markers, Qvec, b);
}
} // namespace mfem
+90 -6
View File
@@ -324,24 +324,108 @@ void DomainLFGradIntegrator::AssembleDevice(const FiniteElementSpace &fes,
const int qorder = 2 * fe.GetOrder();
const Geometry::Type gtype = fe.GetGeomType();
const IntegrationRule *ir = IntRule ? IntRule : &IntRules.Get(gtype, qorder);
const int nq = ir->GetNPoints(), ne = fes.GetMesh()->GetNE();
QuadratureSpace qs(*fes.GetMesh(), *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
DLFGradAssemble(fes, ir, markers, coeff, b);
if (VectorConstantCoefficient *vcQ =
dynamic_cast<VectorConstantCoefficient*>(&Q))
{
Qvec = vcQ->GetVec();
}
else if (VectorQuadratureFunctionCoefficient *vqfQ =
dynamic_cast<VectorQuadratureFunctionCoefficient*>(&Q))
{
const QuadratureFunction &qfun = vqfQ->GetQuadFunction();
MFEM_VERIFY(qfun.Size() == ne*nq,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qfun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different.\n");
qfun.Read();
Qvec.MakeRef(const_cast<QuadratureFunction&>(qfun),0);
}
else
{
const int qvdim = Q.GetVDim();
Vector qvec(qvdim);
Qvec.SetSize(qvdim * nq * ne);
auto C = Reshape(Qvec.HostWrite(), qvdim, nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation& Tr = *fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
const IntegrationPoint &ip = ir->IntPoint(q);
Tr.SetIntPoint(&ip);
Q.Eval(qvec, Tr, ip);
for (int c=0; c < qvdim; ++c)
{
C(c,q,e) = qvec[c];
}
}
}
}
DLFGradAssemble(fes, ir, markers, Qvec, b);
}
void VectorDomainLFGradIntegrator::AssembleDevice(const FiniteElementSpace &fes,
const Array<int> &markers,
Vector &b)
{
const int vdim = fes.GetVDim();
const FiniteElement &fe = *fes.GetFE(0);
const int qorder = 2 * fe.GetOrder();
const Geometry::Type gtype = fe.GetGeomType();
const IntegrationRule *ir = IntRule ? IntRule : &IntRules.Get(gtype, qorder);
const int nq = ir->GetNPoints(), ne = fes.GetMesh()->GetNE(),
ns = fes.GetMesh()->SpaceDimension();
QuadratureSpace qs(*fes.GetMesh(), *ir);
CoefficientVector coeff(Q, qs, CoefficientStorage::COMPRESSED);
DLFGradAssemble(fes, ir, markers, coeff, b);
if (VectorConstantCoefficient *vcQ =
dynamic_cast<VectorConstantCoefficient*>(&Q))
{
Qvec = vcQ->GetVec();
}
else if (QuadratureFunctionCoefficient *qfQ =
dynamic_cast<QuadratureFunctionCoefficient*>(&Q))
{
const QuadratureFunction &qfun = qfQ->GetQuadFunction();
MFEM_VERIFY(qfun.Size() == ne*nq,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qfun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different.\n");
qfun.Read();
Qvec.MakeRef(const_cast<QuadratureFunction&>(qfun),0);
}
else if (VectorQuadratureFunctionCoefficient* vqfQ =
dynamic_cast<VectorQuadratureFunctionCoefficient*>(&Q))
{
const QuadratureFunction &qFun = vqfQ->GetQuadFunction();
MFEM_VERIFY(qFun.Size() == vdim * ns * nq * ne,
"Incompatible QuadratureFunction dimension \n");
MFEM_VERIFY(ir == &qFun.GetSpace()->GetElementIntRule(0),
"IntegrationRule used within integrator and in"
" QuadratureFunction appear to be different");
qFun.Read();
Qvec.MakeRef(const_cast<QuadratureFunction &>(qFun),0);
}
else
{
Vector qvec(vdim);
Qvec.SetSize(vdim * nq * ne);
auto C = Reshape(Qvec.HostWrite(), vdim, nq, ne);
for (int e = 0; e < ne; ++e)
{
ElementTransformation &Tr = *fes.GetElementTransformation(e);
for (int q = 0; q < nq; ++q)
{
const IntegrationPoint &ip = ir->IntPoint(q);
Tr.SetIntPoint(&ip);
Q.Eval(qvec, Tr, ip);
for (int c = 0; c<vdim; ++c) { C(c,q,e) = qvec[c]; }
}
}
}
DLFGradAssemble(fes, ir, markers, Qvec, b);
}
} // namespace mfem
+5 -9
View File
@@ -365,11 +365,8 @@ FiniteElementSpace &LORBase::GetFESpace() const
void LORBase::AssembleSystem(BilinearForm &a_ho, const Array<int> &ess_dofs)
{
if (a)
{
A.Clear();
delete a;
}
A.Clear();
delete a;
if (BatchedLORAssembly::FormIsSupported(a_ho))
{
// Skip forming the space
@@ -471,8 +468,7 @@ void LORDiscretization::FormLORSpace()
fec = fes_ho.FEColl()->Clone(GetLOROrder());
const int vdim = fes_ho.GetVDim();
const Ordering::Type ordering = fes_ho.GetOrdering();
fes = new FiniteElementSpace(mesh, fec, vdim, ordering);
fes = new FiniteElementSpace(mesh, fec, vdim);
SetupProlongationAndRestriction();
}
@@ -517,8 +513,8 @@ void ParLORDiscretization::FormLORSpace()
fec = pfes_ho.FEColl()->Clone(GetLOROrder());
const int vdim = fes_ho.GetVDim();
const Ordering::Type ordering = fes_ho.GetOrdering();
fes = new ParFiniteElementSpace(pmesh, fec, vdim, ordering);
ParFiniteElementSpace *pfes = new ParFiniteElementSpace(pmesh, fec, vdim);
fes = pfes;
SetupProlongationAndRestriction();
}
+1 -1
View File
@@ -95,7 +95,7 @@ protected:
/// Returns the order of the LOR space. 1 for H1 or ND, 0 for L2 or RT.
int GetLOROrder() const;
/// Construct the LOR space (overridden for serial and parallel versions).
/// Construct the LOR space (overriden for serial and parallel versions).
virtual void FormLORSpace() = 0;
/// Construct the LORBase object for the given FE space and refinement type.
+6 -17
View File
@@ -13,9 +13,6 @@
#include "../../general/forall.hpp"
#include "../../fem/pbilinearform.hpp"
#define MFEM_NVTX_COLOR DeepSkyBlue
#include "../../general/nvtx.hpp"
namespace mfem
{
@@ -81,10 +78,8 @@ void BatchedLOR_ADS::Form3DFaceToEdge(Array<int> &face2edge)
}
}
void BatchedLOR_ADS::FormCurlMatrixLocal()
void BatchedLOR_ADS::FormCurlMatrix()
{
NVTX("Discrete Curl");
// The curl matrix maps from LOR edges to LOR faces. Given a quadrilateral
// face (defined by its four edges) f_i = (e_j1, e_j2, e_j3, e_j4), the
// matrix has nonzeros A(i, jk), so there are always exactly four nonzeros
@@ -92,9 +87,10 @@ void BatchedLOR_ADS::FormCurlMatrixLocal()
const int nface_dof = face_fes.GetNDofs();
const int nedge_dof = edge_fes.GetNDofs();
SparseMatrix C_local;
C_local.OverrideSize(nface_dof, nedge_dof);
EnsureCapacity(C_local.GetMemoryI(), nedge_dof+1,
Device::GetDeviceMemoryType());
C_local.GetMemoryI().New(nedge_dof+1, Device::GetDeviceMemoryType());
// Each row always has four nonzeros
const int nnz = 4*nedge_dof;
auto I = C_local.WriteI();
@@ -124,8 +120,8 @@ void BatchedLOR_ADS::FormCurlMatrixLocal()
const auto f2e = Reshape(face2edge.Read(), 4, nface_per_el);
// Fill J and data
EnsureCapacity(C_local.GetMemoryJ(), nnz, Device::GetDeviceMemoryType());
EnsureCapacity(C_local.GetMemoryData(), nnz, Device::GetDeviceMemoryType());
C_local.GetMemoryJ().New(nnz, Device::GetDeviceMemoryType());
C_local.GetMemoryData().New(nnz, Device::GetDeviceMemoryType());
auto J = C_local.WriteJ();
auto V = C_local.WriteData();
@@ -150,11 +146,6 @@ void BatchedLOR_ADS::FormCurlMatrixLocal()
V[i*4 + k] = sgn*sgn_f*sgn_e;
}
});
}
void BatchedLOR_ADS::FormCurlMatrix()
{
FormCurlMatrixLocal();
// Create a block diagonal parallel matrix
OperatorHandle C_diag(Operator::Hypre_ParCSR);
@@ -188,8 +179,6 @@ void BatchedLOR_ADS::FormCurlMatrix()
}
C->CopyRowStarts();
C->CopyColStarts();
C_local.Clear();
}
HypreParMatrix *BatchedLOR_ADS::StealCurlMatrix()
+1 -6
View File
@@ -36,9 +36,7 @@ protected:
ND_FECollection edge_fec; ///< The associated Nedelec collection.
ParFiniteElementSpace edge_fes; ///< The associated Nedelec space.
BatchedLOR_AMS ams; ///< The associated AMS object.
HypreParMatrix *C = nullptr; ///< The discrete curl matrix.
SparseMatrix C_local;
HypreParMatrix *C; ///< The discrete curl matrix.
/// Form the local elementwise discrete curl matrix.
void Form3DFaceToEdge(Array<int> &face2edge);
@@ -66,9 +64,6 @@ public:
/// Form the discrete curl matrix (not part of the public API).
void FormCurlMatrix();
void FormCurlMatrixLocal();
~BatchedLOR_ADS();
};
+20 -34
View File
@@ -13,9 +13,6 @@
#include "../../general/forall.hpp"
#include "../../fem/pbilinearform.hpp"
#define MFEM_NVTX_COLOR DeepSkyBlue
#include "../../general/nvtx.hpp"
namespace mfem
{
@@ -139,9 +136,8 @@ void BatchedLOR_AMS::Form3DEdgeToVertex(Array<int> &edge2vert)
}
}
void BatchedLOR_AMS::FormGradientMatrixLocal()
void BatchedLOR_AMS::FormGradientMatrix()
{
NVTX("Discrete Gradient");
// The gradient matrix maps from LOR vertices to LOR edges. Given an edge
// (defined by its two vertices) e_i = (v_j1, v_j2), the matrix has nonzeros
// A(i, j1) = -1 and A(i, j2) = 1, so there are always exactly two nonzeros
@@ -149,10 +145,10 @@ void BatchedLOR_AMS::FormGradientMatrixLocal()
const int nedge_dof = edge_fes.GetNDofs();
const int nvert_dof = vert_fes.GetNDofs();
SparseMatrix G_local;
G_local.OverrideSize(nedge_dof, nvert_dof);
EnsureCapacity(G_local.GetMemoryI(), nedge_dof+1,
Device::GetDeviceMemoryType());
G_local.GetMemoryI().New(nedge_dof+1, Device::GetDeviceMemoryType());
// Each row always has two nonzeros
const int nnz = 2*nedge_dof;
auto I = G_local.WriteI();
@@ -184,8 +180,8 @@ void BatchedLOR_AMS::FormGradientMatrixLocal()
const auto e2v = Reshape(edge2vertex.Read(), 2, nedge_per_el);
// Fill J and data
EnsureCapacity(G_local.GetMemoryJ(), nnz, Device::GetDeviceMemoryType());
EnsureCapacity(G_local.GetMemoryData(), nnz, Device::GetDeviceMemoryType());
G_local.GetMemoryJ().New(nnz, Device::GetDeviceMemoryType());
G_local.GetMemoryData().New(nnz, Device::GetDeviceMemoryType());
auto J = G_local.WriteJ();
auto V = G_local.WriteData();
@@ -208,11 +204,6 @@ void BatchedLOR_AMS::FormGradientMatrixLocal()
V[i*2 + 0] = -sgn;
V[i*2 + 1] = sgn;
});
}
void BatchedLOR_AMS::FormGradientMatrix()
{
FormGradientMatrixLocal();
// Create a block diagonal parallel matrix
OperatorHandle G_diag(Operator::Hypre_ParCSR);
@@ -246,8 +237,6 @@ void BatchedLOR_AMS::FormGradientMatrix()
}
G->CopyRowStarts();
G->CopyColStarts();
G_local.Clear();
}
template <typename T>
@@ -292,7 +281,7 @@ void BatchedLOR_AMS::FormCoordinateVectors(const Vector &X_vert)
const MemoryClass mc = GetHypreMemoryClass();
bool dev = (mc == MemoryClass::DEVICE);
if (xyz_tvec == nullptr) { xyz_tvec = new Vector(ntdofs*dim); }
xyz_tvec = new Vector(ntdofs*dim);
auto xyz_tv = Reshape(HypreWrite(xyz_tvec->GetMemory()), ntdofs, dim);
const auto xyz_e =
@@ -313,24 +302,21 @@ void BatchedLOR_AMS::FormCoordinateVectors(const Vector &X_vert)
});
// Make x, y, z HypreParVectors point to T-vector data
if (x == nullptr)
{
HYPRE_BigInt glob_size = vert_fes.GlobalTrueVSize();
HYPRE_BigInt *cols = vert_fes.GetTrueDofOffsets();
HYPRE_BigInt glob_size = vert_fes.GlobalTrueVSize();
HYPRE_BigInt *cols = vert_fes.GetTrueDofOffsets();
double *d_x_ptr = xyz_tv + 0*ntdofs;
x = new HypreParVector(vert_fes.GetComm(), glob_size, d_x_ptr, cols, dev);
double *d_y_ptr = xyz_tv + 1*ntdofs;
y = new HypreParVector(vert_fes.GetComm(), glob_size, d_y_ptr, cols, dev);
if (dim == 3)
{
double *d_z_ptr = xyz_tv + 2*ntdofs;
z = new HypreParVector(vert_fes.GetComm(), glob_size, d_z_ptr, cols, dev);
}
else
{
z = NULL;
}
double *d_x_ptr = xyz_tv + 0*ntdofs;
x = new HypreParVector(vert_fes.GetComm(), glob_size, d_x_ptr, cols, dev);
double *d_y_ptr = xyz_tv + 1*ntdofs;
y = new HypreParVector(vert_fes.GetComm(), glob_size, d_y_ptr, cols, dev);
if (dim == 3)
{
double *d_z_ptr = xyz_tv + 2*ntdofs;
z = new HypreParVector(vert_fes.GetComm(), glob_size, d_z_ptr, cols, dev);
}
else
{
z = NULL;
}
}
+3 -8
View File
@@ -33,14 +33,12 @@ protected:
const int order; ///< Polynomial degree.
H1_FECollection vert_fec; ///< The corresponding H1 collection.
ParFiniteElementSpace vert_fes; ///< The corresponding H1 space.
Vector *xyz_tvec = nullptr; ///< Mesh vertex coordinates in true-vector format.
HypreParMatrix *G = nullptr; ///< Discrete gradient matrix.
SparseMatrix G_local;
Vector *xyz_tvec; ///< Mesh vertex coordinates in true-vector format.
HypreParMatrix *G; ///< Discrete gradient matrix.
/// @name Mesh coordinate vectors in HypreParVector format
///@{
HypreParVector *x = nullptr, *y = nullptr, *z = nullptr;
HypreParVector *x, *y, *z;
///@}
/// @name Construct the local (elementwise) discrete gradient
@@ -91,9 +89,6 @@ public:
/// Construct the discrete gradient matrix (not part of the public API).
void FormGradientMatrix();
void FormGradientMatrixLocal();
~BatchedLOR_AMS();
};
+11 -41
View File
@@ -15,8 +15,6 @@
#include <climits>
#include "../pbilinearform.hpp"
#include "../../general/nvtx.hpp"
// Specializations
#include "lor_h1.hpp"
#include "lor_nd.hpp"
@@ -72,13 +70,8 @@ bool BatchedLORAssembly::FormIsSupported(BilinearForm &a)
}
void BatchedLORAssembly::FormLORVertexCoordinates(FiniteElementSpace &fes_ho,
Vector &X_vert,
Vector *evec)
Vector &X_vert)
{
#undef MFEM_NVTX_COLOR
#define MFEM_NVTX_COLOR DeepSkyBlue
NVTX("LOR Coordinates");
Mesh &mesh_ho = *fes_ho.GetMesh();
mesh_ho.EnsureNodes();
@@ -94,33 +87,20 @@ void BatchedLORAssembly::FormLORVertexCoordinates(FiniteElementSpace &fes_ho,
const Operator *nodal_restriction =
nodal_fes->GetElementRestriction(ElementDofOrdering::LEXICOGRAPHIC);
Vector *tmp_evec = nullptr;
Vector *nodal_evec;
if (evec)
{
nodal_evec = evec;
nodal_evec->SetSize(nodal_restriction->Height());
}
else
{
tmp_evec = new Vector(nodal_restriction->Height());
nodal_evec = tmp_evec;
}
// Map from nodal L-vector to E-vector
nodal_restriction->Mult(*nodal_gf, *nodal_evec);
Vector nodal_evec(nodal_restriction->Height());
nodal_restriction->Mult(*nodal_gf, nodal_evec);
IntegrationRule ir = GetCollocatedIntRule(fes_ho);
IntegrationRules irs(0, Quadrature1D::GaussLobatto);
Geometry::Type geom = mesh_ho.GetElementGeometry(0);
const IntegrationRule &ir = irs.Get(geom, 2*nd1d - 3);
// Map from nodal E-vector to Q-vector at the LOR vertex points
X_vert.SetSize(dim*ndof_per_el*nel_ho);
const QuadratureInterpolator *quad_interp =
nodal_fes->GetQuadratureInterpolator(ir);
quad_interp->SetOutputLayout(QVectorLayout::byVDIM);
quad_interp->Values(*nodal_evec, X_vert);
delete tmp_evec;
quad_interp->Values(nodal_evec, X_vert);
}
// The following two functions (GetMinElt and GetAndIncrementNnzIndex) are
@@ -394,11 +374,11 @@ void BatchedLORAssembly::SparseIJToCSR(OperatorHandle &A) const
A_mat->OverrideSize(nvdof, nvdof);
EnsureCapacity(A_mat->GetMemoryI(), nvdof+1, Device::GetDeviceMemoryType());
A_mat->GetMemoryI().New(nvdof+1, Device::GetDeviceMemoryType());
int nnz = FillI(*A_mat);
EnsureCapacity(A_mat->GetMemoryJ(), nnz, Device::GetDeviceMemoryType());
EnsureCapacity(A_mat->GetMemoryData(), nnz, Device::GetDeviceMemoryType());
A_mat->GetMemoryJ().New(nnz, Device::GetDeviceMemoryType());
A_mat->GetMemoryData().New(nnz, Device::GetDeviceMemoryType());
FillJAndData(*A_mat);
}
@@ -477,6 +457,7 @@ void BatchedLORAssembly::ParAssemble(
BilinearForm &a, const Array<int> &ess_dofs, OperatorHandle &A)
{
// Assemble the system matrix local to this partition
OperatorHandle A_local;
AssembleWithoutBC(a, A_local);
ParBilinearForm *pa =
@@ -492,9 +473,6 @@ void BatchedLORAssembly::ParAssemble(
void BatchedLORAssembly::Assemble(
BilinearForm &a, const Array<int> ess_dofs, OperatorHandle &A)
{
#undef MFEM_NVTX_COLOR
#define MFEM_NVTX_COLOR NavyBlue
NVTX("LOR Assemble");
#ifdef MFEM_USE_MPI
if (dynamic_cast<ParFiniteElementSpace*>(&fes_ho))
{
@@ -515,12 +493,4 @@ BatchedLORAssembly::BatchedLORAssembly(FiniteElementSpace &fes_ho_)
FormLORVertexCoordinates(fes_ho, X_vert);
}
IntegrationRule GetCollocatedIntRule(FiniteElementSpace &fes)
{
IntegrationRules irs(0, Quadrature1D::GaussLobatto);
const Geometry::Type geom = fes.GetMesh()->GetElementGeometry(0);
const int nd1d = fes.GetMaxElementOrder() + 1;
return irs.Get(geom, 2*nd1d - 3);
}
} // namespace mfem
+5 -33
View File
@@ -13,7 +13,6 @@
#define MFEM_LOR_BATCHED
#include "lor.hpp"
#include "../qspace.hpp"
namespace mfem
{
@@ -53,8 +52,6 @@ protected:
/// nonzero).
Array<int> sparse_mapping;
OperatorHandle A_local; // Cache this
public:
/// Construct the batched assembly object corresponding to @a fes_ho_.
BatchedLORAssembly(FiniteElementSpace &fes_ho_);
@@ -70,12 +67,12 @@ public:
/// Compute the vertices of the LOR mesh and place the result in @a X_vert.
static void FormLORVertexCoordinates(FiniteElementSpace &fes_ho,
Vector &X_vert,
Vector *evec = nullptr);
Vector &X_vert);
/// Return the vertices of the LOR mesh in E-vector format
const Vector &GetLORVertexCoordinates() { return X_vert; }
protected:
/// After assembling the "sparse IJ" format, convert it to CSR.
void SparseIJToCSR(OperatorHandle &A) const;
@@ -119,12 +116,12 @@ public:
/// If the capacity of @a mem is not large enough, delete it and allocate new
/// memory with size @a capacity.
template <typename T>
void EnsureCapacity(Memory<T> &mem, int capacity, MemoryType mt)
void EnsureCapacity(Memory<T> &mem, int capacity)
{
if (mem.Capacity() < capacity)
{
mem.Delete();
mem.New(capacity, mt);
mem.New(capacity, mem.GetMemoryType());
}
}
@@ -146,25 +143,6 @@ static T *GetIntegrator(BilinearForm &a)
return nullptr;
}
IntegrationRule GetCollocatedIntRule(FiniteElementSpace &fes);
template <typename INTEGRATOR>
void ProjectLORCoefficient(BilinearForm &a, CoefficientVector &coeff_vector)
{
INTEGRATOR *i = GetIntegrator<INTEGRATOR>(a);
if (i)
{
// const_cast since Coefficient::Eval is not const...
auto *coeff = const_cast<Coefficient*>(i->GetCoefficient());
if (coeff) { coeff_vector.Project(*coeff); }
else { coeff_vector.SetConstant(1.0); }
}
else
{
coeff_vector.SetConstant(0.0);
}
}
/// Abstract base class for the batched LOR assembly kernels.
class BatchedLORKernel
{
@@ -173,18 +151,12 @@ protected:
Vector &X_vert; ///< Mesh coordinate vector.
Vector &sparse_ij; ///< Local element sparsity matrix data.
Array<int> &sparse_mapping; ///< Local element sparsity pattern.
IntegrationRule ir; ///< Collocated integration rule.
QuadratureSpace qs; ///< Quadrature space for coefficients.
CoefficientVector c1; ///< Coefficient of first integrator.
CoefficientVector c2; ///< Coefficient of second integrator.
BatchedLORKernel(FiniteElementSpace &fes_ho_,
Vector &X_vert_,
Vector &sparse_ij_,
Array<int> &sparse_mapping_)
: fes_ho(fes_ho_), X_vert(X_vert_), sparse_ij(sparse_ij_),
sparse_mapping(sparse_mapping_), ir(GetCollocatedIntRule(fes_ho)),
qs(*fes_ho.GetMesh(), ir), c1(qs, CoefficientStorage::COMPRESSED),
c2(qs, CoefficientStorage::COMPRESSED)
sparse_mapping(sparse_mapping_)
{ }
};
+81 -241
View File
@@ -30,14 +30,8 @@ void BatchedLOR_H1::Assemble2D()
static constexpr int nnz_per_row = 9;
static constexpr int sz_local_mat = nv*nv;
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1)
: Reshape(c1.Read(), nd1d, nd1d, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1)
: Reshape(c2.Read(), nd1d, nd1d, nel_ho);
const double DQ = diffusion_coeff;
const double MQ = mass_coeff;
sparse_ij.SetSize(nnz_per_row*ndof_per_el*nel_ho);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, nd1d, nd1d, nel_ho);
@@ -103,8 +97,6 @@ void BatchedLOR_H1::Assemble2D()
{
for (int iqy=0; iqy<2; ++iqy)
{
const double mq = const_mq ? MQ(0,0,0) : MQ(kx+iqx, ky+iqy, iel_ho);
const double dq = const_dq ? DQ(0,0,0) : DQ(kx+iqx, ky+iqy, iel_ho);
for (int jy=0; jy<2; ++jy)
{
const double bjy = (jy == iqy) ? 1.0 : 0.0;
@@ -141,9 +133,9 @@ void BatchedLOR_H1::Assemble2D()
val += dix*djx*Q(0,iqy,iqx);
val += (dix*djy + diy*djx)*Q(1,iqy,iqx);
val += diy*djy*Q(2,iqy,iqx);
val *= dq;
val *= DQ;
val += mq*bix*biy*bjx*bjy*Q(3,iqy,iqx);
val += MQ*bix*biy*bjx*bjy*Q(3,iqy,iqx);
local_mat(ii_loc, jj_loc) += val;
}
@@ -205,53 +197,14 @@ void BatchedLOR_H1::Assemble2D()
}
}
template<int ORDER>
static void SparseMapping3D(Array<int> &sparse_mapping)
{
static constexpr int nnz_per_row = 27;
static constexpr int nd1d = ORDER + 1;
static constexpr int ndof_per_el = nd1d*nd1d*nd1d;
sparse_mapping.SetSize(nnz_per_row*ndof_per_el);
sparse_mapping = -1;
auto map = Reshape(sparse_mapping.HostReadWrite(), nnz_per_row, ndof_per_el);
for (int iz=0; iz<nd1d; ++iz)
{
const int jz_begin = (iz > 0) ? iz - 1 : 0;
const int jz_end = (iz < ORDER) ? iz + 1 : ORDER;
for (int iy=0; iy<nd1d; ++iy)
{
const int jy_begin = (iy > 0) ? iy - 1 : 0;
const int jy_end = (iy < ORDER) ? iy + 1 : ORDER;
for (int ix=0; ix<nd1d; ++ix)
{
const int jx_begin = (ix > 0) ? ix - 1 : 0;
const int jx_end = (ix < ORDER) ? ix + 1 : ORDER;
const int ii_el = ix + nd1d*(iy + nd1d*iz);
for (int jz=jz_begin; jz<=jz_end; ++jz)
{
for (int jy=jy_begin; jy<=jy_end; ++jy)
{
for (int jx=jx_begin; jx<=jx_end; ++jx)
{
const int jj_off = (jx-ix+1) + 3*(jy-iy+1) + 9*(jz-iz+1);
const int jj_el = jx + nd1d*(jy + nd1d*jz);
map(jj_off, ii_el) = jj_el;
}
}
}
}
}
}
}
template <int ORDER>
void BatchedLOR_H1::Assemble3D()
{
const int nel_ho = fes_ho.GetNE();
const double DQ = diffusion_coeff;
const double MQ = mass_coeff;
static constexpr int nv = 8;
static constexpr int dim = 3;
static constexpr int ddm2 = (dim*(dim+1))/2;
@@ -264,15 +217,6 @@ void BatchedLOR_H1::Assemble3D()
static constexpr int sz_mass_B = sz_mass_A*2;
static constexpr int sz_local_mat = nv*nv;
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1, 1)
: Reshape(c1.Read(), nd1d, nd1d, nd1d, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1, 1)
: Reshape(c2.Read(), nd1d, nd1d, nd1d, nel_ho);
sparse_ij.SetSize(nel_ho*ndof_per_el*nnz_per_row);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, nd1d, nd1d, nd1d, nel_ho);
@@ -344,6 +288,7 @@ void BatchedLOR_H1::Assemble3D()
//MFEM_UNROLL(2)
for (int iqx=0; iqx<2; ++iqx)
{
const double x = iqx;
const double y = iqy;
const double z = iqz;
@@ -373,21 +318,23 @@ void BatchedLOR_H1::Assemble3D()
}
}
//MFEM_UNROLL(2)
for (int iqx=0; iqx<2; ++iqx)
{
//MFEM_UNROLL(2)
for (int jz=0; jz<2; ++jz)
{
// Note loop starts at iz=jz here, taking advantage of
// symmetries.
//MFEM_UNROLL(2)
for (int iz=jz; iz<2; ++iz)
{
//MFEM_UNROLL(2)
for (int iqy=0; iqy<2; ++iqy)
{
//MFEM_UNROLL(2)
for (int iqz=0; iqz<2; ++iqz)
{
const double mq = const_mq ? MQ(0,0,0,0) : MQ(kx+iqx, ky+iqy, kz+iqz, iel_ho);
const double dq = const_dq ? DQ(0,0,0,0) : DQ(kx+iqx, ky+iqy, kz+iqz, iel_ho);
const double biz = (iz == iqz) ? 1.0 : 0.0;
const double giz = (iz == 0) ? -1.0 : 1.0;
@@ -404,21 +351,23 @@ void BatchedLOR_H1::Assemble3D()
const double J23 = J32;
const double J33 = Q(5,iqz,iqy,iqx);
grad_A(0,0,iqy,iz,jz,iqx) += dq*J11*biz*bjz;
grad_A(1,0,iqy,iz,jz,iqx) += dq*J21*biz*bjz;
grad_A(2,0,iqy,iz,jz,iqx) += dq*J31*giz*bjz;
grad_A(0,1,iqy,iz,jz,iqx) += dq*J12*biz*bjz;
grad_A(1,1,iqy,iz,jz,iqx) += dq*J22*biz*bjz;
grad_A(2,1,iqy,iz,jz,iqx) += dq*J32*giz*bjz;
grad_A(0,2,iqy,iz,jz,iqx) += dq*J13*biz*gjz;
grad_A(1,2,iqy,iz,jz,iqx) += dq*J23*biz*gjz;
grad_A(2,2,iqy,iz,jz,iqx) += dq*J33*giz*gjz;
grad_A(0,0,iqy,iz,jz,iqx) += J11*biz*bjz;
grad_A(1,0,iqy,iz,jz,iqx) += J21*biz*bjz;
grad_A(2,0,iqy,iz,jz,iqx) += J31*giz*bjz;
grad_A(0,1,iqy,iz,jz,iqx) += J12*biz*bjz;
grad_A(1,1,iqy,iz,jz,iqx) += J22*biz*bjz;
grad_A(2,1,iqy,iz,jz,iqx) += J32*giz*bjz;
grad_A(0,2,iqy,iz,jz,iqx) += J13*biz*gjz;
grad_A(1,2,iqy,iz,jz,iqx) += J23*biz*gjz;
grad_A(2,2,iqy,iz,jz,iqx) += J33*giz*gjz;
double wdetJ = Q(6,iqz,iqy,iqx);
mass_A(iqy,iz,jz,iqx) += mq*wdetJ*biz*bjz;
mass_A(iqy,iz,jz,iqx) += wdetJ*biz*bjz;
}
//MFEM_UNROLL(2)
for (int jy=0; jy<2; ++jy)
{
//MFEM_UNROLL(2)
for (int iy=0; iy<2; ++iy)
{
const double biy = (iy == iqy) ? 1.0 : 0.0;
@@ -441,12 +390,16 @@ void BatchedLOR_H1::Assemble3D()
}
}
}
//MFEM_UNROLL(2)
for (int jy=0; jy<2; ++jy)
{
//MFEM_UNROLL(2)
for (int jx=0; jx<2; ++jx)
{
//MFEM_UNROLL(2)
for (int iy=0; iy<2; ++iy)
{
//MFEM_UNROLL(2)
for (int ix=0; ix<2; ++ix)
{
const double bix = (ix == iqx) ? 1.0 : 0.0;
@@ -473,7 +426,9 @@ void BatchedLOR_H1::Assemble3D()
val += bix*bjx*grad_B(2,2,iy,jy,iz,jz,iqx);
val += bix*bjx*grad_B(1,2,iy,jy,iz,jz,iqx);
val += bix*bjx*mass_B(iy,jy,iz,jz,iqx);
val *= DQ;
val += MQ*bix*bjx*mass_B(iy,jy,iz,jz,iqx);
local_mat(ii_loc, jj_loc) += val;
}
@@ -514,176 +469,40 @@ void BatchedLOR_H1::Assemble3D()
}
}
});
SparseMapping3D<ORDER>(sparse_mapping);
}
template <>
void BatchedLOR_H1::Assemble3D<1>()
{
static constexpr int nv = 8;
static constexpr int nd1d = 2;
static constexpr int ndof_per_el = 8;
static constexpr int nnz_per_row = 27;
static constexpr int sz_local_mat = nv*nv;
const int nel_ho = fes_ho.GetNE();
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1, 1)
: Reshape(c1.Read(), nd1d, nd1d, nd1d, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1, 1)
: Reshape(c2.Read(), nd1d, nd1d, nd1d, nel_ho);
sparse_ij.SetSize(nel_ho*ndof_per_el*nnz_per_row);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, nd1d, nd1d, nd1d, nel_ho);
const auto X = X_vert.Read();
MFEM_FORALL_3D(iel_ho, nel_ho, 8, 4, 1,
sparse_mapping.SetSize(nnz_per_row*ndof_per_el);
sparse_mapping = -1;
auto map = Reshape(sparse_mapping.HostReadWrite(), nnz_per_row, ndof_per_el);
for (int iz=0; iz<nd1d; ++iz)
{
static constexpr int e[8] = {0,1,3,2,4,5,7,6};
MFEM_SHARED double vx[8], vy[8], vz[8];
const int tidz = MFEM_THREAD_ID(z);
MFEM_SHARED double local_mat_[sz_local_mat];
DeviceTensor<4> local_mat(local_mat_, 2,2,2, nv);
if (tidz == 0)
const int jz_begin = (iz > 0) ? iz - 1 : 0;
const int jz_end = (iz < ORDER) ? iz + 1 : ORDER;
for (int iy=0; iy<nd1d; ++iy)
{
MFEM_FOREACH_THREAD(xyz,x,8)
const int jy_begin = (iy > 0) ? iy - 1 : 0;
const int jy_end = (iy < ORDER) ? iy + 1 : ORDER;
for (int ix=0; ix<nd1d; ++ix)
{
const int z = xyz%2, y = (xyz/2)%2, x = xyz/2/2;
MFEM_FOREACH_THREAD(j,y,nnz_per_row)
const int jx_begin = (ix > 0) ? ix - 1 : 0;
const int jx_end = (ix < ORDER) ? ix + 1 : ORDER;
const int ii_el = ix + nd1d*(iy + nd1d*iz);
for (int jz=jz_begin; jz<=jz_end; ++jz)
{
if (j < 8) { local_mat(z,y,x,j) = 0.0; }
V(j,x,y,z,iel_ho) = 0.0;
if (j == 0)
for (int jy=jy_begin; jy<=jy_end; ++jy)
{
const int i = x + 2*y + 4*z;
const int ei = 3*(e[i] + 8*iel_ho);
vx[i] = X[ei + 0];
vy[i] = X[ei + 1];
vz[i] = X[ei + 2];
for (int jx=jx_begin; jx<=jx_end; ++jx)
{
const int jj_off = (jx-ix+1) + 3*(jy-iy+1) + 9*(jz-iz+1);
const int jj_el = jx + nd1d*(jy + nd1d*jz);
map(jj_off, ii_el) = jj_el;
}
}
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(xyz,x,8)
{
const int qz = xyz%2, qy = (xyz/2)%2, qx = xyz/2/2;
static constexpr double w = 1.0/8.0;
double J_[3*3];
DeviceTensor<2> J(J_, 3,3);
Jacobian3D(qx,qy,qz, vx,vy,vz, J);
const double detJ = Det3D(J);
const double w_detJ = w/detJ;
// adj(J)
double A_[3*3];
DeviceTensor<2> A(A_, 3, 3);
Adjugate3D(J, A);
const double J11 = w_detJ*(A(0,0)*A(0,0)+A(0,1)*A(0,1)+A(0,2)*A(0,2)); // 1,1
const double J21 = w_detJ*(A(0,0)*A(1,0)+A(0,1)*A(1,1)+A(0,2)*A(1,2)); // 2,1
const double J31 = w_detJ*(A(0,0)*A(2,0)+A(0,1)*A(2,1)+A(0,2)*A(2,2)); // 3,1
const double J12 = J21;
const double J22 = w_detJ*(A(1,0)*A(1,0)+A(1,1)*A(1,1)+A(1,2)*A(1,2)); // 2,2
const double J32 = w_detJ*(A(1,0)*A(2,0)+A(1,1)*A(2,1)+A(1,2)*A(2,2)); // 3,2
const double J13 = J31;
const double J23 = J32;
const double J33 = w_detJ*(A(2,0)*A(2,0)+A(2,1)*A(2,1)+A(2,2)*A(2,2)); // 3,3
const double wdetJ = w*detJ;
const double mq = const_mq ? MQ(0,0,0,0) : MQ(qx,qy,qz, iel_ho);
const double dq = const_dq ? DQ(0,0,0,0) : DQ(qx,qy,qz, iel_ho);
MFEM_FOREACH_THREAD(xyz,y,8)
{
const int jz = xyz%2, jy = (xyz/2)%2, jx = xyz/2/2;
const double bjz = (jz == qz) ? 1.0 : 0.0;
const double gjz = (jz == 0) ? -1.0 : 1.0;
const double bjy = (jy == qy) ? 1.0 : 0.0;
const double gjy = (jy == 0) ? -1.0 : 1.0;
const double bjx = (jx == qx) ? 1.0 : 0.0;
const double gjx = (jx == 0) ? -1.0 : 1.0;
const double djx = gjx*bjy*bjz;
const double djy = bjx*gjy*bjz;
const double djz = bjx*bjy*gjz;
const int jj_loc = jx + 2*jy + 4*jz;
MFEM_FOREACH_THREAD(xyz,z,8)
{
const int iz = xyz%2, iy = (xyz/2)%2, ix = xyz/2/2;
const double biz = (iz == qz) ? 1.0 : 0.0;
const double giz = (iz == 0) ? -1.0 : 1.0;
const double biy = (iy == qy) ? 1.0 : 0.0;
const double giy = (iy == 0) ? -1.0 : 1.0;
const double bix = (ix == qx) ? 1.0 : 0.0;
const double gix = (ix == 0) ? -1.0 : 1.0;
const double dix = gix*biy*biz;
const double diy = bix*giy*biz;
const double diz = bix*biy*giz;
const int ii_loc = ix + 2*iy + 4*iz;
// Only store the lower-triangular part of
// the matrix (by symmetry).
if (jj_loc > ii_loc) { continue; }
double grad_grad = 0.0;
grad_grad += dix*djx*J11;
grad_grad += diy*djx*J12;
grad_grad += diz*djx*J13;
grad_grad += dix*djy*J21;
grad_grad += diy*djy*J22;
grad_grad += diz*djy*J23;
grad_grad += dix*djz*J31;
grad_grad += diy*djz*J32;
grad_grad += diz*djz*J33;
const double basis_basis = wdetJ*bix*biy*biz*bjx*bjy*bjz;
const double value = dq*grad_grad + mq*basis_basis;
AtomicAdd(local_mat(iz,iy,ix, jj_loc), value);
} // i
} // j
} // q
MFEM_SYNC_THREAD;
// Assemble the local matrix into the macro-element sparse matrix
// in a format similar to coordinate format.
// The (I,J) arrays are implicit (not stored explicitly).
if (tidz == 0)
{
MFEM_FOREACH_THREAD(xyz,x,8)
{
const int iz = xyz%2, iy = (xyz/2)%2, ix = xyz/2/2;
const int ii_loc = ix + 2*iy + 4*iz;
MFEM_FOREACH_THREAD(jj_loc,y,8)
{
const int jx = jj_loc%2, jy = (jj_loc/2)%2, jz = jj_loc/2/2;
const int jj_off = (jx-ix+1) + 3*(jy-iy+1) + 9*(jz-iz+1);
if (jj_loc <= ii_loc)
{
AtomicAdd(V(jj_off, ix,iy,iz, iel_ho), local_mat(iz,iy,ix, jj_loc));
}
else
{
AtomicAdd(V(jj_off, ix,iy,iz, iel_ho), local_mat(jz,jy,jx, ii_loc));
}
}
}
}
MFEM_SYNC_THREAD;
});
SparseMapping3D<1>(sparse_mapping);
}
}
// Explicit template instantiations
@@ -696,7 +515,7 @@ template void BatchedLOR_H1::Assemble2D<6>();
template void BatchedLOR_H1::Assemble2D<7>();
template void BatchedLOR_H1::Assemble2D<8>();
//template void BatchedLOR_H1::Assemble3D<1>(); // explicitly specialized
template void BatchedLOR_H1::Assemble3D<1>();
template void BatchedLOR_H1::Assemble3D<2>();
template void BatchedLOR_H1::Assemble3D<3>();
template void BatchedLOR_H1::Assemble3D<4>();
@@ -712,8 +531,29 @@ BatchedLOR_H1::BatchedLOR_H1(BilinearForm &a,
Array<int> &sparse_mapping_)
: BatchedLORKernel(fes_ho_, X_vert_, sparse_ij_, sparse_mapping_)
{
ProjectLORCoefficient<MassIntegrator>(a, c1);
ProjectLORCoefficient<DiffusionIntegrator>(a, c2);
MassIntegrator *mass = GetIntegrator<MassIntegrator>(a);
DiffusionIntegrator *diffusion = GetIntegrator<DiffusionIntegrator>(a);
if (mass != nullptr)
{
auto *coeff = dynamic_cast<const ConstantCoefficient*>(mass->GetCoefficient());
mass_coeff = coeff ? coeff->constant : 1.0;
}
else
{
mass_coeff = 0.0;
}
if (diffusion != nullptr)
{
auto *coeff = dynamic_cast<const ConstantCoefficient*>
(diffusion->GetCoefficient());
diffusion_coeff = coeff ? coeff->constant : 1.0;
}
else
{
diffusion_coeff = 0.0;
}
}
} // namespace mfem
+3
View File
@@ -21,6 +21,9 @@ namespace mfem
// classes BatchedLORAssembly and BatchedLORKernel .
class BatchedLOR_H1 : BatchedLORKernel
{
protected:
// TODO: for now only supporting constant coefficients
double mass_coeff, diffusion_coeff;
public:
template <int ORDER> void Assemble2D();
template <int ORDER> void Assemble3D();
+91 -324
View File
@@ -33,14 +33,8 @@ void BatchedLOR_ND::Assemble2D()
static constexpr int nnz_per_row = 7;
static constexpr int sz_local_mat = ne*ne;
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1)
: Reshape(c1.Read(), op1, op1, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1)
: Reshape(c2.Read(), op1, op1, nel_ho);
const double DQ = curl_curl_coeff;
const double MQ = mass_coeff;
sparse_ij.SetSize(nnz_per_row*ndof_per_el*nel_ho);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, o*op1, dim, nel_ho);
@@ -112,8 +106,6 @@ void BatchedLOR_ND::Assemble2D()
{
for (int iqy=0; iqy<2; ++iqy)
{
const double mq = const_mq ? MQ(0,0,0) : MQ(kx+iqx, ky+iqy, iel_ho);
const double dq = const_dq ? DQ(0,0,0) : DQ(kx+iqx, ky+iqy, iel_ho);
// Loop over x,y components. c=0 => x, c=1 => y
for (int cj=0; cj<dim; ++cj)
{
@@ -144,8 +136,8 @@ void BatchedLOR_ND::Assemble2D()
val += byi*bxj*Q(1,iqy,iqx);
val += bxi*byj*Q(1,iqy,iqx);
val += byi*byj*Q(2,iqy,iqx);
val *= mq;
val += dq*curl_i*curl_j*Q(3,iqy,iqx);
val *= MQ;
val += DQ*curl_i*curl_j*Q(3,iqy,iqx);
local_mat(ii_loc, jj_loc) += val;
}
@@ -216,90 +208,6 @@ void BatchedLOR_ND::Assemble2D()
}
}
template<int ORDER>
static void SparseMapping3D(Array<int> &sparse_mapping)
{
static constexpr int nnz_per_row = 33;
static constexpr int dim = 3;
static constexpr int o = ORDER;
static constexpr int op1 = ORDER + 1;
static constexpr int ndof_per_el = dim*o*op1*op1;
sparse_mapping.SetSize(nnz_per_row*ndof_per_el);
sparse_mapping = -1;
auto map = Reshape(sparse_mapping.HostReadWrite(), nnz_per_row, ndof_per_el);
for (int ci=0; ci<dim; ++ci)
{
const int i_off = ci*o*op1*op1;
const int id0 = ci;
const int id1 = (ci+1)%3;
const int id2 = (ci+2)%3;
const int nxi = (ci == 0) ? o : op1;
const int nyi = (ci == 1) ? o : op1;
for (int i0=0; i0<o; ++i0)
{
for (int i1=0; i1<op1; ++i1)
{
for (int i2=0; i2<op1; ++i2)
{
int ii_lex[3];
ii_lex[id0] = i0;
ii_lex[id1] = i1;
ii_lex[id2] = i2;
const int ii_el = i_off + ii_lex[0] + ii_lex[1]*nxi + ii_lex[2]*nxi*nyi;
for (int cj_rel=0; cj_rel<dim; ++cj_rel)
{
const int cj = (ci + cj_rel) % 3;
const int j_off = cj*o*op1*op1;
const int nxj = (cj == 0) ? o : op1;
const int nyj = (cj == 1) ? o : op1;
const int j0_begin = i0;
const int j0_end = (cj_rel == 0) ? i0 : i0 + 1;
const int j1_begin = (i1 > 0) ? i1-1 : i1;
const int j1_end = (cj_rel == 1)
? ((i1 < o) ? i1 : i1-1)
: ((i1 < o) ? i1+1 : i1);
const int j2_begin = (i2 > 0) ? i2-1 : i2;
const int j2_end = (cj_rel == 2)
? ((i2 < o) ? i2 : i2-1)
: ((i2 < o) ? i2+1 : i2);
for (int j0=j0_begin; j0<=j0_end; ++j0)
{
const int d0 = j0 - i0;
for (int j1=j1_begin; j1<=j1_end; ++j1)
{
const int d1 = j1 - i1 + 1;
for (int j2=j2_begin; j2<=j2_end; ++j2)
{
const int d2 = j2 - i2 + 1;
int jj_lex[3];
jj_lex[id0] = j0;
jj_lex[id1] = j1;
jj_lex[id2] = j2;
const int jj_el = j_off + jj_lex[0] + jj_lex[1]*nxj + jj_lex[2]*nxj*nyj;
int jj_off;
if (cj_rel == 0) { jj_off = d1 + 3*d2; }
else if (cj_rel == 1) { jj_off = 9 + d0 + 2*d1 + 4*d2; }
else /* if (cj_rel == 2) */ { jj_off = 21 + d0 + 2*d1 + 6*d2; }
map(jj_off, ii_el) = jj_el;
}
}
}
}
}
}
}
}
}
template <int ORDER>
void BatchedLOR_ND::Assemble3D()
{
@@ -316,19 +224,13 @@ void BatchedLOR_ND::Assemble3D()
static constexpr int nnz_per_row = 33;
static constexpr int sz_local_mat = ne*ne;
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1, 1)
: Reshape(c1.Read(), op1, op1, op1, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1, 1)
: Reshape(c2.Read(), op1, op1, op1, nel_ho);
const double DQ = curl_curl_coeff;
const double MQ = mass_coeff;
sparse_ij.SetSize(nnz_per_row*ndof_per_el*nel_ho);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, o*op1*op1, dim, nel_ho);
const auto X = X_vert.Read();
auto X = X_vert.Read();
// Last thread dimension is lowered to avoid "too many resources" error
MFEM_FORALL_3D(iel_ho, nel_ho, ORDER, ORDER, (ORDER>6)?4:ORDER,
@@ -416,8 +318,6 @@ void BatchedLOR_ND::Assemble3D()
{
for (int iqx=0; iqx<2; ++iqx)
{
const double mq = const_mq ? MQ(0,0,0,0) : MQ(kx+iqx, ky+iqy, kz+iqz, iel_ho);
const double dq = const_dq ? DQ(0,0,0,0) : DQ(kx+iqx, ky+iqy, kz+iqz, iel_ho);
// Loop over x,y,z components. 0 => x, 1 => y, 2 => z
for (int cj=0; cj<dim; ++cj)
{
@@ -491,7 +391,7 @@ void BatchedLOR_ND::Assemble3D()
basis_basis += Q(4,iqz,iqy,iqx)*(basis_i[1]*basis_j[2] + basis_i[2]*basis_j[1]);
basis_basis += Q(5,iqz,iqy,iqx)*basis_i[2]*basis_j[2];
const double val = dq*curl_curl + mq*basis_basis;
const double val = DQ*curl_curl + MQ*basis_basis;
local_mat(ii_loc, jj_loc) += val;
}
@@ -572,231 +472,78 @@ void BatchedLOR_ND::Assemble3D()
}
}
});
SparseMapping3D<ORDER>(sparse_mapping);
}
template<>
void BatchedLOR_ND::Assemble3D<1>()
{
static constexpr int ne = 12; // number of edges in hexahedron
static constexpr int dim = 3;
static constexpr int o = 1;
static constexpr int op1 = 2;
static constexpr int ndof_per_el = dim*o*op1*op1;
static constexpr int nnz_per_row = 33;
static constexpr int sz_local_mat = ne*ne;
const int nel_ho = fes_ho.GetNE();
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1, 1)
: Reshape(c1.Read(), op1, op1, op1, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1, 1)
: Reshape(c2.Read(), op1, op1, op1, nel_ho);
sparse_ij.SetSize(nnz_per_row*ndof_per_el*nel_ho);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, o*op1*op1, dim, nel_ho);
const auto X = X_vert.Read();
MFEM_FORALL_3D(iel_ho, nel_ho, 8, 1, 4,
sparse_mapping.SetSize(nnz_per_row*ndof_per_el);
sparse_mapping = -1;
auto map = Reshape(sparse_mapping.HostReadWrite(), nnz_per_row, ndof_per_el);
for (int ci=0; ci<dim; ++ci)
{
MFEM_FOREACH_THREAD(iz,z,o) // 1
const int i_off = ci*o*op1*op1;
const int id0 = ci;
const int id1 = (ci+1)%3;
const int id2 = (ci+2)%3;
const int nxi = (ci == 0) ? o : op1;
const int nyi = (ci == 1) ? o : op1;
for (int i0=0; i0<o; ++i0)
{
MFEM_FOREACH_THREAD(iy,y,op1) // 2
for (int i1=0; i1<op1; ++i1)
{
MFEM_FOREACH_THREAD(ix,x,op1) // 2
for (int i2=0; i2<op1; ++i2)
{
for (int c=0; c<dim; ++c)
int ii_lex[3];
ii_lex[id0] = i0;
ii_lex[id1] = i1;
ii_lex[id2] = i2;
const int ii_el = i_off + ii_lex[0] + ii_lex[1]*nxi + ii_lex[2]*nxi*nyi;
for (int cj_rel=0; cj_rel<dim; ++cj_rel)
{
for (int j=0; j<nnz_per_row; ++j)
const int cj = (ci + cj_rel) % 3;
const int j_off = cj*o*op1*op1;
const int nxj = (cj == 0) ? o : op1;
const int nyj = (cj == 1) ? o : op1;
const int j0_begin = i0;
const int j0_end = (cj_rel == 0) ? i0 : i0 + 1;
const int j1_begin = (i1 > 0) ? i1-1 : i1;
const int j1_end = (cj_rel == 1)
? ((i1 < o) ? i1 : i1-1)
: ((i1 < o) ? i1+1 : i1);
const int j2_begin = (i2 > 0) ? i2-1 : i2;
const int j2_end = (cj_rel == 2)
? ((i2 < o) ? i2 : i2-1)
: ((i2 < o) ? i2+1 : i2);
for (int j0=j0_begin; j0<=j0_end; ++j0)
{
V(j,ix+iy*op1+iz*op1*op1,c,iel_ho) = 0.0;
const int d0 = j0 - i0;
for (int j1=j1_begin; j1<=j1_end; ++j1)
{
const int d1 = j1 - i1 + 1;
for (int j2=j2_begin; j2<=j2_end; ++j2)
{
const int d2 = j2 - i2 + 1;
int jj_lex[3];
jj_lex[id0] = j0;
jj_lex[id1] = j1;
jj_lex[id2] = j2;
const int jj_el = j_off + jj_lex[0] + jj_lex[1]*nxj + jj_lex[2]*nxj*nyj;
int jj_off;
if (cj_rel == 0) { jj_off = d1 + 3*d2; }
else if (cj_rel == 1) { jj_off = 9 + d0 + 2*d1 + 4*d2; }
else /* if (cj_rel == 2) */ { jj_off = 21 + d0 + 2*d1 + 6*d2; }
map(jj_off, ii_el) = jj_el;
}
}
}
}
}
}
}
MFEM_SYNC_THREAD;
MFEM_SHARED double local_mat_[sz_local_mat];
DeviceTensor<4> local_mat(local_mat_, 3,4, 3,4);
/// should be optimized
for (int i=0; i<sz_local_mat; ++i) { local_mat[i] = 0.0; }
MFEM_SHARED double vx[8], vy[8], vz[8];
/// should be optimized
LORVertexCoordinates3D<1>(X, iel_ho, 0,0,0, vx,vy,vz);
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(xyz,x,8)
{
const int qz = xyz%2, qy = (xyz/2)%2, qx = xyz/2/2;
static constexpr double w = 1.0/8.0;
double J_[3*3];
DeviceTensor<2> J(J_, 3,3);
Jacobian3D(qx,qy,qz, vx,vy,vz, J);
const double detJ = Det3D(J);
const double w_detJ = w/detJ;
// adj(J)
double A_[3*3];
DeviceTensor<2> A(A_, 3,3);
Adjugate3D(J, A);
const double Q0 = w_detJ*(A(0,0)*A(0,0)+A(0,1)*A(0,1)+A(0,2)*A(0,2)); // 1,1
const double Q1 = w_detJ*(A(0,0)*A(1,0)+A(0,1)*A(1,1)+A(0,2)*A(1,2)); // 2,1
const double Q2 = w_detJ*(A(0,0)*A(2,0)+A(0,1)*A(2,1)+A(0,2)*A(2,2)); // 3,1
const double Q3 = w_detJ*(A(1,0)*A(1,0)+A(1,1)*A(1,1)+A(1,2)*A(1,2)); // 2,2
const double Q4 = w_detJ*(A(1,0)*A(2,0)+A(1,1)*A(2,1)+A(1,2)*A(2,2)); // 3,2
const double Q5 = w_detJ*(A(2,0)*A(2,0)+A(2,1)*A(2,1)+A(2,2)*A(2,2)); // 3,3
// w J^T J / det(J)
const double Q6 = w_detJ*(J(0,0)*J(0,0)+J(1,0)*J(1,0)+J(2,0)*J(2,0)); // 1,1
const double Q7 = w_detJ*(J(0,0)*J(0,1)+J(1,0)*J(1,1)+J(2,0)*J(2,1)); // 2,1
const double Q8 = w_detJ*(J(0,0)*J(0,2)+J(1,0)*J(1,2)+J(2,0)*J(2,2)); // 3,1
const double Q9 = w_detJ*(J(0,1)*J(0,1)+J(1,1)*J(1,1)+J(2,1)*J(2,1)); // 2,2
const double Q10 = w_detJ*(J(0,1)*J(0,2)+J(1,1)*J(1,2)+J(2,1)*J(2,2)); // 3,2
const double Q11 = w_detJ*(J(0,2)*J(0,2)+J(1,2)*J(1,2)+J(2,2)*J(2,2)); // 3,3
const double mq = const_mq ? MQ(0,0,0,0) : MQ(qx,qy,qz, iel_ho);
const double dq = const_dq ? DQ(0,0,0,0) : DQ(qx,qy,qz, iel_ho);
// Loop over x,y,z components. 0 => x, 1 => y, 2 => z
MFEM_FOREACH_THREAD(cj,y,dim)
{
const double jq1 = (cj == 0) ? qy : ((cj == 1) ? qz : qx);
const double jq2 = (cj == 0) ? qz : ((cj == 1) ? qx : qy);
const int jd_0 = cj;
const int jd_1 = (cj + 1)%3;
const int jd_2 = (cj + 2)%3;
MFEM_FOREACH_THREAD(bj,z,4) // 4 edges in each dim
{
const int bj1 = bj%2;
const int bj2 = bj/2;
double curl_j[3];
curl_j[jd_0] = 0.0;
curl_j[jd_1] = ((bj1 == 0) ? jq1 - 1 : -jq1)*((bj2 == 0) ? 1 : -1);
curl_j[jd_2] = ((bj2 == 0) ? 1 - jq2 : jq2)*((bj1 == 0) ? 1 : -1);
double basis_j[3];
basis_j[jd_0] = ((bj1 == 0) ? 1 - jq1 : jq1)*((bj2 == 0) ? 1 - jq2 : jq2);
basis_j[jd_1] = 0.0;
basis_j[jd_2] = 0.0;
const int jj_loc = bj + 4*cj;
for (int ci=0; ci<dim; ++ci)
{
const double iq1 = (ci == 0) ? qy : ((ci == 1) ? qz : qx);
const double iq2 = (ci == 0) ? qz : ((ci == 1) ? qx : qy);
const int id_0 = ci, id_1 = (ci + 1)%3, id_2 = (ci + 2)%3;
for (int bi=0; bi<4; ++bi)
{
const int bi1 = bi%2, bi2 = bi/2;
double curl_i[3];
curl_i[id_0] = 0.0;
curl_i[id_1] = ((bi1 == 0) ? iq1 - 1 : -iq1)*((bi2 == 0) ? 1 : -1);
curl_i[id_2] = ((bi2 == 0) ? 1 - iq2 : iq2)*((bi1 == 0) ? 1 : -1);
double basis_i[3];
basis_i[id_0] = ((bi1 == 0) ? 1 - iq1 : iq1)*((bi2 == 0) ? 1 - iq2 : iq2);
basis_i[id_1] = 0.0;
basis_i[id_2] = 0.0;
const int ii_loc = bi + 4*ci;
// Only store the lower-triangular part of
// the matrix (by symmetry).
if (jj_loc > ii_loc) { continue; }
double curl_curl = 0.0;
curl_curl += Q6*curl_i[0]*curl_j[0];
curl_curl += Q7*(curl_i[0]*curl_j[1] + curl_i[1]*curl_j[0]);
curl_curl += Q8*(curl_i[0]*curl_j[2] + curl_i[2]*curl_j[0]);
curl_curl += Q9*curl_i[1]*curl_j[1];
curl_curl += Q10*(curl_i[1]*curl_j[2] + curl_i[2]*curl_j[1]);
curl_curl += Q11*curl_i[2]*curl_j[2];
double basis_basis = 0.0;
basis_basis += Q0*basis_i[0]*basis_j[0];
basis_basis += Q1*(basis_i[0]*basis_j[1] + basis_i[1]*basis_j[0]);
basis_basis += Q2*(basis_i[0]*basis_j[2] + basis_i[2]*basis_j[0]);
basis_basis += Q3*basis_i[1]*basis_j[1];
basis_basis += Q4*(basis_i[1]*basis_j[2] + basis_i[2]*basis_j[1]);
basis_basis += Q5*basis_i[2]*basis_j[2];
const double val = dq*curl_curl + mq*basis_basis;
AtomicAdd(local_mat(ci,bi, cj,bj), val);
} // bi
} // ci
} // bj
} // cj
} // q
MFEM_SYNC_THREAD;
// Assemble the local matrix into the macro-element sparse matrix
// The nonzeros of the macro-element sparse matrix are ordered as
// follows:
//
// The axes are ordered relative to the direction of the basis
// vector, e.g. for x-vectors, the axes are (x,y,z), for
// y-vectors the axes are (y,z,x), and for z-vectors the axes are
// (z,x,y).
//
// The nonzeros are then given in "rotated lexicographic"
// ordering, according to these axes.
if (MFEM_THREAD_ID(x) == 0 && MFEM_THREAD_ID(y) == 0)
{
MFEM_FOREACH_THREAD(ii_loc,z,ne)
{
const int ci = ii_loc/4, bi = ii_loc%4;
const int i0 = 0, i1 = bi%2, i2 = bi/2;
const int id0 = ci, id1 = (ci+1)%3, id2 = (ci+2)%3;
int ii_lex[3];
ii_lex[id0] = i0, ii_lex[id1] = i1, ii_lex[id2] = i2;
const int nx = (ci == 0) ? o : op1, ny = (ci == 1) ? o : op1;
const int ii = ii_lex[0] + (ii_lex[1])*nx + (ii_lex[2])*nx*ny;
for (int jj_loc=0; jj_loc<ne; ++jj_loc)
{
const int cj = jj_loc/4, bj = jj_loc%4;
// add 3 to take modulus (rather than remainder) when
// (cj - ci) is negative
const int cj_rel = (3 + cj - ci)%3;
const int jd0 = cj_rel, jd1 = (cj_rel+1)%3, jd2 = (cj_rel+2)%3;
int jj_rel[3];
jj_rel[jd0] = 0, jj_rel[jd1] = bj%2, jj_rel[jd2] = bj/2;
const int d0 = jj_rel[0] - i0;
const int d1 = 1 + jj_rel[1] - i1;
const int d2 = 1 + jj_rel[2] - i2;
const int jj_off = (cj_rel == 0) ? d1 + 3*d2 :
(cj_rel == 1) ? 9 + d0 + 2*d1 + 4*d2 :
(cj_rel == 2) ? 21 + d0 + 2*d1 + 6*d2 : -1;
// Symmetry
const double val = (jj_loc <= ii_loc)
? local_mat(ci,bi, cj,bj)
: local_mat(cj,bj, ci,bi);
AtomicAdd(V(jj_off, ii, ci, iel_ho), val);
}
}
}
});
SparseMapping3D<1>(sparse_mapping);
}
}
// Explicit template instantiations
@@ -809,7 +556,7 @@ template void BatchedLOR_ND::Assemble2D<6>();
template void BatchedLOR_ND::Assemble2D<7>();
template void BatchedLOR_ND::Assemble2D<8>();
//template void BatchedLOR_ND::Assemble3D<1>(); // explicitly specialized
template void BatchedLOR_ND::Assemble3D<1>();
template void BatchedLOR_ND::Assemble3D<2>();
template void BatchedLOR_ND::Assemble3D<3>();
template void BatchedLOR_ND::Assemble3D<4>();
@@ -825,8 +572,28 @@ BatchedLOR_ND::BatchedLOR_ND(BilinearForm &a,
Array<int> &sparse_mapping_)
: BatchedLORKernel(fes_ho_, X_vert_, sparse_ij_, sparse_mapping_)
{
ProjectLORCoefficient<VectorFEMassIntegrator>(a, c1);
ProjectLORCoefficient<CurlCurlIntegrator>(a, c2);
VectorFEMassIntegrator *mass = GetIntegrator<VectorFEMassIntegrator>(a);
if (mass != nullptr)
{
auto *coeff = dynamic_cast<const ConstantCoefficient*>(mass->GetCoefficient());
mass_coeff = coeff ? coeff->constant : 1.0;
}
else
{
mass_coeff = 0.0;
}
CurlCurlIntegrator *diffusion = GetIntegrator<CurlCurlIntegrator>(a);
if (diffusion != nullptr)
{
auto *coeff = dynamic_cast<const ConstantCoefficient*>
(diffusion->GetCoefficient());
curl_curl_coeff = coeff ? coeff->constant : 1.0;
}
else
{
curl_curl_coeff = 0.0;
}
}
} // namespace mfem
+2
View File
@@ -21,6 +21,8 @@ namespace mfem
// classes BatchedLORAssembly and BatchedLORKernel .
class BatchedLOR_ND : BatchedLORKernel
{
protected:
double mass_coeff, curl_curl_coeff;
public:
template <int ORDER> void Assemble2D();
template <int ORDER> void Assemble3D();
+91 -279
View File
@@ -33,14 +33,8 @@ void BatchedLOR_RT::Assemble2D()
static constexpr int nnz_per_row = 7;
static constexpr int sz_local_mat = ne*ne;
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1)
: Reshape(c1.Read(), op1, op1, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1)
: Reshape(c2.Read(), op1, op1, nel_ho);
const double DQ = div_div_coeff;
const double MQ = mass_coeff;
sparse_ij.SetSize(nnz_per_row*ndof_per_el*nel_ho);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, o*op1, dim, nel_ho);
@@ -108,8 +102,6 @@ void BatchedLOR_RT::Assemble2D()
{
for (int iqy=0; iqy<2; ++iqy)
{
const double mq = const_mq ? MQ(0,0,0) : MQ(kx+iqx, ky+iqy, iel_ho);
const double dq = const_dq ? DQ(0,0,0) : DQ(kx+iqx, ky+iqy, iel_ho);
// Loop over x,y components. c=0 => x, c=1 => y
for (int cj=0; cj<dim; ++cj)
{
@@ -140,8 +132,8 @@ void BatchedLOR_RT::Assemble2D()
val += byi*bxj*Q(1,iqy,iqx);
val += bxi*byj*Q(1,iqy,iqx);
val += byi*byj*Q(2,iqy,iqx);
val *= mq;
val += dq*div_j*div_i*Q(3,iqy,iqx);
val *= MQ;
val += DQ*div_j*div_i*Q(3,iqy,iqx);
local_mat(ii_loc, jj_loc) += val;
}
@@ -233,87 +225,6 @@ void BatchedLOR_RT::Assemble2D()
}
}
template<int ORDER>
static void SparseMapping3D(Array<int> &sparse_mapping)
{
static constexpr int nnz_per_row = 11;
static constexpr int dim = 3;
static constexpr int o = ORDER;
static constexpr int op1 = ORDER + 1;
static constexpr int ndof_per_el = dim*o*o*op1;
sparse_mapping.SetSize(nnz_per_row*ndof_per_el);
sparse_mapping = -1;
auto map = Reshape(sparse_mapping.HostReadWrite(), nnz_per_row, ndof_per_el);
for (int ci=0; ci<dim; ++ci)
{
const int i_off = ci*o*o*op1;
const int id0 = ci;
const int id1 = (ci+1)%3;
const int id2 = (ci+2)%3;
const int nxi = (ci == 0) ? op1 : o;
const int nyi = (ci == 1) ? op1 : o;
for (int i0=0; i0<op1; ++i0)
{
for (int i1=0; i1<o; ++i1)
{
for (int i2=0; i2<o; ++i2)
{
int ii_lex[3];
ii_lex[id0] = i0;
ii_lex[id1] = i1;
ii_lex[id2] = i2;
const int ii_el = i_off + ii_lex[0] + ii_lex[1]*nxi + ii_lex[2]*nxi*nyi;
for (int cj_rel=0; cj_rel<dim; ++cj_rel)
{
const int cj = (ci + cj_rel) % 3;
const int j_off = cj*o*o*op1;
const int nxj = (cj == 0) ? op1 : o;
const int nyj = (cj == 1) ? op1 : o;
const int j0_begin = (i0 > 0) ? i0-1 : i0;
const int j0_end = (cj_rel == 0)
? ((i0 < o) ? i0+1 : i0)
: ((i0 < o) ? i0 : i0-1);
const int j1_begin = i1;
const int j1_end = (cj_rel == 1) ? i1+1 : i1;
const int j2_begin = i2;
const int j2_end = (cj_rel == 2) ? i2+1 : i2;
for (int j0=j0_begin; j0<=j0_end; ++j0)
{
const int d0 = 1 + j0 - i0;
for (int j1=j1_begin; j1<=j1_end; ++j1)
{
const int d1 = j1 - i1;
for (int j2=j2_begin; j2<=j2_end; ++j2)
{
const int d2 = j2 - i2;
int jj_lex[3];
jj_lex[id0] = j0;
jj_lex[id1] = j1;
jj_lex[id2] = j2;
const int jj_el = j_off + jj_lex[0] + jj_lex[1]*nxj + jj_lex[2]*nxj*nyj;
int jj_off;
if (cj_rel == 0) { jj_off = d0; }
else if (cj_rel == 1) { jj_off = 3 + d0 + 2*d1; }
else /* if (cj_rel == 2) */ { jj_off = 7 + d0 + 2*d2; }
map(jj_off, ii_el) = jj_el;
}
}
}
}
}
}
}
}
}
template <int ORDER>
void BatchedLOR_RT::Assemble3D()
{
@@ -330,19 +241,13 @@ void BatchedLOR_RT::Assemble3D()
static constexpr int nnz_per_row = 11;
static constexpr int sz_local_mat = nf*nf;
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1, 1)
: Reshape(c1.Read(), op1, op1, op1, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1, 1)
: Reshape(c2.Read(), op1, op1, op1, nel_ho);
const double DQ = div_div_coeff;
const double MQ = mass_coeff;
sparse_ij.SetSize(nnz_per_row*ndof_per_el*nel_ho);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, o*o*op1, dim, nel_ho);
const auto X = X_vert.Read();
auto X = X_vert.Read();
// Last thread dimension is lowered to avoid "too many resources" error
MFEM_FORALL_3D(iel_ho, nel_ho, ORDER, ORDER, (ORDER>6)?4:ORDER,
@@ -418,8 +323,6 @@ void BatchedLOR_RT::Assemble3D()
{
for (int iqx=0; iqx<2; ++iqx)
{
const double mq = const_mq ? MQ(0,0,0,0) : MQ(kx+iqx, ky+iqy, kz+iqz, iel_ho);
const double dq = const_dq ? DQ(0,0,0,0) : DQ(kx+iqx, ky+iqy, kz+iqz, iel_ho);
// Loop over x,y,z components. 0 => x, 1 => y, 2 => z
for (int cj=0; cj<dim; ++cj)
{
@@ -473,7 +376,7 @@ void BatchedLOR_RT::Assemble3D()
basis_basis += Q(4,iqz,iqy,iqx)*(basis_i[1]*basis_j[2] + basis_i[2]*basis_j[1]);
basis_basis += Q(5,iqz,iqy,iqx)*basis_i[2]*basis_j[2];
const double val = dq*div_div + mq*basis_basis;
const double val = DQ*div_div + MQ*basis_basis;
// const double val = 1.0;
local_mat(ii_loc, jj_loc) += val;
@@ -555,185 +458,76 @@ void BatchedLOR_RT::Assemble3D()
}
}
});
SparseMapping3D<ORDER>(sparse_mapping);
}
template <>
void BatchedLOR_RT::Assemble3D<1>()
{
static constexpr int o = 1;
static constexpr int op1 = 2;
static constexpr int nf = 6; // number of faces in hexahedron
static constexpr int dim = 3;
static constexpr int ndof_per_el = dim*o*o*op1;
static constexpr int nnz_per_row = 11;
static constexpr int sz_local_mat = nf*nf;
const int nel_ho = fes_ho.GetNE();
const bool const_mq = c1.Size() == 1;
const auto MQ = const_mq
? Reshape(c1.Read(), 1, 1, 1, 1)
: Reshape(c1.Read(), op1, op1, op1, nel_ho);
const bool const_dq = c2.Size() == 1;
const auto DQ = const_dq
? Reshape(c2.Read(), 1, 1, 1, 1)
: Reshape(c2.Read(), op1, op1, op1, nel_ho);
sparse_ij.SetSize(nnz_per_row*ndof_per_el*nel_ho);
auto V = Reshape(sparse_ij.Write(), nnz_per_row, o*o*op1, dim, nel_ho);
const auto X = X_vert.Read();
MFEM_FORALL_3D(iel_ho, nel_ho, 8, 1, 1,
sparse_mapping.SetSize(nnz_per_row*ndof_per_el);
sparse_mapping = -1;
auto map = Reshape(sparse_mapping.HostReadWrite(), nnz_per_row, ndof_per_el);
for (int ci=0; ci<dim; ++ci)
{
MFEM_FOREACH_THREAD(j,x,nnz_per_row)
const int i_off = ci*o*o*op1;
const int id0 = ci;
const int id1 = (ci+1)%3;
const int id2 = (ci+2)%3;
const int nxi = (ci == 0) ? op1 : o;
const int nyi = (ci == 1) ? op1 : o;
for (int i0=0; i0<op1; ++i0)
{
for (int ix = 0; ix < op1; ++ix)
for (int i1=0; i1<o; ++i1)
{
for (int c = 0; c < dim; ++c)
for (int i2=0; i2<o; ++i2)
{
V(j,ix,c,iel_ho) = 0.0;
int ii_lex[3];
ii_lex[id0] = i0;
ii_lex[id1] = i1;
ii_lex[id2] = i2;
const int ii_el = i_off + ii_lex[0] + ii_lex[1]*nxi + ii_lex[2]*nxi*nyi;
for (int cj_rel=0; cj_rel<dim; ++cj_rel)
{
const int cj = (ci + cj_rel) % 3;
const int j_off = cj*o*o*op1;
const int nxj = (cj == 0) ? op1 : o;
const int nyj = (cj == 1) ? op1 : o;
const int j0_begin = (i0 > 0) ? i0-1 : i0;
const int j0_end = (cj_rel == 0)
? ((i0 < o) ? i0+1 : i0)
: ((i0 < o) ? i0 : i0-1);
const int j1_begin = i1;
const int j1_end = (cj_rel == 1) ? i1+1 : i1;
const int j2_begin = i2;
const int j2_end = (cj_rel == 2) ? i2+1 : i2;
for (int j0=j0_begin; j0<=j0_end; ++j0)
{
const int d0 = 1 + j0 - i0;
for (int j1=j1_begin; j1<=j1_end; ++j1)
{
const int d1 = j1 - i1;
for (int j2=j2_begin; j2<=j2_end; ++j2)
{
const int d2 = j2 - i2;
int jj_lex[3];
jj_lex[id0] = j0;
jj_lex[id1] = j1;
jj_lex[id2] = j2;
const int jj_el = j_off + jj_lex[0] + jj_lex[1]*nxj + jj_lex[2]*nxj*nyj;
int jj_off;
if (cj_rel == 0) { jj_off = d0; }
else if (cj_rel == 1) { jj_off = 3 + d0 + 2*d1; }
else /* if (cj_rel == 2) */ { jj_off = 7 + d0 + 2*d2; }
map(jj_off, ii_el) = jj_el;
}
}
}
}
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(xyz,x,8)
{
const int qz = xyz%2, qy = (xyz/2)%2, qx = xyz/2/2;
static constexpr double w = 1.0/8.0;
double local_mat_[sz_local_mat];
DeviceTensor<4> local_mat(local_mat_, 3,2, 3,2);
for (int i=0; i<sz_local_mat; ++i) { local_mat[i] = 0.0; }
double vx[8], vy[8], vz[8];
LORVertexCoordinates3D<1>(X, iel_ho, 0,0,0, vx, vy, vz);
double J_[3*3];
DeviceTensor<2> J(J_, 3,3);
Jacobian3D(qx,qy,qz, vx,vy,vz, J);
const double detJ = Det3D(J);
const double w_detJ = w/detJ;
const double Q0 = w_detJ*(J(0,0)*J(0,0)+J(1,0)*J(1,0)+J(2,0)*J(2,0)); // 1,1
const double Q1 = w_detJ*(J(0,1)*J(0,0)+J(1,1)*J(1,0)+J(2,1)*J(2,0)); // 2,1
const double Q2 = w_detJ*(J(0,2)*J(0,0)+J(1,2)*J(1,0)+J(2,2)*J(2,0)); // 3,1
const double Q3 = w_detJ*(J(0,1)*J(0,1)+J(1,1)*J(1,1)+J(2,1)*J(2,1)); // 2,2
const double Q4 = w_detJ*(J(0,2)*J(0,1)+J(1,2)*J(1,1)+J(2,2)*J(2,1)); // 3,2
const double Q5 = w_detJ*(J(0,2)*J(0,2)+J(1,2)*J(1,2)+J(2,2)*J(2,2)); // 3,3
const double Q6 = w_detJ;
const double mq = const_mq ? MQ(0,0,0,0) : MQ(qx,qy,qz, iel_ho);
const double dq = const_dq ? DQ(0,0,0,0) : DQ(qx,qy,qz, iel_ho);
// Loop over x,y,z components. 0 => x, 1 => y, 2 => z
for (int cj=0; cj<dim; ++cj)//MFEM_FOREACH_THREAD(cj,y,dim)
{
const int jq0 = (cj == 0) ? qx : (cj == 1) ? qy : qz;
const int jd0 = cj, jd1 = (cj + 1)%3, jd2 = (cj + 2)%3;
for (int bj=0; bj<2; ++bj)//MFEM_FOREACH_THREAD(bj,z,2) // 2 faces in each dim
{
const double div_j = (bj == 0) ? -1.0 : 1.0;
double basis_j[3];
basis_j[jd0] = (bj == jq0) ? 1.0 : 0.0;
basis_j[jd1] = 0.0;
basis_j[jd2] = 0.0;
const int jj_loc = bj + 2*cj;
for (int ci=0; ci<dim; ++ci)
{
const double iq0 = (ci == 0) ? qx : ((ci == 1) ? qy : qz);
const int id0 = ci, id1 = (ci + 1)%3, id2 = (ci + 2)%3;
for (int bi=0; bi<2; ++bi)
{
const double div_i = (bi == 0) ? -1.0 : 1.0;
double basis_i[3];
basis_i[id0] = (bi == iq0) ? 1.0 : 0.0;
basis_i[id1] = 0.0;
basis_i[id2] = 0.0;
const int ii_loc = bi + 2*ci;
// Only store the lower-triangular part of
// the matrix (by symmetry).
if (jj_loc > ii_loc) { continue; }
const double div_div = Q6*div_i*div_j;
double basis_basis = 0.0;
basis_basis += Q0*basis_i[0]*basis_j[0];
basis_basis += Q1*(basis_i[0]*basis_j[1] + basis_i[1]*basis_j[0]);
basis_basis += Q2*(basis_i[0]*basis_j[2] + basis_i[2]*basis_j[0]);
basis_basis += Q3*basis_i[1]*basis_j[1];
basis_basis += Q4*(basis_i[1]*basis_j[2] + basis_i[2]*basis_j[1]);
basis_basis += Q5*basis_i[2]*basis_j[2];
const double val = dq*div_div + mq*basis_basis;
local_mat(ci,bi, cj,bj) += val;
} // bi
} // ci
} // bj
} // cj
// Assemble the local matrix into the macro-element sparse matrix
// The nonzeros of the macro-element sparse matrix are ordered as
// follows:
//
// The axes are ordered relative to the direction of the basis
// vector, e.g. for x-vectors, the axes are (x,y,z), for
// y-vectors the axes are (y,z,x), and for z-vectors the axes are
// (z,x,y).
//
// The nonzeros are then given in "rotated lexicographic"
// ordering, according to these axes.
for (int ii_loc=0; ii_loc<nf; ++ii_loc)
{
const int ci = ii_loc/2, bi = ii_loc%2;
const int id0 = ci, id1 = (ci+1)%3, id2 = (ci+2)%3;
const int i0 = bi, i1 = 0, i2 = 0;
int ii_lex[3];
ii_lex[id0] = i0, ii_lex[id1] = i1, ii_lex[id2] = i2;
const int nx = (ci == 0) ? op1 : o;
const int ny = (ci == 1) ? op1 : o;
const int ii = ii_lex[0] + ii_lex[1]*nx + ii_lex[2]*nx*ny;
for (int jj_loc=0; jj_loc<nf; ++jj_loc)
{
const int cj = jj_loc/2, bj = jj_loc%2;
// add 3 to take modulus (rather than remainder) when
// (cj - ci) is negative
const int cj_rel = (3 + cj - ci)%3;
const int jd0 = cj_rel, jd1 = (cj_rel+1)%3, jd2 = (cj_rel+2)%3;
int jj_rel[3];
jj_rel[jd0] = bj, jj_rel[jd1] = 0, jj_rel[jd2] = 0;
const int d0 = jj_rel[0] - i0 + 1;
const int d1 = jj_rel[1] - i1;
const int d2 = jj_rel[2] - i2;
const int jj_off = (cj_rel == 0) ?d0 :
(cj_rel == 1) ? 3 + d0 + 2*d1 :
(cj_rel == 2) ? 7 + d0 + 2*d2 : -1;
// Symmetry
const double val = (jj_loc <= ii_loc)
? local_mat(ci,bi, cj,bj)
: local_mat(cj,bj, ci,bi);
AtomicAdd(V(jj_off, ii, ci, iel_ho), val);
} // jj_loc
} // ii_loc
} // q
});
SparseMapping3D<1>(sparse_mapping);
}
}
// Explicit template instantiations
@@ -746,7 +540,7 @@ template void BatchedLOR_RT::Assemble2D<6>();
template void BatchedLOR_RT::Assemble2D<7>();
template void BatchedLOR_RT::Assemble2D<8>();
//template void BatchedLOR_RT::Assemble3D<1>(); // explicitly specialized
template void BatchedLOR_RT::Assemble3D<1>();
template void BatchedLOR_RT::Assemble3D<2>();
template void BatchedLOR_RT::Assemble3D<3>();
template void BatchedLOR_RT::Assemble3D<4>();
@@ -762,8 +556,26 @@ BatchedLOR_RT::BatchedLOR_RT(BilinearForm &a,
Array<int> &sparse_mapping_)
: BatchedLORKernel(fes_ho_, X_vert_, sparse_ij_, sparse_mapping_)
{
ProjectLORCoefficient<VectorFEMassIntegrator>(a, c1);
ProjectLORCoefficient<DivDivIntegrator>(a, c2);
if (VectorFEMassIntegrator *mass = GetIntegrator<VectorFEMassIntegrator>(a))
{
auto *coeff = dynamic_cast<const ConstantCoefficient*>(mass->GetCoefficient());
mass_coeff = coeff ? coeff->constant : 1.0;
}
else
{
mass_coeff = 0.0;
}
if (DivDivIntegrator *divdiv = GetIntegrator<DivDivIntegrator>(a))
{
auto *coeff = dynamic_cast<const ConstantCoefficient*>
(divdiv->GetCoefficient());
div_div_coeff = coeff ? coeff->constant : 1.0;
}
else
{
div_div_coeff = 0.0;
}
}
} // namespace mfem
+2
View File
@@ -21,6 +21,8 @@ namespace mfem
// classes BatchedLORAssembly and BatchedLORKernel .
class BatchedLOR_RT : BatchedLORKernel
{
protected:
double mass_coeff, div_div_coeff;
public:
template <int ORDER> void Assemble2D();
template <int ORDER> void Assemble3D();
+1 -1
View File
@@ -19,7 +19,7 @@ namespace mfem
/*!
* @brief Interface for mortar element assembly.
* The MortarIntegrator interface is used for performing Petrov-Galerkin
* The MortarIntegrator interface is used for performing Pertrov-Galerkin
* finite element assembly on intersections between elements.
* The quadrature rules are to be generated by a cut algorithm (e.g.,
* mfem::Cut). The quadrature rules are defined in the respective trial and test
+2 -2
View File
@@ -157,14 +157,14 @@ public:
/// Return a (read-only) list of all essential true dofs.
const Array<int> &GetEssentialTrueDofs() const { return ess_tdof_list; }
/// Compute the energy corresponding to the state @a x.
/// Compute the enery corresponding to the state @a x.
/** In general, @a x may have non-homogeneous essential boundary values.
The state @a x must be a "GridFunction size" vector, i.e. its size must
be fes->GetVSize(). */
double GetGridFunctionEnergy(const Vector &x) const;
/// Compute the energy corresponding to the state @a x.
/// Compute the enery corresponding to the state @a x.
/** In general, @a x may have non-homogeneous essential boundary values.
The state @a x must be a true-dof vector. */
+1 -4
View File
@@ -16,9 +16,6 @@
#include "fem.hpp"
#include "../general/sort_pairs.hpp"
#define MFEM_NVTX_COLOR Crimson
#include "../general/nvtx.hpp"
namespace mfem
{
@@ -127,7 +124,6 @@ void ParBilinearForm::pAllocMat()
void ParBilinearForm::ParallelRAP(SparseMatrix &loc_A, OperatorHandle &A,
bool steal_loc_A)
{
NVTX("RAP");
ParFiniteElementSpace &pfespace = *ParFESpace();
// Create a block diagonal parallel matrix
@@ -407,6 +403,7 @@ void ParBilinearForm::FormLinearSystem(
P.MultTranspose(b, true_B);
R.Mult(x, true_X);
p_mat.EliminateBC(p_mat_e, ess_tdof_list, true_X, true_B);
R.EnsureMultTranspose();
R.MultTranspose(true_B, b);
hybridization->ReduceRHS(true_B, B);
X.SetSize(B.Size());
+9 -1
View File
@@ -129,7 +129,7 @@ void ParFiniteElementSpace::ParInit(ParMesh *pm)
ApplyLDofSigns(*elem_dof);
}
// Check for shared triangular faces with interior Nedelec DoFs
// Check for shared trianglular faces with interior Nedelec DoFs
CheckNDSTriaDofs();
}
@@ -959,6 +959,10 @@ void ParFiniteElementSpace::Build_Dof_TrueDof_Matrix() const // matrix P
SparseMatrix Pdiag;
P->GetDiag(Pdiag);
R = Transpose(Pdiag);
// The following call ensures that the action of the transpose of P is
// performed fast when HYPRE is built for GPUs.
P->EnsureMultTranspose();
}
HypreParMatrix *ParFiniteElementSpace::GetPartialConformingInterpolation()
@@ -2624,6 +2628,10 @@ int ParFiniteElementSpace
{
*P_ = MakeVDimHypreMatrix(pmatrix, ndofs, num_true_dofs,
dof_offs, tdof_offs);
// The following call ensures that the action of the transpose of *P_ is
// performed fast when HYPRE is built for GPUs.
(*P_)->EnsureMultTranspose();
}
// clean up possible remaining messages in the queue to avoid receiving

Some files were not shown because too many files have changed in this diff Show More