Compare commits

...
Author SHA1 Message Date
Jakub Červený 94db67a915 Removed blobs, added .gitignore for this directory. 2020-01-03 11:09:09 +01:00
Jakub Červený 00e47ecf81 Cleanup, test code. 2019-12-11 16:58:40 +01:00
Jakub Červený d4271203a0 Improved patch-based version works (sampling regular grid with context). 2019-12-11 16:43:29 +01:00
Jakub Červený eb98055c9d Better implementation of sampling master faces, works now. 2019-12-06 14:43:45 +01:00
Jakub Červený fae164ab3a Added function GetMasterRestriction to sample points on slave faces. 2019-12-05 18:09:18 +01:00
Jakub Červený b3d7ed1889 Slave faces and orientations should work now. 2019-11-29 16:34:48 +01:00
Jakub Červený bd549ce8e4 Conforming neighbor faces work, with derivatives. 2019-11-25 21:37:35 +01:00
Jakub Červený 39dfda881b Debugging patch-based drl4amr. 2019-11-22 20:33:18 +01:00
Jakub Červený e0ff49b50b WIP Drl4Amr::GetLocalImage, conforming faces 2019-11-19 14:50:23 +01:00
Jakub Červený 0cd45041b4 WIP rasterizing local image (element + neighbor edges) 2019-11-14 16:26:41 +01:00
camierjs d8d8f3002b Cleanup 2019-11-01 16:10:28 -07:00
Jakub Cerveny 1d15781516 Better GetImage works. 2019-11-01 17:52:06 +01:00
Jakub Červený b0f0c5df5c WIP better rasterization (Drl4Amr::GetImage()). 2019-11-01 14:57:33 +01:00
camierjs bdf46c9629 Python cleanup @ MacOS 2019-10-31 22:43:07 -07:00
camierjs 454137a10a MacOs python package clean 2019-10-31 22:10:56 -07:00
camierjs 6a343bd456 Get back the image to python 2019-10-31 17:12:31 -07:00
camierjs 895ad83f06 Vis width & height 2019-10-31 08:25:48 -07:00
camierjs dd3bc1928a GetImage simplifications & projection 2019-10-30 19:12:59 -07:00
camierjs 96f0b44dc9 Cleanup 2019-10-30 18:32:06 -07:00
camierjs 6b7569ad69 MacOS fix 2019-10-30 18:22:50 -07:00
camierjs bb7ef48e11 NOTMAC makefile fix 2019-10-30 17:04:15 -07:00
camierjs accded0c4d Intermediate mesh before image output 2019-10-30 16:03:17 -07:00
camierjs dec3f7890a Refinement per element id 2019-10-30 10:30:48 -07:00
camierjs 797f4f2b24 Interpolation x vs xcoeff + norm 2019-10-29 19:05:36 -07:00
camierjs ddb33ce7f0 x0 setup 2019-10-29 17:49:56 -07:00
camierjs 332dcc6ca8 Wip init x0 2019-10-29 11:17:31 -07:00
camierjs 0fc14255cc Add main.cpp and makefile c++/python tests 2019-10-29 10:53:34 -07:00
camierjs 3b9a8be840 Compute, Refine & Update API 2019-10-29 10:25:16 -07:00
camierjs f53439cce3 Cleanup miniapps/drl4amr, switch to ctypes & cdll 2019-10-29 09:35:16 -07:00
camierjs 6b6f0c643a drl4amr examples launching ex6 2019-10-28 18:55:30 -07:00
Aaron Fisher c55c80d17b Merge pull request #923 from mfem/sundials-interface
Sundials 5.0 Interface [sundials-interface]
2019-10-17 14:11:47 -07:00
Aaron Fisher ddac9ed690 Updated the default locations for packages in the mk files. 2019-10-16 15:59:03 -07:00
Aaron Fisher d38e7b55d9 Merge branch 'master' into sundials-interface 2019-10-16 12:39:20 -07:00
Veselin Dobrev 7a869c30c6 In class TimeDependentOperator, add a new property: evaluation
mode, see TimeDependentOperator::SetEvalMode() for description.
This new property is now used to support IMEX time integration
in class ARKStepSolver. With this addition the method
TimeDependentOperator::SUNImplicitMult() is no longer needed
and was removed.
2019-10-11 23:37:49 -07:00
Veselin Dobrev 7c350fa8dd Minor formatting edits. 2019-10-03 22:43:48 -07:00
Veselin Dobrev 5560274fd6 Fix memory leaks in SUNDIALS examples 9/10/10p.
The fix in example 9 uses the approach used in the regular
example 9 to fix the same leak (see PR #816).

Fix a memory leak in SUNDIALS example 9 using the same fix
that was used
2019-10-03 22:10:40 -07:00
Veselin Dobrev ceb91142ed Adjust the command line parameters in the CMake build system
for testing the SUNDIALS example 10/10p to match the ones in
the GNU make files.
2019-10-03 20:42:43 -07:00
Veselin Dobrev 8cc3eb04ae In the SUNDIALS example 10/10p:
- Add smaller tolerances for the linear/nonlinear solves in
  CVODE/ARKODE -- this is necessary due to the use of very large
  relative+absolute tolerances (1e-1) for the time accuracy.
- Restore a few removed sample runs using CVODE+ADAMS and adjust
  their time step.
- Revert the time step for two CVODE+BDF sample runs to the values
  used in the master branch.

In the INSTALL file, update the MFEM version where only SUNDIALS
v5.0.0+ is supported: MFEM v4.0.2 --> MFEM v4.1.
2019-10-03 20:08:51 -07:00
Veselin Dobrev 959eea2830 Some small tweaks:
- Apply 'make style'
- Add a note in the CVODE/ARKODE/KINSOL Init() methods that
  re-initializing may purge some user-set options
- When updating the operator of the CVODE/AKODE/KINSOL objects in
  their respective SetOperator() methods, use an all-reduce to
  make sure all ranks do this in a consistent manner.
2019-10-03 16:45:50 -07:00
Jean-Sylvain CAMIER 632857c15d Merge pull request #1073 from mfem/dcpo-fix
Fix Pconf allows CUDA => DEVICE_MASK
2019-09-23 09:13:14 -07:00
David J. Gardner ebac5cd124 correct input to ARKStepSetTableNum 2019-09-20 13:47:40 -07:00
camierjs 4edd165ff5 Fix Pconf allows CUDA => DEVICE_MASK 2019-09-17 08:47:11 -07:00
Jean-Sylvain CAMIER 3f1adff17d Merge pull request #1050 from mfem/raja-wrap
[GPU] Raja Cuda Wrap 2D & 3D using Kernel Policies [raja-wrap]
2019-09-04 12:55:25 -07:00
camierjs 4f4ecce2ac Merge branch 'master' into raja-wrap 2019-09-04 12:02:33 -07:00
Jean-Sylvain CAMIER 34bd444257 Merge pull request #1038 from mfem/dmpi
[GPU] Aware MPI kernels [dmpi]
2019-09-04 11:54:46 -07:00
camierjs ec9001e6e9 Revert to Raja Kernels 2019-08-30 17:50:33 -07:00
camierjs b478d4c89e Merge branch 'master' into raja-wrap 2019-08-30 17:41:10 -07:00
Tzanio Kolev dc68d860f6 Merge branch 'master' into sundials-interface 2019-08-29 15:16:26 -07:00
camierjs cde8f8530e Merge branch 'raja-wrap' of github.com:mfem/mfem into raja-wrap 2019-08-29 11:13:12 -07:00
camierjs 7a4b1c0b2f MFEM_USE_RAJA_FORALL_ND try 2019-08-29 11:12:39 -07:00
camierjs 1ad1a70293 Revert MFEM_GPU_CHECK undef check 2019-08-29 10:41:49 -07:00
Tzanio 25ea5e9208 Merge branch 'master' into dmpi
Conflicts:
	general/device.cpp
2019-08-28 18:40:39 -07:00
Tzanio Kolev f8a3a379d1 Merge pull request #1009 from mfem/rocm
HIP support
2019-08-28 18:24:40 -07:00
Tzanio 42ddb401af Mentioned the RAJA improvements in CHANGELOG 2019-08-28 18:15:54 -07:00
camierjs 3ab9217f02 Makefile combination verifications 2019-08-28 18:00:17 -07:00
camierjs 64f83231a0 Touch to relaunch to re-build in AppVeyor.
Last build failed: "Build execution time has reached the maximum allowed time for your plan".
2019-08-28 17:17:47 -07:00
camierjs f0e36f71fc Raja forall with 2D & 3D kernel policies 2019-08-28 17:05:22 -07:00
camierjs 8353494bd0 RajaCudaWrap3D using KernelPolicy, RajaCudaWrap2D still WIP 2019-08-28 14:28:09 -07:00
camierjs 4bb27ed051 Move GPU prefix calls to MFEM_GPU 2019-08-28 10:02:38 -07:00
camierjs 2a6227e6f6 Cleanup old comment and remove mfem::out message 2019-08-28 09:48:47 -07:00
camierjs 9f1e923b02 ifndef MFEM_CUDA_CHECK => MFEM_GPU_CHECK 2019-08-26 10:11:20 -07:00
Tzanio Kolev a4617d76b3 Merge pull request #1043 from goxberry/bugfix-makefile-occa-dev
makefile: add missing parenthesis to conditional
2019-08-26 09:43:06 -07:00
Tzanio Kolev 6973800a20 Merge pull request #1030 from mfem/array_const_int_fix
A few small tweaks to support Arrays of const T objects [array_const_int_fix]
2019-08-26 09:42:57 -07:00
Geoffrey M Oxberry e2147e51e2 makefile: add missing parenthesis to conditional
This commit adds a missing parenthesis to a conditional in the MFEM
makefile that checks if PREFIX is set when MFEM_USE_OCCA equals YES.
2019-08-25 00:30:15 -07:00
Tzanio ed0c911ee3 Small updates 2019-08-23 16:14:14 -07:00
camierjs f67d502138 Revert to use engine kernels 2019-08-23 12:33:51 -07:00
David J. Gardner 2849ea1fb4 fix linear solve output in ex10p 2019-08-22 22:53:01 -07:00
David J. Gardner 2c8a87dc5d fix linear solve output in ex10 2019-08-22 22:12:22 -07:00
David J. Gardner e08232488e clarify what UseMFEMMassLinearSolver and UseSundialsMassLinearSolver attach 2019-08-22 11:24:45 -07:00
David J. Gardner 3dd964eecf clarify what UseMFEMLinearSolver and UseSundialsLinearSolver attach 2019-08-22 10:31:50 -07:00
Veselin Dobrev 185ed1d4e0 Small updates in comments. 2019-08-21 22:27:30 -07:00
Tzanio d3accd9ee8 Small change in INSTALL 2019-08-21 18:42:46 -07:00
camierjs 454c1b41a3 makefile hip targets updates and warning fix 2019-08-21 18:25:01 -07:00
camierjs 6f105f2c04 INSTALL and makefile updates 2019-08-21 18:12:48 -07:00
camierjs 994ae7ed78 Update Device::Configure backend priority list 2019-08-21 17:46:23 -07:00
camierjs 48985c6447 Remove mfem::out put 2019-08-21 17:45:00 -07:00
Tzanio 5b9ca638ec minor 2019-08-21 17:19:34 -07:00
camierjs c2995218e9 global and shared keywords 2019-08-20 18:12:23 -07:00
camierjs 56a59f6725 examples/ex1p.cpp device option 2019-08-20 18:09:01 -07:00
camierjs ffafd43bf4 Remove warnings from CudaConformingProlongationOperator and CudaGroupCommunicator 2019-08-20 17:52:13 -07:00
Aaron Fisher 8d8e71dbbc Improved the SUNDIALS interface documentation a bit. 2019-08-20 15:44:45 -07:00
camierjs 04ce993757 Cleanup warnings 2019-08-20 15:33:38 -07:00
camierjs eae9341fe8 Cleanup WIP 2019-08-20 13:17:09 -07:00
camierjs 64c3341df9 CudaConformingProlongationOperator and GroupCommunicator 2019-08-19 18:16:22 -07:00
camierjs a9b99f3601 Device selection 2019-08-19 12:28:14 -07:00
camierjs 931d19578f Cleanup 2019-08-16 17:36:42 -07:00
camierjs f3e6b4d9a2 GPU Aware MPI through DeviceConformingProlongationOperator 2019-08-16 11:33:08 -07:00
camierjs 5f14a12c48 gpu_aware_mpi, but last Mult 2019-08-15 19:36:25 -07:00
camierjs ba7fc29455 auto send_buf = ext_buf.Write() + send_offset 2019-08-15 19:18:49 -07:00
camierjs b868a45da6 Cleanup 2019-08-15 19:09:24 -07:00
David J. Gardner 3830906b62 fix to set maa before KINInit() 2019-08-15 18:13:30 -07:00
camierjs 2bf9cca31d CUDA MPI with buffers 2019-08-15 18:09:10 -07:00
David J. Gardner 269dac7933 add wrapper for KINSetMAA 2019-08-15 17:42:04 -07:00
David J. Gardner 02ed1504a3 update comment 2019-08-15 17:34:09 -07:00
David J. Gardner d9973bf879 remove resize flag from sundials base class 2019-08-15 17:33:54 -07:00
David J. Gardner 9e868eff4f update arkstep reinit/resize 2019-08-15 17:16:17 -07:00
David J. Gardner 4038b2e212 update cvode reinit/resize 2019-08-15 17:16:17 -07:00
David J. Gardner 83a3a40302 remove extra includes, add comments 2019-08-15 17:16:16 -07:00
David J. Gardner 5b897053ba note which methods must be called after SetOperator 2019-08-15 17:16:16 -07:00
David J. Gardner 7b2bf5a88f update kinsol setoperator for reinitializing/resizing 2019-08-15 17:16:10 -07:00
David J. Gardner f6bcd67b58 attach MFEM ls if prec is non-null 2019-08-15 13:34:14 -07:00
David J. Gardner 9ecc568e13 fix setting FuncNormTol 2019-08-15 13:30:49 -07:00
David J. Gardner f4ff8099e8 free A if non-null and attaching new ls 2019-08-15 13:30:18 -07:00
Veselin Dobrev 18f0cf0873 Merge pull request #714 from najlkin/pr6
Minor improvements of BilinearForm and MixedBilinearForm [najlkin:pr6]
2019-08-13 19:00:12 -07:00
Veselin Dobrev 418271e61e Merge pull request #917 from mfem/hypreparmat-copyconstr
Add copy constructor (deep copy data) to HypreParMatrix
2019-08-13 18:59:26 -07:00
Veselin Dobrev dff56a705b Merge pull request #992 from mfem/variable-names-fix
Fix variable names beginning with underscore [variable-names-fix]
2019-08-13 18:58:01 -07:00
Veselin Dobrev 4e9246faf5 Merge pull request #998 from mfem/bugfix-eval3d
Bugfix eval3d (and Diffusion)
2019-08-13 18:57:16 -07:00
Veselin Dobrev a35279d17f Merge pull request #1003 from mfem/bugfix/get-bdrelem-face-trans-dev
Setting element number in transformation objects [bugfix/get-bdrelem-face-trans-dev]
2019-08-13 18:56:23 -07:00
David J. Gardner 05226e6e79 add mass matrix mult wraper if using ARKode mass matrix support 2019-08-09 16:24:09 -07:00
David J. Gardner c4f88ab02a attach MFEM linear solver by default 2019-08-09 15:32:45 -07:00
David J. Gardner cef852ec19 wrap long lines 2019-08-09 14:58:06 -07:00
David J. Gardner 5700636721 revise CVODE wrapper to use Init(f) rather than Init(f, t, x) 2019-08-09 12:14:23 -07:00
David J. Gardner 1822da30b5 remove trailing whitespace 2019-08-09 12:04:45 -07:00
David J. Gardner e6994f5d66 remove ARKStep Create utility function 2019-08-09 12:01:16 -07:00
David J. Gardner b87916cead fix typo in comments 2019-08-09 11:36:59 -07:00
David J. Gardner c86634bc11 revise ARKStep wrapper to use Init(f) rather than Init(f, t, x) 2019-08-09 11:33:04 -07:00
Aaron Fisher 1c4bee5def Remove the t parameter from SUNMassSetup and SUNImplicitSetup. 2019-08-08 15:08:53 -07:00
Aaron Fisher e3cb07a4ec Swapped the locations of the x,b parameters in SUNImplicitSolve and SUNMassSolve. Also updated the build system defaults to point to the proper place in for SUNDIALS5.0. 2019-08-08 14:47:35 -07:00
Aaron Fisher 4f5fb640df Updated the CHANGELOG/INSTALL for the upgrade to SUNDIALS 5.0. 2019-08-08 10:24:40 -07:00
Veselin Dobrev d967d12b39 A few tweaks to support Array<const int> objects. 2019-08-08 07:03:02 -04:00
David J. Gardner 754d9b62cf add SUN prefix to implicit setup/solve and mass setup/solve 2019-08-05 17:05:29 -07:00
David J. Gardner 6f8a71f961 remove f2 from ode class, add SUNImplicitMult for IMEX problem 2019-08-05 16:56:54 -07:00
David J. Gardner acf0be8304 remove old LS interface 2019-08-05 15:13:48 -07:00
David J. Gardner 6e7b82d403 update examples 10, 10p, and 16 to new LS interface 2019-08-05 14:55:05 -07:00
camierjs 00cab9eb6c WIP direct CUDA MPI 2019-08-02 18:15:00 -07:00
camierjs 6414a14c07 dmpi converges 2019-08-02 10:04:50 -07:00
camierjs 0d1bb17f99 Direct MPI from Engines 2019-08-01 15:11:06 -07:00
Veselin Dobrev 8bb5309cbc In MixedBilinearForm::Assemble, use local loop variables everywhere. 2019-07-30 16:58:19 -07:00
Tzanio f8f261f32d minor 2019-07-26 23:08:01 -07:00
Jan Nikl 4f8678c9e4 Fixed for-init scoping in bilinearform.cpp 2019-07-25 20:25:53 +02:00
camierjs 850c049386 std namespace fix 2019-07-25 09:35:36 -07:00
camierjs e754227754 Review updates 2019-07-25 09:31:53 -07:00
camierjs 762613713c rocm => hip 2019-07-24 17:55:53 -07:00
camierjs 4ab197b21b general/rocm.hpp double defines fix 2019-07-24 11:46:08 -07:00
camierjs 9323158b73 config/defaults.mk fix 2019-07-23 18:15:06 -07:00
camierjs 486cffbba9 Merge branch 'master' into rocm 2019-07-23 13:55:29 -07:00
camierjs 5bf6cad024 GPU, CUDA, ROCM calls 2019-07-23 13:29:14 -07:00
Veselin Dobrev ef67de8c61 Merge branch 'master' into najlkin/pr6 2019-07-19 20:44:21 -07:00
Veselin Dobrev 4d0eca20e4 Update a comment in class MixedBilinearForm. 2019-07-19 20:35:36 -07:00
Stowell, Mark L 2395e9765d Setting element number in transformation objects 2019-07-14 15:09:40 -07:00
Tzanio Kolev 369fc7c903 Merge pull request #1001 from mfem/hypre-appveyor-update
In .appveyor.yml, update the source location for hypre [hypre-appveyor-update]
2019-07-12 20:16:15 -07:00
Veselin Dobrev 4c8a94e32e In .appveyor.yml, fix the hypre build for the new .tar.gz source. 2019-07-12 18:39:35 -07:00
Veselin Dobrev f30c5b0a5c In .appveyor.yml, update the source location for hypre. 2019-07-12 18:17:15 -07:00
Veselin Dobrev 50d95d0618 Various small fixes in SUNDIALS-related code. 2019-07-12 17:58:32 -07:00
Veselin Dobrev d33da44075 Fix a copy-paste bug in the SUNDIALS versions of ex9 and ex9p. 2019-07-11 19:20:15 -07:00
Veselin Dobrev deb6286d9e Some tweaks in the doxygen documentation of class TimeDependentOperator. 2019-07-11 17:51:10 -07:00
Veselin Dobrev aa27304d17 Some tweaks of doxygen documentation in linalg/sundials.hpp. 2019-07-10 21:55:21 -07:00
Veselin Dobrev 43a5097176 Fix errors in 'make test' in the SUNDIALS versions of ex9/ex9p.
Add a version check for SUNDIALS v5.0.0 in linalg/sundials.hpp.
2019-07-10 20:00:39 -07:00
Michael Franco 36fa400f27 Only allocate as much memory as needed in 3D case 2019-07-10 16:13:05 -07:00
Michael Franco 4a40065cfc Fix Eval3D bug issue #997 2019-07-10 16:09:58 -07:00
Veselin Dobrev 41bc2aec88 Fix doxygen warnings.
Fix build warnings when using the option -Wall.

In class SundialsLinearSolver, move the default implementations
of the methods ODELinSys and ODEMassSys to the .cpp file.
2019-07-08 20:21:13 -07:00
Veselin Dobrev 40795c21c6 Merge pull request #981 from mfem/out-of-source-build-fix
Fix an issue with the out-of-source build [out-of-source-build-fix]
2019-07-03 14:10:46 -07:00
Veselin Dobrev 4e4c2bd638 Merge pull request #980 from mfem/bugfix-const-correctness
Fix const correctness in eliminate* functions [bugfix-const-correctness]
2019-07-03 14:09:53 -07:00
Veselin Dobrev 5b24fd6676 Rename some variables beginning with '_' to instead end
with '_'. Some of these variables were causing issues
under cygwin.
2019-07-03 14:00:14 -07:00
Veselin Dobrev 588d254043 In class TimeDependentOperator, move the default implementations
of the virtual methods to the .cpp file.
2019-07-02 19:02:48 -07:00
Veselin Dobrev 100b2077fa Apply 'make style' 2019-07-02 17:55:01 -07:00
David J. Gardner b67e5b1b99 remove old comment 2019-06-28 10:46:05 -07:00
David J. Gardner 590583c60e Merge branch 'master' into sundials-interface
Conflicts:
  linalg/ode.hpp
2019-06-28 10:43:05 -07:00
Veselin Dobrev 878a82cf00 Merge pull request #963 from mfem/lor-mesh-bugfix
Fix bug when creating low-order refined meshes in parallel [lor-mesh-bugfix]
2019-06-25 16:14:05 -07:00
Veselin Dobrev ae8f987156 Merge pull request #960 from mfem/lininteg-quadrature-order
Increase default quadrature in vector linear forms [lininteg-quadrature-order]
2019-06-25 16:13:30 -07:00
Veselin Dobrev a67bc51147 Merge pull request #959 from mfem/fix-quadratic-cubit-hex
Fix reading quadratic hex meshes from cubit [fix-quadratic-cubit-hex]
2019-06-25 16:12:51 -07:00
Veselin Dobrev cec95e06d3 Merge pull request #736 from mfem/rectangular-parallel-operator
Implement Operator::FormDiscreteOperator() [rectangular-parallel-operator]
2019-06-25 16:11:40 -07:00
Michael Franco 29d2fdee65 Fix const correctness in eliminate* functions 2019-06-25 11:34:45 -07:00
Veselin Dobrev 24970c1e01 Fix an issue with the out-of-source build when the build path
contains tokens like 'linux' or 'unix'.

The fix is to define the complete path to the config file as a
string macro instead of trying to concatenate the build path with
'/config/_config.hpp' and then stringify it.
2019-06-24 20:43:55 -07:00
Veselin Dobrev 318967d694 Fix issues uncovered by the regression tests. 2019-06-20 12:34:28 -07:00
Will Pazner b5649e8599 Merge branch 'lor-mesh-bugfix' of github.com:mfem/mfem into lor-mesh-bugfix 2019-06-17 10:34:23 -07:00
Will Pazner 357efe7524 Fix bug when creating LOR mesh in parallel
Boundary elements were inserted into mesh partitions that
legitimately had no boundaries.
2019-06-17 10:34:02 -07:00
Tzanio Kolev 3320cb796c Merge pull request #941 from mfem/const-array-iterators
Add const iterators to Array class
2019-06-16 19:48:52 +02:00
Tzanio Kolev 142f31e48a Merge pull request #940 from mfem/bugfix-duplicate-flop-count
Fix issue #938
2019-06-16 19:48:22 +02:00
Will Pazner 8ba7663222 Increase default quadrature in vector linear forms
To ensure h^{p+1} convergence, the quadrature used for the
right-hand side needs to be sufficiently accurate. This makes the
vector versions of LinearFormIntegrator use the same degree of
exactness as the scalar versions.
2019-06-15 14:08:32 -07:00
Julian Andrej 62df58dd37 Fix reading quadratic hex meshes from cubit 2019-06-14 15:57:36 -07:00
Andrew T. Barker 50b28d2396 Operator: fix error from mfem 4.0 merge. 2019-06-12 15:52:06 -07:00
Andrew T. Barker e3da15847d Merge remote-tracking branch 'origin/master' into rectangular-parallel-operator
Conflicts:
	linalg/operator.cpp
2019-06-05 15:38:30 -07:00
Socratis 791634a7fc Make sure rowstarts and colstarts are copied and communicators are cloned. 2019-06-04 11:13:50 -07:00
Will Pazner 57b9d5a0b3 Move global flop_count to mfem::internal namespace 2019-05-30 16:51:34 -07:00
camierjs e15d1a6b9c ex1.cpp benchmarks 2019-05-30 16:50:48 -07:00
Will Pazner 089c3bfce1 Add const iterators to Array class 2019-05-30 16:43:10 -07:00
Will Pazner 0aa0e36139 Fix issue #938
Global variable defined in tconfig.hpp would cause duplicate
symbol errors when linking.
2019-05-30 16:38:09 -07:00
Will Pazner 363c82277a Temporary fix for flop_count non-static global 2019-05-29 14:15:49 -07:00
camierjs 3fce1070c6 ex1 @ ROCm converges 2019-05-29 12:32:27 -07:00
camierjs 689fb17113 Vector::operator* tries 2019-05-28 21:07:04 -07:00
camierjs f1c9da0164 First ex1 run on Radeon Instinct MI25 2019-05-28 20:46:08 -07:00
camierjs 4f7447f053 First pass toward ROCm 2019-05-28 20:24:55 -07:00
Veselin Dobrev edfb62d8c2 Merge pull request #932 from mfem/hypre-errors-dev
Flexible handling of hypre errors [hypre-errors-dev]
2019-05-28 12:42:55 -07:00
Tzanio f417834319 Expanding the comments for the cases when error_mode = IGNORE_HYPRE_ERRORS;
is used.

Adding it to AMS, which could also have this issue (though may be rare).
2019-05-26 13:15:34 -07:00
Veselin Dobrev daa301f6f7 Add support for more flexible handling of hypre errors in
class HypreSolver.

The default is still to abort on hypre errors, except in some
special cases -- see the documentation of the new method
HypreSolver::SetErrorMode() for details.
2019-05-26 12:00:00 -07:00
Veselin Dobrev edbe2affc7 Update version numbers to 4.0.1 -- a new development version. 2019-05-25 08:20:15 -07:00
Tzanio 4d900b0c5f Preparing for v4.0 release 2019-05-24 18:50:01 -07:00
Tzanio Kolev dd6d3c642a Merge pull request #913 from mfem/memory-dev
Add Memory class [memory-dev]
2019-05-24 16:31:55 -07:00
Tzanio b14e78d5fb Updated CHANGELOG and the documentation in doc/ 2019-05-24 16:26:55 -07:00
Veselin Dobrev 31238435af Small fix in the doxygen comments in class Memory. 2019-05-24 16:16:39 -07:00
camierjs 32f1a33dd2 Small renaming in SmemPAMassApply3D kernel 2019-05-24 15:57:13 -07:00
Tzanio 32d7e036e7 Renamed
MemoryType GetSuitableMemoryType(MemoryClass mc);

to

  MemoryType GetMemoryType(MemoryClass mc);
2019-05-24 15:54:24 -07:00
Tzanio 2c1d07c127 Merge branch 'memory-dev' of github.com:mfem/mfem into memory-dev 2019-05-24 15:41:51 -07:00
Tzanio 342fb5058f Renamed MFEM_FORALL_IF -> MFEM_FORALL_SWITCH 2019-05-24 15:41:47 -07:00
camierjs c78de477bf Remove !MFEM_USE_SUBVECTOR_KERNELS code sections 2019-05-24 15:33:08 -07:00
Tzanio 0401024513 Minor. 2019-05-24 15:32:24 -07:00
Tzanio b3c3c5cf4c Merge branch 'memory-dev' of github.com:mfem/mfem into memory-dev 2019-05-24 14:23:56 -07:00
Tzanio 1a69dcff78 Replaced FIXMEs with TODO or NOTE 2019-05-24 14:23:39 -07:00
camierjs bd08fa9592 Merge branch 'memory-dev' of github.com:mfem/mfem into memory-dev 2019-05-24 14:21:12 -07:00
camierjs 283264ade4 {Read,Write,ReadWrite}Access => {Read,Write,ReadWrite}
Add shortcut for Host{Read,Write,ReadWrite}
2019-05-24 14:08:45 -07:00
Tzanio bee66cbcef Merge branch 'memory-dev' of github.com:mfem/mfem into memory-dev 2019-05-24 14:02:56 -07:00
Tzanio e36d1ea8fd Several change to (hopefully) simplify the interface:
* The parameter of Vector::UseDevice(bool) no longer has a default value. All
  the calls to UseDevice() in operator.cpp bilinearform_ext.cpp gridfunc.?pp and
  linearform.?pp have been replaced with UseDevice(true).

* Added shortcuts for the device flags of the Memory objects inside the Vector
  and Array classes:

    bool Array::UseDevice()
    bool Vector::UseDevice()

* The internal device flag in Memory::FlagMask is now called USE_DEVICE (it was
  previously called EXEC_FLAG). The accessor function for this flag have been
  renamed:

    bool Memory::GetExecFlag()     -> bool Memory::UseDevice()
    void Memory::SetExecFlag(bool) -> void Memory::UseDevice(bool)

* Further renamed:

    Memory::SyncWith         -> Memory::Sync
    Memory::SyncAliasToBase  -> Memory::SyncAlias
    Memory::SyncAliasToBase_ -> Memory::SyncAlias_

* Replaced FIXME with TODO in general/mem_manager.cpp
2019-05-24 14:01:15 -07:00
camierjs 48ad7cf360 Revert ATTR and constexpr comments 2019-05-24 12:38:06 -07:00
Tzanio 87ceaf15d3 Merge branch 'memory-dev' of github.com:mfem/mfem into memory-dev 2019-05-24 12:12:57 -07:00
camierjs 913be3d6fe Merge branch 'memory-dev' of github.com:mfem/mfem into memory-dev 2019-05-24 11:00:25 -07:00
camierjs b3c18e6d5e Introduce MFEM_FOREACH_THREAD, MFEM_THREAD_ID and MFEM_THREAD_SIZE
Add ATTR to MFEM_ATTR_SHARED
2019-05-24 10:59:40 -07:00
Veselin Dobrev 75065b0ea1 Addressing some feedback from the PR. 2019-05-24 10:25:34 -07:00
Tzanio 52696ab6ce Added an internal "hpc" target to the makefile which builds with MPI and
all currently available backends.

We may choose to advertise this later, but for now it is mostly for
developers and testing.
2019-05-24 09:28:02 -07:00
Veselin Dobrev b69fd1e038 Some tweaks and additions to the MemoryManager class.
Modify the Device class to require the creation of an object in
order to use backends other then Backend::CPU. At destriction,
this object will call the Destroy() method of the MemoryManager to
deallocate any remaining registered device pointers.

In class Device, remove the method Disable() and make the method
Enable() private.

Use a global Array<double> as the buffer used by the cuda functions
for minimum and dot product.
2019-05-23 23:21:27 -07:00
camierjs 0afec3491a Fem diffusion and mass kernels w/o Bt and Gt 2019-05-23 14:22:09 -07:00
David J. Gardner 567e9c39a9 remove orig files 2019-05-23 09:56:31 -07:00
Veselin Dobrev 798ded1f55 Merge branch 'master' into memory-dev 2019-05-22 23:53:04 -07:00
Tzanio Kolev bc92fc1a9f Merge pull request #922 from mfem/mesh-ext-additions
Improve the integration of some of the new GPU classes with existing classes [mesh-ext-additions]
2019-05-22 19:15:14 -07:00
Tzanio bfb9540bbd Updated README.html files 2019-05-22 19:10:54 -07:00
Tzanio 08f1bf7ab7 Updated CHANGELOG 2019-05-22 18:38:52 -07:00
Tzanio 26f137ae88 Merge branch 'mesh-ext-additions' of github.com:mfem/mfem into mesh-ext-additions 2019-05-22 17:37:10 -07:00
Veselin Dobrev 5e8117e112 Remove bilininteg_ext.cpp from the CMake build system. 2019-05-22 17:34:52 -07:00
Veselin Dobrev 504f01aa75 Merge branch 'master' into mesh-ext-additions
Moved implementation from fem/bilininteg_ext.cpp into
fem/bilininteg_diffusion.cpp and fem/bilininteg_mass.cpp

Update the layouts used in the shared memory kernels.
2019-05-22 17:31:04 -07:00
Tzanio 7bb68afa73 Patch from Stefano for HYPRE_MIXEDINT. 2019-05-22 16:59:25 -07:00
Tzanio b128777209 Another minor 2019-05-22 16:14:49 -07:00
Tzanio 387682795f Merge branch 'mesh-ext-additions' of github.com:mfem/mfem into mesh-ext-additions 2019-05-22 16:03:05 -07:00
Tzanio 14c2ea6dd2 minor 2019-05-22 15:44:31 -07:00
camierjs 73feb74bc1 Avoid cudaErrorCudartUnloading error while freeing cuda memory at exit. 2019-05-22 11:36:50 -07:00
Tzanio c3999ba78b A few shortcuts for Vector + Memory 2019-05-22 08:16:12 -07:00
Veselin Dobrev cdbcc9e5e1 Integrate class DofToQuad with the FiniteElement class.
Rename class ElemRestriction to ElementRestriction and integrate
it with class FiniteElementSpace; it is accessible with the method
FiniteElementSpace::GetElementRestriction().

Rename class XTMesh to GeometricFactors and remove invJ from the
possible geometric factors.

Introduce class QuadratureInterpolator (created and owned by class
FiniteElementSpace) that interpolates E-vectors to quadrature points,
see FiniteElementSpace::GetQuadratureInterpolator().

Switch the layouts of the E-vectors and Q-vectors to have the
local (element) DOFs and quadrature points, respectively, as the
fastest changing index.
2019-05-22 07:19:44 -07:00
Tzanio 41cf43c88c Minor 2019-05-22 03:44:49 -07:00
Tzanio Kolev cdcc96f074 Merge pull request #921 from mfem/shared-kernels
Mass + diffusion shared memory kernels [shared-kernels]
2019-05-21 20:28:27 -07:00
Tzanio e745fad2cf Small edits 2019-05-21 20:19:04 -07:00
Tzanio 3d88be85bc Removed MFEM_USE_MM -- it is no longer necessary 2019-05-21 19:18:50 -07:00
Tzanio 3a860f124c Added a call to Device::Enable at the end of Device::Configure.
Updated examples to use only Device::Configure (no more calls to Enable/Disable)
2019-05-21 19:17:06 -07:00
camierjs 00e5a1b736 make style and makefile revert pathnames 2019-05-21 18:23:27 -07:00
camierjs be5206688e Mass + diffusion shared memory kernels 2019-05-21 17:47:45 -07:00
Tzanio Kolev 274bc26e0e Merge pull request #756 from mfem/stefanozampini/small-improvements
Stefanozampini/small improvements
2019-05-21 16:11:36 -07:00
Tzanio 0b2a456137 Switching to beam-tet.mesh as the default in Example 19 2019-05-21 16:04:45 -07:00
Veselin Dobrev c0245d4865 Addressing some PR feedback. 2019-05-21 15:05:05 -07:00
Tzanio Kolev 902f34a4fd Merge pull request #918 from mfem/inf-reciprocal
XL compiler O3 -qnostrict (1/inf=nan) work-around [inf-reciprocal]
2019-05-21 07:58:26 -07:00
Veselin Dobrev 534bf51466 Addressing some of the PR feedback. 2019-05-21 02:19:05 -07:00
Tzanio Kolev e51f058263 Merge branch 'master' into memory-dev 2019-05-20 22:28:22 -07:00
Tzanio 4355a8951f Small clarification 2019-05-20 16:46:51 -07:00
camierjs f6610b15a7 XL compiler O3 -qnostrict (1/inf=nan) work-around 2019-05-20 14:57:29 -07:00
Julian Andrej 99471634bc Add copy constructor (deep copy data) to HypreParMatrix 2019-05-20 12:36:25 -07:00
Tzanio Kolev 5808fd93c5 Merge pull request #742 from mfem/gzdata-collection-dev
Added the gzstream capability to the data collection classes [gzdata-collection-dev]
2019-05-20 11:41:06 -07:00
Tzanio 6fd69230ff Fix uninitialised value found by valgrind 2019-05-20 11:40:11 -07:00
Tzanio Kolev 727125fa1d Merge pull request #911 from mfem/ex22-rename
Renamed Example 22 to Example 21 [ex22-rename]
2019-05-20 10:43:10 -07:00
Tzanio Kolev e70bf08ae8 Merge pull request #878 from mfem/vfe-dev
ExchangeFaceNbrData for VectorFE [vfe-dev]
2019-05-20 10:05:25 -07:00
Tzanio Kolev f18497b675 Merge pull request #908 from mfem/compiler-warnings
Fix nvcc warnings in make all [compiler-warnings]
2019-05-20 10:04:47 -07:00
Tzanio Kolev e4ccdeb697 Merge pull request #715 from najlkin/pr7
Fixed Mesh::GetBdrElementTransformation() collisions with the functions for faces [najlkin:pr7]
2019-05-20 09:58:47 -07:00
Tzanio Kolev 879b328504 Merge pull request #912 from mfem/occa-omp
Enable OpenMP foralls with all OMP_MASK'ed devices [occa-omp]
2019-05-20 09:54:25 -07:00
Veselin Dobrev 9cef797ebd Bugfix in Memory::SyncWith() 2019-05-19 10:33:07 -07:00
Veselin Dobrev c38115aab5 Comment out unused private variable in examples/petsc/ex10p.cpp to
suppress a warning.
2019-05-18 20:35:21 -07:00
Veselin Dobrev e332d2713f Address some FIXME comments. 2019-05-18 16:29:06 -07:00
Veselin Dobrev f8d1f5d557 Make sure GridFunction::MakeRef marks itself and its base vector
for execution on the mfem::Device.

Refine the logic in Vector::operator=.

Update the comment for Memory::SyncWith.
2019-05-18 00:40:43 -07:00
Veselin Dobrev 36188eafe5 Comment out some delete statements that can lead to double
deletion, e.g. in laghos.
2019-05-17 22:53:16 -07:00
Veselin Dobrev 0ec3fb3a46 Replace '#if 0' comments inside MFEM_FORALL macros with C++ style
comments.

These were generating warnings and also seem to break compilation
with Visual Studio.
2019-05-17 19:58:19 -07:00
Veselin Dobrev 31f2ce99cc Some small tweaks and additions related to the classes Memory and
MemoryManager.
2019-05-17 16:26:20 -07:00
Will Pazner a1e4d0ae90 Fix bug when creating LOR mesh in parallel
Boundary elements were inserted into mesh partitions that
legitimately had no boundaries.
2019-05-16 22:12:52 -07:00
camierjs 6cdead168e Add flags to compute J, invJ, detJ, X 2019-05-16 18:20:58 -07:00
camierjs c89a99b864 Update mesh/CMakeLists.txt 2019-05-16 16:12:26 -07:00
camierjs cb3daf3c5a Rename to XTMesh and cleanup 2019-05-16 16:10:45 -07:00
Veselin Dobrev 40378a046b Introduce a new Memory class for handling host + device allocations
and transfers.

The Memory class is now used by some MFEM classes (like Array and
Vector) which can be used on the Device. Such classes now provide
methods to access the underlying Memory object, e.g. GetMemory.

Updated ex1/ex1p and ex6/ex6p to not need to enable/disable the
Device at specific points -- the Device is now enabled just at the
start. Also, the same examples can now run on Device (e.g. -d cuda)
without the partial assembly option (-pa) -- full assembly will
be still done on CPU but the sparse matrix action and vector
operations will be done using the Device.

Reverted changes in class DenseMatrix related to using the Device.
At this point, DenseMatrix operations are only used for small matrices
and using the Device in this case is not a good option.
2019-05-16 14:53:39 -07:00
camierjs bce17bca45 Enable OpenMP foralls with all OMP_MASK'ed devices 2019-05-16 11:49:19 -07:00
camierjs 4110071899 Remove commented lines of unused variables 2019-05-16 10:21:06 -07:00
Tzanio fbf12be2cf Renamed Example 22 to Example 21 to close the gap in numbering before the
mfem-4.0 release.
2019-05-15 13:15:51 -07:00
David J. Gardner c5ced79ac7 update 16p to use new TimeDependentOperator methods 2019-05-15 12:37:03 -07:00
Tzanio 5b00c3d0e6 Merge branch 'master' into stefanozampini/small-improvements
Conflicts:
	fem/bilininteg.hpp
2019-05-15 11:38:18 -07:00
Tzanio ceb8f71e38 minor 2019-05-15 11:35:02 -07:00
Tzanio fa14a82fc8 make style 2019-05-15 11:26:29 -07:00
Stefano Zampini c5586d5e5a examples/petsc/ex10p: added PetscPreconditionerFactory example of usage
added matrix free tests
2019-05-15 11:29:06 +03:00
Stefano Zampini 90a83cb398 PetscSolver::SetPreconditionerFactory : prevent from segfaulting 2019-05-15 11:29:06 +03:00
Stefano Zampini 69981d62eb Fix for the -snes_mf_operator case
The rational here is that since MFEM has only one matrix returned by the GetGradient method,
it is that matrix that have be used to construct the preconditioner
2019-05-15 11:29:06 +03:00
Stefano Zampini 11da55e772 Fix deprecated function from PETSc 3.12 2019-05-15 11:29:06 +03:00
Tzanio 41d09fde1b make style 2019-05-14 20:58:37 -07:00
camierjs 714bb88be2 Fix nvcc warnings in make all 2019-05-13 11:23:43 -07:00
David J. Gardner cec74a9c2d update arkode to work with the new mass methods 2019-05-10 17:07:55 -07:00
David J. Gardner e209394abd update arkode to work with the new ls methods 2019-05-10 16:58:41 -07:00
David J. Gardner 98c9710b02 update cvode to work with new methods 2019-05-10 16:57:01 -07:00
David J. Gardner 1a52284203 add sundials specific methods to timedependent operator 2019-05-10 16:48:49 -07:00
David J. Gardner 957a01d81e add method to resize arkode 2019-05-10 15:58:12 -07:00
David J. Gardner 3ba5af74c3 update CVODE and ARKStep init to support reinitialization 2019-05-10 15:22:59 -07:00
David J. Gardner d9bdd0b9f9 add back MFEM steppers in ex9/9p 2019-05-09 16:07:38 -07:00
David J. Gardner 1ebac08b89 add back mfem steppers to ex16 2019-05-09 15:56:39 -07:00
David J. Gardner e3b44216ef minor ex10/10p updates 2019-05-09 15:34:46 -07:00
David J. Gardner 02c9f681ac fix cv and ark in ex10 2019-05-09 15:18:02 -07:00
David J. Gardner 07bc8ced69 update sundials ex10p 2019-05-09 15:00:42 -07:00
Yohann Dudouit 103f631925 Attempt to put GeometryExtension in the Mesh. 2019-05-09 14:32:18 -07:00
David J. Gardner f2ad3f1fc9 remove zeroing kin_pp
When connecting as SUNLinearSolver the initial guess is zeroed out
before calling the solve routine so this is not needed any more.
2019-05-08 10:09:43 -07:00
David J. Gardner e562b8a0c9 update ex10 2019-05-07 15:27:29 -07:00
David J. Gardner 1513847ff1 update KINSOL interface, add utility functions, clean up 2019-05-05 21:34:54 -07:00
Tzanio Kolev 21cdc4d8a3 Merge pull request #885 from mfem/bugfix/csr-mat-sum
Handle special case where B_offd is empty [bugfix/csr-mat-sum]
2019-05-01 20:58:28 -07:00
Tzanio Kolev b12684d9c6 Merge pull request #887 from mfem/mat-add-doc
Augmenting comments related to hypre_ParCSRMatrixSum [mat-add-doc]
2019-05-01 20:51:55 -07:00
Stowell, Mark L 675b137cbe Adding comments related to hypre_ParCSRMatrixSum 2019-04-25 22:12:07 -07:00
Stowell, Mark L 83db3da392 Handle special case where B_offd is empty 2019-04-24 16:31:45 -07:00
Tzanio Kolev c097bda546 Merge pull request #870 from mfem/4.0-rc2-docs
Updated documentation to mention GPU support [4.0-rc2-docs]
2019-04-24 14:02:40 -07:00
Tzanio Kolev 56608342e8 Merge pull request #875 from mfem/okina-ex6fix
Okina ex6fix [okina-ex6fix]
2019-04-24 14:02:05 -07:00
Tzanio Kolev a1c1da8e9c Merge pull request #882 from mfem/autotest-devices
Update device sample runs [autotest-devices]
2019-04-24 14:01:50 -07:00
Tzanio 72968077c6 CUDA driver no longer needed in top-level CMakeList.txt.
See https://github.com/mfem/mfem/pull/862#issuecomment-485171820.
2019-04-24 13:59:45 -07:00
Tzanio 9cebf45288 Merge branch 'master' into okina-ex6fix
Conflicts:
	general/cuda.hpp
2019-04-24 13:58:31 -07:00
Tzanio Kolev ffc2dfc70b Merge pull request #862 from mfem/okina-cmake
Okina-CMake: CUDA, OCCA, RAJA + MPI [okina-cmake]
2019-04-24 13:54:24 -07:00
Tzanio Kolev d996ee2d39 Merge pull request #868 from mfem/okina-nodrv
Remove CU driver calls [okina-nodrv]
2019-04-24 13:54:02 -07:00
Veselin Dobrev 49c25eec31 Fix a typo. 2019-04-24 12:10:25 -07:00
David J. Gardner fbba86c71d update ex9 and ex16 2019-04-24 11:42:55 -07:00
camierjs 27f8d46aae Add the -dev option to launch devices tests 2019-04-24 10:14:10 -07:00
Tzanio 4a64afedc2 doxygen fix 2019-04-24 09:00:16 -07:00
Tzanio 03d36aa518 Set the RC2 date to today 2019-04-24 07:07:45 -07:00
Veselin Dobrev b2f154112c Update the script config/sample-runs.sh to filter out device runs. 2019-04-23 21:47:50 -07:00
Tzanio 247f9c7445 Small edits 2019-04-23 21:39:21 -07:00
Tzanio 2baece3ab1 Adjusted documentation, mentioned in CHANGELOG 2019-04-23 21:32:43 -07:00
Veselin Dobrev 5b5769dea3 In class SparseMatrix:
* Add methods BuildTranspose() and ResetTranspose() that control
  the use of the internal transpose matrix, At.
* The method AddMultTranspose() will always use At, if it is built.
  If At is not build and the Device is enabled, an error will be
  generated pointing to BuildTranspose().
* Introduce separate non-const and const versions of the methods
  GetI(), GetJ(), and GetData().
* Made the method ActualWidth() const.
* Some edits in the documentation, the code formatting, and the
  error messages.

In the method BilinearForm::FormLinearSystem(), call the method
SparseMatrix::BuildTranspose() for the nonconforming prolongation
matrix, when necessary.
2019-04-23 17:48:40 -07:00
Tzanio f617acf414 Removed a comment about the CUDA driver (no longer needed). 2019-04-22 18:30:14 -07:00
Veselin Dobrev 68fbe31aa1 In config/defaults.mk, use '=' to set {OCCA,RAJA}_DIR.
This makes it easier to configure mfem by copying defaults.mk to
user.mk and editting it: if using '?=', the value given in user.mk
will not overwrite the one from defaults.mk.
2019-04-22 16:51:13 -07:00
Veselin Dobrev 8142e822d8 Fix non-CUDA builds. 2019-04-22 16:41:09 -07:00
Veselin Dobrev a7303349e0 Add checks that the memory manager is enabled when CUDA is enabled. 2019-04-22 16:26:46 -07:00
Veselin Dobrev 10b3988449 Reworked the macro MFEM_CUDA_CHECK:
* It always performs the error check, no just in debug mode.
* All 'cuda*' runtime calls are now wrapped with this macro.
2019-04-22 15:17:04 -07:00
Veselin Dobrev bc876e1c64 Remove the CUDA driver from the GNU make build system. 2019-04-22 15:14:18 -07:00
Tzanio 634d7e8de5 minor styling 2019-04-22 12:29:40 -07:00
Tzanio 8888cfb04b make style 2019-04-22 12:25:42 -07:00
Tzanio Kolev 1a8c36e92f Merge pull request #871 from mfem/bugfix/windows
Bugfix/windows
2019-04-21 21:44:57 -07:00
Pratyuksh Bansal 1a8fada58c Added conversion of local dofs for vector elements to AssembleSharedFaces 2019-04-21 16:56:59 +02:00
Pratyuksh Bansal 15639cec41 Add conversion of local dofs for vector elements in ExchangeFaceNbrData 2019-04-21 16:55:24 +02:00
Veselin Dobrev aa027c2b8d In the CMake build system, define CUDA_ARCH in defaults.cmake. 2019-04-19 21:06:57 -07:00
Veselin Dobrev 8a8d9419d1 Some small modifications in the build systems. 2019-04-19 20:51:16 -07:00
camierjs 6e82bd6ada Update new nodes before rebalancing 2019-04-19 17:01:07 -07:00
Tzanio c9d80fc64f minor 2019-04-19 15:58:20 -07:00
Tzanio c9762fe73e Update INSTALL: CMake + CUDA build, list Homebrew/Science as deprecated. 2019-04-19 15:40:04 -07:00
Tzanio 41f45474b2 Merge branch 'okina-cmake' of github.com:mfem/mfem into okina-cmake 2019-04-19 15:39:38 -07:00
Tzanio 4313a7b00f Small adjustemnt in XSDKDefaults.cmake 2019-04-19 15:29:03 -07:00
camierjs 3e4301a4cb Install the okl files 2019-04-19 14:57:04 -07:00
camierjs 6991239cc0 XSDKDefaults tweaks 2019-04-19 14:36:28 -07:00
camierjs 9655ceaaef Merge branch 'okina-cmake' of github.com:mfem/mfem into okina-cmake 2019-04-19 14:07:38 -07:00
camierjs b92b3acc1a Add CMAKE_CUDA_STANDARD/REQUIRED/EXTENSIONS 2019-04-19 14:06:50 -07:00
Tzanio 226ccf8db7 RC2-related changes in CHANGELOG 2019-04-19 13:15:01 -07:00
Tzanio 4ffe4a4beb styling 2019-04-19 12:40:49 -07:00
Tzanio dbaa40a116 Renamed tA to At 2019-04-19 12:31:58 -07:00
camierjs 8ad78ae156 Add MPI + CUDA support 2019-04-19 12:00:59 -07:00
camierjs 118a4dcde4 MFEM_CUDA_CHECK fix 2019-04-19 11:39:46 -07:00
camierjs 3e34d4b99a Cleanup, style and remove all CUdevice, CUcontext & CUstream 2019-04-19 11:00:23 -07:00
camierjs 06ee78f67c Switch to use Device::IsEnable 2019-04-19 10:59:01 -07:00
camierjs be8eac8997 Keep legacy AddMultTranspose code for non-accelerated runs 2019-04-19 10:16:02 -07:00
camierjs 890579e228 Cleanup and add a self-transposed sparse matrix that is used 2019-04-19 09:48:15 -07:00
camierjs ceb8bb1417 Remove AtomicAdd from sparsemat AddMultTranspose 2019-04-18 18:21:06 -07:00
Veselin Dobrev a06fe30a73 Fix the "check" CMake target for Visual Studio. 2019-04-18 16:07:54 -07:00
David J. Gardner 16e20eb471 use sundials mat and ls NewEmpty functions 2019-04-18 11:01:48 -07:00
jonesholger 2bb423434c Update .appveyor.yml 2019-04-17 23:02:30 -07:00
jonesholger 1bf5b9098f Update .appveyor.yml 2019-04-17 22:29:28 -07:00
jonesholger ed431414c2 Update .appveyor.yml 2019-04-17 22:17:43 -07:00
Holger Jones 0bbe93c26f Need to specify release config 2019-04-17 21:52:48 -07:00
Holger Jones d41d992798 modify test target to RUN_TESTS, which is known to work under windows 2019-04-17 21:31:32 -07:00
Holger Jones 881cc50cfd fix to bring in platform specific rmdir; lowered pts threshold in inversetransform test 2019-04-17 21:21:08 -07:00
Veselin Dobrev fb7be12a77 Merge pull request #863 from mfem/cmake-unit-tests-fix
Fix a bug in the CMake file for the unit tests [cmake-unit-tests-fix]
2019-04-17 19:43:59 -07:00
Veselin Dobrev 4e48ebc0cf Merge pull request #837 from rcarson3/hypre-dep-dev
Update hypre version and point users to the LLNL repository for hypre [rcarson3:hypre-dep-dev]
2019-04-17 19:41:39 -07:00
Tzanio ac12259cba Mention that MFEM_USE_LEGACY_OPENMP is deprecated in INSTALL. 2019-04-17 19:25:23 -07:00
Tzanio 99fbdcdf73 Mentioned GPU classes in doc/CodeDocumentation.dox 2019-04-17 18:12:55 -07:00
David J. Gardner 4ee8c71150 remove constructor taking sun_mem 2019-04-17 18:05:07 -07:00
David J. Gardner 57fb37f6ab fix typo 2019-04-17 18:04:46 -07:00
David J. Gardner ab8993cee6 remove unneeded variable 2019-04-17 18:01:55 -07:00
David J. Gardner 4afb8d724f updates for IMEX support 2019-04-17 18:01:01 -07:00
Tzanio aaead6c866 Updated README and CONTRIBUTING to mention GPUs 2019-04-17 18:00:21 -07:00
camierjs d3a1685cb6 Remove CU driver calls 2019-04-17 17:36:43 -07:00
David J. Gardner df3cea4cbc fix default ode opt 2019-04-17 16:54:01 -07:00
camierjs 16ca9883eb Use CMake 3.8 CUDA native support to compile MFEM + examples 2019-04-17 16:15:26 -07:00
Veselin Dobrev 2a0c8f25d3 Fix a bug in tests/unit/CMakeLists.txt that prevented building of
the unit tests with CMake.
2019-04-16 22:38:30 -07:00
Tzanio Kolev e9691ba40e Merge pull request #859 from mfem/install-okl-fix
Fix error messages when installing *.okl files [install-okl-fix]
2019-04-16 21:45:36 -07:00
camierjs 21edf56417 First cmake MFEM_USE_CUDA/OCCA/RAJA pass 2019-04-16 18:54:01 -07:00
David J. Gardner e780dffb94 get ex16p working with CVODE 2019-04-16 18:23:29 -07:00
David J. Gardner 289e57a247 simplify input check 2019-04-16 18:21:52 -07:00
David J. Gardner 3f2bdc0586 update constructors/destructors 2019-04-16 18:17:10 -07:00
David J. Gardner ffc4024147 fix naming conflicts, add destroy/free functions 2019-04-16 18:16:40 -07:00
David J. Gardner 7e75a67cd8 fix comments, remove extra break 2019-04-16 13:35:16 -07:00
David J. Gardner 565aed1e0d remove operator from LS base class 2019-04-16 13:33:41 -07:00
Tzanio 80a7cbaafe Hypre clarification 2019-04-16 10:36:16 -07:00
Veselin Dobrev 90b5e07681 In the main makefile, install *.okl files in a separate loop to
avoid error messages.
2019-04-15 19:47:20 -07:00
camierjs e9de1fcf7b Remove '>' in front of devices tests 2019-04-15 16:05:18 -07:00
Tzanio 99759b6e7c Updated hypre's URL 2019-04-13 22:36:19 -07:00
David J. Gardner e2455eb460 minor update to error message 2019-04-12 17:45:38 -07:00
David J. Gardner 5f2afd1719 fix error checks, clean up ex9, running with new interface 2019-04-12 17:41:54 -07:00
David J. Gardner 5ab614c896 update ex9p 2019-04-12 17:06:34 -07:00
David J. Gardner d5b9e221ae update sunmat wrap, uncomment linsys fn, add printinfo 2019-04-12 17:03:56 -07:00
David J. Gardner 7682a5d42a remove temp files 2019-04-12 11:23:22 -07:00
David J. Gardner 29458ab6ae Merge branch 'master' into sundials-interface 2019-04-12 11:09:44 -07:00
David J. Gardner 657ece56e6 update sundials files with new interface prototype 2019-04-12 11:08:32 -07:00
Stefano Zampini 6c129fc80d rename SparseMatrix::Chop -> SparseMatrix::Threshold 2019-04-10 11:33:03 +03:00
Veselin Dobrev a63cbf6841 Update .travis.yml
Link `hypre-2.10.0b` as `hypre`.
2019-04-07 14:19:20 -07:00
rcarson3 311f538fd5 Update hypre dependency version and point to github source
The hypre dependency version has been updated to now be the latest available on the LLNL github repository for hypre. The INSTALL file has also been updated to point people to the github page to have them download/clone the repository. Next, the build make/cmake files have also been updated to reflect that the hypre directory is now just hypre rather than hypre-2.10.0b.
2019-04-04 15:34:26 -07:00
Tzanio a4f7d19221 Styling 2019-03-30 20:21:02 -07:00
Tzanio 4c4619153c Added Operator::GetOutputRestriction() 2019-03-30 19:56:26 -07:00
David J. Gardner c767ba78f7 make backups of original interfaces 2019-03-26 10:44:52 -07:00
David J. Gardner 29f229c0cb fixes for approch 1 2019-03-26 10:44:22 -07:00
David J. Gardner 9b7b02bcdc rename _1 files to prevent building 2019-03-26 10:41:48 -07:00
David J. Gardner a4df3089ba updates based on feedback from Dan
Finish out native ARKode mass matrix support. Fix some typos.
2019-03-25 11:30:55 -07:00
David J. Gardner 0e0dc504ec initial update to use ARKStep mass matrix 2019-03-22 18:35:04 -07:00
David J. Gardner 7d8471cc89 update first approch to include ARKStep 2019-03-22 18:33:29 -07:00
David J. Gardner 8732d80050 outline one approach to sundials interfacing 2019-03-22 15:29:51 -07:00
Andrew T. Barker 38462c6175 Operator: separate FormSystemOperator() and FormDiscreteOperator()
FormSystemOperator() is for square operators (eg BilinearForm), potentially
with boundary conditions, while FormDiscreteOperator is for rectangular
operators, eg. matrix-free discrete gradient.
2019-03-21 09:30:11 -07:00
Stefano Zampini 0c1318dfd3 FiniteElementForGeometry can return NULL
This fixes the segfault but this should be handled better
2019-03-19 11:18:03 +03:00
Stefano Zampini c7eeca7c51 WIP: specify partitioning for NCMesh 2019-03-18 11:34:24 +03:00
Stefano Zampini e93b207273 ParNCMesh::GetConformingSharedStructures relax checks when elements are present 2019-03-18 11:34:24 +03:00
Stefano Zampini dde32310a0 PetscNonlinearSolver: expose update method 2019-03-14 21:17:40 +03:00
Stefano Zampini 3a47471713 make config: allow specifying a compiler to compile get_hypre_version
this fixes configs in supercomputers when login nodes != backend nodes
2019-03-14 21:17:40 +03:00
Stefano Zampini f4b3269a41 BDDC: add support for approximate solvers and scalar spaces 2019-03-14 21:17:40 +03:00
Stefano Zampini fb98bbc443 PetscLinearSolver: change default wrap flag to true 2019-03-14 21:17:40 +03:00
Stefano Zampini 3ed3353645 SparseMatrix: added Chop method to remove zeros from CSR of the matrix 2019-03-14 21:17:40 +03:00
Stefano Zampini c70e1dc9c8 MFEMInitializePetsc: added a couple of variations 2019-03-14 21:17:40 +03:00
Stefano Zampini aebe7919ba Mesh::FindPoints: fix for non-conforming meshes
It may happen that the closest element has a slave face with the actual owner of the point
2019-03-14 21:17:40 +03:00
Stefano Zampini 4a29f90147 HypreSolver: error when setup or solve fail 2019-03-14 21:17:40 +03:00
Stefano Zampini c11a4f4376 Petsc: add support for Operator::ANY_TYPE 2019-03-14 21:17:40 +03:00
Stefano Zampini 1b212fd2e5 prevent Convert_Array_IS from segfaulting 2019-03-14 21:17:40 +03:00
Stefano Zampini 5a7061c1ff PetscParMatrix: clarify constructor 2019-03-14 21:17:40 +03:00
Stefano Zampini f5de5a11bc VectorDeltaCoefficient: added a couple of setters 2019-03-14 21:17:40 +03:00
Stefano Zampini 0cc2429f20 FiniteElementSpace: prevent GetFE() from segfaulting 2019-03-14 21:17:40 +03:00
Stefano Zampini a0749535a3 Assume all build-* folder are build directories for VPATH builds 2019-03-14 21:17:40 +03:00
Stefano Zampini 70653ee1e5 BilinearIntegrators: made all parameters (scalars and coefficients) protected to make them accessible to derived class
For all public integrators, make coefficients usage consistent and store a pointer instead of a reference
This affected Convection, Derivative and *ProductInterpolator integrators
2019-03-14 21:17:40 +03:00
Stefano Zampini 7155d89824 Fix bug in ParGridFunction::ProjectDiscCoefficient
Calling parallel assemble is conceptually wrong, since a GridFunction
represents also vdofs. See https://github.com/mfem/mfem/issues/443

Suggested-by: Veselin Dobrev <dobrev@llnl.gov>
2019-03-14 21:17:40 +03:00
Stefano Zampini 8c88f1bdbe Add missing typecasts to PetscObject for PetscParVector and PetscParMatrix 2019-03-14 21:17:40 +03:00
Stefano Zampini eab5053902 {Vector|Matrix}ArrayCoefficient: customizable ownership of scalar coefficients 2019-03-14 21:17:40 +03:00
Stefano Zampini c4d20e38bd VectorMassIntegrators: make coefficients accessible to derived classes 2019-03-14 21:17:40 +03:00
Andrew T. Barker 48d25ce1cd Operator: FormParallelOperator -> FormSystemOperator 2019-02-18 14:53:25 -08:00
Tzanio 5f9ac7cde1 make style 2019-02-16 21:23:47 -08:00
Aaron Fisher 111f2158cb Added the gzstream capability to the data collection classes. 2019-02-12 13:05:32 -08:00
Andrew T. Barker 64d80b8fa3 Operator: implement FormParallelOperator() 2019-01-29 12:27:40 -08:00
Jan Nikl ac67858e76 Fixed Mesh::GetBdrElementTransformation collisions with the functions for faces. 2019-01-04 11:37:00 +01:00
Jan Nikl 4a7c708f8a Added some descriptions of the methods in BilinearForm and MixedBilinearForm. 2019-01-04 10:52:57 +01:00
Jan Nikl 2587806b28 Added AssembleElementMatrix() and AssembleBdrElementMatrix() to MixedBilinearForm. 2019-01-04 10:48:09 +01:00
Jan Nikl d3085cc755 Added AssembleElementMatrix() and AssembleBdrElementMatrix() versions not returning the VDofs used. 2019-01-04 10:47:19 +01:00
Jan Nikl c18e4e6162 Added ComputeElementMatrix() and ComputeBdrElementMatrix() to MixedBilinearForm. 2019-01-04 07:04:57 +01:00
Jan Nikl 59decf0c0e Added ComputeBdrElementMatrix() to BilinearForm. 2019-01-04 06:52:02 +01:00
Jan Nikl 7d26633521 Added boundary attribute markers for boundary integrators in MixedBilinearForm. 2019-01-03 23:28:10 +01:00
Jan Nikl c2ec29cd8e Renamed the lists of integrators in MixedBilinearForm to agree with BilinearForm. 2019-01-03 23:05:28 +01:00
Jan Nikl e5e0f0d507 Added boundary trace face integrators to MixedBilinearForm. 2019-01-03 23:02:45 +01:00
158 changed files with 13258 additions and 7357 deletions
+9 -8
View File
@@ -26,24 +26,25 @@ install:
- cd ..
# Install hypre
- ps: Start-FileDownload 'https://computation.llnl.gov/project/linear_solvers/download/hypre-2.10.0b.tar.gz'
- 7z x hypre-2.10.0b.tar.gz -so | 7z x -si -ttar > nul
- cd hypre-2.10.0b
- cmake -Hsrc -Bbuild -DMPI_C_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include" -DMPI_C_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include"
# - cmake -Hsrc -Bbuild -DCMAKE_BUILD_TYPE=Release -DMPI_C_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include" -DMPI_C_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include"
- ps: Start-FileDownload 'https://github.com/hypre-space/hypre/archive/V2-10-0b.tar.gz'
- 7z x V2-10-0b.tar.gz -so | 7z x -si -ttar > nul
- cd hypre-2-10-0b
- cmake -H. -Bbuild -DHYPRE_USING_FEI=OFF -DMPI_C_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include" -DMPI_C_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include"
- cmake --build build
- cmake --build build --target install
- cd ..
# MFEM
before_build:
- cmake -H. -DCMAKE_INSTALL_PREFIX=install -Bbuild_parallel -DMFEM_USE_MPI=TRUE -DMFEM_USE_METIS_5=TRUE -DMPI_CXX_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include" -DHYPRE_LIBRARIES=%cd%\hypre-2.10.0b\src\hypre\lib\HYPRE.lib -DHYPRE_INCLUDE_DIRS=%cd%\hypre-2.10.0b\src\hypre\include -DHYPRE_VERSION=21000 -DMETIS_LIBRARIES=%cd%\metis-5.1.0\build\libmetis\Debug\metis.lib -DMETIS_INCLUDE_DIRS=%cd%\metis-5.1.0\include
- cmake -H. -DCMAKE_INSTALL_PREFIX=install -Bbuild_serial -DMFEM_USE_MPI=FALSE -DMFEM_USE_METIS_5=TRUE -DMPI_CXX_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include" -DHYPRE_LIBRARIES=%cd%\hypre-2.10.0b\src\hypre\lib\HYPRE.lib -DHYPRE_INCLUDE_DIRS=%cd%\hypre-2.10.0b\src\hypre\include -DHYPRE_VERSION=21000 -DMETIS_LIBRARIES=%cd%\metis-5.1.0\build\libmetis\Debug\metis.lib -DMETIS_INCLUDE_DIRS=%cd%\metis-5.1.0\include
- cmake -H. -DCMAKE_INSTALL_PREFIX=install -Bbuild_parallel -DMFEM_USE_MPI=TRUE -DMFEM_USE_METIS_5=TRUE -DMPI_CXX_LIBRARIES="C:\Program Files (x86)\Microsoft SDKs\MPI\Lib\x86\msmpi.lib" -DMPI_CXX_INCLUDE_PATH="C:\Program Files (x86)\Microsoft SDKs\MPI\Include" -DHYPRE_LIBRARIES=%cd%\hypre-2-10-0b\hypre\lib\HYPRE.lib -DHYPRE_INCLUDE_DIRS=%cd%\hypre-2-10-0b\hypre\include -DHYPRE_VERSION=21000 -DMETIS_LIBRARIES=%cd%\metis-5.1.0\build\libmetis\Debug\metis.lib -DMETIS_INCLUDE_DIRS=%cd%\metis-5.1.0\include
- cmake -H. -DCMAKE_INSTALL_PREFIX=install -Bbuild_serial -DMFEM_USE_MPI=FALSE
build_script:
- cmake --build build_parallel
- cmake --build build_serial
- cmake --build build_serial --target exec
after_build:
# - cmake --build build_parallel --target check
- cmake --build build_serial --target check
- cmake --build build_serial --target RUN_TESTS
+6 -3
View File
@@ -82,9 +82,9 @@ examples/ex20.dat
examples/ex20p_?????.dat
examples/gnuplot_ex20.inp
examples/gnuplot_ex20p.inp
examples/ex22*.mesh
examples/ex22*.sol
examples/ex22p_*.*
examples/ex21*.mesh
examples/ex21*.sol
examples/ex21p_*.*
examples/sundials/ex9
examples/sundials/ex1[06]
@@ -183,3 +183,6 @@ miniapps/nurbs/Example1*
# Unit test binary and outputs
tests/unit/output_meshes
tests/unit/unit_tests
# VPATH builds
build-*/*
+1
View File
@@ -205,6 +205,7 @@ install:
else
echo "Reusing cached hypre-2.10.0b/";
fi;
ln -s hypre-2.10.0b hypre;
else
echo "Serial build, not using hypre";
fi
+65 -33
View File
@@ -8,23 +8,30 @@
http://mfem.org
Version 4.0-RC1, Apr 11, 2019
=============================
Version 4.0.1 (development)
===========================
- Improved RAJA backend
- Improved multi-GPU MPI communication.
Requirements and Limitations
----------------------------
- This is a release candidate for mfem-4.0.
- Use at your own risk -- not everything will work, the API may change.
- We are looking for feedback from friendly users.
- Unlike previous MFEM releases, this version requires a C++11 compiler.
GPU support
-----------
- Added initial support for AMD GPUs based on HIP: a C++ runtime API and kernel
language that can run on both AMD and NVIDIA hardware. The list of current
backends is: "occa-cuda", "raja-cuda", "cuda", "hip", "occa-omp", "raja-omp",
"omp", "occa-cpu", "raja-cpu", and "cpu".
- GPU-related limitations:
* NVCC is not supported in the CMake build system yet.
* Element batching is currently ignored.
* Full-assembly (on device), element assembly, and matrix-free bilinear forms
are not supported yet.
* FunctionCoefficients do not currently work on GPUs.
* Partial assembly kernels are not implemented yet for simplices.
Miscellaneous
-------------
- Upgraded the SUNDIALS interface to utilize SUNDIALS version 5.0. This
necessitated a complete rework of the interface and requires changes at
the application level. Example usage of this new interface can be found
in the examples/sundials directory.
Version 4.0, released on May 24, 2019
=====================================
Unlike previous MFEM releases, this version requires a C++11 compiler.
GPU support
-----------
@@ -35,7 +42,7 @@ GPU support
seamlessly with a new lightweight device/host memory manager. The kernels can
be implemented either in OCCA, or as a simple wrapper around for-loops, which
can then be dispatched to RAJA and native backends. See the files forall.hpp
and mem_manager.hpp in the general/ directory.
and mem_manager.hpp in the general/ directory for more details.
- Several of the MFEM example codes (ex1, ex1p, ex6, and ex6p) can now take
advantage of GPU acceleration with the backend selectable at runtime. Many of
@@ -43,26 +50,44 @@ GPU support
bilinear forms) have been extended to take advantage of kernel acceleration by
simply replacing loops with the MFEM_FORALL() macro.
- In addition to pure CUDA, the library currently supports OCCA, RAJA and OpenMP
kernels, which could be mixed and matched in different parts of the same
application. We plan on adding support for more programming models and devices
in the future, without the need for significant modifications in user code.
The list of current backends is: "occa-cuda", "raja-cuda", "cuda", "occa-omp",
"raja-omp", "omp", "occa-cpu", "raja-cpu", and "cpu".
- In addition to native CUDA kernels, the library currently supports OCCA, RAJA
and OpenMP kernels, which could be mixed and matched in different parts of the
same application. We plan on adding support for more programming models and
devices in the future, without the need for significant modifications in user
code. The list of current backends is: "occa-cuda", "raja-cuda", "cuda",
"occa-omp", "raja-omp", "omp", "occa-cpu", "raja-cpu", and "cpu".
- GPU-related limitations:
* Hypre preconditioners are not yet available in GPU mode, and in particular
hypre must be built in CPU mode.
* Only constant coefficients are currently supported on GPUs.
* Optimized element assembly, and matrix-free bilinear forms are not
implemented yet. Element batching is currently ignored.
* In device mode, full assembly is performed on the host (but the matvec
action is performed on the device).
* Partial assembly kernels are not implemented yet for simplices.
Discretization improvements
---------------------------
- Partial assembled finite element operators are now available in the core
library, based on the new classes PABilinearFormExtension, ElementRestriction,
DofToQuad and GeometricFactors (associated with the classes BilinearForm,
FiniteElementSpace, FiniteElement and Mesh, respectively). The kernels for
partial assembled Setup/Assembly and Action/Mult are implemented in the
BilinearFormIntegrator methods AssemblePA and AddMultPA.
- Added support for a general "low-order refined"-to-"high-order" transfer of
GridFunction data from a "low-order refined" (LOR) space defined on a refined
mesh to a "high-order" (HO) finite element space defined on a coarse mesh. See
the new classes InterpolationGridTransfer and L2ProjectionGridTransfer and the
new LOR Transfer miniapp: miniapps/tools/lor-transfer.cpp.
- Added support for derefinement of vector (RT + ND) spaces.
- Added element flux, and flux energy computation in class ElasticityIntegrator,
allowing for the use of Zienkiewicz-Zhu type error estimators with the
integrator. For an illustration of this addition, see the new Example 22.
integrator. For an illustration of this addition, see the new Example 21.
- Added support for derefinement of vector (RT + ND) spaces.
- Added a variety of coefficients which are sums or products of existing
coefficients as well as grid function coefficients which return the
@@ -74,13 +99,13 @@ Support for wedge elements and meshes with mixed element types
type PRISM) which have two triangular faces and three quadrilateral faces.
Several examples of such meshes can be found in the data/ directory.
- Added H1 and L2 finite elements of arbitrary order for Wedge elements.
- Added support for mixed meshes containing triangles and quadrilaterals in 2D
or tetrahedra, wedges, and hexahedra in 3D. This includes support for uniform
refinement of such meshes. Several examples of such meshes can be found in the
data/ directory.
- Added H1 and L2 finite elements of arbitrary order for Wedge elements.
- Added support for reading and writing linear and quadratic meshes containing
wedge elements in VTK mesh format. Several examples of such meshes can be
found in the data/ directory.
@@ -101,6 +126,10 @@ Other meshing improvements
This guarantees that the shape regularity of the elements will be preserved
under refinement.
- The TMOP mesh optimization algorithms were extended to support user-defined
space-dependent limiting terms. Improved the TMOP objective functions by more
accurate normalization of the different terms.
- Added support for parallel communication groups on non-conforming meshes.
- Improved parallel partitioning of non-conforming meshes. If the coarse mesh
@@ -114,10 +143,6 @@ Other meshing improvements
- Added support for reading linear and quadratic 2D quadrilateral and triangular
Cubit meshes.
- The TMOP mesh optimization algorithms were extended to support user-defined
space-dependent limiting terms. Improved the TMOP objective functions by more
accurate normalization of the different terms.
New and updated examples and miniapps
-------------------------------------
- Added a new meshing miniapp, Toroid, which can produce a variety of torus
@@ -133,7 +158,7 @@ New and updated examples and miniapps
from a Hamiltonian. The example demonstrates the use of the variable order,
symplectic integration algorithm implemented in class SIAVSolver.
- Added a new example, Example 22/22p, that illustrates the use of AMR to solve
- Added a new example, Example 21/21p, that illustrates the use of AMR to solve
a linear elasticity problem. This is an extension of Example 2/2p.
New and improved solvers and preconditioners
@@ -145,17 +170,24 @@ New and improved solvers and preconditioners
Miscellaneous
-------------
- Added unit tests based on the Catch++ library.
- Added unit tests based on the Catch++ library in the test/ directory.
- Renamed the option MFEM_USE_OPENMP to MFEM_USE_LEGACY_OPENMP. This legacy
option is deprecated and planned for removal in a future release. The original
option name, MFEM_USE_OPENMP, is now used to enable the new OpenMP backends in
the new kernels.
- In SparseMatrix added the option to perform MultTranspose() by matvec with
computed and stored transpose matrix. This is required for deterministic
results when using devices such as CUDA and OpenMP.
- Altered the way FGMRES counts its iterations so that it matches GMRES.
- Various other simplifications, extensions, and bugfixes in the code.
- Construct abstract parallel rectangular truedof-to-truedof operators via
Operator::FormDiscreteOperator().
API changes
-----------
- In multiple places, use Geometry::Type instead of int, where appropriate.
+59 -8
View File
@@ -50,7 +50,7 @@ project(mfem NONE)
# Current version of MFEM, see also `makefile`.
# mfem_VERSION = (string)
# MFEM_VERSION = (int) [automatically derived from mfem_VERSION]
set(${PROJECT_NAME}_VERSION 3.4.1)
set(${PROJECT_NAME}_VERSION 4.0.1)
# Prohibit in-source build
if (${PROJECT_SOURCE_DIR} STREQUAL ${PROJECT_BINARY_DIR})
@@ -86,6 +86,13 @@ include("${CMAKE_CURRENT_SOURCE_DIR}/config/XSDKDefaults.cmake")
# Enable languages.
enable_language(CXX)
if (MFEM_USE_CUDA)
# MFEM_USE_CUDA requires CMake 3.8 or newer (for direct CUDA support)
cmake_minimum_required(VERSION 3.8 FATAL_ERROR)
enable_language(CUDA)
message(STATUS "Using CUDA architecture: ${CUDA_ARCH}")
endif()
if (XSDK_ENABLE_C)
enable_language(C)
endif()
@@ -266,6 +273,31 @@ if (MFEM_USE_PUMI)
endif()
endif()
# CUDA
if (MFEM_USE_CUDA)
set(CMAKE_CUDA_STANDARD 11)
set(CMAKE_CUDA_STANDARD_REQUIRED ON)
set(CMAKE_CUDA_EXTENSIONS OFF)
set(CMAKE_CUDA_FLAGS "-arch=${CUDA_ARCH} --expt-extended-lambda"
CACHE STRING "CUDA flags set for MFEM" FORCE)
if (MFEM_USE_MPI)
set(CUDA_CCBIN_COMPILER ${MPI_CXX_COMPILER})
else()
set(CUDA_CCBIN_COMPILER ${CMAKE_CXX_COMPILER})
endif()
string(APPEND CMAKE_CUDA_FLAGS " -ccbin ${CUDA_CCBIN_COMPILER}")
endif()
# OCCA
if (MFEM_USE_OCCA)
find_package(OCCA REQUIRED)
endif()
# RAJA
if (MFEM_USE_RAJA)
find_package(RAJA REQUIRED)
endif()
# MFEM_TIMER_TYPE
if (NOT DEFINED MFEM_TIMER_TYPE)
if (APPLE)
@@ -291,7 +323,7 @@ endif()
# be before SuiteSparse.
set(MFEM_TPLS MPI_CXX OPENMP BLAS LAPACK METIS HYPRE SuiteSparse SUNDIALS PETSC
MESQUITE SuperLUDist STRUMPACK AXOM CONDUIT GECKO GNUTLS NETCDF MPFR PUMI
POSIXCLOCKS MFEMBacktrace ZLIB)
POSIXCLOCKS MFEMBacktrace ZLIB OCCA RAJA)
# Add all *_FOUND libraries in the variable TPL_LIBRARIES.
set(TPL_LIBRARIES "")
set(TPL_INCLUDE_DIRS "")
@@ -327,6 +359,13 @@ set(MFEM_SOURCE_DIRS general linalg mesh fem)
foreach(DIR IN LISTS MFEM_SOURCE_DIRS)
add_subdirectory(${DIR})
endforeach()
if (MFEM_USE_CUDA)
foreach(file IN LISTS SOURCES)
set_property(SOURCE ${file} PROPERTY LANGUAGE CUDA)
endforeach()
endif()
add_subdirectory(config)
set(MASTER_HEADERS
${PROJECT_SOURCE_DIR}/mfem.hpp
@@ -337,6 +376,11 @@ set(CMAKE_INSTALL_RPATH_USE_LINK_PATH ON CACHE BOOL "")
set(CMAKE_INSTALL_RPATH "${_lib_path}" CACHE PATH "")
set(CMAKE_INSTALL_NAME_DIR "${_lib_path}" CACHE PATH "")
set(MFEM_SOURCE_DIR ${CMAKE_CURRENT_SOURCE_DIR} CACHE PATH
"The MFEM source directory" FORCE)
set(MFEM_INSTALL_DIR ${CMAKE_INSTALL_PREFIX} CACHE PATH
"The MFEM install directory" FORCE)
# Declaring the library
add_library(mfem ${SOURCES} ${HEADERS} ${MASTER_HEADERS})
# message(STATUS "TPL_LIBRARIES = ${TPL_LIBRARIES}")
@@ -351,11 +395,11 @@ endif()
set_target_properties(mfem PROPERTIES VERSION "${mfem_VERSION}")
set_target_properties(mfem PROPERTIES SOVERSION "${mfem_VERSION}")
# If building out-of-source, define MFEM_BUILD_DIR to point to the build
# directory.
# If building out-of-source, define MFEM_CONFIG_FILE to point to the config file
# inside the build directory.
if (NOT ("${PROJECT_SOURCE_DIR}" STREQUAL "${PROJECT_BINARY_DIR}"))
target_compile_definitions(mfem PRIVATE
"MFEM_BUILD_DIR=${PROJECT_BINARY_DIR}")
"MFEM_CONFIG_FILE=\"${PROJECT_BINARY_DIR}/config/_config.hpp\"")
endif()
# Generate configuration file in the build directory: config/_config.hpp.
@@ -371,7 +415,7 @@ if (NOT ("${PROJECT_SOURCE_DIR}" STREQUAL "${PROJECT_BINARY_DIR}"))
"Writing substitute header --> \"${Header}\"")
file(WRITE "${PROJECT_BINARY_DIR}/${Header}"
"// Auto-generated file.
#define MFEM_BUILD_DIR ${PROJECT_BINARY_DIR}
#define MFEM_CONFIG_FILE \"${PROJECT_BINARY_DIR}/config/_config.hpp\"
#include \"${PROJECT_SOURCE_DIR}/${Header}\"
")
# This version will be installed in the top include directory:
@@ -434,12 +478,12 @@ endif()
# Add 'check' target - quick test
if (NOT MFEM_USE_MPI)
add_custom_target(check
${CMAKE_CTEST_COMMAND} -R '^ex1_ser' -C ${CMAKE_CFG_INTDIR}
${CMAKE_CTEST_COMMAND} -R \"^ex1_ser\" -C ${CMAKE_CFG_INTDIR}
USES_TERMINAL)
add_dependencies(check ex1)
else()
add_custom_target(check
${CMAKE_CTEST_COMMAND} -R '^ex1p' -C ${CMAKE_CFG_INTDIR}
${CMAKE_CTEST_COMMAND} -R \"^ex1p\" -C ${CMAKE_CFG_INTDIR}
USES_TERMINAL)
add_dependencies(check ex1p)
endif()
@@ -484,6 +528,13 @@ install(DIRECTORY ${MFEM_SOURCE_DIRS}
DESTINATION ${INSTALL_INCLUDE_DIR}/mfem
FILES_MATCHING PATTERN "*.hpp")
# Install the okl files
if (MFEM_USE_OCCA)
install(DIRECTORY ${MFEM_SOURCE_DIRS}
DESTINATION ${INSTALL_INCLUDE_DIR}/mfem
FILES_MATCHING PATTERN "*.okl")
endif()
# Install ${HEADERS}
# ---
# foreach (HDR ${HEADERS})
+10
View File
@@ -142,6 +142,16 @@ Origin](#developers-certificate-of-origin-11) at the end of this file.*
+ [`HypreParMatrix`](http://mfem.github.io/doxygen/html/classmfem_1_1HypreParMatrix.html) and [`HypreParVector`](http://mfem.github.io/doxygen/html/classmfem_1_1HypreParVector.html)
+ [`HypreSolver`](http://mfem.github.io/doxygen/html/classmfem_1_1HypreSolver.html) and other [hypre classes](http://mfem.github.io/doxygen/html/hypre_8hpp.html)
- GPU and multi-core CPU support is based on device kernels supporting different
backends (CUDA, OCCA, RAJA, OpenMP, etc.) and an internal lightweight
device/host memory manager.
- The main device-relevant classes and sources are:
+ [`Device`](http://mfem.github.io/doxygen/html/device_8hpp.html)
+ [`MemoryManager`](http://mfem.github.io/doxygen/html/mem_manager_8hpp.html)
+ the [`MFEM_FORALL`](http://mfem.github.io/doxygen/html/forall_8hpp.html) macro
+ the [`cuda.hpp`](http://mfem.github.io/doxygen/html/cuda_8hpp.html) and [`occa.hpp`](http://mfem.github.io/doxygen/html/occa_8hpp.html) files
- The `general/` directory contains C++ classes that serve as utilities for
communication, error handling, arrays, (Boolean) tables, timing, etc.
+63 -30
View File
@@ -13,22 +13,31 @@ of MFEM is a (modern) C++ compiler, such as g++. The parallel version of MFEM
requires an MPI C++ compiler, as well as the following external libraries:
- hypre (a library of high-performance preconditioners)
http://www.llnl.gov/CASC/hypre
https://github.com/hypre-space/hypre
- METIS (a family of multilevel partitioning algorithms)
http://glaros.dtc.umn.edu/gkhome/metis/metis/overview
The hypre dependency can be downloaded as a tarball from GitHub or from the
project webpage https://www.llnl.gov/casc/hypre. For example, the 2.16.0 release
of hypre is available at
https://github.com/hypre-space/hypre/archive/v2.16.0.tar.gz
The METIS dependency can be disabled but that is not generally recommended, see
the option MFEM_USE_METIS.
MFEM also includes support for devices such as GPUs, and programming models such
as CUDA, OCCA, OpenMP and RAJA.
as CUDA, HIP, OCCA, OpenMP and RAJA.
- Starting with version 4.0, MFEM requires a C++11 compiler
- CUDA support requires an NVIDIA GPU and an installation of the CUDA Toolkit
https://developer.nvidia.com/cuda-toolkit
- HIP support requires an AMD GPU and an installation of the ROCm software stack
https://rocm.github.io/ROCmInstall.html#installing-from-amd-rocm-repositories
- OCCA support requires the OCCA library
https://libocca.org
@@ -48,7 +57,7 @@ following package managers:
- Spack, https://github.com/spack/spack
- OpenHPC, http://openhpc.community
- Homebrew/Science, https://github.com/Homebrew/homebrew-science
- Homebrew/Science, https://github.com/Homebrew/homebrew-science (deprecated)
We also recommend downloading and building the MFEM-based GLVis visualization
tool which can be used to visualize the meshes and solution in MFEM's examples
@@ -60,15 +69,19 @@ Serial build:
make serial -j 4
Parallel build:
(download hypre 2.10.0b and METIS 4 from above URLs)
(download hypre and METIS 4 from above URLs)
(build METIS 4 in ../metis-4.0 relative to mfem/)
(build hypre 2.10.0b in ../hypre-2.10.0b relative to mfem/)
(build hypre in ../hypre relative to mfem/)
make parallel -j 4
CUDA build:
make cuda -j 4
(build for a specific compute capability: 'make cuda -j 4 CUDA_ARCH=sm_30')
HIP build:
make hip -j 4
(build for a specific AMD GPU chip: 'make hip -j 4 HIP_ARCH=gfx900')
Example codes (serial/parallel, depending on the build):
cd examples
make -j 4
@@ -87,13 +100,19 @@ Serial build:
make -j 4 (assuming "UNIX Makefiles" generator)
Parallel build:
(download hypre 2.10.0b and METIS 4 from above URLs)
(download hypre and METIS 4 from above URLs)
(build METIS 4 in ../metis-4.0 relative to mfem/)
(build hypre 2.10.0b in ../hypre-2.10.0b relative to mfem/)
(build hypre in ../hypre relative to mfem/)
mkdir <mfem-build-dir> ; cd <mfem-build-dir>
cmake <mfem-source-dir> -DMFEM_USE_MPI=YES
make -j 4
CUDA build:
(this build requires CMake 3.8 or newer)
mkdir <mfem-build-dir> ; cd <mfem-build-dir>
cmake <mfem-source-dir> -DMFEM_USE_CUDA=YES
make -j 4
Example codes (serial/parallel, depending on the build):
make examples -j 4
@@ -149,14 +168,18 @@ Note that re-configuration is only needed to change the currently configured
options. Several shortcut targets combining (re-)configuration and compilation
are also defined:
make serial -> Builds serial optimized version of the library
make parallel -> Builds parallel optimized version of the library
make debug -> Builds serial debug version of the library
make pdebug -> Builds parallel debug version of the library
make cuda -> Builds serial cuda optimized version of the library
make pcuda -> Builds parallel cuda optimized version of the library
make cudebug -> Builds serial cuda debug version of the library
make pcudebug -> Builds parallel cuda debug version of the library
make serial -> Builds serial optimized version of the library
make parallel -> Builds parallel optimized version of the library
make debug -> Builds serial debug version of the library
make pdebug -> Builds parallel debug version of the library
make cuda -> Builds serial cuda optimized version of the library
make pcuda -> Builds parallel cuda optimized version of the library
make cudebug -> Builds serial cuda debug version of the library
make pcudebug -> Builds parallel cuda debug version of the library
make hip -> Builds serial hip optimized version of the library
make phip -> Builds parallel hip optimized version of the library
make hipdebug -> Builds serial hip debug version of the library
make phipdebug -> Builds parallel hip debug version of the library
Note that any of the above shortcuts accept configuration options, either at the
command line or through a user configuration file.
@@ -278,6 +301,7 @@ MFEM_THREAD_SAFE = YES/NO
MFEM_USE_LEGACY_OPENMP = YES/NO
Enable (basic) experimental OpenMP support. Requires MFEM_THREAD_SAFE.
This option is deprecated.
MFEM_USE_OPENMP = YES/NO
Enable the OpenMP backend.
@@ -391,28 +415,32 @@ MFEM_USE_PUMI = YES/NO
models and effectively supports automated adaptive analysis. PUMI enables
support for parallel unstructured mesh modifications in MFEM.
MFEM_USE_MM = YES/NO
Enables support for the MFEM's memory manager (MM), which is required to
support devices with different memory spaces.
MFEM_USE_CUDA = YES/NO
Enables support for CUDA devices in MFEM. CUDA is a parallel computing
platform and programming model for general computing on graphical processing
units (GPUs). This option requires MFEM_USE_MM. The variable CUDA_ARCH is
used to specify the CUDA compute capability used during compilation (by
default, CUDA_ARCH=sm_60). When enabled, this option uses the CUDA_* build
options, see below.
units (GPUs). The variable CUDA_ARCH is used to specify the CUDA compute
capability used during compilation (by default, CUDA_ARCH=sm_60). When
enabled, this option uses the CUDA_* build options, see below.
MFEM_USE_HIP = YES/NO
Enables support for AMD devices in MFEM. HIP is a heterogeneous-compute
interface for portability developed by AMD that can target both AMD and
NVIDIA GPUs. The variable HIP_ARCH is used to specify the AMD GPU processor
used during compilation (by default, HIP_ARCH=gfx900). When enabled, this
option uses the HIP_* build options, see below.
MFEM_USE_RAJA = YES/NO
Enable support for the RAJA performance portability layer in MFEM. RAJA
provides a portable abstraction for loops, supporting different programming
model backends. When using the RAJA CUDA backend, MFEM_USE_MM is required.
model backends. When using RAJA built with CUDA support, CUDA support must be
also enabled in MFEM, i.e. MFEM_USE_CUDA=YES must be set.
MFEM_USE_OCCA = YES/NO
Enables support for the OCCA library in MFEM. OCCA is an open-source library
which aims to make it easy to program different types of devices (e.g. CPU,
GPU, FPGA) by providing an unified API for interacting with JIT-compiled
backends. When using the OCCA CUDA backend, MFEM_USE_MM is required.
backends. In order to use the OCCA CUDA backend, CUDA support must be enabled
in MFEM as well, i.e. MFEM_USE_CUDA=YES must be set.
MFEM_BUILD_TAG = (any value)
An optional tag to characterize the build. Exported to config/config.mk.
@@ -435,7 +463,7 @@ directory and use the string @MFEM_DIR@, e.g. HYPRE_OPT = -I@MFEM_DIR@/../hypre.
The specific libraries and their options are:
- HYPRE, required for the parallel build, i.e. when MFEM_USE_MPI = YES.
URL: http://www.llnl.gov/CASC/hypre
URL: https://github.com/hypre-space/hypre and https://www.llnl.gov/casc/hypre
Options: HYPRE_OPT, HYPRE_LIB.
- METIS, used when MFEM_USE_METIS = YES. If using METIS 5, set
@@ -460,6 +488,7 @@ The specific libraries and their options are:
- SUNDIALS (optional), used when MFEM_USE_SUNDIALS = YES.
Beginning with MFEM v3.3, SUNDIALS v2.7.0 is supported.
Beginning with MFEM v3.3.2, SUNDIALS v3.0.0 is also supported.
Beginning with MFEM v4.1, only SUNDIALS v5.0.0+ is supported.
If MFEM_USE_MPI is enabled, we expect that SUNDIALS is built with support for
both MPI and hypre.
URL: http://computation.llnl.gov/projects/sundials/sundials-software
@@ -533,6 +562,10 @@ The specific libraries and their options are:
URL: https://developer.nvidia.com/cuda-toolkit
Options: CUDA_CXX, CUDA_ARCH, CUDA_OPT, CUDA_LIB.
- HIP, used when MFEM_USE_HIP = YES.
URL: https://rocm.github.io/ROCmInstall.html
Options: HIP_CXX, HIP_ARCH, HIP_OPT, HIP_LIB.
- OCCA, used when MFEM_USE_OCCA = YES.
URL: https://libocca.org
Options: OCCA_DIR, OCCA_OPT, OCCA_LIB.
@@ -645,6 +678,8 @@ Configuration variables (CMake)
===============================
See the configuration file config/defaults.cmake for the default settings.
Note: the option MFEM_USE_CUDA requires CMake version 3.8 or newer!
Non-standard CMake variables for compilers:
CXX - If set, overwrite the auto-detected C++ compiler, serial build
MPICXX - If set, overwrite the auto-detected MPI C++ compiler, parallel build
@@ -675,13 +710,9 @@ MFEM_USE_NETCDF
MFEM_USE_MPFR
MFEM_USE_GZSTREAM
MFEM_USE_PUMI
The following GNU make options are not supported with CMake yet:
MFEM_USE_CUDA
MFEM_USE_OCCA
MFEM_USE_RAJA
MFEM_USE_MM
The following options are CMake specific:
@@ -728,6 +759,8 @@ The CMake build system adds auto-detection for the following packages/libraries:
- LIBUNWIND
- POSIXCLOCKS
- PUMI
- OCCA
- RAJA
The following built-in CMake packages are also used:
+17 -16
View File
@@ -8,9 +8,9 @@
http://mfem.org
MFEM is a modular parallel C++ library for finite element methods. Its goal is
to enable the research and development of scalable finite element discretization
and solver algorithms through general finite element abstractions, accurate and
flexible visualization, and tight integration with the hypre library.
to enable high-performance scalable finite element discretization research and
application development on a wide variety of platforms, ranging from laptops to
supercomputers.
* For building instructions, see the file INSTALL, or type "make help".
@@ -39,23 +39,24 @@ conforming and non-conforming (AMR) adaptive refinement. Arbitrary element
transformations, allowing for high-order mesh elements with curved boundaries,
are also supported.
MFEM is commonly used as a "finite element to linear algebra translator", since
it can take a problem described in terms of finite element-type objects, and
produce the corresponding linear algebra vectors and sparse matrices. In order
to facilitate this, MFEM uses compressed sparse row (CSR) sparse matrix storage
and includes simple smoothers and Krylov solvers, such as PCG, MINRES and GMRES,
as well as support for sequential sparse direct solvers from the SuiteSparse
When used as a "finite element to linear algebra translator", MFEM can take a
problem described in terms of finite element-type objects, and produce the
corresponding linear algebra vectors and fully or partially assembled operators,
e.g. in the form of global sparse matrices or matrix-free operators. The library
includes simple smoothers and Krylov solvers, such as PCG, MINRES and GMRES, as
well as support for sequential sparse direct solvers from the SuiteSparse
library. Nonlinear solvers (the Newton method), eigensolvers (LOBPCG), and
several explicit and implicit Runge-Kutta time integrators are also available.
MFEM supports MPI-based parallelism throughout the library, and can readily be
used as a scalable unstructured finite element problem generator. MFEM-based
applications require minimal changes to transition from a serial to a
high-performing parallel version of the code, where they can take advantage of
the integrated scalable linear solvers from the hypre library. Comprehensive
support for other external packages, e.g. PETSc and SUNDIALS is also included,
giving access to many additional linear and nonlinear solvers, preconditioners,
time integrators, etc.
used as a scalable unstructured finite element problem generator. As of version
4.0, MFEM offers initial support for GPU acceleration, and programming models,
such as CUDA, OCCA, RAJA and OpenMP. MFEM-based applications require minimal
changes to switch from a serial to a high-performing MPI-parallel version of the
code, where they can take advantage of the integrated linear solvers from the
hypre library. Comprehensive support for other external packages, e.g. PETSc
and SUNDIALS is also included, giving access to many additional linear and
nonlinear solvers, preconditioners, time integrators, etc.
For examples of using MFEM, see the examples/ and miniapps/ directories, as well
as the OpenGL visualization tool GLVis which is available at http://glvis.org.
+23 -4
View File
@@ -74,7 +74,7 @@
IF (NOT COMMAND PRINT_VAR)
FUNCTION(PRINT_VAR VAR_NAME)
MESSAGE("-- " "${VAR_NAME} = '${${VAR_NAME}}'")
MESSAGE(STATUS "${VAR_NAME} = '${${VAR_NAME}}'")
ENDFUNCTION()
ENDIF()
@@ -166,14 +166,14 @@ IF (USE_XSDK_DEFAULTS)
ENDIF()
XSDK_HANDLE_LANG_DEFAULTS(Fortran FC "FFLAGS;FCFLAGS")
ENDIF()
# Set XSDK defaults for other CMake variables
IF ("${BUILD_SHARED_LIBS}" STREQUAL "")
MESSAGE("-- " "XSDK: Setting default BUILD_SHARED_LIBS=TRUE")
SET(BUILD_SHARED_LIBS TRUE CACHE BOOL "Set by default in XSDK mode")
ENDIF()
IF ("${CMAKE_BUILD_TYPE}" STREQUAL "")
MESSAGE("-- " "XSDK: Setting default CMAKE_BUILD_TYPE=DEBUG")
SET(CMAKE_BUILD_TYPE DEBUG CACHE STRING "Set by default in XSDK mode")
@@ -181,6 +181,13 @@ IF (USE_XSDK_DEFAULTS)
ENDIF()
##################################################################################
#
# MFEM-specific additions: set TPL MFEM_USE_* defaults
#
##################################################################################
IF (DEFINED TPL_ENABLE_MPI)
SET(MFEM_USE_MPI ${TPL_ENABLE_MPI} CACHE BOOL "Enable MPI parallel build" FORCE)
ENDIF()
@@ -252,3 +259,15 @@ ENDIF()
IF (DEFINED TPL_ENABLE_PUMI)
SET(MFEM_USE_PUMI ${TPL_ENABLE_PUMI} CACHE BOOL "Enable PUMI" FORCE)
ENDIF()
IF (DEFINED TPL_ENABLE_CUDA)
SET(MFEM_USE_CUDA ${TPL_ENABLE_CUDA} CACHE BOOL "Enable CUDA" FORCE)
ENDIF()
IF (DEFINED TPL_ENABLE_OCCA)
SET(MFEM_USE_OCCA ${TPL_ENABLE_OCCA} CACHE BOOL "Enable OCCA" FORCE)
ENDIF()
IF (DEFINED TPL_ENABLE_RAJA)
SET(MFEM_USE_RAJA ${TPL_ENABLE_RAJA} CACHE BOOL "Enable RAJA" FORCE)
ENDIF()
+3
View File
@@ -41,6 +41,9 @@ set(MFEM_USE_MPFR @MFEM_USE_MPFR@)
set(MFEM_USE_SIDRE @MFEM_USE_SIDRE@)
set(MFEM_USE_CONDUIT @MFEM_USE_CONDUIT@)
set(MFEM_USE_PUMI @MFEM_USE_PUMI@)
set(MFEM_USE_CUDA @MFEM_USE_CUDA@)
set(MFEM_USE_OCCA @MFEM_USE_OCCA@)
set(MFEM_USE_RAJA @MFEM_USE_RAJA@)
set(MFEM_CXX_COMPILER "@CMAKE_CXX_COMPILER@")
set(MFEM_CXX_FLAGS "@CMAKE_CXX_FLAGS@")
+16
View File
@@ -30,6 +30,12 @@
#define MFEM_VERSION_MINOR (((MFEM_VERSION)/100)%100)
#define MFEM_VERSION_PATCH ((MFEM_VERSION)%100)
// MFEM source directory.
#define MFEM_SOURCE_DIR "@MFEM_SOURCE_DIR@"
// MFEM install directory.
#define MFEM_INSTALL_DIR "@MFEM_INSTALL_DIR@"
// Description of the git commit used to build MFEM.
#cmakedefine MFEM_GIT_STRING "@MFEM_GIT_STRING@"
@@ -104,6 +110,16 @@
// Enable MFEM functionality based on the PUMI library
#cmakedefine MFEM_USE_PUMI
// Build the GPU/CUDA-enabled version of the MFEM library.
// Requires a CUDA compiler (nvcc).
#cmakedefine MFEM_USE_CUDA
// Enable MFEM functionality based on the RAJA library
#cmakedefine MFEM_USE_RAJA
// Enable MFEM functionality based on the OCCA library
#cmakedefine MFEM_USE_OCCA
// Which library functions to use in class StopWatch for measuring time.
// For a list of the available options, see INSTALL.
// If not defined, an option is selected automatically.
+19
View File
@@ -0,0 +1,19 @@
# Copyright (c) 2010, Lawrence Livermore National Security, LLC. Produced at the
# Lawrence Livermore National Laboratory. LLNL-CODE-443211. All Rights reserved.
# See file COPYRIGHT for details.
#
# This file is part of the MFEM library. For more information and source code
# availability see http://mfem.org.
#
# MFEM is free software; you can redistribute it and/or modify it under the
# terms of the GNU Lesser General Public License (as published by the Free
# Software Foundation) version 2.1 dated February 1999.
# Defines the following variables:
# - OCCA_FOUND
# - OCCA_LIBRARIES
# - OCCA_INCLUDE_DIRS
include(MfemCmakeUtilities)
mfem_find_package(OCCA OCCA OCCA_DIR "include" "occa.hpp" "lib" "occa"
"Paths to headers required by OCCA." "Libraries required by OCCA.")
+30
View File
@@ -0,0 +1,30 @@
# Copyright (c) 2010, Lawrence Livermore National Security, LLC. Produced at the
# Lawrence Livermore National Laboratory. LLNL-CODE-443211. All Rights reserved.
# See file COPYRIGHT for details.
#
# This file is part of the MFEM library. For more information and source code
# availability see http://mfem.org.
#
# MFEM is free software; you can redistribute it and/or modify it under the
# terms of the GNU Lesser General Public License (as published by the Free
# Software Foundation) version 2.1 dated February 1999.
# Defines the following variables:
# - RAJA_FOUND
# - RAJA_LIBRARIES
# - RAJA_INCLUDE_DIRS
include(MfemCmakeUtilities)
mfem_find_package(RAJA RAJA RAJA_DIR "include" "RAJA/RAJA.hpp" "lib" "RAJA"
"Paths to headers required by RAJA." "Libraries required by RAJA.")
if (NOT RAJA_CONFIG_CMAKE)
set(RAJA_CONFIG_CMAKE "${RAJA_DIR}/share/raja/cmake/raja-config.cmake")
endif()
if (EXISTS "${RAJA_CONFIG_CMAKE}")
include("${RAJA_CONFIG_CMAKE}")
if (ENABLE_CUDA AND NOT MFEM_USE_CUDA)
message(FATAL_ERROR
"RAJA is built with CUDA: MFEM_USE_CUDA=YES is required")
endif()
endif()
@@ -232,10 +232,12 @@ function(mfem_find_package Name Prefix DirVar IncSuffixes Header LibSuffixes
# If we have the TPL_ versions of _INCLUDE_DIRS and _LIBRARIES then set the
# standard ${Prefix} versions
if (TPL_${Prefix}_INCLUDE_DIRS)
set(${Prefix}_INCLUDE_DIRS ${TPL_${Prefix}_INCLUDE_DIRS} CACHE STRING "TPL_${Prefix}_INCLUDE_DIRS was found." FORCE)
set(${Prefix}_INCLUDE_DIRS ${TPL_${Prefix}_INCLUDE_DIRS} CACHE STRING
"TPL_${Prefix}_INCLUDE_DIRS was found." FORCE)
endif()
if (TPL_${Prefix}_LIBRARIES)
set(${Prefix}_LIBRARIES ${TPL_${Prefix}_LIBRARIES} CACHE STRING "TPL_${Prefix}_LIBRARIES was found." FORCE)
set(${Prefix}_LIBRARIES ${TPL_${Prefix}_LIBRARIES} CACHE STRING
"TPL_${Prefix}_LIBRARIES was found." FORCE)
endif()
# Quick return
@@ -718,7 +720,7 @@ function(mfem_export_mk_files)
MFEM_USE_MEMALLOC MFEM_USE_SUNDIALS MFEM_USE_MESQUITE MFEM_USE_SUITESPARSE
MFEM_USE_SUPERLU MFEM_USE_STRUMPACK MFEM_USE_GECKO MFEM_USE_GNUTLS
MFEM_USE_NETCDF MFEM_USE_PETSC MFEM_USE_MPFR MFEM_USE_SIDRE
MFEM_USE_CONDUIT MFEM_USE_PUMI)
MFEM_USE_CONDUIT MFEM_USE_PUMI MFEM_USE_CUDA MFEM_USE_OCCA MFEM_USE_RAJA)
foreach(var ${CONFIG_MK_BOOL_VARS})
if (${var})
set(${var} YES)
@@ -726,6 +728,7 @@ function(mfem_export_mk_files)
set(${var} NO)
endif()
endforeach()
# TODO: Add support for MFEM_USE_CUDA=YES
set(MFEM_CXX ${CMAKE_CXX_COMPILER})
set(MFEM_CPPFLAGS "")
string(STRIP "${CMAKE_CXX_FLAGS_${BUILD_TYPE}} ${CMAKE_CXX_FLAGS}"
+3 -6
View File
@@ -10,18 +10,15 @@
// Software Foundation) version 2.1 dated February 1999.
// Support out-of-source builds: if MFEM_BUILD_DIR is defined, load the config
// file MFEM_BUILD_DIR/config/_config.hpp.
// Support out-of-source builds: if MFEM_CONFIG_FILE is defined, include it.
//
// Otherwise, use the local file: _config.hpp.
#ifndef MFEM_CONFIG_HPP
#define MFEM_CONFIG_HPP
#ifdef MFEM_BUILD_DIR
#define MFEM_QUOTE(a) #a
#define MFEM_MAKE_PATH(x,y) MFEM_QUOTE(x/y)
#include MFEM_MAKE_PATH(MFEM_BUILD_DIR,config/_config.hpp)
#ifdef MFEM_CONFIG_FILE
#include MFEM_CONFIG_FILE
#else
#include "_config.hpp"
#endif
+5 -4
View File
@@ -121,19 +121,20 @@
// Enable MFEM functionality based on the PUMI library
// #define MFEM_USE_PUMI
// Build the GPU/CUDA-enabled version of the MFEM library.
// Build the NVIDIA GPU/CUDA-enabled version of the MFEM library.
// Requires a CUDA compiler (nvcc).
// #define MFEM_USE_CUDA
// Build the AMD GPU/HIP-enabled version of the MFEM library.
// Requires a HIP compiler (hipcc).
// #define MFEM_USE_HIP
// Enable functionality based on the RAJA library.
// #define MFEM_USE_RAJA
// Enable functionality based on the OCCA library.
// #define MFEM_USE_OCCA
// Enable MFEM's internal Memory Manager (needed e.g. for MFEM_USE_CUDA)
// #define MFEM_USE_MM
// Version of HYPRE used for building MFEM.
// #define MFEM_HYPRE_VERSION @MFEM_HYPRE_VERSION@
+1 -1
View File
@@ -42,9 +42,9 @@ MFEM_USE_SIDRE = @MFEM_USE_SIDRE@
MFEM_USE_CONDUIT = @MFEM_USE_CONDUIT@
MFEM_USE_PUMI = @MFEM_USE_PUMI@
MFEM_USE_CUDA = @MFEM_USE_CUDA@
MFEM_USE_HIP = @MFEM_USE_HIP@
MFEM_USE_RAJA = @MFEM_USE_RAJA@
MFEM_USE_OCCA = @MFEM_USE_OCCA@
MFEM_USE_MM = @MFEM_USE_MM@
# Compiler, compile options, and link options
MFEM_CXX = @MFEM_CXX@
+11 -2
View File
@@ -42,6 +42,9 @@ option(MFEM_USE_MPFR "Enable MPFR usage." OFF)
option(MFEM_USE_SIDRE "Enable Axom/Sidre usage" OFF)
option(MFEM_USE_CONDUIT "Enable Conduit usage" OFF)
option(MFEM_USE_PUMI "Enable PUMI" OFF)
option(MFEM_USE_CUDA "Enable CUDA" OFF)
option(MFEM_USE_OCCA "Enable OCCA" OFF)
option(MFEM_USE_RAJA "Enable RAJA" OFF)
set(MFEM_MPI_NP 4 CACHE STRING "Number of processes used for MPI tests")
@@ -59,13 +62,16 @@ option(MFEM_ENABLE_MINIAPPS "Build all of the miniapps" OFF)
# set(CXX g++)
# set(MPICXX mpicxx)
# Set the target CUDA architecture
set(CUDA_ARCH "sm_60" CACHE STRING "Target CUDA architecture.")
set(MFEM_DIR ${CMAKE_CURRENT_SOURCE_DIR})
# The *_DIR paths below will be the first place searched for the corresponding
# headers and library. If these fail, then standard cmake search is performed.
# Note: if the variables are already in the cache, they are not overwritten.
set(HYPRE_DIR "${MFEM_DIR}/../hypre-2.10.0b/src/hypre" CACHE PATH
set(HYPRE_DIR "${MFEM_DIR}/../hypre/src/hypre" CACHE PATH
"Path to the hypre library.")
# If hypre was compiled to depend on BLAS and LAPACK:
# set(HYPRE_REQUIRED_PACKAGES "BLAS" "LAPACK" CACHE STRING
@@ -75,7 +81,7 @@ set(METIS_DIR "${MFEM_DIR}/../metis-4.0" CACHE PATH "Path to the METIS library."
set(LIBUNWIND_DIR "" CACHE PATH "Path to Libunwind.")
set(SUNDIALS_DIR "${MFEM_DIR}/../sundials-3.0.0" CACHE PATH
set(SUNDIALS_DIR "${MFEM_DIR}/../sundials-5.0.0/instdir" CACHE PATH
"Path to the SUNDIALS library.")
# The following may be necessary, if SUNDIALS was built with KLU:
# set(SUNDIALS_REQUIRED_PACKAGES "SuiteSparse/KLU/AMD/BTF/COLAMD/config"
@@ -154,6 +160,9 @@ set(Axom_REQUIRED_PACKAGES "Conduit/relay" CACHE STRING
set(PUMI_DIR "${MFEM_DIR}/../pumi-2.1.0" CACHE STRING
"Directory where PUMI is installed")
set(OCCA_DIR "${MFEM_DIR}/../occa" CACHE PATH "Path to OCCA")
set(RAJA_DIR "${MFEM_DIR}/../raja" CACHE PATH "Path to RAJA")
set(BLAS_INCLUDE_DIRS "" CACHE STRING "Path to BLAS headers.")
set(BLAS_LIBRARIES "" CACHE STRING "The BLAS library.")
set(LAPACK_INCLUDE_DIRS "" CACHE STRING "Path to LAPACK headers.")
+21 -11
View File
@@ -46,6 +46,14 @@ CUDA_FLAGS = -x=cu --expt-extended-lambda -arch=$(CUDA_ARCH)
CUDA_XCOMPILER = -Xcompiler=
CUDA_XLINKER = -Xlinker=
# HIP configuration options
HIP_CXX = hipcc
# The HIP_ARCH option specifies the AMD GPU processor, similar to CUDA_ARCH. For
# example: gfx600 (tahiti), gfx700 (kaveri), gfx701 (hawaii), gfx801 (carrizo),
# gfx900, gfx1010, etc.
HIP_ARCH = gfx900
HIP_FLAGS = --amdgpu-target=$(HIP_ARCH)
ifneq ($(NOTMAC),)
AR = ar
ARFLAGS = cruv
@@ -122,9 +130,9 @@ MFEM_USE_SIDRE = NO
MFEM_USE_CONDUIT = NO
MFEM_USE_PUMI = NO
MFEM_USE_CUDA = NO
MFEM_USE_HIP = NO
MFEM_USE_RAJA = NO
MFEM_USE_OCCA = NO
MFEM_USE_MM = NO
# Compile and link options for zlib.
ZLIB_DIR =
@@ -136,7 +144,7 @@ LIBUNWIND_OPT = -g
LIBUNWIND_LIB = $(if $(NOTMAC),-lunwind -ldl,)
# HYPRE library configuration (needed to build the parallel version)
HYPRE_DIR = @MFEM_DIR@/../hypre-2.10.0b/src/hypre
HYPRE_DIR = @MFEM_DIR@/../hypre/src/hypre
HYPRE_OPT = -I$(HYPRE_DIR)/include
HYPRE_LIB = -L$(HYPRE_DIR)/lib -lHYPRE
@@ -174,9 +182,9 @@ OPENMP_LIB =
POSIX_CLOCKS_LIB = -lrt
# SUNDIALS library configuration
SUNDIALS_DIR = @MFEM_DIR@/../sundials-3.0.0
SUNDIALS_DIR = @MFEM_DIR@/../sundials-5.0.0/instdir
SUNDIALS_OPT = -I$(SUNDIALS_DIR)/include
SUNDIALS_LIB = -Wl,-rpath,$(SUNDIALS_DIR)/lib -L$(SUNDIALS_DIR)/lib\
SUNDIALS_LIB = -Wl,-rpath,$(SUNDIALS_DIR)/lib64 -L$(SUNDIALS_DIR)/lib64\
-lsundials_arkode -lsundials_cvode -lsundials_nvecserial -lsundials_kinsol
ifeq ($(MFEM_USE_MPI),YES)
@@ -201,7 +209,7 @@ SUITESPARSE_LIB = -Wl,-rpath,$(SUITESPARSE_DIR)/lib -L$(SUITESPARSE_DIR)/lib\
# SuperLU library configuration
SUPERLU_DIR = @MFEM_DIR@/../SuperLU_DIST_5.1.0
SUPERLU_OPT = -I$(SUPERLU_DIR)/SRC
SUPERLU_LIB = -Wl,-rpath,$(SUPERLU_DIR)/SRC -L$(SUPERLU_DIR)/SRC -lsuperlu_dist
SUPERLU_LIB = -Wl,-rpath,$(SUPERLU_DIR)/lib -L$(SUPERLU_DIR)/lib -lsuperlu_dist_5.1.0
# SCOTCH library configuration (required by STRUMPACK <= v2.1.0, optional in
# STRUMPACK >= v2.2.0)
@@ -300,19 +308,21 @@ PUMI_OPT = -I$(PUMI_DIR)/include
PUMI_LIB = -L$(PUMI_DIR)/lib -lpumi -lcrv -lma -lmds -lapf -lpcu -lgmi -lparma\
-llion -lmth -lapf_zoltan -lspr
# CUDA library configuration. Since we compile and link with nvcc (when CUDA is
# enabled) we only need to explicitly link with the CUDA driver, libcuda.*,
# which is usually in a system path.
# CUDA library configuration (currently not needed)
CUDA_OPT =
CUDA_LIB = $(if $(NOTMAC),,-L/usr/local/cuda/lib) -lcuda
CUDA_LIB =
# HIP library configuration (currently not needed)
HIP_OPT =
HIP_LIB =
# OCCA library configuration
OCCA_DIR ?= @MFEM_DIR@/../occa
OCCA_DIR = @MFEM_DIR@/../occa
OCCA_OPT = -I$(OCCA_DIR)/include
OCCA_LIB = $(XLINKER)-rpath,$(OCCA_DIR)/lib -L$(OCCA_DIR)/lib -locca
# RAJA library configuration
RAJA_DIR ?= @MFEM_DIR@/../raja
RAJA_DIR = @MFEM_DIR@/../raja
RAJA_OPT = -I$(RAJA_DIR)/include
ifdef CUB_DIR
RAJA_OPT += -I$(CUB_DIR)
+2 -1
View File
@@ -36,6 +36,7 @@ CONFIG_MK = config.mk
all: header config-mk
MPI = $(MFEM_USE_MPI:NO=)
GHV_CXX ?= $(MFEM_CXX)
GHV = get_hypre_version
GHV_FLAGS = $(subst @MFEM_DIR@,$(if $(MFEM_DIR),$(MFEM_DIR),..),$(HYPRE_OPT))
SMX = $(if $(MFEM_USE_PUMI:NO=),MFEM_USE_SIMMETRIX)
@@ -44,7 +45,7 @@ SMX_FILE = $(subst @MFEM_DIR@,$(if $(MFEM_DIR),$(MFEM_DIR),..),$(SMX_PATH))
$(GHV): $(SRC)$(GHV).cpp
$(call mfem-info, Determining HYPRE version ...)
$(MFEM_CXX) ${GHV_FLAGS} $(SRC)$(GHV).cpp -o $(GHV)
$(GHV_CXX) ${GHV_FLAGS} $(SRC)$(GHV).cpp -o $(GHV)
$(GHV).out: $(GHV)
./$(GHV) > $(GHV).out
.INTERMEDIATE: $(GHV) $(GHV).out
+19 -1
View File
@@ -18,6 +18,8 @@ run_prefix=""
run_vg="valgrind --leak-check=full --show-reachable=yes --track-origins=yes"
run_suffix="-no-vis"
skip_gen_meshes="yes"
# filter-out device runs ("no") or non-device runs ("yes"):
device_runs="no"
cur_dir="${PWD}"
mfem_dir="$(cd "$(dirname "$0")"/.. && pwd)"
mfem_build_dir=""
@@ -148,6 +150,11 @@ function extract_sample_runs()
if [ "$skip_gen_meshes" == "yes" ]; then
runs=`printf "%s" "$runs" | grep -v ".* -m .*\.gen"`
fi
if [ "$device_runs" == "yes" ]; then
runs=`printf "%s" "$runs" | grep ".* -d .*"`
else
runs=`printf "%s" "$runs" | grep -v ".* -d .*"`
fi
IFS=$'\n'
runs=(${runs})
IFS="${old_IFS}"
@@ -169,6 +176,9 @@ function help_message()
-g <dir> <pattern>
Specify explicitly a group (dir + file pattern) to run; This
option can be used multiple times to define multiple groups
-dev configure only sample runs using devices.
To test with a parallel build, the parallel (-p|-par) option
should be set first on the command line.
-v Enable valgrind
-o <dir> [${output_dir:-"<empty>: output goes to stdout"}]
If not empty, save output to files inside <dir>
@@ -253,7 +263,7 @@ case "$1" in
-h|-help)
opt_help="yes"
;;
-p|-parallel)
-p|-par)
mfem_config="MFEM_USE_MPI=YES MFEM_DEBUG=NO"
;;
-g)
@@ -264,6 +274,10 @@ case "$1" in
groups=("${groups[@]}" "${test_group}")
shift 2
;;
-dev)
device_runs="yes"
mfem_config+=" MFEM_USE_CUDA=YES MFEM_USE_OCCA=YES MFEM_USE_RAJA=YES MFEM_USE_OPENMP=YES"
;;
-v)
valgrind="yes"
;;
@@ -294,6 +308,10 @@ case "$1" in
-n)
run_prefix="echo"
;;
-*)
echo "unknown option: '$1'"
exit 1
;;
*=*)
eval $1
;;
+2 -3
View File
@@ -43,15 +43,14 @@
#define MFEM_ALIGN_SIZE(size,type) \
MFEM_ROUNDUP(size,(MFEM_SIMD_SIZE)/sizeof(type))
#ifdef MFEM_COUNT_FLOPS
namespace mfem
{
namespace internal
{
long long flop_count;
extern long long flop_count;
}
}
#ifdef MFEM_COUNT_FLOPS
#define MFEM_FLOPS_RESET() (mfem::internal::flop_count = 0)
#define MFEM_FLOPS_ADD(cnt) (mfem::internal::flop_count += (cnt))
#define MFEM_FLOPS_GET() (mfem::internal::flop_count)
+1 -1
View File
@@ -38,7 +38,7 @@ PROJECT_NAME = "MFEM"
# could be handy for archiving the generated documentation or if some version
# control system is used.
PROJECT_NUMBER = v3.4.1
PROJECT_NUMBER = v4.0.1
# Using the PROJECT_BRIEF tag one can provide an optional one line description
# for a project that appears at the top of each page and should give viewer a
+8 -2
View File
@@ -35,6 +35,12 @@ namespace mfem {
* - HypreParMatrix and HypreParVector
* - HypreSolver and other \link hypre.hpp hypre classes\endlink
*
* <H3>Main GPU classes</H3>
* - Device
* - Memory
* - MemoryManager
* - MFEM_FORALL macro in forall.hpp
*
* <H3>Example codes</H3>
* - <a class="el" href="examples_2ex1_8cpp_source.html">Example 1</a>: nodal H1 FEM for the Laplace problem
* - <a class="el" href="examples_2ex1p_8cpp_source.html">Example 1p</a>: parallel nodal H1 FEM for the Laplace problem
@@ -73,8 +79,8 @@ namespace mfem {
* - <a class="el" href="ex19p_8cpp_source.html">Example 19p</a>: parallel incompressible nonlinear elasticity
* - <a class="el" href="ex20_8cpp_source.html">Example 20</a>: symplectic ODE integration
* - <a class="el" href="ex20p_8cpp_source.html">Example 20p</a>: parallel symplectic ODE integration
* - <a class="el" href="ex22_8cpp_source.html">Example 22</a>: adaptive mesh refinement for linear elasticity
* - <a class="el" href="ex22p_8cpp_source.html">Example 22p</a>: parallel adaptive mesh refinement for linear elasticity
* - <a class="el" href="ex21_8cpp_source.html">Example 21</a>: adaptive mesh refinement for linear elasticity
* - <a class="el" href="ex21p_8cpp_source.html">Example 21p</a>: parallel adaptive mesh refinement for linear elasticity
*
* <H4>SUNDIALS Examples</H4>
* - Variants of Examples
Binary file not shown.

After

Width:  |  Height:  |  Size: 134 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 73 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 128 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 66 KiB

+2 -2
View File
@@ -27,7 +27,7 @@ list(APPEND ALL_EXE_SRCS
ex18.cpp
ex19.cpp
ex20.cpp
ex22.cpp
ex21.cpp
)
if (MFEM_USE_MPI)
@@ -52,7 +52,7 @@ if (MFEM_USE_MPI)
ex18p.cpp
ex19p.cpp
ex20p.cpp
ex22p.cpp
ex21p.cpp
)
endif()
+245 -164
View File
File diff suppressed because one or more lines are too long
+21 -25
View File
@@ -26,12 +26,12 @@
// ex1 -m ../data/mobius-strip.mesh -o -1 -sc
//
// Device sample runs:
// > ex1 -pa -d cuda
// > ex1 -pa -d raja-cuda
// > ex1 -pa -d occa-cuda
// > ex1 -pa -d raja-omp
// > ex1 -pa -d occa-omp
// > ex1 -m ../data/beam-hex.mesh -pa -d cuda
// ex1 -pa -d cuda
// ex1 -pa -d raja-cuda
// ex1 -pa -d occa-cuda
// ex1 -pa -d raja-omp
// ex1 -pa -d occa-omp
// ex1 -m ../data/beam-hex.mesh -pa -d cuda
//
// Description: This example code demonstrates the use of MFEM to define a
// simple finite element discretization of the Laplace problem
@@ -62,7 +62,7 @@ int main(int argc, char *argv[])
int order = 1;
bool static_cond = false;
bool pa = false;
const char *device = "cpu";
const char *device_config = "cpu";
bool visualization = true;
OptionsParser args(argc, argv);
@@ -75,7 +75,7 @@ int main(int argc, char *argv[])
"--no-static-condensation", "Enable static condensation.");
args.AddOption(&pa, "-pa", "--partial-assembly", "-no-pa",
"--no-partial-assembly", "Enable Partial Assembly.");
args.AddOption(&device, "-d", "--device",
args.AddOption(&device_config, "-d", "--device",
"Device configuration string, see Device::Configure().");
args.AddOption(&visualization, "-vis", "--visualization", "-no-vis",
"--no-visualization",
@@ -88,13 +88,18 @@ int main(int argc, char *argv[])
}
args.PrintOptions(cout);
// 2. Read the mesh from the given mesh file. We can handle triangular,
// 2. Enable hardware devices such as GPUs, and programming models such as
// CUDA, OCCA, RAJA and OpenMP based on command line options.
Device device(device_config);
device.Print();
// 3. Read the mesh from the given mesh file. We can handle triangular,
// quadrilateral, tetrahedral, hexahedral, surface and volume meshes with
// the same code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
// 3. Refine the mesh to increase the resolution. In this example we do
// 4. Refine the mesh to increase the resolution. In this example we do
// 'ref_levels' of uniform refinement. We choose 'ref_levels' to be the
// largest number that gives a final mesh with no more than 50,000
// elements.
@@ -107,7 +112,7 @@ int main(int argc, char *argv[])
}
}
// 4. Define a finite element space on the mesh. Here we use continuous
// 5. Define a finite element space on the mesh. Here we use continuous
// Lagrange finite elements of the specified order. If order < 1, we
// instead use an isoparametric/isogeometric space.
FiniteElementCollection *fec;
@@ -128,7 +133,7 @@ int main(int argc, char *argv[])
cout << "Number of finite element unknowns: "
<< fespace->GetTrueVSize() << endl;
// 5. Determine the list of true (i.e. conforming) essential boundary dofs.
// 6. Determine the list of true (i.e. conforming) essential boundary dofs.
// In this example, the boundary conditions are defined by marking all
// the boundary attributes from the mesh as essential (Dirichlet) and
// converting them to a list of true dofs.
@@ -140,7 +145,7 @@ int main(int argc, char *argv[])
fespace->GetEssentialTrueDofs(ess_bdr, ess_tdof_list);
}
// 6. Set up the linear form b(.) which corresponds to the right-hand side of
// 7. Set up the linear form b(.) which corresponds to the right-hand side of
// the FEM linear system, which in this case is (1,phi_i) where phi_i are
// the basis functions in the finite element fespace.
LinearForm *b = new LinearForm(fespace);
@@ -148,12 +153,6 @@ int main(int argc, char *argv[])
b->AddDomainIntegrator(new DomainLFIntegrator(one));
b->Assemble();
// 7. Set device config parameters from the command line options and switch
// to working on the device.
Device::Configure(device);
Device::Print();
Device::Enable();
// 8. Define the solution vector x as a finite element grid function
// corresponding to fespace. Initialize x with initial guess of zero,
// which satisfies the boundary conditions.
@@ -203,10 +202,7 @@ int main(int argc, char *argv[])
// 12. Recover the solution as a finite element grid function.
a->RecoverFEMSolution(X, *b, x);
// 13. Switch back to the host.
Device::Disable();
// 14. Save the refined mesh and the solution. This output can be viewed later
// 13. Save the refined mesh and the solution. This output can be viewed later
// using GLVis: "glvis -m refined.mesh -g sol.gf".
ofstream mesh_ofs("refined.mesh");
mesh_ofs.precision(8);
@@ -215,7 +211,7 @@ int main(int argc, char *argv[])
sol_ofs.precision(8);
x.Save(sol_ofs);
// 15. Send the solution by socket to a GLVis server.
// 14. Send the solution by socket to a GLVis server.
if (visualization)
{
char vishost[] = "localhost";
@@ -225,7 +221,7 @@ int main(int argc, char *argv[])
sol_sock << "solution\n" << *mesh << x << flush;
}
// 16. Free the used memory.
// 15. Free the used memory.
delete a;
delete b;
delete fespace;
+1 -1
View File
@@ -144,7 +144,7 @@ void InitialDeformation(const Vector &x, Vector &y);
int main(int argc, char *argv[])
{
// 1. Parse command-line options
const char *mesh_file = "../data/beam-hex.mesh";
const char *mesh_file = "../data/beam-tet.mesh";
int ref_levels = 0;
int order = 2;
bool visualization = true;
+1 -1
View File
@@ -150,7 +150,7 @@ int main(int argc, char *argv[])
MPI_Comm_rank(MPI_COMM_WORLD, &myid);
// 2. Parse command-line options
const char *mesh_file = "../data/beam-hex.mesh";
const char *mesh_file = "../data/beam-tet.mesh";
int ser_ref_levels = 0;
int par_ref_levels = 0;
int order = 2;
+19 -23
View File
@@ -26,9 +26,9 @@
// mpirun -np 4 ex1p -m ../data/mobius-strip.mesh -o -1 -sc
//
// Device sample runs:
// > mpirun -np 4 ex1p -pa -d cuda
// > mpirun -np 4 ex1p -pa -d occa-cuda
// > mpirun -np 4 ex1p -pa -d raja-omp
// mpirun -np 4 ex1p -pa -d cuda
// mpirun -np 4 ex1p -pa -d occa-cuda
// mpirun -np 4 ex1p -pa -d raja-omp
//
// Description: This example code demonstrates the use of MFEM to define a
// simple finite element discretization of the Laplace problem
@@ -65,7 +65,7 @@ int main(int argc, char *argv[])
int order = 1;
bool static_cond = false;
bool pa = false;
const char *device = "cpu";
const char *device_config = "cpu";
bool visualization = true;
OptionsParser args(argc, argv);
@@ -78,7 +78,7 @@ int main(int argc, char *argv[])
"--no-static-condensation", "Enable static condensation.");
args.AddOption(&pa, "-pa", "--partial-assembly", "-no-pa",
"--no-partial-assembly", "Enable Partial Assembly.");
args.AddOption(&device, "-d", "--device",
args.AddOption(&device_config, "-d", "--device",
"Device configuration string, see Device::Configure().");
args.AddOption(&visualization, "-vis", "--visualization", "-no-vis",
"--no-visualization",
@@ -98,13 +98,18 @@ int main(int argc, char *argv[])
args.PrintOptions(cout);
}
// 3. Read the (serial) mesh from the given mesh file on all processors. We
// 3. Enable hardware devices such as GPUs, and programming models such as
// CUDA, OCCA, RAJA and OpenMP based on command line options.
Device device(device_config);
if (myid == 0) { device.Print(); }
// 4. Read the (serial) mesh from the given mesh file on all processors. We
// can handle triangular, quadrilateral, tetrahedral, hexahedral, surface
// and volume meshes with the same code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
// 4. Refine the serial mesh on all processors to increase the resolution. In
// 5. Refine the serial mesh on all processors to increase the resolution. In
// this example we do 'ref_levels' of uniform refinement. We choose
// 'ref_levels' to be the largest number that gives a final mesh with no
// more than 10,000 elements.
@@ -117,7 +122,7 @@ int main(int argc, char *argv[])
}
}
// 5. Define a parallel mesh by a partitioning of the serial mesh. Refine
// 6. Define a parallel mesh by a partitioning of the serial mesh. Refine
// this mesh further in parallel to increase the resolution. Once the
// parallel mesh is defined, the serial mesh can be deleted.
ParMesh *pmesh = new ParMesh(MPI_COMM_WORLD, *mesh);
@@ -130,7 +135,7 @@ int main(int argc, char *argv[])
}
}
// 6. Define a parallel finite element space on the parallel mesh. Here we
// 7. Define a parallel finite element space on the parallel mesh. Here we
// use continuous Lagrange finite elements of the specified order. If
// order < 1, we instead use an isoparametric/isogeometric space.
FiniteElementCollection *fec;
@@ -157,7 +162,7 @@ int main(int argc, char *argv[])
cout << "Number of finite element unknowns: " << size << endl;
}
// 7. Determine the list of true (i.e. parallel conforming) essential
// 8. Determine the list of true (i.e. parallel conforming) essential
// boundary dofs. In this example, the boundary conditions are defined
// by marking all the boundary attributes from the mesh as essential
// (Dirichlet) and converting them to a list of true dofs.
@@ -169,7 +174,7 @@ int main(int argc, char *argv[])
fespace->GetEssentialTrueDofs(ess_bdr, ess_tdof_list);
}
// 8. Set up the parallel linear form b(.) which corresponds to the
// 9. Set up the parallel linear form b(.) which corresponds to the
// right-hand side of the FEM linear system, which in this case is
// (1,phi_i) where phi_i are the basis functions in fespace.
ParLinearForm *b = new ParLinearForm(fespace);
@@ -177,12 +182,6 @@ int main(int argc, char *argv[])
b->AddDomainIntegrator(new DomainLFIntegrator(one));
b->Assemble();
// 9. Set device config parameters from the command line options and switch
// to working on the device.
Device::Configure(device);
if (myid == 0) { Device::Print(); }
Device::Enable();
// 10. Define the solution vector x as a parallel finite element grid function
// corresponding to fespace. Initialize x with initial guess of zero,
// which satisfies the boundary conditions.
@@ -225,10 +224,7 @@ int main(int argc, char *argv[])
// local finite element solution on each processor.
a->RecoverFEMSolution(X, *b, x);
// 15. Switch back to the host.
Device::Disable();
// 16. Save the refined mesh and the solution in parallel. This output can
// 15. Save the refined mesh and the solution in parallel. This output can
// be viewed later using GLVis: "glvis -np <np> -m mesh -g sol".
{
ostringstream mesh_name, sol_name;
@@ -244,7 +240,7 @@ int main(int argc, char *argv[])
x.Save(sol_ofs);
}
// 17. Send the solution by socket to a GLVis server.
// 16. Send the solution by socket to a GLVis server.
if (visualization)
{
char vishost[] = "localhost";
@@ -255,7 +251,7 @@ int main(int argc, char *argv[])
sol_sock << "solution\n" << *pmesh << x << flush;
}
// 18. Free the used memory.
// 17. Free the used memory.
delete a;
delete b;
delete fespace;
+14 -14
View File
@@ -1,16 +1,16 @@
// MFEM Example 22
// MFEM Example 21
//
// Compile with: make ex22
// Compile with: make ex21
//
// Sample runs: ex22
// ex22 -o 3
// ex22 -m ../data/beam-quad.mesh
// ex22 -m ../data/beam-quad.mesh -o 3
// ex22 -m ../data/beam-quad.mesh -o 3 -f 1
// ex22 -m ../data/beam-tet.mesh
// ex22 -m ../data/beam-tet.mesh -o 2
// ex22 -m ../data/beam-hex.mesh
// ex22 -m ../data/beam-hex.mesh -o 2
// Sample runs: ex21
// ex21 -o 3
// ex21 -m ../data/beam-quad.mesh
// ex21 -m ../data/beam-quad.mesh -o 3
// ex21 -m ../data/beam-quad.mesh -o 3 -f 1
// ex21 -m ../data/beam-tet.mesh
// ex21 -m ../data/beam-tet.mesh -o 2
// ex21 -m ../data/beam-hex.mesh
// ex21 -m ../data/beam-hex.mesh -o 2
//
// Description: This is a version of Example 2 with a simple adaptive mesh
// refinement loop. The problem being solved is again the linear
@@ -287,11 +287,11 @@ int main(int argc, char *argv[])
}
{
ofstream mesh_ref_out("ex22_reference.mesh");
ofstream mesh_ref_out("ex21_reference.mesh");
mesh_ref_out.precision(16);
mesh.Print(mesh_ref_out);
ofstream mesh_out("ex22_deformed.mesh");
ofstream mesh_out("ex21_deformed.mesh");
mesh_out.precision(16);
GridFunction nodes(&fespace), *nodes_p = &nodes;
mesh.GetNodes(nodes);
@@ -301,7 +301,7 @@ int main(int argc, char *argv[])
mesh.Print(mesh_out);
mesh.SwapNodes(nodes_p, own_nodes);
ofstream x_out("ex22_displacement.sol");
ofstream x_out("ex21_displacement.sol");
x_out.precision(16);
x.Save(x_out);
}
+14 -14
View File
@@ -1,15 +1,15 @@
// MFEM Example 22
// MFEM Example 21
//
// Compile with: make ex22p
// Compile with: make ex21p
//
// Sample runs: mpirun -np 4 ex22p
// mpirun -np 4 ex22p -o 3
// mpirun -np 4 ex22p -m ../data/beam-quad.mesh
// mpirun -np 4 ex22p -m ../data/beam-quad.mesh -o 3
// mpirun -np 4 ex22p -m ../data/beam-tet.mesh
// mpirun -np 4 ex22p -m ../data/beam-tet.mesh -o 2
// mpirun -np 4 ex22p -m ../data/beam-hex.mesh
// mpirun -np 4 ex22p -m ../data/beam-hex.mesh -o 2
// Sample runs: mpirun -np 4 ex21p
// mpirun -np 4 ex21p -o 3
// mpirun -np 4 ex21p -m ../data/beam-quad.mesh
// mpirun -np 4 ex21p -m ../data/beam-quad.mesh -o 3
// mpirun -np 4 ex21p -m ../data/beam-tet.mesh
// mpirun -np 4 ex21p -m ../data/beam-tet.mesh -o 2
// mpirun -np 4 ex21p -m ../data/beam-hex.mesh
// mpirun -np 4 ex21p -m ../data/beam-hex.mesh -o 2
//
// Description: This is a version of Example 2p with a simple adaptive mesh
// refinement loop. The problem being solved is again the linear
@@ -330,7 +330,7 @@ int main(int argc, char *argv[])
x.Update();
}
// 22. Inform also the bilinear and linear forms that the space has
// 21. Inform also the bilinear and linear forms that the space has
// changed.
a.Update();
b.Update();
@@ -338,9 +338,9 @@ int main(int argc, char *argv[])
{
ostringstream mref_name, mesh_name, sol_name;
mref_name << "ex22p_reference_mesh." << setfill('0') << setw(6) << myid;
mesh_name << "ex22p_deformed_mesh." << setfill('0') << setw(6) << myid;
sol_name << "ex22p_displacement." << setfill('0') << setw(6) << myid;
mref_name << "ex21p_reference_mesh." << setfill('0') << setw(6) << myid;
mesh_name << "ex21p_deformed_mesh." << setfill('0') << setw(6) << myid;
sol_name << "ex21p_displacement." << setfill('0') << setw(6) << myid;
ofstream mesh_ref_out(mref_name.str().c_str());
mesh_ref_out.precision(16);
+14 -15
View File
@@ -16,9 +16,9 @@
// ex6 -m ../data/amr-quad.mesh
//
// Device sample runs:
// > ex6 -pa -d cuda
// > ex6 -pa -d occa-cuda
// > ex6 -pa -d raja-omp
// ex6 -pa -d cuda
// ex6 -pa -d occa-cuda
// ex6 -pa -d raja-omp
//
// Description: This is a version of Example 1 with a simple adaptive mesh
// refinement loop. The problem being solved is again the Laplace
@@ -49,7 +49,7 @@ int main(int argc, char *argv[])
const char *mesh_file = "../data/star.mesh";
int order = 1;
bool pa = false;
const char *device = "cpu";
const char *device_config = "cpu";
bool visualization = true;
OptionsParser args(argc, argv);
@@ -59,7 +59,7 @@ int main(int argc, char *argv[])
"Finite element order (polynomial degree).");
args.AddOption(&pa, "-pa", "--partial-assembly", "-no-pa",
"--no-partial-assembly", "Enable Partial Assembly.");
args.AddOption(&device, "-d", "--device",
args.AddOption(&device_config, "-d", "--device",
"Device configuration string, see Device::Configure().");
args.AddOption(&visualization, "-vis", "--visualization", "-no-vis",
"--no-visualization",
@@ -72,14 +72,19 @@ int main(int argc, char *argv[])
}
args.PrintOptions(cout);
// 2. Read the mesh from the given mesh file. We can handle triangular,
// 2. Enable hardware devices such as GPUs, and programming models such as
// CUDA, OCCA, RAJA and OpenMP based on command line options.
Device device(device_config);
device.Print();
// 3. Read the mesh from the given mesh file. We can handle triangular,
// quadrilateral, tetrahedral, hexahedral, surface and volume meshes with
// the same code.
Mesh mesh(mesh_file, 1, 1);
int dim = mesh.Dimension();
int sdim = mesh.SpaceDimension();
// 3. Since a NURBS mesh can currently only be refined uniformly, we need to
// 4. Since a NURBS mesh can currently only be refined uniformly, we need to
// convert it to a piecewise-polynomial curved mesh. First we refine the
// NURBS mesh a bit more and then project the curvature to quadratic Nodes.
if (mesh.NURBSext)
@@ -91,15 +96,11 @@ int main(int argc, char *argv[])
mesh.SetCurvature(2);
}
// 4. Define a finite element space on the mesh. The polynomial order is
// 5. Define a finite element space on the mesh. The polynomial order is
// one (linear) by default, but this can be changed on the command line.
H1_FECollection fec(order, dim);
FiniteElementSpace fespace(&mesh, &fec);
// 5. Set device config parameters from the command line options.
Device::Configure(device);
Device::Print();
// 6. As in Example 1, we set up bilinear and linear forms corresponding to
// the Laplace problem -\Delta u = 1. We don't assemble the discrete
// problem yet, this will be done in the main loop.
@@ -168,8 +169,7 @@ int main(int argc, char *argv[])
x.ProjectBdrCoefficient(zero, ess_bdr);
fespace.GetEssentialTrueDofs(ess_bdr, ess_tdof_list);
// 15. Switch to the device and assemble the stiffness matrix.
Device::Enable();
// 15. Assemble the stiffness matrix.
a.Assemble();
// 16. Create the linear system: eliminate boundary conditions, constrain
@@ -204,7 +204,6 @@ int main(int argc, char *argv[])
// 18. After solving the linear system, reconstruct the solution as a
// finite element GridFunction. Constrained nodes are interpolated
// from true DOFs (it may therefore happen that x.Size() >= X.Size()).
Device::Disable();
a.RecoverFEMSolution(X, b, x);
// 19. Send solution by socket to the GLVis server.
+18 -19
View File
@@ -16,9 +16,9 @@
// mpirun -np 4 ex6p -m ../data/amr-quad.mesh
//
// Device sample runs:
// > mpirun -np 4 ex6p -pa -d cuda
// > mpirun -np 4 ex6p -pa -d occa-cuda
// > mpirun -np 4 ex6p -pa -d raja-omp
// mpirun -np 4 ex6p -pa -d cuda
// mpirun -np 4 ex6p -pa -d occa-cuda
// mpirun -np 4 ex6p -pa -d raja-omp
//
// Description: This is a version of Example 1 with a simple adaptive mesh
// refinement loop. The problem being solved is again the Laplace
@@ -55,7 +55,7 @@ int main(int argc, char *argv[])
const char *mesh_file = "../data/star.mesh";
int order = 1;
bool pa = false;
const char *device = "cpu";
const char *device_config = "cpu";
bool visualization = true;
OptionsParser args(argc, argv);
@@ -65,7 +65,7 @@ int main(int argc, char *argv[])
"Finite element order (polynomial degree).");
args.AddOption(&pa, "-pa", "--partial-assembly", "-no-pa",
"--no-partial-assembly", "Enable Partial Assembly.");
args.AddOption(&device, "-d", "--device",
args.AddOption(&device_config, "-d", "--device",
"Device configuration string, see Device::Configure().");
args.AddOption(&visualization, "-vis", "--visualization", "-no-vis",
"--no-visualization",
@@ -85,14 +85,19 @@ int main(int argc, char *argv[])
args.PrintOptions(cout);
}
// 3. Read the (serial) mesh from the given mesh file on all processors. We
// 3. Enable hardware devices such as GPUs, and programming models such as
// CUDA, OCCA, RAJA and OpenMP based on command line options.
Device device(device_config);
if (myid == 0) { device.Print(); }
// 4. Read the (serial) mesh from the given mesh file on all processors. We
// can handle triangular, quadrilateral, tetrahedral, hexahedral, surface
// and volume meshes with the same code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
int sdim = mesh->SpaceDimension();
// 4. Refine the serial mesh on all processors to increase the resolution.
// 5. Refine the serial mesh on all processors to increase the resolution.
// Also project a NURBS mesh to a piecewise-quadratic curved mesh. Make
// sure that the mesh is non-conforming.
if (mesh->NURBSext)
@@ -102,7 +107,7 @@ int main(int argc, char *argv[])
}
mesh->EnsureNCMesh();
// 5. Define a parallel mesh by partitioning the serial mesh.
// 6. Define a parallel mesh by partitioning the serial mesh.
// Once the parallel mesh is defined, the serial mesh can be deleted.
ParMesh pmesh(MPI_COMM_WORLD, *mesh);
delete mesh;
@@ -112,15 +117,11 @@ int main(int argc, char *argv[])
Array<int> ess_bdr(pmesh.bdr_attributes.Max());
ess_bdr = 1;
// 6. Define a finite element space on the mesh. The polynomial order is
// 7. Define a finite element space on the mesh. The polynomial order is
// one (linear) by default, but this can be changed on the command line.
H1_FECollection fec(order, dim);
ParFiniteElementSpace fespace(&pmesh, &fec);
// 7. Set device config parameters from the command line options.
Device::Configure(device);
if (myid == 0) { Device::Print(); }
// 8. As in Example 1p, we set up bilinear and linear forms corresponding to
// the Laplace problem -\Delta u = 1. We don't assemble the discrete
// problem yet, this will be done in the main loop.
@@ -200,11 +201,10 @@ int main(int argc, char *argv[])
fespace.GetEssentialTrueDofs(ess_bdr, ess_tdof_list);
b.Assemble();
// 15. Switch to the device and assemble the stiffness matrix. Note that
// MFEM doesn't care at this point that the mesh is nonconforming and
// parallel. The FE space is considered 'cut' along hanging
// edges/faces, and also across processor boundaries.
Device::Enable();
// 15. Assemble the stiffness matrix. Note that MFEM doesn't care at this
// point that the mesh is nonconforming and parallel. The FE space is
// considered 'cut' along hanging edges/faces, and also across
// processor boundaries.
a.Assemble();
// 16. Create the parallel linear system: eliminate boundary conditions.
@@ -232,7 +232,6 @@ int main(int argc, char *argv[])
// 18. Switch back to the host and extract the parallel grid function
// corresponding to the finite element approximation X. This is the
// local solution on each processor.
Device::Disable();
a.RecoverFEMSolution(X, b, x);
// 19. Send the solution by socket to a GLVis server.
+7 -3
View File
@@ -22,9 +22,9 @@ MFEM_LIB_FILE = mfem_is_not_built
-include $(CONFIG_MK)
SEQ_EXAMPLES = ex1 ex2 ex3 ex4 ex5 ex6 ex7 ex8 ex9 ex10 ex14 ex15 ex16 ex17\
ex18 ex19 ex20 ex22
ex18 ex19 ex20 ex21 drl4amr
PAR_EXAMPLES = ex1p ex2p ex3p ex4p ex5p ex6p ex7p ex8p ex9p ex10p ex11p ex12p\
ex13p ex14p ex15p ex16p ex17p ex18p ex19p ex20p ex22p
ex13p ex14p ex15p ex16p ex17p ex18p ex19p ex20p ex21p
ifeq ($(MFEM_USE_MPI),NO)
EXAMPLES = $(SEQ_EXAMPLES)
@@ -71,6 +71,10 @@ ifeq ($(MFEM_USE_MPI),YES)
ex18p: $(SRC)ex18.hpp
endif
drl4amr:
python drl4amr.py build
g++ -pthread -shared -Wl,-z,relro build/temp.linux-x86_64-2.7/drl4amr.o -L/usr/lib64 -lpython2.7 -o build/lib.linux-x86_64-2.7/drl4amr.so -L.. -lmfem
MFEM_TESTS = EXAMPLES
include $(MFEM_TEST_MK)
test: $(SUBDIRS_TEST)
@@ -125,4 +129,4 @@ clean-exec:
@rm -f vortex-mesh.* vortex.mesh vortex-?-init.* vortex-?-final.*
@rm -f deformation.* pressure.*
@rm -f ex20.dat ex20p_?????.dat gnuplot_ex20.inp gnuplot_ex20p.inp
@rm -f ex22*.mesh ex22*.sol ex22p_*.*
@rm -f ex21*.mesh ex21*.sol ex21p_*.*
+58 -9
View File
@@ -27,8 +27,11 @@
// method HyperelasticOperator::ImplicitSolve is the only
// requirement for high-order implicit (SDIRK) time integration.
// If using PETSc to solve the nonlinear problem, use the option
// file provided (rc_ex10p) that customizes the
// Newton-Krylov method.
// files provided (see rc_ex10p, rc_ex10p_mf, rc_ex10p_mfop) that
// customize the Newton-Krylov method.
// When option --jfnk is used, PETSc will use a Jacobian-free
// Newton-Krylov method, using a user-defined preconditioner
// constructed with the PetscPreconditionerFactory class.
//
// We recommend viewing examples 2 and 9 before viewing this
// example.
@@ -86,12 +89,15 @@ protected:
Solver *J_solver;
/// Preconditioner for the Jacobian solve in the Newton method
Solver *J_prec;
/// Preconditioner factory for JFNK
PetscPreconditionerFactory *J_factory;
mutable Vector z; // auxiliary vector
public:
HyperelasticOperator(ParFiniteElementSpace &f, Array<int> &ess_bdr,
double visc, double mu, double K, bool use_petsc);
double visc, double mu, double K,
bool use_petsc, bool petsc_use_jfnk);
/// Compute the right-hand side of the ODE system.
virtual void Mult(const Vector &vx, Vector &dvx_dt) const;
@@ -136,8 +142,21 @@ public:
virtual Operator &GetGradient(const Vector &k) const;
virtual ~ReducedSystemOperator();
};
/** Auxiliary class to provide preconditioners for matrix-free methods */
class PreconditionerFactory : public PetscPreconditionerFactory
{
private:
// const ReducedSystemOperator& op; // unused for now (generates warning)
public:
PreconditionerFactory(const ReducedSystemOperator& op_, const string& name_)
: PetscPreconditionerFactory(name_) /* , op(op_) */ {}
virtual mfem::Solver* NewPreconditioner(const mfem::OperatorHandle&);
virtual ~PreconditionerFactory() {}
};
/** Function representing the elastic energy density for the given hyperelastic
model+deformation. Used in HyperelasticOperator::GetElasticEnergyDensity. */
@@ -187,6 +206,7 @@ int main(int argc, char *argv[])
int vis_steps = 1;
bool use_petsc = true;
const char *petscrc_file = "";
bool petsc_use_jfnk = false;
OptionsParser args(argc, argv);
args.AddOption(&mesh_file, "-m", "--mesh",
@@ -221,6 +241,9 @@ int main(int argc, char *argv[])
"Use or not PETSc to solve the nonlinear system.");
args.AddOption(&petscrc_file, "-petscopts", "--petscopts",
"PetscOptions file to use.");
args.AddOption(&petsc_use_jfnk, "-jfnk", "--jfnk", "-no-jfnk",
"--no-jfnk",
"Use JFNK with user-defined preconditioner factory.");
args.Parse();
if (!args.Good())
{
@@ -344,7 +367,8 @@ int main(int argc, char *argv[])
// 9. Initialize the hyperelastic operator, the GLVis visualization and print
// the initial energies.
HyperelasticOperator *oper = new HyperelasticOperator(fespace, ess_bdr, visc,
mu, K, use_petsc);
mu, K, use_petsc,
petsc_use_jfnk);
socketstream vis_v, vis_w;
if (visualization)
@@ -520,7 +544,7 @@ Operator &ReducedSystemOperator::GetGradient(const Vector &k) const
add(*v, dt, k, w);
add(*x, dt, w, z);
localJ->Add(dt*dt, H->GetLocalGradient(z));
// if we are using PETSc, the HypreParCSR jacobian will be converted to
// if we are using PETSc, the HypreParCSR Jacobian will be converted to
// PETSc's AIJ on the fly
Jacobian = M->ParallelAssemble(localJ);
delete localJ;
@@ -537,7 +561,8 @@ ReducedSystemOperator::~ReducedSystemOperator()
HyperelasticOperator::HyperelasticOperator(ParFiniteElementSpace &f,
Array<int> &ess_bdr, double visc,
double mu, double K, bool use_petsc)
double mu, double K, bool use_petsc,
bool use_petsc_factory)
: TimeDependentOperator(2*f.TrueVSize(), 0.0), fespace(f),
M(&fespace), S(&fespace), H(&fespace),
viscosity(visc), M_solver(f.GetComm()),
@@ -590,6 +615,8 @@ HyperelasticOperator::HyperelasticOperator(ParFiniteElementSpace &f,
J_minres->SetPreconditioner(*J_prec);
J_solver = J_minres;
J_factory = NULL;
newton_solver.iterative_mode = false;
newton_solver.SetSolver(*J_solver);
newton_solver.SetOperator(*reduced_oper);
@@ -600,12 +627,20 @@ HyperelasticOperator::HyperelasticOperator(ParFiniteElementSpace &f,
}
else
{
// if using PETSc, we create the same solver (NEWTON+MINRES+Jacobi)
// if using PETSc, we create the same solver (Newton + MINRES + Jacobi)
// by command line options (see rc_ex10p)
J_solver = NULL;
J_prec = NULL;
J_factory = NULL;
pnewton_solver = new PetscNonlinearSolver(f.GetComm(),
*reduced_oper);
// we can setup a factory to construct a "physics-based" preconditioner
if (use_petsc_factory)
{
J_factory = new PreconditionerFactory(*reduced_oper, "JFNK preconditioner");
pnewton_solver->SetPreconditionerFactory(J_factory);
}
pnewton_solver->SetPrintLevel(1); // print Newton iterations
pnewton_solver->SetRelTol(rel_tol);
pnewton_solver->SetAbsTol(0.0);
@@ -691,12 +726,26 @@ HyperelasticOperator::~HyperelasticOperator()
{
delete J_solver;
delete J_prec;
delete J_factory;
delete reduced_oper;
delete model;
delete Mmat;
delete pnewton_solver;
}
// This method gets called every time we need a preconditioner "oh"
// contains the PetscParMatrix that wraps the operator constructed in
// the GetGradient() method (see also PetscSolver::SetJacobianType()).
// In this example, we just return a customizable PetscPreconditioner
// using that matrix. However, the OperatorHandle argument can be
// ignored, and any "physics-based" solver can be constructed since we
// have access to the HyperElasticOperator class.
Solver* PreconditionerFactory::NewPreconditioner(const mfem::OperatorHandle& oh)
{
PetscParMatrix *pP;
oh.Get(pP);
return new PetscPreconditioner(*pP,"jfnk_");
}
double ElasticEnergyCoefficient::Eval(ElementTransformation &T,
const IntegrationPoint &ip)
@@ -710,8 +759,8 @@ double ElasticEnergyCoefficient::Eval(ElementTransformation &T,
void InitialDeformation(const Vector &x, Vector &y)
{
// set the initial configuration to be the same as the reference, stress
// free, configuration
// set the initial configuration to be the same as the reference,
// stress free, configuration
y = x;
}
+7
View File
@@ -84,6 +84,10 @@ EX9_E_ARGS := -m ../../data/periodic-hexagon.mesh --usepetsc --petscopts r
EX9_ES_ARGS := -m ../../data/periodic-hexagon.mesh --usepetsc --petscopts rc_ex9p_expl --no-step
EX9_IS_ARGS := -m ../../data/periodic-hexagon.mesh --usepetsc --petscopts rc_ex9p_impl --implicit -tf 0.5
EX10_ARGS := -m ../../data/beam-quad.mesh --usepetsc --petscopts rc_ex10p -tf 30 -s 3 -rs 2 -dt 3
EX10_MF_ARGS := -m ../../data/beam-quad.mesh --usepetsc --petscopts rc_ex10p_mf -tf 6 -s 3 -rs 0 -dt 3
EX10_MFOP_ARGS := -m ../../data/beam-quad.mesh --usepetsc --petscopts rc_ex10p_mfop -tf 6 -s 3 -rs 0 -dt 3
EX10_JFNK_ARGS := -m ../../data/beam-quad.mesh --usepetsc --petscopts rc_ex10p_jfnk --jfnk -tf 6 -s 3 -rs 0 -dt 3
ex1p-test-par: ex1p
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX1_ARGS_W))
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX1_ARGS_P))
@@ -107,6 +111,9 @@ ex9p-test-par: ex9p
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX9_IS_ARGS))
ex10p-test-par: ex10p
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX10_ARGS))
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX10_MF_ARGS))
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX10_MFOP_ARGS))
@$(call mfem-test,$<, $(RUN_MPI), $(TESTNAME),$(EX10_JFNK_ARGS))
# Testing: "test" target and mfem-test* variables are defined in config/test.mk
+5
View File
@@ -0,0 +1,5 @@
# matrix-free Jacobian action, preconditioner constructed using PetscPreconditionerFactory
-snes_monitor
-snes_mf_operator
-ksp_type minres
-jfnk_pc_type jacobi
+4
View File
@@ -0,0 +1,4 @@
# matrix free -> no preconditioner
-snes_monitor
-snes_mf
-ksp_type minres
+5
View File
@@ -0,0 +1,5 @@
# matrix-free Jacobian action, preconditioner constructed from the matrix obtained by the GetGradient() method
-snes_monitor
-snes_mf_operator
-ksp_type minres
-pc_type jacobi
+4 -4
View File
@@ -42,12 +42,12 @@ add_mfem_examples(SUNDIALS_EXAMPLES_SRCS ${PFX} "" test_sundials)
# ctest -R sundials
# Command line options for the tests.
# Example 9: test explicit CVODE time stepping
set(EX9_COMMON_OPTS -m ../../data/periodic-hexagon.mesh -p 0 -s 11)
# Example 9: test CVODE with CV_ADAMS (non-stiff implicit) time stepping
set(EX9_COMMON_OPTS -m ../../data/periodic-hexagon.mesh -p 0 -s 7)
set(EX9_TEST_OPTS ${EX9_COMMON_OPTS} -r 2 -dt 0.0018 -vs 25)
set(EX9P_TEST_OPTS ${EX9_COMMON_OPTS} -rp 1 -dt 0.0009 -vs 50)
# Example 10: test implicit CVODE time stepping
set(EX10_COMMON_OPTS -m ../../data/beam-quad.mesh -o 2 -s 5 -dt 0.15 -vs 10)
# Example 10: test CVODE with CV_BDF (stiff implicit) time stepping
set(EX10_COMMON_OPTS -m ../../data/beam-quad.mesh -o 2 -s 5 -dt 0.15 -tf 6 -vs 10)
set(EX10_TEST_OPTS ${EX10_COMMON_OPTS} -r 2)
set(EX10P_TEST_OPTS ${EX10_COMMON_OPTS} -rp 1)
# Example 16: use the default options
+204 -210
View File
@@ -4,16 +4,16 @@
// Compile with: make ex10
//
// Sample runs:
// ex10 -m ../../data/beam-quad.mesh -r 2 -o 2 -s 5 -dt 0.15 -vs 10
// ex10 -m ../../data/beam-tri.mesh -r 2 -o 2 -s 7 -dt 0.3 -vs 5
// ex10 -m ../../data/beam-hex.mesh -r 1 -o 2 -s 5 -dt 0.2 -vs 5
// ex10 -m ../../data/beam-quad.mesh -r 2 -o 2 -s 12 -dt 0.15 -vs 10
// ex10 -m ../../data/beam-tri.mesh -r 2 -o 2 -s 16 -dt 0.3 -vs 5
// ex10 -m ../../data/beam-hex.mesh -r 1 -o 2 -s 12 -dt 0.2 -vs 5
// ex10 -m ../../data/beam-tri.mesh -r 2 -o 2 -s 2 -dt 3 -nls kinsol
// ex10 -m ../../data/beam-quad.mesh -r 2 -o 2 -s 2 -dt 3 -nls kinsol
// ex10 -m ../../data/beam-hex.mesh -r 1 -o 2 -s 2 -dt 3 -nls kinsol
// ex10 -m ../../data/beam-quad.mesh -r 2 -o 2 -s 15 -dt 5e-3 -vs 60
// ex10 -m ../../data/beam-tri.mesh -r 2 -o 2 -s 16 -dt 0.01 -vs 30
// ex10 -m ../../data/beam-hex.mesh -r 1 -o 2 -s 15 -dt 0.01 -vs 30
// ex10 -m ../../data/beam-quad-amr.mesh -r 2 -o 2 -s 5 -dt 0.15 -vs 10
// ex10 -m ../../data/beam-quad.mesh -r 2 -o 2 -s 14 -dt 0.15 -vs 10
// ex10 -m ../../data/beam-tri.mesh -r 2 -o 2 -s 17 -dt 0.01 -vs 30
// ex10 -m ../../data/beam-hex.mesh -r 1 -o 2 -s 14 -dt 0.15 -vs 10
// ex10 -m ../../data/beam-quad-amr.mesh -r 2 -o 2 -s 12 -dt 0.15 -vs 10
//
// Description: This examples solves a time dependent nonlinear elasticity
// problem of the form dv/dt = H(x) + S v, dx/dt = v, where H is a
@@ -53,7 +53,6 @@ using namespace std;
using namespace mfem;
class ReducedSystemOperator;
class SundialsJacSolver;
/** After spatial discretization, the hyperelastic model can be written as a
* system of ODEs:
@@ -92,12 +91,17 @@ protected:
mutable Vector z; // auxiliary vector
SparseMatrix *grad_H;
SparseMatrix *Jacobian;
double saved_gamma; // saved gamma value from implicit setup
public:
/// Solver type to use in the ImplicitSolve() method, used by SDIRK methods.
enum NonlinearSolverType
{
NEWTON = 0, ///< Use MFEM's plain NewtonSolver
KINSOL = 1 ///< Use SUNDIALS' KINSOL (through MFEM's class KinSolver)
KINSOL = 1 ///< Use SUNDIALS' KINSOL (through MFEM's class KINSolver)
};
HyperelasticOperator(FiniteElementSpace &f, Array<int> &ess_bdr,
@@ -106,15 +110,41 @@ public:
/// Compute the right-hand side of the ODE system.
virtual void Mult(const Vector &vx, Vector &dvx_dt) const;
/** Solve the Backward-Euler equation: k = f(x + dt*k, t), for the unknown k.
This is the only requirement for high-order SDIRK implicit integration.*/
virtual void ImplicitSolve(const double dt, const Vector &x, Vector &k);
/** Connect the Jacobian linear system solver (SundialsJacSolver) used by
SUNDIALS' CVODE and ARKODE time integrators to the internal objects
created by HyperelasticOperator. This method is called by the InitSystem
method of SundialsJacSolver. */
void InitSundialsJacSolver(SundialsJacSolver &sjsolv);
/// Custom Jacobian system solver for the SUNDIALS time integrators.
/** For the ODE system represented by HyperelasticOperator
M dv/dt = -(H(x) + S*v)
dx/dt = v,
this class facilitates the solution of linear systems of the form
(M + γS) yv + γJ yx = M bv, J=(dH/dx)(x)
- γ yv + yx = bx
for given bv, bx, x, and γ = GetTimeStep(). */
/** Linear solve applicable to the SUNDIALS format.
Solves (Mass - dt J) y = Mass b, where in our case:
Mass = | M 0 | J = | -S -grad_H | y = | v_hat | b = | b_v |
| 0 I | | I 0 | | x_hat | | b_x |
The result replaces the rhs b.
We substitute x_hat = b_x + dt v_hat and solve
(M + dt S + dt^2 grad_H) v_hat = M b_v - dt grad_H b_x. */
/** Setup the linear system. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSetup(const Vector &y, const Vector &fy,
int jok, int *jcur, double gamma);
/** Solve the linear system. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSolve(const Vector &b, Vector &x, double tol);
double ElasticEnergy(const Vector &x) const;
double KineticEnergy(const Vector &v) const;
@@ -152,53 +182,6 @@ public:
virtual ~ReducedSystemOperator();
};
/// Custom Jacobian system solver for the SUNDIALS time integrators.
/** For the ODE system represented by HyperelasticOperator
M dv/dt = -(H(x) + S*v)
dx/dt = v,
this class facilitates the solution of linear systems of the form
(M + γS) yv + γJ yx = M bv, J=(dH/dx)(x)
- γ yv + yx = bx
for given bv, bx, x, and γ = GetTimeStep(). */
class SundialsJacSolver : public SundialsODELinearSolver
{
private:
BilinearForm *M, *S;
NonlinearForm *H;
SparseMatrix *grad_H, *Jacobian;
Solver *J_solver;
public:
SundialsJacSolver()
: M(), S(), H(), grad_H(), Jacobian(), J_solver() { }
/// Connect the solver to the objects created inside HyperelasticOperator.
void SetOperators(BilinearForm &M_, BilinearForm &S_,
NonlinearForm &H_, Solver &solver)
{
M = &M_; S = &S_; H = &H_; J_solver = &solver;
}
/** Linear solve applicable to the SUNDIALS format.
Solves (Mass - dt J) y = Mass b, where in our case:
Mass = | M 0 | J = | -S -grad_H | y = | v_hat | b = | b_v |
| 0 I | | I 0 | | x_hat | | b_x |
The result replaces the rhs b.
We substitute x_hat = b_x + dt v_hat and solve
(M + dt S + dt^2 grad_H) v_hat = M b_v - dt grad_H b_x. */
int InitSystem(void *sundials_mem);
int SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred, int &jac_cur,
Vector &v_temp1, Vector &v_temp2, Vector &v_temp3);
int SolveSystem(void *sundials_mem, Vector &b, const Vector &weight,
const Vector &y_cur, const Vector &f_cur);
int FreeSystem(void *sundials_mem);
};
/** Function representing the elastic energy density for the given hyperelastic
model+deformation. Used in HyperelasticOperator::GetElasticEnergyDensity. */
@@ -243,6 +226,12 @@ int main(int argc, char *argv[])
// Relative and absolute tolerances for CVODE and ARKODE.
const double reltol = 1e-1, abstol = 1e-1;
// Since this example uses the loose tolerances defined above, it is
// necessary to lower the linear solver tolerance for CVODE which is relative
// to the above tolerances.
const double cvode_eps_lin = 1e-4;
// Similarly, the nonlinear tolerance for ARKODE needs to be tightened.
const double arkode_eps_nonlin = 1e-6;
OptionsParser args(argc, argv);
args.AddOption(&mesh_file, "-m", "--mesh",
@@ -252,15 +241,24 @@ int main(int argc, char *argv[])
args.AddOption(&order, "-o", "--order",
"Order (degree) of the finite elements.");
args.AddOption(&ode_solver_type, "-s", "--ode-solver",
"ODE solver: 1 - Backward Euler, 2 - SDIRK2, 3 - SDIRK3,\n\t"
" 4 - CVODE implicit, approximate Jacobian,\n\t"
" 5 - CVODE implicit, specified Jacobian,\n\t"
" 6 - ARKODE implicit, approximate Jacobian,\n\t"
" 7 - ARKODE implicit, specified Jacobian,\n\t"
" 11 - Forward Euler, 12 - RK2,\n\t"
" 13 - RK3 SSP, 14 - RK4,\n\t"
" 15 - CVODE (adaptive order) explicit,\n\t"
" 16 - ARKODE default (4th order) explicit.");
"ODE solver:\n\t"
"1 - Backward Euler,\n\t"
"2 - SDIRK2, L-stable\n\t"
"3 - SDIRK3, L-stable\n\t"
"4 - Implicit Midpoint,\n\t"
"5 - SDIRK2, A-stable,\n\t"
"6 - SDIRK3, A-stable,\n\t"
"7 - Forward Euler,\n\t"
"8 - RK2,\n\t"
"9 - RK3 SSP,\n\t"
"10 - RK4,\n\t"
"11 - CVODE implicit BDF, approximate Jacobian,\n\t"
"12 - CVODE implicit BDF, specified Jacobian,\n\t"
"13 - CVODE implicit ADAMS, approximate Jacobian,\n\t"
"14 - CVODE implicit ADAMS, specified Jacobian,\n\t"
"15 - ARKODE implicit, approximate Jacobian,\n\t"
"16 - ARKODE implicit, specified Jacobian,\n\t"
"17 - ARKODE explicit, 4th order.");
args.AddOption(&nls, "-nls", "--nonlinear-solver",
"Nonlinear systems solver: "
"\"newton\" (plain Newton) or \"kinsol\" (KINSOL).");
@@ -287,72 +285,19 @@ int main(int argc, char *argv[])
}
args.PrintOptions(cout);
// check for vaild ODE solver option
if (ode_solver_type < 1 || ode_solver_type > 17)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
return 1;
}
// 2. Read the mesh from the given mesh file. We can handle triangular,
// quadrilateral, tetrahedral and hexahedral meshes with the same code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
// 3. Define the ODE solver used for time integration. Several implicit
// singly diagonal implicit Runge-Kutta (SDIRK) methods, as well as
// explicit Runge-Kutta methods are available.
ODESolver *ode_solver;
CVODESolver *cvode = NULL;
ARKODESolver *arkode = NULL;
SundialsJacSolver *sjsolver = NULL;
switch (ode_solver_type)
{
// Implicit L-stable methods
case 1: ode_solver = new BackwardEulerSolver; break;
case 2: ode_solver = new SDIRK23Solver(2); break;
case 3: ode_solver = new SDIRK33Solver; break;
case 4:
case 5:
cvode = new CVODESolver(CV_BDF, CV_NEWTON);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
if (ode_solver_type == 5)
{
sjsolver = new SundialsJacSolver;
cvode->SetLinearSolver(*sjsolver);
}
ode_solver = cvode; break;
case 6:
case 7:
arkode = new ARKODESolver(ARKODESolver::IMPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 7)
{
// Custom Jacobian inversion.
sjsolver = new SundialsJacSolver;
arkode->SetLinearSolver(*sjsolver);
}
ode_solver = arkode; break;
// Explicit methods
case 11: ode_solver = new ForwardEulerSolver; break;
case 12: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 13: ode_solver = new RK3SSPSolver; break;
case 14: ode_solver = new RK4Solver; break;
case 15:
cvode = new CVODESolver(CV_ADAMS, CV_FUNCTIONAL);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 16:
arkode = new ARKODESolver(ARKODESolver::IMPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
// Implicit A-stable methods (not L-stable)
case 22: ode_solver = new ImplicitMidpointSolver; break;
case 23: ode_solver = new SDIRK23Solver; break;
case 24: ode_solver = new SDIRK34Solver; break;
default:
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
delete mesh;
return 3;
}
// 3. Setup the nonlinear solver
map<string,HyperelasticOperator::NonlinearSolverType> nls_map;
nls_map["newton"] = HyperelasticOperator::NEWTON;
nls_map["kinsol"] = HyperelasticOperator::KINSOL;
@@ -439,11 +384,82 @@ int main(int argc, char *argv[])
cout << "initial kinetic energy (KE) = " << ke0 << endl;
cout << "initial total energy (TE) = " << (ee0 + ke0) << endl;
// 8. Define the ODE solver used for time integration. Several implicit
// singly diagonal implicit Runge-Kutta (SDIRK) methods, as well as
// explicit Runge-Kutta methods are available.
double t = 0.0;
oper.SetTime(t);
ode_solver->Init(oper);
// 8. Perform time-integration (looping over the time iterations, ti, with a
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKStepSolver *arkode = NULL;
switch (ode_solver_type)
{
// Implicit L-stable methods
case 1: ode_solver = new BackwardEulerSolver; break;
case 2: ode_solver = new SDIRK23Solver(2); break;
case 3: ode_solver = new SDIRK33Solver; break;
// Implicit A-stable methods (not L-stable)
case 4: ode_solver = new ImplicitMidpointSolver; break;
case 5: ode_solver = new SDIRK23Solver; break;
case 6: ode_solver = new SDIRK34Solver; break;
// Explicit methods
case 7: ode_solver = new ForwardEulerSolver; break;
case 8: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 9: ode_solver = new RK3SSPSolver; break;
case 10: ode_solver = new RK4Solver; break;
// CVODE BDF
case 11:
case 12:
cvode = new CVODESolver(CV_BDF);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
CVodeSetEpsLin(cvode->GetMem(), cvode_eps_lin);
cvode->SetMaxStep(dt);
if (ode_solver_type == 11)
{
cvode->UseSundialsLinearSolver();
}
ode_solver = cvode; break;
// CVODE Adams
case 13:
case 14:
cvode = new CVODESolver(CV_ADAMS);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
CVodeSetEpsLin(cvode->GetMem(), cvode_eps_lin);
cvode->SetMaxStep(dt);
if (ode_solver_type == 13)
{
cvode->UseSundialsLinearSolver();
}
ode_solver = cvode; break;
// ARKStep Implicit methods
case 15:
case 16:
arkode = new ARKStepSolver(ARKStepSolver::IMPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
ARKStepSetNonlinConvCoef(arkode->GetMem(), arkode_eps_nonlin);
arkode->SetMaxStep(dt);
if (ode_solver_type == 15)
{
arkode->UseSundialsLinearSolver();
}
ode_solver = arkode; break;
// ARKStep Explicit methods
case 17:
arkode = new ARKStepSolver(ARKStepSolver::EXPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
}
// Initialize MFEM integrators, SUNDIALS integrators are initialized above
if (ode_solver_type < 11) { ode_solver->Init(oper); }
// 9. Perform time-integration (looping over the time iterations, ti, with a
// time-step dt).
bool last_step = false;
for (int ti = 1; !last_step; ti++)
@@ -478,7 +494,7 @@ int main(int argc, char *argv[])
}
}
// 9. Save the displaced mesh, the velocity and elastic energy.
// 10. Save the displaced mesh, the velocity and elastic energy.
{
v.SetFromTrueVector(); x.SetFromTrueVector();
GridFunction *nodes = &x;
@@ -497,9 +513,8 @@ int main(int argc, char *argv[])
w.Save(ee_ofs);
}
// 10. Free the used memory.
// 11. Free the used memory.
delete ode_solver;
delete sjsolver;
delete mesh;
return 0;
@@ -579,81 +594,14 @@ ReducedSystemOperator::~ReducedSystemOperator()
}
int SundialsJacSolver::InitSystem(void *sundials_mem)
{
TimeDependentOperator *td_oper = GetTimeDependentOperator(sundials_mem);
HyperelasticOperator *he_oper;
// During development, we use dynamic_cast<> to ensure the setup is correct:
he_oper = dynamic_cast<HyperelasticOperator*>(td_oper);
MFEM_VERIFY(he_oper, "operator is not HyperelasticOperator");
// When the implementation is finalized, we can switch to static_cast<>:
// he_oper = static_cast<HyperelasticOperator*>(td_oper);
he_oper->InitSundialsJacSolver(*this);
return 0;
}
int SundialsJacSolver::SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred,
int &jac_cur, Vector &v_temp1,
Vector &v_temp2, Vector &v_temp3)
{
int sc = y_pred.Size() / 2;
const Vector x(y_pred.GetData() + sc, sc);
double dt = GetTimeStep(sundials_mem);
// J = M + dt*(S + dt*grad(H))
delete Jacobian;
Jacobian = Add(1.0, M->SpMat(), dt, S->SpMat());
grad_H = dynamic_cast<SparseMatrix *>(&H->GetGradient(x));
Jacobian->Add(dt * dt, *grad_H);
J_solver->SetOperator(*Jacobian);
jac_cur = 1;
return 0;
}
int SundialsJacSolver::SolveSystem(void *sundials_mem, Vector &b,
const Vector &weight, const Vector &y_cur,
const Vector &f_cur)
{
int sc = b.Size() / 2;
// Vector x(y_cur.GetData() + sc, sc);
Vector b_v(b.GetData() + 0, sc);
Vector b_x(b.GetData() + sc, sc);
Vector rhs(sc);
double dt = GetTimeStep(sundials_mem);
// rhs = M b_v - dt*grad(H) b_x
grad_H->Mult(b_x, rhs);
rhs *= -dt;
M->AddMult(b_v, rhs);
J_solver->iterative_mode = false;
J_solver->Mult(rhs, b_v);
b_x.Add(dt, b_v);
return 0;
}
int SundialsJacSolver::FreeSystem(void *sundials_mem)
{
delete Jacobian;
return 0;
}
HyperelasticOperator::HyperelasticOperator(FiniteElementSpace &f,
Array<int> &ess_bdr, double visc,
double mu, double K,
NonlinearSolverType nls_type)
: TimeDependentOperator(2*f.GetTrueVSize(), 0.0), fespace(f),
M(&fespace), S(&fespace), H(&fespace),
viscosity(visc), z(height/2)
viscosity(visc), z(height/2),
grad_H(NULL), Jacobian(NULL)
{
const double rel_tol = 1e-8;
const int skip_zero_entries = 0;
@@ -702,23 +650,24 @@ HyperelasticOperator::HyperelasticOperator(FiniteElementSpace &f,
if (nls_type == KINSOL)
{
KinSolver *kinsolver = new KinSolver(KIN_NONE, true);
kinsolver->SetMaxSetupCalls(4);
KINSolver *kinsolver = new KINSolver(KIN_NONE, true);
newton_solver = kinsolver;
newton_solver->SetOperator(*reduced_oper);
newton_solver->SetMaxIter(200);
newton_solver->SetRelTol(rel_tol);
newton_solver->SetPrintLevel(0);
kinsolver->SetMaxSetupCalls(4);
}
else
{
newton_solver = new NewtonSolver();
newton_solver->SetOperator(*reduced_oper);
newton_solver->SetMaxIter(10);
newton_solver->SetRelTol(rel_tol);
newton_solver->SetPrintLevel(-1);
}
newton_solver->SetSolver(*J_solver);
newton_solver->iterative_mode = false;
newton_solver->SetOperator(*reduced_oper);
}
void HyperelasticOperator::Mult(const Vector &vx, Vector &dvx_dt) const
@@ -768,9 +717,53 @@ void HyperelasticOperator::ImplicitSolve(const double dt,
add(v, dt, dv_dt, dx_dt);
}
void HyperelasticOperator::InitSundialsJacSolver(SundialsJacSolver &sjsolv)
int HyperelasticOperator::SUNImplicitSetup(const Vector &y,
const Vector &fy, int jok, int *jcur,
double gamma)
{
sjsolv.SetOperators(M, S, H, *J_solver);
int sc = y.Size() / 2;
const Vector x(y.GetData() + sc, sc);
// J = M + dt*(S + dt*grad(H))
if (Jacobian) { delete Jacobian; }
Jacobian = Add(1.0, M.SpMat(), gamma, S.SpMat());
grad_H = dynamic_cast<SparseMatrix *>(&H.GetGradient(x));
Jacobian->Add(gamma * gamma, *grad_H);
// Set Jacobian solve operator
J_solver->SetOperator(*Jacobian);
// Indicate that the Jacobian was updated
*jcur = 1;
// Save gamma for use in solve
saved_gamma = gamma;
// Return success
return 0;
}
int HyperelasticOperator::SUNImplicitSolve(const Vector &b, Vector &x,
double tol)
{
int sc = b.Size() / 2;
Vector b_v(b.GetData() + 0, sc);
Vector b_x(b.GetData() + sc, sc);
Vector x_v(x.GetData() + 0, sc);
Vector x_x(x.GetData() + sc, sc);
Vector rhs(sc);
// rhs = M b_v - dt*grad(H) b_x
grad_H->Mult(b_x, rhs);
rhs *= -saved_gamma;
M.AddMult(b_v, rhs);
J_solver->iterative_mode = false;
J_solver->Mult(rhs, x_v);
add(b_x, saved_gamma, x_v, x_x);
return 0;
}
double HyperelasticOperator::ElasticEnergy(const Vector &x) const
@@ -792,6 +785,7 @@ void HyperelasticOperator::GetElasticEnergyDensity(
HyperelasticOperator::~HyperelasticOperator()
{
delete Jacobian;
delete newton_solver;
delete J_solver;
delete J_prec;
+219 -229
View File
@@ -4,16 +4,16 @@
// Compile with: make ex10p
//
// Sample runs:
// mpirun -np 4 ex10p -m ../../data/beam-quad.mesh -rp 1 -o 2 -s 5 -dt 0.15 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-tri.mesh -rp 1 -o 2 -s 7 -dt 0.25 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-hex.mesh -rp 0 -o 2 -s 5 -dt 0.15 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-quad.mesh -rp 1 -o 2 -s 12 -dt 0.15 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-tri.mesh -rp 1 -o 2 -s 16 -dt 0.25 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-hex.mesh -rp 0 -o 2 -s 12 -dt 0.15 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-tri.mesh -rp 1 -o 2 -s 2 -dt 3 -nls kinsol
// mpirun -np 4 ex10p -m ../../data/beam-quad.mesh -rp 1 -o 2 -s 2 -dt 3 -nls kinsol
// mpirun -np 4 ex10p -m ../../data/beam-hex.mesh -rs 1 -o 2 -s 2 -dt 3 -nls kinsol
// mpirun -np 4 ex10p -m ../../data/beam-quad.mesh -rp 1 -o 2 -s 15 -dt 3e-3 -vs 120
// mpirun -np 4 ex10p -m ../../data/beam-tri.mesh -rp 1 -o 2 -s 16 -dt 5e-3 -vs 60
// mpirun -np 4 ex10p -m ../../data/beam-hex.mesh -rp 0 -o 2 -s 15 -dt 5e-3 -vs 60
// mpirun -np 4 ex10p -m ../../data/beam-quad-amr.mesh -rp 1 -o 2 -s 5 -dt 0.15 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-quad.mesh -rp 1 -o 2 -s 14 -dt 0.15 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-tri.mesh -rp 1 -o 2 -s 17 -dt 5e-3 -vs 60
// mpirun -np 4 ex10p -m ../../data/beam-hex.mesh -rp 0 -o 2 -s 14 -dt 0.15 -vs 10
// mpirun -np 4 ex10p -m ../../data/beam-quad-amr.mesh -rp 1 -o 2 -s 12 -dt 0.15 -vs 10
//
// Description: This examples solves a time dependent nonlinear elasticity
// problem of the form dv/dt = H(x) + S v, dx/dt = v, where H is a
@@ -53,7 +53,6 @@ using namespace std;
using namespace mfem;
class ReducedSystemOperator;
class SundialsJacSolver;
/** After spatial discretization, the hyperelastic model can be written as a
* system of ODEs:
@@ -94,12 +93,17 @@ protected:
mutable Vector z; // auxiliary vector
const SparseMatrix *local_grad_H;
HypreParMatrix *Jacobian;
double saved_gamma; // saved gamma value from implicit setup
public:
/// Solver type to use in the ImplicitSolve() method, used by SDIRK methods.
enum NonlinearSolverType
{
NEWTON = 0, ///< Use MFEM's plain NewtonSolver
KINSOL = 1 ///< Use SUNDIALS' KINSOL (through MFEM's class KinSolver)
KINSOL = 1 ///< Use SUNDIALS' KINSOL (through MFEM's class KINSolver)
};
HyperelasticOperator(ParFiniteElementSpace &f, Array<int> &ess_bdr,
@@ -108,15 +112,41 @@ public:
/// Compute the right-hand side of the ODE system.
virtual void Mult(const Vector &vx, Vector &dvx_dt) const;
/** Solve the Backward-Euler equation: k = f(x + dt*k, t), for the unknown k.
This is the only requirement for high-order SDIRK implicit integration.*/
virtual void ImplicitSolve(const double dt, const Vector &x, Vector &k);
/** Connect the Jacobian linear system solver (SundialsJacSolver) used by
SUNDIALS' CVODE and ARKODE time integrators to the internal objects
created by HyperelasticOperator. This method is called by the InitSystem
method of SundialsJacSolver. */
void InitSundialsJacSolver(SundialsJacSolver &sjsolv);
/// Custom Jacobian system solver for the SUNDIALS time integrators.
/** For the ODE system represented by HyperelasticOperator
M dv/dt = -(H(x) + S*v)
dx/dt = v,
this class facilitates the solution of linear systems of the form
(M + γS) yv + γJ yx = M bv, J=(dH/dx)(x)
- γ yv + yx = bx
for given bv, bx, x, and γ = GetTimeStep(). */
/** Linear solve applicable to the SUNDIALS format.
Solves (Mass - dt J) y = Mass b, where in our case:
Mass = | M 0 | J = | -S -grad_H | y = | v_hat | b = | b_v |
| 0 I | | I 0 | | x_hat | | b_x |
The result replaces the rhs b.
We substitute x_hat = b_x + dt v_hat and solve
(M + dt S + dt^2 grad_H) v_hat = M b_v - dt grad_H b_x. */
/** Setup the linear system. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSetup(const Vector &y, const Vector &fy,
int jok, int *jcur, double gamma);
/** Solve the linear system. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSolve(const Vector &b, Vector &x, double tol);
double ElasticEnergy(const ParGridFunction &x) const;
double KineticEnergy(const ParGridFunction &v) const;
@@ -157,57 +187,6 @@ public:
virtual ~ReducedSystemOperator();
};
/// Custom Jacobian system solver for the SUNDIALS time integrators.
/** For the ODE system represented by HyperelasticOperator
M dv/dt = -(H(x) + S*v)
dx/dt = v,
this class facilitates the solution of linear systems of the form
(M + γS) yv + γJ yx = M bv, J=(dH/dx)(x)
- γ yv + yx = bx
for given bv, bx, x, and γ = GetTimeStep(). */
class SundialsJacSolver : public SundialsODELinearSolver
{
private:
ParBilinearForm *M, *S;
ParNonlinearForm *H;
const SparseMatrix *local_grad_H;
HypreParMatrix *Jacobian;
Solver *J_solver;
const Array<int> *ess_tdof_list;
public:
SundialsJacSolver()
: M(), S(), H(), local_grad_H(), Jacobian(), J_solver() { }
/// Connect the solver to the objects created inside HyperelasticOperator.
void SetOperators(ParBilinearForm &M_, ParBilinearForm &S_,
ParNonlinearForm &H_, Solver &solver,
const Array<int> &ess_tdof_list_)
{
M = &M_; S = &S_; H = &H_; J_solver = &solver;
ess_tdof_list = &ess_tdof_list_;
}
/** Linear solve applicable to the SUNDIALS format.
Solves (Mass - dt J) y = Mass b, where in our case:
Mass = | M 0 | J = | -S -grad_H | y = | v_hat | b = | b_v |
| 0 I | | I 0 | | x_hat | | b_x |
The result replaces the rhs b.
We substitute x_hat = b_x + dt v_hat and solve
(M + dt S + dt^2 grad_H) v_hat = M b_v - dt grad_H b_x. */
int InitSystem(void *sundials_mem);
int SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred, int &jac_cur,
Vector &v_temp1, Vector &v_temp2, Vector &v_temp3);
int SolveSystem(void *sundials_mem, Vector &b, const Vector &weight,
const Vector &y_cur, const Vector &f_cur);
int FreeSystem(void *sundials_mem);
};
/** Function representing the elastic energy density for the given hyperelastic
model+deformation. Used in HyperelasticOperator::GetElasticEnergyDensity. */
@@ -259,6 +238,12 @@ int main(int argc, char *argv[])
// Relative and absolute tolerances for CVODE and ARKODE.
const double reltol = 1e-1, abstol = 1e-1;
// Since this example uses the loose tolerances defined above, it is
// necessary to lower the linear solver tolerance for CVODE which is relative
// to the above tolerances.
const double cvode_eps_lin = 1e-4;
// Similarly, the nonlinear tolerance for ARKODE needs to be tightened.
const double arkode_eps_nonlin = 1e-6;
OptionsParser args(argc, argv);
args.AddOption(&mesh_file, "-m", "--mesh",
@@ -270,15 +255,24 @@ int main(int argc, char *argv[])
args.AddOption(&order, "-o", "--order",
"Order (degree) of the finite elements.");
args.AddOption(&ode_solver_type, "-s", "--ode-solver",
"ODE solver: 1 - Backward Euler, 2 - SDIRK2, 3 - SDIRK3,\n\t"
" 4 - CVODE implicit, approximate Jacobian,\n\t"
" 5 - CVODE implicit, specified Jacobian,\n\t"
" 6 - ARKODE implicit, approximate Jacobian,\n\t"
" 7 - ARKODE implicit, specified Jacobian,\n\t"
" 11 - Forward Euler, 12 - RK2,\n\t"
" 13 - RK3 SSP, 14 - RK4,\n\t"
" 15 - CVODE (adaptive order) explicit,\n\t"
" 16 - ARKODE default (4th order) explicit.");
"ODE solver:\n\t"
"1 - Backward Euler,\n\t"
"2 - SDIRK2, L-stable\n\t"
"3 - SDIRK3, L-stable\n\t"
"4 - Implicit Midpoint,\n\t"
"5 - SDIRK2, A-stable,\n\t"
"6 - SDIRK3, A-stable,\n\t"
"7 - Forward Euler,\n\t"
"8 - RK2,\n\t"
"9 - RK3 SSP,\n\t"
"10 - RK4,\n\t"
"11 - CVODE implicit BDF, approximate Jacobian,\n\t"
"12 - CVODE implicit BDF, specified Jacobian,\n\t"
"13 - CVODE implicit ADAMS, approximate Jacobian,\n\t"
"14 - CVODE implicit ADAMS, specified Jacobian,\n\t"
"15 - ARKODE implicit, approximate Jacobian,\n\t"
"16 - ARKODE implicit, specified Jacobian,\n\t"
"17 - ARKODE explicit, 4th order.");
args.AddOption(&nls, "-nls", "--nonlinear-solver",
"Nonlinear systems solver: "
"\"newton\" (plain Newton) or \"kinsol\" (KINSOL).");
@@ -312,76 +306,24 @@ int main(int argc, char *argv[])
args.PrintOptions(cout);
}
// check for vaild ODE solver option
if (ode_solver_type < 1 || ode_solver_type > 17)
{
if (myid == 0)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
}
MPI_Finalize();
return 1;
}
// 3. Read the serial mesh from the given mesh file on all processors. We can
// handle triangular, quadrilateral, tetrahedral and hexahedral meshes
// with the same code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
// 4. Define the ODE solver used for time integration. Several implicit
// singly diagonal implicit Runge-Kutta (SDIRK) methods, as well as
// explicit Runge-Kutta methods are available.
ODESolver *ode_solver;
CVODESolver *cvode = NULL;
ARKODESolver *arkode = NULL;
SundialsJacSolver *sjsolver = NULL;
switch (ode_solver_type)
{
// Implicit L-stable methods
case 1: ode_solver = new BackwardEulerSolver; break;
case 2: ode_solver = new SDIRK23Solver(2); break;
case 3: ode_solver = new SDIRK33Solver; break;
case 4:
case 5:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_BDF, CV_NEWTON);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
if (ode_solver_type == 5)
{
sjsolver = new SundialsJacSolver;
cvode->SetLinearSolver(*sjsolver); // Custom Jacobian inversion.
}
ode_solver = cvode; break;
case 6:
case 7:
arkode = new ARKODESolver(MPI_COMM_WORLD, ARKODESolver::IMPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 7)
{
sjsolver = new SundialsJacSolver;
arkode->SetLinearSolver(*sjsolver); // Custom Jacobian inversion.
}
ode_solver = arkode; break;
// Explicit methods
case 11: ode_solver = new ForwardEulerSolver; break;
case 12: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 13: ode_solver = new RK3SSPSolver; break;
case 14: ode_solver = new RK4Solver; break;
case 15:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_ADAMS, CV_FUNCTIONAL);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 16:
arkode = new ARKODESolver(MPI_COMM_WORLD, ARKODESolver::EXPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
// Implicit A-stable methods (not L-stable)
case 22: ode_solver = new ImplicitMidpointSolver; break;
case 23: ode_solver = new SDIRK23Solver; break;
case 24: ode_solver = new SDIRK34Solver; break;
default:
if (myid == 0)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
}
delete mesh;
MPI_Finalize();
return 3;
}
// 4. Nonlinear solver
map<string,HyperelasticOperator::NonlinearSolverType> nls_map;
nls_map["newton"] = HyperelasticOperator::NEWTON;
nls_map["kinsol"] = HyperelasticOperator::KINSOL;
@@ -391,7 +333,6 @@ int main(int argc, char *argv[])
{
cout << "Unknown type of nonlinear solver: " << nls << endl;
}
delete ode_solver;
delete mesh;
MPI_Finalize();
return 4;
@@ -495,11 +436,82 @@ int main(int argc, char *argv[])
cout << "initial total energy (TE) = " << (ee0 + ke0) << endl;
}
// 10. Define the ODE solver used for time integration. Several implicit
// singly diagonal implicit Runge-Kutta (SDIRK) methods, as well as
// explicit Runge-Kutta methods are available.
double t = 0.0;
oper.SetTime(t);
ode_solver->Init(oper);
// 10. Perform time-integration
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKStepSolver *arkode = NULL;
switch (ode_solver_type)
{
// Implicit L-stable methods
case 1: ode_solver = new BackwardEulerSolver; break;
case 2: ode_solver = new SDIRK23Solver(2); break;
case 3: ode_solver = new SDIRK33Solver; break;
// Implicit A-stable methods (not L-stable)
case 4: ode_solver = new ImplicitMidpointSolver; break;
case 5: ode_solver = new SDIRK23Solver; break;
case 6: ode_solver = new SDIRK34Solver; break;
// Explicit methods
case 7: ode_solver = new ForwardEulerSolver; break;
case 8: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 9: ode_solver = new RK3SSPSolver; break;
case 10: ode_solver = new RK4Solver; break;
// CVODE BDF
case 11:
case 12:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_BDF);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
CVodeSetEpsLin(cvode->GetMem(), cvode_eps_lin);
cvode->SetMaxStep(dt);
if (ode_solver_type == 11)
{
cvode->UseSundialsLinearSolver();
}
ode_solver = cvode; break;
// CVODE Adams
case 13:
case 14:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_ADAMS);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
CVodeSetEpsLin(cvode->GetMem(), cvode_eps_lin);
cvode->SetMaxStep(dt);
if (ode_solver_type == 13)
{
cvode->UseSundialsLinearSolver();
}
ode_solver = cvode; break;
// ARKStep Implicit methods
case 15:
case 16:
arkode = new ARKStepSolver(MPI_COMM_WORLD, ARKStepSolver::IMPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
ARKStepSetNonlinConvCoef(arkode->GetMem(), arkode_eps_nonlin);
arkode->SetMaxStep(dt);
if (ode_solver_type == 15)
{
arkode->UseSundialsLinearSolver();
}
ode_solver = arkode; break;
// ARKStep Explicit methods
case 17:
arkode = new ARKStepSolver(MPI_COMM_WORLD, ARKStepSolver::EXPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
}
// Initialize MFEM integrators, SUNDIALS integrators are initialized above
if (ode_solver_type < 11) { ode_solver->Init(oper); }
// 11. Perform time-integration
// (looping over the time iterations, ti, with a time-step dt).
bool last_step = false;
for (int ti = 1; !last_step; ti++)
@@ -538,7 +550,7 @@ int main(int argc, char *argv[])
}
}
// 11. Save the displaced mesh, the velocity and elastic energy.
// 12. Save the displaced mesh, the velocity and elastic energy.
{
v_gf.SetFromTrueVector(); x_gf.SetFromTrueVector();
GridFunction *nodes = &x_gf;
@@ -563,9 +575,8 @@ int main(int argc, char *argv[])
w_gf.Save(ee_ofs);
}
// 12. Free the used memory.
// 13. Free the used memory.
delete ode_solver;
delete sjsolver;
delete pmesh;
MPI_Finalize();
@@ -653,92 +664,14 @@ ReducedSystemOperator::~ReducedSystemOperator()
}
int SundialsJacSolver::InitSystem(void *sundials_mem)
{
TimeDependentOperator *td_oper = GetTimeDependentOperator(sundials_mem);
HyperelasticOperator *he_oper;
// During development, we use dynamic_cast<> to ensure the setup is correct:
he_oper = dynamic_cast<HyperelasticOperator*>(td_oper);
MFEM_VERIFY(he_oper, "operator is not HyperelasticOperator");
// When the implementation is finalized, we can switch to static_cast<>:
// he_oper = static_cast<HyperelasticOperator*>(td_oper);
he_oper->InitSundialsJacSolver(*this);
return 0;
}
int SundialsJacSolver::SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred,
int &jac_cur, Vector &v_temp1,
Vector &v_temp2, Vector &v_temp3)
{
int sc = y_pred.Size() / 2;
const Vector x(y_pred.GetData() + sc, sc);
double dt = GetTimeStep(sundials_mem);
// J = M + dt*(S + dt*grad(H))
delete Jacobian;
SparseMatrix *localJ = Add(1.0, M->SpMat(), dt, S->SpMat());
local_grad_H = &H->GetLocalGradient(x);
localJ->Add(dt*dt, *local_grad_H);
Jacobian = M->ParallelAssemble(localJ);
delete localJ;
HypreParMatrix *Je = Jacobian->EliminateRowsCols(*ess_tdof_list);
delete Je;
J_solver->SetOperator(*Jacobian);
jac_cur = 1;
return 0;
}
int SundialsJacSolver::SolveSystem(void *sundials_mem, Vector &b,
const Vector &weight, const Vector &y_cur,
const Vector &f_cur)
{
int sc = b.Size() / 2;
ParFiniteElementSpace *fes = H->ParFESpace();
// Vector x(y_cur.GetData() + sc, sc);
Vector b_v(b.GetData() + 0, sc);
Vector b_x(b.GetData() + sc, sc);
Vector rhs(sc);
double dt = GetTimeStep(sundials_mem);
// We can assume that b_v and b_x have zeros at essential tdofs.
// rhs = M b_v - dt*grad(H) b_x
ParGridFunction lb_x(fes), lrhs(fes);
lb_x.Distribute(b_x);
local_grad_H->Mult(lb_x, lrhs);
lrhs.ParallelAssemble(rhs);
rhs *= -dt;
M->TrueAddMult(b_v, rhs);
rhs.SetSubVector(*ess_tdof_list, 0.0);
J_solver->iterative_mode = false;
J_solver->Mult(rhs, b_v);
b_x.Add(dt, b_v);
return 0;
}
int SundialsJacSolver::FreeSystem(void *sundials_mem)
{
delete Jacobian;
return 0;
}
HyperelasticOperator::HyperelasticOperator(ParFiniteElementSpace &f,
Array<int> &ess_bdr, double visc,
double mu, double K,
NonlinearSolverType nls_type)
: TimeDependentOperator(2*f.TrueVSize(), 0.0), fespace(f),
M(&fespace), S(&fespace), H(&fespace),
viscosity(visc), M_solver(f.GetComm()), z(height/2)
viscosity(visc), M_solver(f.GetComm()), z(height/2),
local_grad_H(NULL), Jacobian(NULL)
{
const double rel_tol = 1e-8;
const int skip_zero_entries = 0;
@@ -788,23 +721,24 @@ HyperelasticOperator::HyperelasticOperator(ParFiniteElementSpace &f,
if (nls_type == KINSOL)
{
KinSolver *kinsolver = new KinSolver(f.GetComm(), KIN_NONE, true);
kinsolver->SetMaxSetupCalls(4);
KINSolver *kinsolver = new KINSolver(f.GetComm(), KIN_NONE, true);
newton_solver = kinsolver;
newton_solver->SetOperator(*reduced_oper);
newton_solver->SetMaxIter(200);
newton_solver->SetRelTol(rel_tol);
newton_solver->SetPrintLevel(0);
kinsolver->SetMaxSetupCalls(4);
}
else
{
newton_solver = new NewtonSolver(f.GetComm());
newton_solver->SetOperator(*reduced_oper);
newton_solver->SetMaxIter(10);
newton_solver->SetRelTol(rel_tol);
newton_solver->SetPrintLevel(-1);
}
newton_solver->SetSolver(*J_solver);
newton_solver->iterative_mode = false;
newton_solver->SetOperator(*reduced_oper);
}
void HyperelasticOperator::Mult(const Vector &vx, Vector &dvx_dt) const
@@ -858,9 +792,64 @@ void HyperelasticOperator::ImplicitSolve(const double dt,
add(v, dt, dv_dt, dx_dt);
}
void HyperelasticOperator::InitSundialsJacSolver(SundialsJacSolver &sjsolv)
int HyperelasticOperator::SUNImplicitSetup(const Vector &y,
const Vector &fy, int jok, int *jcur,
double gamma)
{
sjsolv.SetOperators(M, S, H, *J_solver, ess_tdof_list);
int sc = y.Size() / 2;
const Vector x(y.GetData() + sc, sc);
// J = M + dt*(S + dt*grad(H))
if (Jacobian) { delete Jacobian; }
SparseMatrix *localJ = Add(1.0, M.SpMat(), gamma, S.SpMat());
local_grad_H = &H.GetLocalGradient(x);
localJ->Add(gamma*gamma, *local_grad_H);
Jacobian = M.ParallelAssemble(localJ);
delete localJ;
HypreParMatrix *Je = Jacobian->EliminateRowsCols(ess_tdof_list);
delete Je;
// Set Jacobian solve operator
J_solver->SetOperator(*Jacobian);
// Indicate that the Jacobian was updated
*jcur = 1;
// Save gamma for use in solve
saved_gamma = gamma;
// Return success
return 0;
}
int HyperelasticOperator::SUNImplicitSolve(const Vector &b, Vector &x,
double tol)
{
int sc = b.Size() / 2;
ParFiniteElementSpace *fes = H.ParFESpace();
Vector b_v(b.GetData() + 0, sc);
Vector b_x(b.GetData() + sc, sc);
Vector x_v(x.GetData() + 0, sc);
Vector x_x(x.GetData() + sc, sc);
Vector rhs(sc);
// We can assume that b_v and b_x have zeros at essential tdofs.
// rhs = M b_v - dt*grad(H) b_x
ParGridFunction lb_x(fes), lrhs(fes);
lb_x.Distribute(b_x);
local_grad_H->Mult(lb_x, lrhs);
lrhs.ParallelAssemble(rhs);
rhs *= -saved_gamma;
M.TrueAddMult(b_v, rhs);
rhs.SetSubVector(ess_tdof_list, 0.0);
J_solver->iterative_mode = false;
J_solver->Mult(rhs, x_v);
add(b_x, saved_gamma, x_v, x_x);
return 0;
}
double HyperelasticOperator::ElasticEnergy(const ParGridFunction &x) const
@@ -886,6 +875,7 @@ void HyperelasticOperator::GetElasticEnergyDensity(
HyperelasticOperator::~HyperelasticOperator()
{
delete Jacobian;
delete newton_solver;
delete J_solver;
delete J_prec;
+124 -165
View File
@@ -7,9 +7,9 @@
// ex16 -m ../../data/inline-tri.mesh
// ex16 -m ../../data/disc-nurbs.mesh -tf 2
// ex16 -s 12 -a 0.0 -k 1.0
// ex16 -s 1 -a 1.0 -k 0.0 -dt 1e-4 -tf 5e-2 -vs 25
// ex16 -s 2 -a 0.5 -k 0.5 -o 4 -dt 1e-4 -tf 2e-2 -vs 25
// ex16 -s 3 -dt 1.0e-4 -tf 4.0e-2 -vs 40
// ex16 -s 8 -a 1.0 -k 0.0 -dt 1e-4 -tf 5e-2 -vs 25
// ex16 -s 9 -a 0.5 -k 0.5 -o 4 -dt 1e-4 -tf 2e-2 -vs 25
// ex16 -s 10 -dt 1.0e-4 -tf 4.0e-2 -vs 40
// ex16 -m ../../data/fichera-q2.mesh
// ex16 -m ../../data/escher.mesh
// ex16 -m ../../data/beam-tet.mesh -tf 10 -dt 0.1
@@ -58,7 +58,6 @@ protected:
SparseMatrix Mmat, Kmat;
SparseMatrix *T; // T = M + dt K
double current_dt;
CGSolver M_solver; // Krylov solver for inverting the mass matrix M
DSmoother M_prec; // Preconditioner for the mass matrix M
@@ -75,13 +74,30 @@ public:
const Vector &u);
virtual void Mult(const Vector &u, Vector &du_dt) const;
/** Solve the Backward-Euler equation: k = f(u + dt*k, t), for the unknown k.
This is the only requirement for high-order SDIRK implicit integration.*/
virtual void ImplicitSolve(const double dt, const Vector &u, Vector &k);
/** Solve the system (M + dt K) y = M b. The result y replaces the input b.
This method is used by the implicit SUNDIALS solvers. */
void SundialsSolve(const double dt, Vector &b);
/// Custom Jacobian system solver for the SUNDIALS time integrators.
/** For the ODE system represented by ConductionOperator
M du/dt = -K(u),
this class facilitates the solution of linear systems of the form
(M + γK) y = M b,
for given b, u (not used), and γ = GetTimeStep(). */
/** Setup the system (M + dt K) x = M b. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSetup(const Vector &x, const Vector &fx,
int jok, int *jcur, double gamma);
/** Solve the system (M + dt K) x = M b. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSolve(const Vector &b, Vector &x, double tol);
/// Update the diffusion BilinearForm K using the given true-dof vector `u`.
void SetParameters(const Vector &u);
@@ -89,33 +105,6 @@ public:
virtual ~ConductionOperator();
};
/// Custom Jacobian system solver for the SUNDIALS time integrators.
/** For the ODE system represented by ConductionOperator
M du/dt = -K(u),
this class facilitates the solution of linear systems of the form
(M + γK) y = M b,
for given b, u (not used), and γ = GetTimeStep(). */
class SundialsJacSolver : public SundialsODELinearSolver
{
private:
ConductionOperator *oper;
public:
SundialsJacSolver() : oper(NULL) { }
int InitSystem(void *sundials_mem);
int SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred, int &jac_cur,
Vector &v_temp1, Vector &v_temp2, Vector &v_temp3);
int SolveSystem(void *sundials_mem, Vector &b, const Vector &weight,
const Vector &y_cur, const Vector &f_cur);
int FreeSystem(void *sundials_mem);
};
double InitialTemperature(const Vector &x);
int main(int argc, char *argv[])
@@ -124,7 +113,7 @@ int main(int argc, char *argv[])
const char *mesh_file = "../../data/star.mesh";
int ref_levels = 2;
int order = 2;
int ode_solver_type = 11; // 11 = CVODE implicit
int ode_solver_type = 9; // CVODE implicit BDF
double t_final = 0.5;
double dt = 1.0e-2;
double alpha = 1.0e-2;
@@ -147,12 +136,19 @@ int main(int argc, char *argv[])
args.AddOption(&order, "-o", "--order",
"Order (degree) of the finite elements.");
args.AddOption(&ode_solver_type, "-s", "--ode-solver",
"ODE solver:\n"
"\t 1/11 - CVODE (explicit/implicit),\n"
"\t 2/12 - ARKODE (default explicit/implicit),\n"
"\t 3 - ARKODE (Fehlberg-6-4-5)\n"
"\t 4 - Forward Euler, 5 - RK2, 6 - RK3 SSP, 7 - RK4,\n"
"\t 8 - Backward Euler, 9 - SDIRK23, 10 - SDIRK33.");
"ODE solver:\n\t"
"1 - Forward Euler,\n\t"
"2 - RK2,\n\t"
"3 - RK3 SSP,\n\t"
"4 - RK4,\n\t"
"5 - Backward Euler,\n\t"
"6 - SDIRK 2,\n\t"
"7 - SDIRK 3,\n\t"
"8 - CVODE (implicit Adams),\n\t"
"9 - CVODE (implicit BDF),\n\t"
"10 - ARKODE (default explicit),\n\t"
"11 - ARKODE (explicit Fehlberg-6-4-5),\n\t"
"12 - ARKODE (default impicit).");
args.AddOption(&t_final, "-tf", "--t-final",
"Final time; start time is 0.");
args.AddOption(&dt, "-dt", "--time-step",
@@ -175,6 +171,11 @@ int main(int argc, char *argv[])
args.PrintUsage(cout);
return 1;
}
if (ode_solver_type < 1 || ode_solver_type > 12)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
return 3;
}
args.PrintOptions(cout);
// 2. Read the mesh from the given mesh file. We can handle triangular,
@@ -182,61 +183,7 @@ int main(int argc, char *argv[])
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
// 3. Define the ODE solver used for time integration. Several
// SUNDIALS solvers are available, as well as included both
// explicit and implicit MFEM ODE solvers.
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKODESolver *arkode = NULL;
SundialsJacSolver sun_solver; // Used by the implicit SUNDIALS ode solvers.
switch (ode_solver_type)
{
// SUNDIALS solvers
case 1:
cvode = new CVODESolver(CV_ADAMS, CV_FUNCTIONAL);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 11:
cvode = new CVODESolver(CV_BDF, CV_NEWTON);
cvode->SetLinearSolver(sun_solver);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 2:
case 3:
arkode = new ARKODESolver(ARKODESolver::EXPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 3) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
case 12:
arkode = new ARKODESolver(ARKODESolver::IMPLICIT);
arkode->SetLinearSolver(sun_solver);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
// Other MFEM explicit methods
case 4: ode_solver = new ForwardEulerSolver; break;
case 5: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 6: ode_solver = new RK3SSPSolver; break;
case 7: ode_solver = new RK4Solver; break;
// MFEM implicit L-stable methods
case 8: ode_solver = new BackwardEulerSolver; break;
case 9: ode_solver = new SDIRK23Solver(2); break;
case 10: ode_solver = new SDIRK33Solver; break;
default:
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
delete mesh;
return 3;
}
// Since we want to update the diffusion coefficient after every time step,
// we need to use the "one-step" mode of the SUNDIALS solvers.
if (cvode) { cvode->SetStepMode(CV_ONE_STEP); }
if (arkode) { arkode->SetStepMode(ARK_ONE_STEP); }
// 4. Refine the mesh to increase the resolution. In this example we do
// 3. Refine the mesh to increase the resolution. In this example we do
// 'ref_levels' of uniform refinement, where 'ref_levels' is a
// command-line parameter.
for (int lev = 0; lev < ref_levels; lev++)
@@ -244,7 +191,7 @@ int main(int argc, char *argv[])
mesh->UniformRefinement();
}
// 5. Define the vector finite element space representing the current and the
// 4. Define the vector finite element space representing the current and the
// initial temperature, u_ref.
H1_FECollection fe_coll(order, dim);
FiniteElementSpace fespace(mesh, &fe_coll);
@@ -254,14 +201,14 @@ int main(int argc, char *argv[])
GridFunction u_gf(&fespace);
// 6. Set the initial conditions for u. All boundaries are considered
// 5. Set the initial conditions for u. All boundaries are considered
// natural.
FunctionCoefficient u_0(InitialTemperature);
u_gf.ProjectCoefficient(u_0);
Vector u;
u_gf.GetTrueDofs(u);
// 7. Initialize the conduction operator and the visualization.
// 6. Initialize the conduction operator and the visualization.
ConductionOperator oper(fespace, alpha, kappa, u);
u_gf.SetFromTrueDofs(u);
@@ -307,13 +254,65 @@ int main(int argc, char *argv[])
}
}
// 7. Define the ODE solver used for time integration.
double t = 0.0;
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKStepSolver *arkode = NULL;
switch (ode_solver_type)
{
// MFEM explicit methods
case 1: ode_solver = new ForwardEulerSolver; break;
case 2: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 3: ode_solver = new RK3SSPSolver; break;
case 4: ode_solver = new RK4Solver; break;
// MFEM implicit L-stable methods
case 5: ode_solver = new BackwardEulerSolver; break;
case 6: ode_solver = new SDIRK23Solver(2); break;
case 7: ode_solver = new SDIRK33Solver; break;
// CVODE
case 8:
cvode = new CVODESolver(CV_ADAMS);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 9:
cvode = new CVODESolver(CV_BDF);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
// ARKODE
case 10:
case 11:
arkode = new ARKStepSolver(ARKStepSolver::EXPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 11) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
case 12:
arkode = new ARKStepSolver(ARKStepSolver::IMPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
}
// Initialize MFEM integrators, SUNDIALS integrators are initialized above
if (ode_solver_type < 8) { ode_solver->Init(oper); }
// Since we want to update the diffusion coefficient after every time step,
// we need to use the "one-step" mode of the SUNDIALS solvers.
if (cvode) { cvode->SetStepMode(CV_ONE_STEP); }
if (arkode) { arkode->SetStepMode(ARK_ONE_STEP); }
// 8. Perform time-integration (looping over the time iterations, ti, with a
// time-step dt).
cout << "Integrating the ODE ..." << endl;
tic_toc.Clear();
tic_toc.Start();
ode_solver->Init(oper);
double t = 0.0;
bool last_step = false;
for (int ti = 1; !last_step; ti++)
@@ -371,7 +370,7 @@ int main(int argc, char *argv[])
ConductionOperator::ConductionOperator(FiniteElementSpace &f, double al,
double kap, const Vector &u)
: TimeDependentOperator(f.GetTrueVSize(), 0.0), fespace(f), M(NULL), K(NULL),
T(NULL), current_dt(0.0), z(height)
T(NULL), z(height)
{
const double rel_tol = 1e-8;
@@ -417,32 +416,14 @@ void ConductionOperator::ImplicitSolve(const double dt,
// Solve the equation:
// du_dt = M^{-1}*[-K(u + dt*du_dt)]
// for du_dt
if (!T)
{
T = Add(1.0, Mmat, dt, Kmat);
current_dt = dt;
T_solver.SetOperator(*T);
}
MFEM_VERIFY(dt == current_dt, ""); // SDIRK methods use the same dt
if (T) { delete T; }
T = Add(1.0, Mmat, dt, Kmat);
T_solver.SetOperator(*T);
Kmat.Mult(u, z);
z.Neg();
T_solver.Mult(z, du_dt);
}
void ConductionOperator::SundialsSolve(const double dt, Vector &b)
{
// Solve the system (M + dt K) y = M b. The result y replaces the input b.
if (!T || dt != current_dt)
{
delete T;
T = Add(1.0, Mmat, dt, Kmat);
current_dt = dt;
T_solver.SetOperator(*T);
}
Mmat.Mult(b, z);
T_solver.Mult(z, b);
}
void ConductionOperator::SetParameters(const Vector &u)
{
GridFunction u_alpha_gf(&fespace);
@@ -460,8 +441,26 @@ void ConductionOperator::SetParameters(const Vector &u)
K->AddDomainIntegrator(new DiffusionIntegrator(u_coeff));
K->Assemble();
K->FormSystemMatrix(ess_tdof_list, Kmat);
delete T;
T = NULL; // re-compute T on the next ImplicitSolve or SundialsSolve
}
int ConductionOperator::SUNImplicitSetup(const Vector &x,
const Vector &fx, int jok, int *jcur,
double gamma)
{
// Setup the ODE Jacobian T = M + gamma K.
if (T) { delete T; }
T = Add(1.0, Mmat, gamma, Kmat);
T_solver.SetOperator(*T);
*jcur = 1;
return (0);
}
int ConductionOperator::SUNImplicitSolve(const Vector &b, Vector &x, double tol)
{
// Solve the system A x = z => (M - gamma K) x = M b.
Mmat.Mult(b, z);
T_solver.Mult(z, x);
return (0);
}
ConductionOperator::~ConductionOperator()
@@ -471,46 +470,6 @@ ConductionOperator::~ConductionOperator()
delete K;
}
int SundialsJacSolver::InitSystem(void *sundials_mem)
{
TimeDependentOperator *td_oper = GetTimeDependentOperator(sundials_mem);
// During development, we use dynamic_cast<> to ensure the setup is correct:
oper = dynamic_cast<ConductionOperator*>(td_oper);
MFEM_VERIFY(oper, "operator is not ConductionOperator");
// When the implementation is finalized, we can switch to static_cast<>:
// oper = static_cast<ConductionOperator*>(td_oper);
return 0;
}
int SundialsJacSolver::SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred,
int &jac_cur, Vector &v_temp1,
Vector &v_temp2, Vector &v_temp3)
{
jac_cur = 1;
return 0;
}
int SundialsJacSolver::SolveSystem(void *sundials_mem, Vector &b,
const Vector &weight, const Vector &y_cur,
const Vector &f_cur)
{
oper->SundialsSolve(GetTimeStep(sundials_mem), b);
return 0;
}
int SundialsJacSolver::FreeSystem(void *sundials_mem)
{
return 0;
}
double InitialTemperature(const Vector &x)
{
if (x.Norml2() < 0.5)
+116 -161
View File
@@ -8,9 +8,9 @@
// mpirun -np 4 ex16p -m ../../data/inline-tri.mesh
// mpirun -np 4 ex16p -m ../../data/disc-nurbs.mesh -tf 2
// mpirun -np 4 ex16p -s 12 -a 0.0 -k 1.0
// mpirun -np 4 ex16p -s 1 -a 1.0 -k 0.0 -dt 4e-6 -tf 2e-2 -vs 50
// mpirun -np 8 ex16p -s 2 -a 0.5 -k 0.5 -o 4 -dt 8e-6 -tf 2e-2 -vs 50
// mpirun -np 4 ex16p -s 3 -dt 2.0e-4 -tf 4.0e-2
// mpirun -np 4 ex16p -s 8 -a 1.0 -k 0.0 -dt 4e-6 -tf 2e-2 -vs 50
// mpirun -np 8 ex16p -s 9 -a 0.5 -k 0.5 -o 4 -dt 8e-6 -tf 2e-2 -vs 50
// mpirun -np 4 ex16p -s 10 -dt 2.0e-4 -tf 4.0e-2
// mpirun -np 16 ex16p -m ../../data/fichera-q2.mesh
// mpirun -np 16 ex16p -m ../../data/escher-p2.mesh
// mpirun -np 8 ex16p -m ../../data/beam-tet.mesh -tf 10 -dt 0.1
@@ -77,13 +77,19 @@ public:
const Vector &u);
virtual void Mult(const Vector &u, Vector &du_dt) const;
/** Solve the Backward-Euler equation: k = f(u + dt*k, t), for the unknown k.
This is the only requirement for high-order SDIRK implicit integration.*/
virtual void ImplicitSolve(const double dt, const Vector &u, Vector &k);
/** Solve the system (M + dt K) y = M b. The result y replaces the input b.
This method is used by the implicit SUNDIALS solvers. */
void SundialsSolve(const double dt, Vector &b);
/** Setup the system (M + dt K) x = M b. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSetup(const Vector &x, const Vector &fx,
int jok, int *jcur, double gamma);
/** Solve the system (M + dt K) x = M b. This method is used by the implicit
SUNDIALS solvers. */
virtual int SUNImplicitSolve(const Vector &b, Vector &x, double tol);
/// Update the diffusion BilinearForm K using the given true-dof vector `u`.
void SetParameters(const Vector &u);
@@ -91,33 +97,6 @@ public:
virtual ~ConductionOperator();
};
/// Custom Jacobian system solver for the SUNDIALS time integrators.
/** For the ODE system represented by ConductionOperator
M du/dt = -K(u),
this class facilitates the solution of linear systems of the form
(M + γK) y = M b,
for given b, u (not used), and γ = GetTimeStep(). */
class SundialsJacSolver : public SundialsODELinearSolver
{
private:
ConductionOperator *oper;
public:
SundialsJacSolver() : oper(NULL) { }
int InitSystem(void *sundials_mem);
int SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred, int &jac_cur,
Vector &v_temp1, Vector &v_temp2, Vector &v_temp3);
int SolveSystem(void *sundials_mem, Vector &b, const Vector &weight,
const Vector &y_cur, const Vector &f_cur);
int FreeSystem(void *sundials_mem);
};
double InitialTemperature(const Vector &x);
int main(int argc, char *argv[])
@@ -133,7 +112,7 @@ int main(int argc, char *argv[])
int ser_ref_levels = 2;
int par_ref_levels = 1;
int order = 2;
int ode_solver_type = 11; // 11 = CVODE implicit
int ode_solver_type = 9; // CVODE implicit BDF
double t_final = 0.5;
double dt = 1.0e-2;
double alpha = 1.0e-2;
@@ -158,12 +137,19 @@ int main(int argc, char *argv[])
args.AddOption(&order, "-o", "--order",
"Order (degree) of the finite elements.");
args.AddOption(&ode_solver_type, "-s", "--ode-solver",
"ODE solver:\n"
"\t 1/11 - CVODE (explicit/implicit),\n"
"\t 2/12 - ARKODE (default explicit/implicit),\n"
"\t 3 - ARKODE (Fehlberg-6-4-5)\n"
"\t 4 - Forward Euler, 5 - RK2, 6 - RK3 SSP, 7 - RK4,\n"
"\t 8 - Backward Euler, 9 - SDIRK23, 10 - SDIRK33.");
"ODE solver:\n\t"
"1 - Forward Euler,\n\t"
"2 - RK2,\n\t"
"3 - RK3 SSP,\n\t"
"4 - RK4,\n\t"
"5 - Backward Euler,\n\t"
"6 - SDIRK 2,\n\t"
"7 - SDIRK 3,\n\t"
"8 - CVODE (implicit Adams),\n\t"
"9 - CVODE (implicit BDF),\n\t"
"10 - ARKODE (default explicit),\n\t"
"11 - ARKODE (explicit Fehlberg-6-4-5),\n\t"
"12 - ARKODE (default impicit).");
args.AddOption(&t_final, "-tf", "--t-final",
"Final time; start time is 0.");
args.AddOption(&dt, "-dt", "--time-step",
@@ -193,67 +179,24 @@ int main(int argc, char *argv[])
args.PrintOptions(cout);
}
// check for vaild ODE solver option
if (ode_solver_type < 1 || ode_solver_type > 12)
{
if (myid == 0)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
}
MPI_Finalize();
return 1;
}
// 3. Read the serial mesh from the given mesh file on all processors. We can
// handle triangular, quadrilateral, tetrahedral and hexahedral meshes
// with the same code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
// 4. Define the ODE solver used for time integration. Several
// SUNDIALS solvers are available, as well as included both
// explicit and implicit MFEM ODE solvers.
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKODESolver *arkode = NULL;
SundialsJacSolver sun_solver; // Used by the implicit SUNDIALS ode solvers.
switch (ode_solver_type)
{
// SUNDIALS solvers
case 1:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_ADAMS, CV_FUNCTIONAL);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 11:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_BDF, CV_NEWTON);
cvode->SetLinearSolver(sun_solver);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 2:
case 3:
arkode = new ARKODESolver(MPI_COMM_WORLD, ARKODESolver::EXPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 3) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
case 12:
arkode = new ARKODESolver(MPI_COMM_WORLD, ARKODESolver::IMPLICIT);
arkode->SetLinearSolver(sun_solver);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
// Other MFEM explicit methods
case 4: ode_solver = new ForwardEulerSolver; break;
case 5: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 6: ode_solver = new RK3SSPSolver; break;
case 7: ode_solver = new RK4Solver; break;
// MFEM implicit L-stable methods
case 8: ode_solver = new BackwardEulerSolver; break;
case 9: ode_solver = new SDIRK23Solver(2); break;
case 10: ode_solver = new SDIRK33Solver; break;
default:
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
delete mesh;
return 3;
}
// Since we want to update the diffusion coefficient after every time step,
// we need to use the "one-step" mode of the SUNDIALS solvers.
if (cvode) { cvode->SetStepMode(CV_ONE_STEP); }
if (arkode) { arkode->SetStepMode(ARK_ONE_STEP); }
// 5. Refine the mesh in serial to increase the resolution. In this example
// 4. Refine the mesh in serial to increase the resolution. In this example
// we do 'ser_ref_levels' of uniform refinement, where 'ser_ref_levels' is
// a command-line parameter.
for (int lev = 0; lev < ser_ref_levels; lev++)
@@ -261,7 +204,7 @@ int main(int argc, char *argv[])
mesh->UniformRefinement();
}
// 6. Define a parallel mesh by a partitioning of the serial mesh. Refine
// 5. Define a parallel mesh by a partitioning of the serial mesh. Refine
// this mesh further in parallel to increase the resolution. Once the
// parallel mesh is defined, the serial mesh can be deleted.
ParMesh *pmesh = new ParMesh(MPI_COMM_WORLD, *mesh);
@@ -271,7 +214,7 @@ int main(int argc, char *argv[])
pmesh->UniformRefinement();
}
// 7. Define the vector finite element space representing the current and the
// 6. Define the vector finite element space representing the current and the
// initial temperature, u_ref.
H1_FECollection fe_coll(order, dim);
ParFiniteElementSpace fespace(pmesh, &fe_coll);
@@ -284,14 +227,14 @@ int main(int argc, char *argv[])
ParGridFunction u_gf(&fespace);
// 8. Set the initial conditions for u. All boundaries are considered
// 7. Set the initial conditions for u. All boundaries are considered
// natural.
FunctionCoefficient u_0(InitialTemperature);
u_gf.ProjectCoefficient(u_0);
Vector u;
u_gf.GetTrueDofs(u);
// 9. Initialize the conduction operator and the VisIt visualization.
// 8. Initialize the conduction operator and the VisIt visualization.
ConductionOperator oper(fespace, alpha, kappa, u);
u_gf.SetFromTrueDofs(u);
@@ -350,6 +293,60 @@ int main(int argc, char *argv[])
}
}
// 9. Define the ODE solver used for time integration.
double t = 0.0;
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKStepSolver *arkode = NULL;
switch (ode_solver_type)
{
// MFEM explicit methods
case 1: ode_solver = new ForwardEulerSolver; break;
case 2: ode_solver = new RK2Solver(0.5); break; // midpoint method
case 3: ode_solver = new RK3SSPSolver; break;
case 4: ode_solver = new RK4Solver; break;
// MFEM implicit L-stable methods
case 5: ode_solver = new BackwardEulerSolver; break;
case 6: ode_solver = new SDIRK23Solver(2); break;
case 7: ode_solver = new SDIRK33Solver; break;
// CVODE
case 8:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_ADAMS);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 9:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_BDF);
cvode->Init(oper);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
// ARKODE
case 10:
case 11:
arkode = new ARKStepSolver(MPI_COMM_WORLD, ARKStepSolver::EXPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 11) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
case 12:
arkode = new ARKStepSolver(MPI_COMM_WORLD, ARKStepSolver::IMPLICIT);
arkode->Init(oper);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
ode_solver = arkode; break;
}
// Initialize MFEM integrators, SUNDIALS integrators are initialized above
if (ode_solver_type < 8) { ode_solver->Init(oper); }
// Since we want to update the diffusion coefficient after every time step,
// we need to use the "one-step" mode of the SUNDIALS solvers.
if (cvode) { cvode->SetStepMode(CV_ONE_STEP); }
if (arkode) { arkode->SetStepMode(ARK_ONE_STEP); }
// 10. Perform time-integration (looping over the time iterations, ti, with a
// time-step dt).
if (myid == 0)
@@ -358,8 +355,6 @@ int main(int argc, char *argv[])
}
tic_toc.Clear();
tic_toc.Start();
ode_solver->Init(oper);
double t = 0.0;
bool last_step = false;
for (int ti = 1; !last_step; ti++)
@@ -428,7 +423,7 @@ int main(int argc, char *argv[])
ConductionOperator::ConductionOperator(ParFiniteElementSpace &f, double al,
double kap, const Vector &u)
: TimeDependentOperator(f.GetTrueVSize(), 0.0), fespace(f), M(NULL), K(NULL),
T(NULL), current_dt(0.0),
T(NULL),
M_solver(f.GetComm()), T_solver(f.GetComm()), z(height)
{
const double rel_tol = 1e-8;
@@ -476,30 +471,32 @@ void ConductionOperator::ImplicitSolve(const double dt,
// Solve the equation:
// du_dt = M^{-1}*[-K(u + dt*du_dt)]
// for du_dt
if (!T)
{
T = Add(1.0, Mmat, dt, Kmat);
current_dt = dt;
T_solver.SetOperator(*T);
}
MFEM_VERIFY(dt == current_dt, ""); // SDIRK methods use the same dt
if (T) { delete T; }
T = Add(1.0, Mmat, dt, Kmat);
T_solver.SetOperator(*T);
Kmat.Mult(u, z);
z.Neg();
T_solver.Mult(z, du_dt);
}
void ConductionOperator::SundialsSolve(const double dt, Vector &b)
int ConductionOperator::SUNImplicitSetup(const Vector &x,
const Vector &fx, int jok, int *jcur,
double gamma)
{
// Solve the system (M + dt K) y = M b. The result y replaces the input b.
if (!T || dt != current_dt)
{
delete T;
T = Add(1.0, Mmat, dt, Kmat);
current_dt = dt;
T_solver.SetOperator(*T);
}
// Setup the ODE Jacobian T = M + gamma K.
if (T) { delete T; }
T = Add(1.0, Mmat, gamma, Kmat);
T_solver.SetOperator(*T);
*jcur = 1;
return (0);
}
int ConductionOperator::SUNImplicitSolve(const Vector &b, Vector &x, double tol)
{
// Solve the system A x = z => (M - gamma K) x = M b.
Mmat.Mult(b, z);
T_solver.Mult(z, b);
T_solver.Mult(z, x);
return (0);
}
void ConductionOperator::SetParameters(const Vector &u)
@@ -519,8 +516,6 @@ void ConductionOperator::SetParameters(const Vector &u)
K->AddDomainIntegrator(new DiffusionIntegrator(u_coeff));
K->Assemble(0); // keep sparsity pattern of M and K the same
K->FormSystemMatrix(ess_tdof_list, Kmat);
delete T;
T = NULL; // re-compute T on the next ImplicitSolve or SundialsSolve
}
ConductionOperator::~ConductionOperator()
@@ -530,46 +525,6 @@ ConductionOperator::~ConductionOperator()
delete K;
}
int SundialsJacSolver::InitSystem(void *sundials_mem)
{
TimeDependentOperator *td_oper = GetTimeDependentOperator(sundials_mem);
// During development, we use dynamic_cast<> to ensure the setup is correct:
oper = dynamic_cast<ConductionOperator*>(td_oper);
MFEM_VERIFY(oper, "operator is not ConductionOperator");
// When the implementation is finalized, we can switch to static_cast<>:
// oper = static_cast<ConductionOperator*>(td_oper);
return 0;
}
int SundialsJacSolver::SetupSystem(void *sundials_mem, int conv_fail,
const Vector &y_pred, const Vector &f_pred,
int &jac_cur, Vector &v_temp1,
Vector &v_temp2, Vector &v_temp3)
{
jac_cur = 1;
return 0;
}
int SundialsJacSolver::SolveSystem(void *sundials_mem, Vector &b,
const Vector &weight, const Vector &y_cur,
const Vector &f_cur)
{
oper->SundialsSolve(GetTimeStep(sundials_mem), b);
return 0;
}
int SundialsJacSolver::FreeSystem(void *sundials_mem)
{
return 0;
}
double InitialTemperature(const Vector &x)
{
if (x.Norml2() < 0.5)
+74 -63
View File
@@ -4,14 +4,14 @@
// Compile with: make ex9
//
// Sample runs:
// ex9 -m ../../data/periodic-segment.mesh -p 0 -r 2 -s 11 -dt 0.005
// ex9 -m ../../data/periodic-square.mesh -p 1 -r 2 -s 12 -dt 0.005 -tf 9
// ex9 -m ../../data/periodic-hexagon.mesh -p 0 -r 2 -s 11 -dt 0.0018 -vs 25
// ex9 -m ../../data/periodic-hexagon.mesh -p 0 -r 2 -s 13 -dt 0.01 -vs 15
// ex9 -m ../../data/amr-quad.mesh -p 1 -r 2 -s 13 -dt 0.002 -tf 9
// ex9 -m ../../data/star-q3.mesh -p 1 -r 2 -s 13 -dt 0.005 -tf 9
// ex9 -m ../../data/disc-nurbs.mesh -p 1 -r 3 -s 11 -dt 0.005 -tf 9
// ex9 -m ../../data/periodic-cube.mesh -p 0 -r 2 -s 12 -dt 0.02 -tf 8 -o 2
// ex9 -m ../../data/periodic-segment.mesh -p 0 -r 2 -s 7 -dt 0.005
// ex9 -m ../../data/periodic-square.mesh -p 1 -r 2 -s 8 -dt 0.005 -tf 9
// ex9 -m ../../data/periodic-hexagon.mesh -p 0 -r 2 -s 7 -dt 0.0018 -vs 25
// ex9 -m ../../data/periodic-hexagon.mesh -p 0 -r 2 -s 9 -dt 0.01 -vs 15
// ex9 -m ../../data/amr-quad.mesh -p 1 -r 2 -s 9 -dt 0.002 -tf 9
// ex9 -m ../../data/star-q3.mesh -p 1 -r 2 -s 9 -dt 0.005 -tf 9
// ex9 -m ../../data/disc-nurbs.mesh -p 1 -r 3 -s 7 -dt 0.005 -tf 9
// ex9 -m ../../data/periodic-cube.mesh -p 0 -r 2 -s 8 -dt 0.02 -tf 8 -o 2
//
// Description: This example code solves the time-dependent advection equation
// du/dt + v.grad(u) = 0, where v is a given fluid velocity, and
@@ -109,11 +109,15 @@ int main(int argc, char *argv[])
args.AddOption(&order, "-o", "--order",
"Order (degree) of the finite elements.");
args.AddOption(&ode_solver_type, "-s", "--ode-solver",
"ODE solver: 1 - Forward Euler,\n\t"
" 2 - RK2 SSP, 3 - RK3 SSP, 4 - RK4, 6 - RK6,\n\t"
" 11 - CVODE (adaptive order) explicit,\n\t"
" 12 - ARKODE default (4th order) explicit,\n\t"
" 13 - ARKODE RK8.");
"ODE solver:\n\t"
"1 - Forward Euler,\n\t"
"2 - RK2 SSP,\n\t"
"3 - RK3 SSP,\n\t"
"4 - RK4,\n\t"
"6 - RK6,\n\t"
"7 - CVODE (adaptive order implicit Adams),\n\t"
"8 - ARKODE default (4th order) explicit,\n\t"
"9 - ARKODE RK8.");
args.AddOption(&t_final, "-tf", "--t-final",
"Final time; start time is 0.");
args.AddOption(&dt, "-dt", "--time-step",
@@ -135,65 +139,41 @@ int main(int argc, char *argv[])
args.PrintUsage(cout);
return 1;
}
// check for vaild ODE solver option
if (ode_solver_type < 1 || ode_solver_type > 9)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
return 3;
}
args.PrintOptions(cout);
// 2. Read the mesh from the given mesh file. We can handle geometrically
// periodic meshes in this code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
Mesh mesh(mesh_file, 1, 1);
int dim = mesh.Dimension();
// 3. Define the ODE solver used for time integration. Several explicit
// Runge-Kutta methods are available.
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKODESolver *arkode = NULL;
switch (ode_solver_type)
{
case 1: ode_solver = new ForwardEulerSolver; break;
case 2: ode_solver = new RK2Solver(1.0); break;
case 3: ode_solver = new RK3SSPSolver; break;
case 4: ode_solver = new RK4Solver; break;
case 6: ode_solver = new RK6Solver; break;
case 11:
cvode = new CVODESolver(CV_ADAMS, CV_FUNCTIONAL);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 12:
case 13:
arkode = new ARKODESolver(ARKODESolver::EXPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 13) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
default:
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
delete mesh;
return 3;
}
// 4. Refine the mesh to increase the resolution. In this example we do
// 3. Refine the mesh to increase the resolution. In this example we do
// 'ref_levels' of uniform refinement, where 'ref_levels' is a
// command-line parameter. If the mesh is of NURBS type, we convert it to
// a (piecewise-polynomial) high-order mesh.
for (int lev = 0; lev < ref_levels; lev++)
{
mesh->UniformRefinement();
mesh.UniformRefinement();
}
if (mesh->NURBSext)
if (mesh.NURBSext)
{
mesh->SetCurvature(max(order, 1));
mesh.SetCurvature(max(order, 1));
}
mesh->GetBoundingBox(bb_min, bb_max, max(order, 1));
mesh.GetBoundingBox(bb_min, bb_max, max(order, 1));
// 5. Define the discontinuous DG finite element space of the given
// 4. Define the discontinuous DG finite element space of the given
// polynomial order on the refined mesh.
DG_FECollection fec(order, dim);
FiniteElementSpace fes(mesh, &fec);
FiniteElementSpace fes(&mesh, &fec);
cout << "Number of unknowns: " << fes.GetVSize() << endl;
// 6. Set up and assemble the bilinear and linear forms corresponding to the
// 5. Set up and assemble the bilinear and linear forms corresponding to the
// DG discretization. The DGTraceIntegrator involves integrals over mesh
// interior faces.
VectorFunctionCoefficient velocity(dim, velocity_function);
@@ -220,7 +200,7 @@ int main(int argc, char *argv[])
k.Finalize(skip_zeros);
b.Assemble();
// 7. Define the initial conditions, save the corresponding grid function to
// 6. Define the initial conditions, save the corresponding grid function to
// a file and (optionally) save data in the VisIt format and initialize
// GLVis visualization.
GridFunction u(&fes);
@@ -229,7 +209,7 @@ int main(int argc, char *argv[])
{
ofstream omesh("ex9.mesh");
omesh.precision(precision);
mesh->Print(omesh);
mesh.Print(omesh);
ofstream osol("ex9-init.gf");
osol.precision(precision);
u.Save(osol);
@@ -243,14 +223,14 @@ int main(int argc, char *argv[])
if (binary)
{
#ifdef MFEM_USE_SIDRE
dc = new SidreDataCollection("Example9", mesh);
dc = new SidreDataCollection("Example9", &mesh);
#else
MFEM_ABORT("Must build with MFEM_USE_SIDRE=YES for binary output.");
#endif
}
else
{
dc = new VisItDataCollection("Example9", mesh);
dc = new VisItDataCollection("Example9", &mesh);
dc->SetPrecision(precision);
}
dc->RegisterField("solution", &u);
@@ -275,7 +255,7 @@ int main(int argc, char *argv[])
else
{
sout.precision(precision);
sout << "solution\n" << *mesh << u;
sout << "solution\n" << mesh << u;
sout << "pause\n";
sout << flush;
cout << "GLVis visualization paused."
@@ -283,15 +263,46 @@ int main(int argc, char *argv[])
}
}
// 8. Define the time-dependent evolution operator describing the ODE
// right-hand side, and perform time-integration (looping over the time
// iterations, ti, with a time-step dt).
// 7. Define the time-dependent evolution operator describing the ODE
// right-hand side, and define the ODE solver used for time integration.
FE_Evolution adv(m.SpMat(), k.SpMat(), b);
double t = 0.0;
adv.SetTime(t);
ode_solver->Init(adv);
// Create the time integrator
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKStepSolver *arkode = NULL;
switch (ode_solver_type)
{
case 1: ode_solver = new ForwardEulerSolver; break;
case 2: ode_solver = new RK2Solver(1.0); break;
case 3: ode_solver = new RK3SSPSolver; break;
case 4: ode_solver = new RK4Solver; break;
case 6: ode_solver = new RK6Solver; break;
case 7:
cvode = new CVODESolver(CV_ADAMS);
cvode->Init(adv);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
cvode->UseSundialsLinearSolver();
ode_solver = cvode; break;
case 8:
case 9:
arkode = new ARKStepSolver(ARKStepSolver::EXPLICIT);
arkode->Init(adv);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 9) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
}
// Initialize MFEM integrators, SUNDIALS integrators are initialized above
if (ode_solver_type < 7) { ode_solver->Init(adv); }
// 8. Perform time-integration (looping over the time iterations, ti,
// with a time-step dt).
bool done = false;
for (int ti = 0; !done; )
{
@@ -309,7 +320,7 @@ int main(int argc, char *argv[])
if (visualization)
{
sout << "solution\n" << *mesh << u << flush;
sout << "solution\n" << mesh << u << flush;
}
if (visit)
+67 -56
View File
@@ -4,14 +4,14 @@
// Compile with: make ex9p
//
// Sample runs:
// mpirun -np 4 ex9p -m ../../data/periodic-segment.mesh -p 1 -rp 1 -s 11 -dt 0.0025
// mpirun -np 4 ex9p -m ../../data/periodic-square.mesh -p 1 -rp 1 -s 12 -dt 0.0025 -tf 9
// mpirun -np 4 ex9p -m ../../data/periodic-hexagon.mesh -p 0 -rp 1 -s 11 -dt 0.0009 -vs 25
// mpirun -np 4 ex9p -m ../../data/periodic-hexagon.mesh -p 0 -rp 1 -s 13 -dt 0.005 -vs 15
// mpirun -np 4 ex9p -m ../../data/amr-quad.mesh -p 1 -rp 1 -s 13 -dt 0.001 -tf 9
// mpirun -np 4 ex9p -m ../../data/star-q3.mesh -p 1 -rp 1 -s 13 -dt 0.0025 -tf 9
// mpirun -np 4 ex9p -m ../../data/disc-nurbs.mesh -p 1 -rp 2 -s 11 -dt 0.0025 -tf 9
// mpirun -np 4 ex9p -m ../../data/periodic-cube.mesh -p 0 -rp 1 -s 12 -dt 0.01 -tf 8 -o 2
// mpirun -np 4 ex9p -m ../../data/periodic-segment.mesh -p 1 -rp 1 -s 7 -dt 0.0025
// mpirun -np 4 ex9p -m ../../data/periodic-square.mesh -p 1 -rp 1 -s 8 -dt 0.0025 -tf 9
// mpirun -np 4 ex9p -m ../../data/periodic-hexagon.mesh -p 0 -rp 1 -s 7 -dt 0.0009 -vs 25
// mpirun -np 4 ex9p -m ../../data/periodic-hexagon.mesh -p 0 -rp 1 -s 9 -dt 0.005 -vs 15
// mpirun -np 4 ex9p -m ../../data/amr-quad.mesh -p 1 -rp 1 -s 9 -dt 0.001 -tf 9
// mpirun -np 4 ex9p -m ../../data/star-q3.mesh -p 1 -rp 1 -s 9 -dt 0.0025 -tf 9
// mpirun -np 4 ex9p -m ../../data/disc-nurbs.mesh -p 1 -rp 2 -s 7 -dt 0.0025 -tf 9
// mpirun -np 4 ex9p -m ../../data/periodic-cube.mesh -p 0 -rp 1 -s 8 -dt 0.01 -tf 8 -o 2
//
// Description: This example code solves the time-dependent advection equation
// du/dt + v.grad(u) = 0, where v is a given fluid velocity, and
@@ -117,11 +117,15 @@ int main(int argc, char *argv[])
args.AddOption(&order, "-o", "--order",
"Order (degree) of the finite elements.");
args.AddOption(&ode_solver_type, "-s", "--ode-solver",
"ODE solver: 1 - Forward Euler,\n\t"
" 2 - RK2 SSP, 3 - RK3 SSP, 4 - RK4, 6 - RK6,\n\t"
" 11 - CVODE (adaptive order) explicit,\n\t"
" 12 - ARKODE default (4th order) explicit,\n\t"
" 13 - ARKODE RK8.");
"ODE solver:\n\t"
"1 - Forward Euler,\n\t"
"2 - RK2 SSP,\n\t"
"3 - RK3 SSP,\n\t"
"4 - RK4,\n\t"
"6 - RK6,\n\t"
"7 - CVODE (adaptive order implicit Adams),\n\t"
"8 - ARKODE default (4th order) explicit,\n\t"
"9 - ARKODE RK8.");
args.AddOption(&t_final, "-tf", "--t-final",
"Final time; start time is 0.");
args.AddOption(&dt, "-dt", "--time-step",
@@ -151,47 +155,23 @@ int main(int argc, char *argv[])
{
args.PrintOptions(cout);
}
// check for vaild ODE solver option
if (ode_solver_type < 1 || ode_solver_type > 9)
{
if (myid == 0)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
}
MPI_Finalize();
return 3;
}
// 3. Read the serial mesh from the given mesh file on all processors. We can
// handle geometrically periodic meshes in this code.
Mesh *mesh = new Mesh(mesh_file, 1, 1);
int dim = mesh->Dimension();
// 4. Define the ODE solver used for time integration. Several explicit
// Runge-Kutta methods are available.
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKODESolver *arkode = NULL;
switch (ode_solver_type)
{
case 1: ode_solver = new ForwardEulerSolver; break;
case 2: ode_solver = new RK2Solver(1.0); break;
case 3: ode_solver = new RK3SSPSolver; break;
case 4: ode_solver = new RK4Solver; break;
case 6: ode_solver = new RK6Solver; break;
case 11:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_ADAMS, CV_FUNCTIONAL);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
ode_solver = cvode; break;
case 12:
case 13:
arkode = new ARKODESolver(MPI_COMM_WORLD, ARKODESolver::EXPLICIT);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 13) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
default:
if (myid == 0)
{
cout << "Unknown ODE solver type: " << ode_solver_type << '\n';
}
delete mesh;
MPI_Finalize();
return 3;
}
// 5. Refine the mesh in serial to increase the resolution. In this example
// 4. Refine the mesh in serial to increase the resolution. In this example
// we do 'ser_ref_levels' of uniform refinement, where 'ser_ref_levels' is
// a command-line parameter. If the mesh is of NURBS type, we convert it
// to a (piecewise-polynomial) high-order mesh.
@@ -205,7 +185,7 @@ int main(int argc, char *argv[])
}
mesh->GetBoundingBox(bb_min, bb_max, max(order, 1));
// 6. Define the parallel mesh by a partitioning of the serial mesh. Refine
// 5. Define the parallel mesh by a partitioning of the serial mesh. Refine
// this mesh further in parallel to increase the resolution. Once the
// parallel mesh is defined, the serial mesh can be deleted.
ParMesh *pmesh = new ParMesh(MPI_COMM_WORLD, *mesh);
@@ -215,7 +195,7 @@ int main(int argc, char *argv[])
pmesh->UniformRefinement();
}
// 7. Define the parallel discontinuous DG finite element space on the
// 6. Define the parallel discontinuous DG finite element space on the
// parallel refined mesh of the given polynomial order.
DG_FECollection fec(order, dim);
ParFiniteElementSpace *fes = new ParFiniteElementSpace(pmesh, &fec);
@@ -226,7 +206,7 @@ int main(int argc, char *argv[])
cout << "Number of unknowns: " << global_vSize << endl;
}
// 8. Set up and assemble the parallel bilinear and linear forms (and the
// 7. Set up and assemble the parallel bilinear and linear forms (and the
// parallel hypre matrices) corresponding to the DG discretization. The
// DGTraceIntegrator involves integrals over mesh interior faces.
VectorFunctionCoefficient velocity(dim, velocity_function);
@@ -257,7 +237,7 @@ int main(int argc, char *argv[])
HypreParMatrix *K = k->ParallelAssemble();
HypreParVector *B = b->ParallelAssemble();
// 9. Define the initial conditions, save the corresponding grid function to
// 8. Define the initial conditions, save the corresponding grid function to
// a file and (optionally) save data in the VisIt format and initialize
// GLVis visualization.
ParGridFunction *u = new ParGridFunction(fes);
@@ -330,15 +310,46 @@ int main(int argc, char *argv[])
}
}
// 10. Define the time-dependent evolution operator describing the ODE
// right-hand side, and perform time-integration (looping over the time
// iterations, ti, with a time-step dt).
// 9. Define the time-dependent evolution operator describing the ODE
// right-hand side, and define the ODE solver used for time integration.
FE_Evolution adv(*M, *K, *B);
double t = 0.0;
adv.SetTime(t);
ode_solver->Init(adv);
// Create the time integrator
ODESolver *ode_solver = NULL;
CVODESolver *cvode = NULL;
ARKStepSolver *arkode = NULL;
switch (ode_solver_type)
{
case 1: ode_solver = new ForwardEulerSolver; break;
case 2: ode_solver = new RK2Solver(1.0); break;
case 3: ode_solver = new RK3SSPSolver; break;
case 4: ode_solver = new RK4Solver; break;
case 6: ode_solver = new RK6Solver; break;
case 7:
cvode = new CVODESolver(MPI_COMM_WORLD, CV_ADAMS);
cvode->Init(adv);
cvode->SetSStolerances(reltol, abstol);
cvode->SetMaxStep(dt);
cvode->UseSundialsLinearSolver();
ode_solver = cvode; break;
case 8:
case 9:
arkode = new ARKStepSolver(MPI_COMM_WORLD, ARKStepSolver::EXPLICIT);
arkode->Init(adv);
arkode->SetSStolerances(reltol, abstol);
arkode->SetMaxStep(dt);
if (ode_solver_type == 9) { arkode->SetERKTableNum(FEHLBERG_13_7_8); }
ode_solver = arkode; break;
}
// Initialize MFEM integrators, SUNDIALS integrators are initialized above
if (ode_solver_type < 7) { ode_solver->Init(adv); }
// 10. Perform time-integration (looping over the time iterations, ti,
// with a time-step dt).
bool done = false;
for (int ti = 0; !done; )
{
+3 -3
View File
@@ -60,15 +60,15 @@ PARALLEL_NAME := Parallel SUNDIALS example
@$(call mfem-test,$<,, $(SERIAL_NAME))
# Testing: Specific execution options:
# Example 9: test explicit CVODE time stepping
EX9_COMMON_ARGS := -m ../../data/periodic-hexagon.mesh -p 0 -s 11
# Example 9: test CVODE with CV_ADAMS (non-stiff implicit) time stepping
EX9_COMMON_ARGS := -m ../../data/periodic-hexagon.mesh -p 0 -s 7
EX9_ARGS := $(EX9_COMMON_ARGS) -r 2 -dt 0.0018 -vs 25
EX9P_ARGS := $(EX9_COMMON_ARGS) -rp 1 -dt 0.0009 -vs 50
ex9-test-seq: ex9
@$(call mfem-test,$<,, $(SERIAL_NAME),$(EX9_ARGS))
ex9p-test-par: ex9p
@$(call mfem-test,$<, $(RUN_MPI), $(PARALLEL_NAME),$(EX9P_ARGS))
# Example 10: test implicit CVODE time stepping
# Example 10: test CVODE with CV_BDF (stiff implicit) time stepping
EX10_COMMON_ARGS := -m ../../data/beam-quad.mesh -o 2 -s 5 -dt 0.15 -tf 6 -vs 10
EX10_ARGS := $(EX10_COMMON_ARGS) -r 2
EX10P_ARGS := $(EX10_COMMON_ARGS) -rp 1
+2 -2
View File
@@ -13,7 +13,8 @@ set(SRCS
bilinearform.cpp
bilinearform_ext.cpp
bilininteg.cpp
bilininteg_ext.cpp
bilininteg_diffusion.cpp
bilininteg_mass.cpp
coefficient.cpp
datacollection.cpp
eltrans.cpp
@@ -37,7 +38,6 @@ set(HDRS
bilinearform.hpp
bilinearform_ext.hpp
bilininteg.hpp
bilininteg_ext.hpp
coefficient.hpp
datacollection.hpp
eltrans.hpp
+274 -53
View File
@@ -55,7 +55,7 @@ void BilinearForm::AllocMat()
int *I = dof_dof.GetI();
int *J = dof_dof.GetJ();
double *data = mfem::New<double>(I[height]);
double *data = new double[I[height]];
mat = new SparseMatrix(I, J, data, height, height, true, true, true);
*mat = 0.0;
@@ -122,11 +122,7 @@ void BilinearForm::SetAssemblyLevel(AssemblyLevel assembly_level)
switch (assembly)
{
case AssemblyLevel::FULL:
if (Device::IsEnabled())
{
mfem_error("Full assembly not supported yet in device mode!");
// ext = new FABilinearFormExtension(this);
}
// ext = new FABilinearFormExtension(this);
// Use the original BilinearForm implementation for now
break;
case AssemblyLevel::ELEMENT:
@@ -298,6 +294,33 @@ void BilinearForm::ComputeElementMatrix(int i, DenseMatrix &elmat)
}
}
void BilinearForm::ComputeBdrElementMatrix(int i, DenseMatrix &elmat)
{
if (bbfi.Size())
{
const FiniteElement &be = *fes->GetBE(i);
ElementTransformation *eltrans = fes->GetBdrElementTransformation(i);
bbfi[0]->AssembleElementMatrix(be, *eltrans, elmat);
for (int k = 1; k < bbfi.Size(); k++)
{
bbfi[k]->AssembleElementMatrix(be, *eltrans, elemmat);
elmat += elemmat;
}
}
else
{
fes->GetBdrElementVDofs(i, vdofs);
elmat.SetSize(vdofs.Size());
elmat = 0.0;
}
}
void BilinearForm::AssembleElementMatrix(
int i, const DenseMatrix &elmat, int skip_zeros)
{
AssembleElementMatrix(i, elmat, vdofs, skip_zeros);
}
void BilinearForm::AssembleElementMatrix(
int i, const DenseMatrix &elmat, Array<int> &vdofs, int skip_zeros)
{
@@ -320,6 +343,12 @@ void BilinearForm::AssembleElementMatrix(
}
}
void BilinearForm::AssembleBdrElementMatrix(
int i, const DenseMatrix &elmat, int skip_zeros)
{
AssembleBdrElementMatrix(i, elmat, vdofs, skip_zeros);
}
void BilinearForm::AssembleBdrElementMatrix(
int i, const DenseMatrix &elmat, Array<int> &vdofs, int skip_zeros)
{
@@ -344,11 +373,6 @@ void BilinearForm::AssembleBdrElementMatrix(
void BilinearForm::Assemble(int skip_zeros)
{
if (Device::IsEnabled() && (assembly != AssemblyLevel::PARTIAL))
{
mfem_error("Chosen assembly level not supported yet in device mode!");
}
if (ext)
{
ext->Assemble();
@@ -588,14 +612,14 @@ void BilinearForm::FormLinearSystem(const Array<int> &ess_tdof_list, Vector &x,
Vector &b, OperatorHandle &A, Vector &X,
Vector &B, int copy_interior)
{
const SparseMatrix *P = fes->GetConformingProlongation();
if (ext)
{
ext->FormLinearSystem(ess_tdof_list, x, b, A, X, B, copy_interior);
return;
}
const SparseMatrix *P = fes->GetConformingProlongation();
FormSystemMatrix(ess_tdof_list, A);
// Transform the system and perform the elimination in B, based on the
@@ -621,8 +645,8 @@ void BilinearForm::FormLinearSystem(const Array<int> &ess_tdof_list, Vector &x,
{
// A, X and B point to the same data as mat, x and b
EliminateVDofsInRHS(ess_tdof_list, x, b);
X.NewDataAndSize(x.GetData(), x.Size());
B.NewDataAndSize(b.GetData(), b.Size());
X.NewMemoryAndSize(x.GetMemory(), x.Size(), false);
B.NewMemoryAndSize(b.GetMemory(), b.Size(), false);
if (!copy_interior) { X.SetSubVectorComplement(ess_tdof_list, 0.0); }
}
}
@@ -723,6 +747,10 @@ void BilinearForm::RecoverFEMSolution(const Vector &X,
else
{
// X and x point to the same data
// If the validity flags of X's Memory were changed (e.g. if it was
// moved to device memory) then we need to tell x about that.
x.SyncMemory(X);
}
}
else // non-conforming space
@@ -1021,9 +1049,13 @@ MixedBilinearForm::MixedBilinearForm (FiniteElementSpace *tr_fes,
extern_bfs = 1;
// Copy the pointers to the integrators
dom = mbf->dom;
bdr = mbf->bdr;
skt = mbf->skt;
dbfi = mbf->dbfi;
bbfi = mbf->bbfi;
tfbfi = mbf->tfbfi;
btfbfi = mbf->btfbfi;
bbfi_marker = mbf->bbfi_marker;
btfbfi_marker = mbf->btfbfi_marker;
}
double & MixedBilinearForm::Elem (int i, int j)
@@ -1077,22 +1109,42 @@ void MixedBilinearForm::GetBlocks(Array2D<SparseMatrix *> &blocks) const
void MixedBilinearForm::AddDomainIntegrator (BilinearFormIntegrator * bfi)
{
dom.Append (bfi);
dbfi.Append (bfi);
}
void MixedBilinearForm::AddBoundaryIntegrator (BilinearFormIntegrator * bfi)
{
bdr.Append (bfi);
bbfi.Append (bfi);
bbfi_marker.Append(NULL); // NULL marker means apply everywhere
}
void MixedBilinearForm::AddBoundaryIntegrator (BilinearFormIntegrator * bfi,
Array<int> &bdr_marker)
{
bbfi.Append (bfi);
bbfi_marker.Append(&bdr_marker);
}
void MixedBilinearForm::AddTraceFaceIntegrator (BilinearFormIntegrator * bfi)
{
skt.Append (bfi);
tfbfi.Append (bfi);
}
void MixedBilinearForm::AddBdrTraceFaceIntegrator(BilinearFormIntegrator *bfi)
{
btfbfi.Append(bfi);
btfbfi_marker.Append(NULL); // NULL marker means apply everywhere
}
void MixedBilinearForm::AddBdrTraceFaceIntegrator(BilinearFormIntegrator *bfi,
Array<int> &bdr_marker)
{
btfbfi.Append(bfi);
btfbfi_marker.Append(&bdr_marker);
}
void MixedBilinearForm::Assemble (int skip_zeros)
{
int i, k;
Array<int> tr_vdofs, te_vdofs;
ElementTransformation *eltrans;
DenseMatrix elemmat;
@@ -1104,48 +1156,75 @@ void MixedBilinearForm::Assemble (int skip_zeros)
mat = new SparseMatrix(height, width);
}
if (dom.Size())
if (dbfi.Size())
{
for (i = 0; i < test_fes -> GetNE(); i++)
for (int i = 0; i < test_fes -> GetNE(); i++)
{
trial_fes -> GetElementVDofs (i, tr_vdofs);
test_fes -> GetElementVDofs (i, te_vdofs);
eltrans = test_fes -> GetElementTransformation (i);
for (k = 0; k < dom.Size(); k++)
for (int k = 0; k < dbfi.Size(); k++)
{
dom[k] -> AssembleElementMatrix2 (*trial_fes -> GetFE(i),
*test_fes -> GetFE(i),
*eltrans, elemmat);
dbfi[k] -> AssembleElementMatrix2 (*trial_fes -> GetFE(i),
*test_fes -> GetFE(i),
*eltrans, elemmat);
mat -> AddSubMatrix (te_vdofs, tr_vdofs, elemmat, skip_zeros);
}
}
}
if (bdr.Size())
if (bbfi.Size())
{
for (i = 0; i < test_fes -> GetNBE(); i++)
// Which boundary attributes need to be processed?
Array<int> bdr_attr_marker(mesh->bdr_attributes.Size() ?
mesh->bdr_attributes.Max() : 0);
bdr_attr_marker = 0;
for (int k = 0; k < bbfi.Size(); k++)
{
if (bbfi_marker[k] == NULL)
{
bdr_attr_marker = 1;
break;
}
Array<int> &bdr_marker = *bbfi_marker[k];
MFEM_ASSERT(bdr_marker.Size() == bdr_attr_marker.Size(),
"invalid boundary marker for boundary integrator #"
<< k << ", counting from zero");
for (int i = 0; i < bdr_attr_marker.Size(); i++)
{
bdr_attr_marker[i] |= bdr_marker[i];
}
}
for (int i = 0; i < test_fes -> GetNBE(); i++)
{
const int bdr_attr = mesh->GetBdrAttribute(i);
if (bdr_attr_marker[bdr_attr-1] == 0) { continue; }
trial_fes -> GetBdrElementVDofs (i, tr_vdofs);
test_fes -> GetBdrElementVDofs (i, te_vdofs);
eltrans = test_fes -> GetBdrElementTransformation (i);
for (k = 0; k < bdr.Size(); k++)
for (int k = 0; k < bbfi.Size(); k++)
{
bdr[k] -> AssembleElementMatrix2 (*trial_fes -> GetBE(i),
*test_fes -> GetBE(i),
*eltrans, elemmat);
if (bbfi_marker[k] &&
(*bbfi_marker[k])[bdr_attr-1] == 0) { continue; }
bbfi[k] -> AssembleElementMatrix2 (*trial_fes -> GetBE(i),
*test_fes -> GetBE(i),
*eltrans, elemmat);
mat -> AddSubMatrix (te_vdofs, tr_vdofs, elemmat, skip_zeros);
}
}
}
if (skt.Size())
if (tfbfi.Size())
{
FaceElementTransformations *ftr;
Array<int> te_vdofs2;
const FiniteElement *trial_face_fe, *test_fe1, *test_fe2;
int nfaces = mesh->GetNumFaces();
for (i = 0; i < nfaces; i++)
for (int i = 0; i < nfaces; i++)
{
ftr = mesh->GetFaceElementTransformations(i);
trial_fes->GetFaceVDofs(i, tr_vdofs);
@@ -1165,14 +1244,70 @@ void MixedBilinearForm::Assemble (int skip_zeros)
// want to actually make a fake element.
test_fe2 = test_fe1;
}
for (int k = 0; k < skt.Size(); k++)
for (int k = 0; k < tfbfi.Size(); k++)
{
skt[k]->AssembleFaceMatrix(*trial_face_fe, *test_fe1, *test_fe2,
*ftr, elemmat);
tfbfi[k]->AssembleFaceMatrix(*trial_face_fe, *test_fe1, *test_fe2,
*ftr, elemmat);
mat->AddSubMatrix(te_vdofs, tr_vdofs, elemmat, skip_zeros);
}
}
}
if (btfbfi.Size())
{
FaceElementTransformations *ftr;
Array<int> te_vdofs2;
const FiniteElement *trial_face_fe, *test_fe1, *test_fe2;
// Which boundary attributes need to be processed?
Array<int> bdr_attr_marker(mesh->bdr_attributes.Size() ?
mesh->bdr_attributes.Max() : 0);
bdr_attr_marker = 0;
for (int k = 0; k < btfbfi.Size(); k++)
{
if (btfbfi_marker[k] == NULL)
{
bdr_attr_marker = 1;
break;
}
Array<int> &bdr_marker = *btfbfi_marker[k];
MFEM_ASSERT(bdr_marker.Size() == bdr_attr_marker.Size(),
"invalid boundary marker for boundary trace face integrator #"
<< k << ", counting from zero");
for (int i = 0; i < bdr_attr_marker.Size(); i++)
{
bdr_attr_marker[i] |= bdr_marker[i];
}
}
for (int i = 0; i < trial_fes -> GetNBE(); i++)
{
const int bdr_attr = mesh->GetBdrAttribute(i);
if (bdr_attr_marker[bdr_attr-1] == 0) { continue; }
ftr = mesh->GetBdrFaceTransformations(i);
if (ftr)
{
trial_fes->GetFaceVDofs(i, tr_vdofs);
test_fes->GetElementVDofs(ftr->Elem1No, te_vdofs);
trial_face_fe = trial_fes->GetFaceElement(i);
test_fe1 = test_fes->GetFE(ftr->Elem1No);
// The test_fe2 object is really a dummy and not used on the
// boundaries, but we can't dereference a NULL pointer, and we don't
// want to actually make a fake element.
test_fe2 = test_fe1;
for (int k = 0; k < btfbfi.Size(); k++)
{
if (btfbfi_marker[k] &&
(*btfbfi_marker[k])[bdr_attr-1] == 0) { continue; }
btfbfi[k]->AssembleFaceMatrix(*trial_face_fe, *test_fe1, *test_fe2,
*ftr, elemmat);
mat->AddSubMatrix(te_vdofs, tr_vdofs, elemmat, skip_zeros);
}
}
}
}
}
void MixedBilinearForm::ConformingAssemble()
@@ -1201,8 +1336,93 @@ void MixedBilinearForm::ConformingAssemble()
width = mat->Width();
}
void MixedBilinearForm::ComputeElementMatrix(int i, DenseMatrix &elmat)
{
if (dbfi.Size())
{
const FiniteElement &trial_fe = *trial_fes->GetFE(i);
const FiniteElement &test_fe = *test_fes->GetFE(i);
ElementTransformation *eltrans = test_fes->GetElementTransformation(i);
dbfi[0]->AssembleElementMatrix2(trial_fe, test_fe, *eltrans, elmat);
for (int k = 1; k < dbfi.Size(); k++)
{
dbfi[k]->AssembleElementMatrix2(trial_fe, test_fe, *eltrans, elemmat);
elmat += elemmat;
}
}
else
{
trial_fes->GetElementVDofs(i, trial_vdofs);
test_fes->GetElementVDofs(i, test_vdofs);
elmat.SetSize(test_vdofs.Size(), trial_vdofs.Size());
elmat = 0.0;
}
}
void MixedBilinearForm::ComputeBdrElementMatrix(int i, DenseMatrix &elmat)
{
if (bbfi.Size())
{
const FiniteElement &trial_be = *trial_fes->GetBE(i);
const FiniteElement &test_be = *test_fes->GetBE(i);
ElementTransformation *eltrans = test_fes->GetBdrElementTransformation(i);
bbfi[0]->AssembleElementMatrix2(trial_be, test_be, *eltrans, elmat);
for (int k = 1; k < bbfi.Size(); k++)
{
bbfi[k]->AssembleElementMatrix2(trial_be, test_be, *eltrans, elemmat);
elmat += elemmat;
}
}
else
{
trial_fes->GetBdrElementVDofs(i, trial_vdofs);
test_fes->GetBdrElementVDofs(i, test_vdofs);
elmat.SetSize(test_vdofs.Size(), trial_vdofs.Size());
elmat = 0.0;
}
}
void MixedBilinearForm::AssembleElementMatrix(
int i, const DenseMatrix &elmat, int skip_zeros)
{
AssembleElementMatrix(i, elmat, trial_vdofs, test_vdofs, skip_zeros);
}
void MixedBilinearForm::AssembleElementMatrix(
int i, const DenseMatrix &elmat, Array<int> &trial_vdofs,
Array<int> &test_vdofs, int skip_zeros)
{
trial_fes->GetElementVDofs(i, trial_vdofs);
test_fes->GetElementVDofs(i, test_vdofs);
if (mat == NULL)
{
mat = new SparseMatrix(height, width);
}
mat->AddSubMatrix(test_vdofs, trial_vdofs, elmat, skip_zeros);
}
void MixedBilinearForm::AssembleBdrElementMatrix(
int i, const DenseMatrix &elmat, int skip_zeros)
{
AssembleBdrElementMatrix(i, elmat, trial_vdofs, test_vdofs, skip_zeros);
}
void MixedBilinearForm::AssembleBdrElementMatrix(
int i, const DenseMatrix &elmat, Array<int> &trial_vdofs,
Array<int> &test_vdofs, int skip_zeros)
{
trial_fes->GetBdrElementVDofs(i, trial_vdofs);
test_fes->GetBdrElementVDofs(i, test_vdofs);
if (mat == NULL)
{
mat = new SparseMatrix(height, width);
}
mat->AddSubMatrix(test_vdofs, trial_vdofs, elmat, skip_zeros);
}
void MixedBilinearForm::EliminateTrialDofs (
Array<int> &bdr_attr_is_ess, const Vector &sol, Vector &rhs )
const Array<int> &bdr_attr_is_ess, const Vector &sol, Vector &rhs )
{
int i, j, k;
Array<int> tr_vdofs, cols_marker (trial_fes -> GetVSize());
@@ -1225,12 +1445,12 @@ void MixedBilinearForm::EliminateTrialDofs (
}
void MixedBilinearForm::EliminateEssentialBCFromTrialDofs (
Array<int> &marked_vdofs, const Vector &sol, Vector &rhs)
const Array<int> &marked_vdofs, const Vector &sol, Vector &rhs)
{
mat -> EliminateCols (marked_vdofs, &sol, &rhs);
}
void MixedBilinearForm::EliminateTestDofs (Array<int> &bdr_attr_is_ess)
void MixedBilinearForm::EliminateTestDofs (const Array<int> &bdr_attr_is_ess)
{
int i, j, k;
Array<int> te_vdofs;
@@ -1264,9 +1484,10 @@ MixedBilinearForm::~MixedBilinearForm()
if (!extern_bfs)
{
int i;
for (i = 0; i < dom.Size(); i++) { delete dom[i]; }
for (i = 0; i < bdr.Size(); i++) { delete bdr[i]; }
for (i = 0; i < skt.Size(); i++) { delete skt[i]; }
for (i = 0; i < dbfi.Size(); i++) { delete dbfi[i]; }
for (i = 0; i < bbfi.Size(); i++) { delete bbfi[i]; }
for (i = 0; i < tfbfi.Size(); i++) { delete tfbfi[i]; }
for (i = 0; i < btfbfi.Size(); i++) { delete btfbfi[i]; }
}
}
@@ -1283,7 +1504,7 @@ void DiscreteLinearOperator::Assemble(int skip_zeros)
mat = new SparseMatrix(height, width);
}
if (dom.Size() > 0)
if (dbfi.Size() > 0)
{
for (int i = 0; i < test_fes->GetNE(); i++)
{
@@ -1293,17 +1514,17 @@ void DiscreteLinearOperator::Assemble(int skip_zeros)
dom_fe = trial_fes->GetFE(i);
ran_fe = test_fes->GetFE(i);
dom[0]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, totelmat);
for (int j = 1; j < dom.Size(); j++)
dbfi[0]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, totelmat);
for (int j = 1; j < dbfi.Size(); j++)
{
dom[j]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, elmat);
dbfi[j]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, elmat);
totelmat += elmat;
}
mat->SetSubMatrix(ran_vdofs, dom_vdofs, totelmat, skip_zeros);
}
}
if (skt.Size())
if (tfbfi.Size())
{
const int nfaces = test_fes->GetMesh()->GetNumFaces();
for (int i = 0; i < nfaces; i++)
@@ -1314,10 +1535,10 @@ void DiscreteLinearOperator::Assemble(int skip_zeros)
dom_fe = trial_fes->GetFaceElement(i);
ran_fe = test_fes->GetFaceElement(i);
skt[0]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, totelmat);
for (int j = 1; j < skt.Size(); j++)
tfbfi[0]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, totelmat);
for (int j = 1; j < tfbfi.Size(); j++)
{
skt[j]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, elmat);
tfbfi[j]->AssembleElementMatrix2(*dom_fe, *ran_fe, *T, elmat);
totelmat += elmat;
}
mat->SetSubMatrix(ran_vdofs, dom_vdofs, totelmat, skip_zeros);
+130 -12
View File
@@ -413,9 +413,49 @@ public:
void FreeElementMatrices()
{ delete element_matrices; element_matrices = NULL; }
/// Compute the element matrix of the given element
/** The element matrix is computed by calling the domain integrators
or the one stored internally by a prior call of ComputeElementMatrices()
is returned when available.
*/
void ComputeElementMatrix(int i, DenseMatrix &elmat);
/// Compute the boundary element matrix of the given boundary element
void ComputeBdrElementMatrix(int i, DenseMatrix &elmat);
/// Assemble the given element matrix
/** The element matrix @a elmat is assembled for the element @a i, i.e.
added to the system matrix. The flag @a skip_zeros skips the zero
elements of the matrix, unless they are breaking the symmetry of
the system matrix.
*/
void AssembleElementMatrix(int i, const DenseMatrix &elmat,
int skip_zeros = 1);
/// Assemble the given element matrix
/** The element matrix @a elmat is assembled for the element @a i, i.e.
added to the system matrix. The vdofs of the element are returned
in @a vdofs. The flag @a skip_zeros skips the zero elements of the
matrix, unless they are breaking the symmetry of the system matrix.
*/
void AssembleElementMatrix(int i, const DenseMatrix &elmat,
Array<int> &vdofs, int skip_zeros = 1);
/// Assemble the given boundary element matrix
/** The boundary element matrix @a elmat is assembled for the boundary
element @a i, i.e. added to the system matrix. The flag @a skip_zeros
skips the zero elements of the matrix, unless they are breaking the
symmetry of the system matrix.
*/
void AssembleBdrElementMatrix(int i, const DenseMatrix &elmat,
int skip_zeros = 1);
/// Assemble the given boundary element matrix
/** The boundary element matrix @a elmat is assembled for the boundary
element @a i, i.e. added to the system matrix. The vdofs of the element
are returned in @a vdofs. The flag @a skip_zeros skips the zero elements
of the matrix, unless they are breaking the symmetry of the system matrix.
*/
void AssembleBdrElementMatrix(int i, const DenseMatrix &elmat,
Array<int> &vdofs, int skip_zeros = 1);
@@ -513,16 +553,26 @@ protected:
FiniteElementSpace *trial_fes, ///< Not owned
*test_fes; ///< Not owned
/** @brief Indicates the BilinearFormIntegrator%s stored in #dom, #bdr, and
#skt are owned by another MixedBilinearForm. */
/** @brief Indicates the BilinearFormIntegrator%s stored in #dbfi, #bbfi,
#tfbfi and #btfbfi are owned by another MixedBilinearForm. */
int extern_bfs;
/// Domain integrators.
Array<BilinearFormIntegrator*> dom;
Array<BilinearFormIntegrator*> dbfi;
/// Boundary integrators.
Array<BilinearFormIntegrator*> bdr;
Array<BilinearFormIntegrator*> bbfi;
Array<Array<int>*> bbfi_marker;///< Entries are not owned.
/// Trace face (skeleton) integrators.
Array<BilinearFormIntegrator*> skt;
Array<BilinearFormIntegrator*> tfbfi;
/// Boundary trace face (skeleton) integrators.
Array<BilinearFormIntegrator*> btfbfi;
Array<Array<int>*> btfbfi_marker;///< Entries are not owned.
DenseMatrix elemmat;
Array<int> trial_vdofs, test_vdofs;
private:
/// Copy construction is not supported; body is undefined.
@@ -586,6 +636,10 @@ public:
/// Adds a boundary integrator. Assumes ownership of @a bfi.
void AddBoundaryIntegrator(BilinearFormIntegrator *bfi);
/// Adds a boundary integrator. Assumes ownership of @a bfi.
void AddBoundaryIntegrator (BilinearFormIntegrator * bfi,
Array<int> &bdr_marker);
/** @brief Add a trace face integrator. Assumes ownership of @a bfi.
This type of integrator assembles terms over all faces of the mesh using
@@ -593,14 +647,32 @@ public:
test space. */
void AddTraceFaceIntegrator(BilinearFormIntegrator *bfi);
/// Adds a boundary trace face integrator. Assumes ownership of @a bfi.
void AddBdrTraceFaceIntegrator (BilinearFormIntegrator * bfi);
/// Adds a boundary trace face integrator. Assumes ownership of @a bfi.
void AddBdrTraceFaceIntegrator (BilinearFormIntegrator * bfi,
Array<int> &bdr_marker);
/// Access all integrators added with AddDomainIntegrator().
Array<BilinearFormIntegrator*> *GetDBFI() { return &dom; }
Array<BilinearFormIntegrator*> *GetDBFI() { return &dbfi; }
/// Access all integrators added with AddBoundaryIntegrator().
Array<BilinearFormIntegrator*> *GetBBFI() { return &bdr; }
Array<BilinearFormIntegrator*> *GetBBFI() { return &bbfi; }
/** @brief Access all boundary markers added with AddBoundaryIntegrator().
If no marker was specified when the integrator was added, the
corresponding pointer (to Array<int>) will be NULL. */
Array<Array<int>*> *GetBBFI_Marker() { return &bbfi_marker; }
/// Access all integrators added with AddTraceFaceIntegrator().
Array<BilinearFormIntegrator*> *GetTFBFI() { return &skt; }
Array<BilinearFormIntegrator*> *GetTFBFI() { return &tfbfi; }
/// Access all integrators added with AddBdrTraceFaceIntegrator().
Array<BilinearFormIntegrator*> *GetBTFBFI() { return &btfbfi; }
/** @brief Access all boundary markers added with AddBdrTraceFaceIntegrator().
If no marker was specified when the integrator was added, the
corresponding pointer (to Array<int>) will be NULL. */
Array<Array<int>*> *GetBTFBFI_Marker() { return &btfbfi_marker; }
void operator=(const double a) { *mat = a; }
@@ -613,13 +685,59 @@ public:
MixedBilinearForm becomes an operator on the conforming FE spaces. */
void ConformingAssemble();
void EliminateTrialDofs(Array<int> &bdr_attr_is_ess,
/// Compute the element matrix of the given element
void ComputeElementMatrix(int i, DenseMatrix &elmat);
/// Compute the boundary element matrix of the given boundary element
void ComputeBdrElementMatrix(int i, DenseMatrix &elmat);
/// Assemble the given element matrix
/** The element matrix @a elmat is assembled for the element @a i, i.e.
added to the system matrix. The flag @a skip_zeros skips the zero
elements of the matrix, unless they are breaking the symmetry of
the system matrix.
*/
void AssembleElementMatrix(int i, const DenseMatrix &elmat,
int skip_zeros = 1);
/// Assemble the given element matrix
/** The element matrix @a elmat is assembled for the element @a i, i.e.
added to the system matrix. The vdofs of the element are returned
in @a trial_vdofs and @a test_vdofs. The flag @a skip_zeros skips
the zero elements of the matrix, unless they are breaking the symmetry
of the system matrix.
*/
void AssembleElementMatrix(int i, const DenseMatrix &elmat,
Array<int> &trial_vdofs, Array<int> &test_vdofs,
int skip_zeros = 1);
/// Assemble the given boundary element matrix
/** The boundary element matrix @a elmat is assembled for the boundary
element @a i, i.e. added to the system matrix. The flag @a skip_zeros
skips the zero elements of the matrix, unless they are breaking the
symmetry of the system matrix.
*/
void AssembleBdrElementMatrix(int i, const DenseMatrix &elmat,
int skip_zeros = 1);
/// Assemble the given boundary element matrix
/** The boundary element matrix @a elmat is assembled for the boundary
element @a i, i.e. added to the system matrix. The vdofs of the element
are returned in @a trial_vdofs and @a test_vdofs. The flag @a skip_zeros
skips the zero elements of the matrix, unless they are breaking the
symmetry of the system matrix.
*/
void AssembleBdrElementMatrix(int i, const DenseMatrix &elmat,
Array<int> &trial_vdofs, Array<int> &test_vdofs,
int skip_zeros = 1);
void EliminateTrialDofs(const Array<int> &bdr_attr_is_ess,
const Vector &sol, Vector &rhs);
void EliminateEssentialBCFromTrialDofs(Array<int> &marked_vdofs,
void EliminateEssentialBCFromTrialDofs(const Array<int> &marked_vdofs,
const Vector &sol, Vector &rhs);
virtual void EliminateTestDofs(Array<int> &bdr_attr_is_ess);
virtual void EliminateTestDofs(const Array<int> &bdr_attr_is_ess);
void Update();
@@ -684,7 +802,7 @@ public:
{ AddTraceFaceIntegrator(di); }
/// Access all interpolators added with AddDomainInterpolator().
Array<BilinearFormIntegrator*> *GetDI() { return &dom; }
Array<BilinearFormIntegrator*> *GetDI() { return &dbfi; }
/** @brief Construct the internal matrix representation of the discrete
linear operator. */
+52 -137
View File
@@ -36,16 +36,18 @@ const Operator *BilinearFormExtension::GetRestriction() const
// Data and methods for partially-assembled bilinear forms
PABilinearFormExtension::PABilinearFormExtension(BilinearForm *form) :
BilinearFormExtension(form),
trialFes(a->FESpace()), testFes(a->FESpace()),
localX(trialFes->GetNE() * trialFes->GetFE(0)->GetDof() * trialFes->GetVDim()),
localY( testFes->GetNE() * testFes->GetFE(0)->GetDof() * testFes->GetVDim()),
elem_restrict(new ElemRestriction(*a->FESpace())) { }
PABilinearFormExtension::~PABilinearFormExtension()
PABilinearFormExtension::PABilinearFormExtension(BilinearForm *form)
: BilinearFormExtension(form),
trialFes(a->FESpace()), testFes(a->FESpace())
{
delete elem_restrict;
elem_restrict_lex = trialFes->GetElementRestriction(
ElementDofOrdering::LEXICOGRAPHIC);
if (elem_restrict_lex)
{
localX.SetSize(elem_restrict_lex->Height(), Device::GetMemoryType());
localY.SetSize(elem_restrict_lex->Height(), Device::GetMemoryType());
localY.UseDevice(true); // ensure 'localY = 0.0' is done on device
}
}
void PABilinearFormExtension::Assemble()
@@ -54,7 +56,7 @@ void PABilinearFormExtension::Assemble()
const int integratorCount = integrators.Size();
for (int i = 0; i < integratorCount; ++i)
{
integrators[i]->Assemble(*a->FESpace());
integrators[i]->AssemblePA(*a->FESpace());
}
}
@@ -64,12 +66,13 @@ void PABilinearFormExtension::Update()
height = width = fes->GetVSize();
trialFes = fes;
testFes = fes;
localX.SetSize(trialFes->GetNE() * trialFes->GetFE(0)->GetDof() *
trialFes->GetVDim());
localY.SetSize(testFes->GetNE() * testFes->GetFE(0)->GetDof() *
testFes->GetVDim());
delete elem_restrict;
elem_restrict = new ElemRestriction(*fes);
elem_restrict_lex = trialFes->GetElementRestriction(
ElementDofOrdering::LEXICOGRAPHIC);
if (elem_restrict_lex)
{
localX.SetSize(elem_restrict_lex->Height());
localY.SetSize(elem_restrict_lex->Height());
}
}
void PABilinearFormExtension::FormSystemMatrix(const Array<int> &ess_tdof_list,
@@ -97,140 +100,52 @@ void PABilinearFormExtension::FormLinearSystem(const Array<int> &ess_tdof_list,
void PABilinearFormExtension::Mult(const Vector &x, Vector &y) const
{
Array<BilinearFormIntegrator*> &integrators = *a->GetDBFI();
elem_restrict->Mult(x, localX);
localY = 0.0;
const int iSz = integrators.Size();
for (int i = 0; i < iSz; ++i)
if (elem_restrict_lex)
{
integrators[i]->MultAssembled(localX, localY);
elem_restrict_lex->Mult(x, localX);
localY = 0.0;
for (int i = 0; i < iSz; ++i)
{
integrators[i]->AddMultPA(localX, localY);
}
elem_restrict_lex->MultTranspose(localY, y);
}
else
{
y.UseDevice(true); // typically this is a large vector, so store on device
y = 0.0;
for (int i = 0; i < iSz; ++i)
{
integrators[i]->AddMultPA(x, y);
}
}
elem_restrict->MultTranspose(localY, y);
}
void PABilinearFormExtension::MultTranspose(const Vector &x, Vector &y) const
{
Array<BilinearFormIntegrator*> &integrators = *a->GetDBFI();
elem_restrict->Mult(x, localX);
localY = 0.0;
const int iSz = integrators.Size();
for (int i = 0; i < iSz; ++i)
if (elem_restrict_lex)
{
integrators[i]->MultAssembledTranspose(localX, localY);
}
elem_restrict->MultTranspose(localY, y);
}
ElemRestriction::ElemRestriction(const FiniteElementSpace &f)
: fes(f),
ne(fes.GetNE()),
vdim(fes.GetVDim()),
byvdim(fes.GetOrdering() == Ordering::byVDIM),
ndofs(fes.GetNDofs()),
dof(fes.GetFE(0)->GetDof()),
nedofs(ne*dof),
offsets(ndofs+1),
indices(ne*dof)
{
for (int e = 0; e < ne; ++e)
{
const FiniteElement *fe = fes.GetFE(e);
const TensorBasisElement* el =
dynamic_cast<const TensorBasisElement*>(fe);
if (el) { continue; }
mfem_error("Finite element not supported with partial assembly");
}
const FiniteElement *fe = fes.GetFE(0);
const TensorBasisElement* el = dynamic_cast<const TensorBasisElement*>(fe);
const Array<int> &dof_map = el->GetDofMap();
const bool dof_map_is_identity = (dof_map.Size()==0);
const Table& e2dTable = fes.GetElementToDofTable();
const int* elementMap = e2dTable.GetJ();
// We'll be keeping a count of how many local nodes point to its global dof
for (int i = 0; i <= ndofs; ++i)
{
offsets[i] = 0;
}
for (int e = 0; e < ne; ++e)
{
for (int d = 0; d < dof; ++d)
elem_restrict_lex->Mult(x, localX);
localY = 0.0;
for (int i = 0; i < iSz; ++i)
{
const int gid = elementMap[dof*e + d];
++offsets[gid + 1];
integrators[i]->AddMultTransposePA(localX, localY);
}
elem_restrict_lex->MultTranspose(localY, y);
}
else
{
y.UseDevice(true);
y = 0.0;
for (int i = 0; i < iSz; ++i)
{
integrators[i]->AddMultTransposePA(x, y);
}
}
// Aggregate to find offsets for each global dof
for (int i = 1; i <= ndofs; ++i)
{
offsets[i] += offsets[i - 1];
}
// For each global dof, fill in all local nodes that point to it
for (int e = 0; e < ne; ++e)
{
for (int d = 0; d < dof; ++d)
{
const int did = dof_map_is_identity?d:dof_map[d];
const int gid = elementMap[dof*e + did];
const int lid = dof*e + d;
indices[offsets[gid]++] = lid;
}
}
// We shifted the offsets vector by 1 by using it as a counter
// Now we shift it back.
for (int i = ndofs; i > 0; --i)
{
offsets[i] = offsets[i - 1];
}
offsets[0] = 0;
}
void ElemRestriction::Mult(const Vector& x, Vector& y) const
{
const int vd = vdim;
const bool t = byvdim;
const DeviceArray d_offsets(offsets, ndofs+1);
const DeviceArray d_indices(indices, nedofs);
const DeviceMatrix d_x(x, t?vd:ndofs, t?ndofs:vd);
DeviceMatrix d_y(y, t?vd:nedofs, t?nedofs:vd);
MFEM_FORALL(i, ndofs,
{
const int offset = d_offsets[i];
const int nextOffset = d_offsets[i+1];
for (int c = 0; c < vd; ++c)
{
const double dofValue = d_x(t?c:i,t?i:c);
for (int j = offset; j < nextOffset; ++j)
{
const int idx_j = d_indices[j];
d_y(t?c:idx_j,t?idx_j:c) = dofValue;
}
}
});
}
void ElemRestriction::MultTranspose(const Vector& x, Vector& y) const
{
const int vd = vdim;
const bool t = byvdim;
const DeviceArray d_offsets(offsets, ndofs+1);
const DeviceArray d_indices(indices, nedofs);
const DeviceMatrix d_x(x, t?vd:nedofs, t?nedofs:vd);
DeviceMatrix d_y(y, t?vd:ndofs, t?ndofs:vd);
MFEM_FORALL(i, ndofs,
{
const int offset = d_offsets[i];
const int nextOffset = d_offsets[i + 1];
for (int c = 0; c < vd; ++c)
{
double dofValue = 0;
for (int j = offset; j < nextOffset; ++j)
{
const int idx_j = d_indices[j];
dofValue += d_x(t?c:idx_j,t?idx_j:c);
}
d_y(t?c:i,t?i:c) = dofValue;
}
});
}
} // namespace mfem
+11 -23
View File
@@ -14,32 +14,16 @@
#include "../config/config.hpp"
#include "fespace.hpp"
#include "../general/device.hpp"
namespace mfem
{
class BilinearForm;
/// Element restriction operator
class ElemRestriction: public Operator
{
public:
const FiniteElementSpace &fes;
const int ne;
const int vdim;
const bool byvdim;
const int ndofs;
const int dof;
const int nedofs;
Array<int> offsets;
Array<int> indices;
public:
ElemRestriction(const FiniteElementSpace&);
void Mult(const Vector &x, Vector &y) const;
void MultTranspose(const Vector &x, Vector &y) const;
};
/** @brief Class extending the BilinearForm class to support the different
AssemblyLevel%s. */
class BilinearFormExtension : public Operator
{
protected:
@@ -48,6 +32,9 @@ protected:
public:
BilinearFormExtension(BilinearForm *form);
virtual MemoryClass GetMemoryClass() const
{ return Device::GetMemoryClass(); }
/// Get the finite element space prolongation matrix
virtual const Operator *GetProlongation() const;
@@ -80,6 +67,7 @@ public:
int copy_interior = 0) {}
void Mult(const Vector &x, Vector &y) const {}
void MultTranspose(const Vector &x, Vector &y) const {}
void Update() {}
~FABilinearFormExtension() {}
};
@@ -99,6 +87,7 @@ public:
int copy_interior = 0) {}
void Mult(const Vector &x, Vector &y) const {}
void MultTranspose(const Vector &x, Vector &y) const {}
void Update() {}
~EABilinearFormExtension() {}
};
@@ -106,9 +95,9 @@ public:
class PABilinearFormExtension : public BilinearFormExtension
{
protected:
const FiniteElementSpace *trialFes, *testFes;
const FiniteElementSpace *trialFes, *testFes; // Not owned
mutable Vector localX, localY;
ElemRestriction *elem_restrict;
const Operator *elem_restrict_lex; // Not owned
public:
PABilinearFormExtension(BilinearForm*);
@@ -123,8 +112,6 @@ public:
void Mult(const Vector &x, Vector &y) const;
void MultTranspose(const Vector &x, Vector &y) const;
void Update();
~PABilinearFormExtension();
};
/// Data and methods for matrix-free bilinear forms
@@ -143,6 +130,7 @@ public:
int copy_interior = 0) {}
void Mult(const Vector &x, Vector &y) const {}
void MultTranspose(const Vector &x, Vector &y) const {}
void Update() {}
~MFBilinearFormExtension() {}
};
+55 -102
View File
@@ -19,19 +19,20 @@ using namespace std;
namespace mfem
{
void BilinearFormIntegrator::Assemble(const FiniteElementSpace&)
void BilinearFormIntegrator::AssemblePA(const FiniteElementSpace&)
{
mfem_error ("BilinearFormIntegrator::Assemble (...)\n"
" is not implemented for this class.");
}
void BilinearFormIntegrator::MultAssembled(Vector&, Vector&)
void BilinearFormIntegrator::AddMultPA(const Vector &, Vector &) const
{
mfem_error ("BilinearFormIntegrator::MultAssembled (...)\n"
" is not implemented for this class.");
}
void BilinearFormIntegrator::MultAssembledTranspose(Vector&, Vector&)
void BilinearFormIntegrator::AddMultTransposePA(const Vector &, Vector &) const
{
mfem_error ("BilinearFormIntegrator::MultAssembledTranspose (...)\n"
" is not implemented for this class.");
@@ -378,6 +379,7 @@ void MixedScalarVectorIntegrator::AssembleElementMatrix2(
}
}
void DiffusionIntegrator::AssembleElementMatrix
( const FiniteElement &el, ElementTransformation &Trans,
DenseMatrix &elmat )
@@ -397,29 +399,7 @@ void DiffusionIntegrator::AssembleElementMatrix
#endif
elmat.SetSize(nd);
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int order;
if (el.Space() == FunctionSpace::Pk)
{
order = 2*el.GetOrder() - 2;
}
else
// order = 2*el.GetOrder() - 2; // <-- this seems to work fine too
{
order = 2*el.GetOrder() + dim - 1;
}
if (el.Space() == FunctionSpace::rQk)
{
ir = &RefinedIntRules.Get(el.GetGeomType(), order);
}
else
{
ir = &IntRules.Get(el.GetGeomType(), order);
}
}
const IntegrationRule *ir = IntRule ? IntRule : &GetRule(el, el);
elmat = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
@@ -475,28 +455,7 @@ void DiffusionIntegrator::AssembleElementMatrix2(
#endif
elmat.SetSize(te_nd, tr_nd);
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int order;
if (trial_fe.Space() == FunctionSpace::Pk)
{
order = trial_fe.GetOrder() + test_fe.GetOrder() - 2;
}
else
{
order = trial_fe.GetOrder() + test_fe.GetOrder() + dim - 1;
}
if (trial_fe.Space() == FunctionSpace::rQk)
{
ir = &RefinedIntRules.Get(trial_fe.GetGeomType(), order);
}
else
{
ir = &IntRules.Get(trial_fe.GetGeomType(), order);
}
}
const IntegrationRule *ir = IntRule ? IntRule : &GetRule(trial_fe, test_fe);
elmat = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
@@ -551,29 +510,7 @@ void DiffusionIntegrator::AssembleElementVector(
elvect.SetSize(nd);
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int order;
if (el.Space() == FunctionSpace::Pk)
{
order = 2*el.GetOrder() - 2;
}
else
// order = 2*el.GetOrder() - 2; // <-- this seems to work fine too
{
order = 2*el.GetOrder() + dim - 1;
}
if (el.Space() == FunctionSpace::rQk)
{
ir = &RefinedIntRules.Get(el.GetGeomType(), order);
}
else
{
ir = &IntRules.Get(el.GetGeomType(), order);
}
}
const IntegrationRule *ir = IntRule ? IntRule : &GetRule(el, el);
elvect = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
@@ -733,6 +670,27 @@ double DiffusionIntegrator::ComputeFluxEnergy
return energy;
}
const IntegrationRule &DiffusionIntegrator::GetRule(
const FiniteElement &trial_fe, const FiniteElement &test_fe)
{
int order;
if (trial_fe.Space() == FunctionSpace::Pk)
{
order = trial_fe.GetOrder() + test_fe.GetOrder() - 2;
}
else
{
// order = 2*el.GetOrder() - 2; // <-- this seems to work fine too
order = trial_fe.GetOrder() + test_fe.GetOrder() + trial_fe.GetDim() - 1;
}
if (trial_fe.Space() == FunctionSpace::rQk)
{
return RefinedIntRules.Get(trial_fe.GetGeomType(), order);
}
return IntRules.Get(trial_fe.GetGeomType(), order);
}
void MassIntegrator::AssembleElementMatrix
( const FiniteElement &el, ElementTransformation &Trans,
@@ -748,21 +706,7 @@ void MassIntegrator::AssembleElementMatrix
elmat.SetSize(nd);
shape.SetSize(nd);
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
// int order = 2 * el.GetOrder();
int order = 2 * el.GetOrder() + Trans.OrderW();
if (el.Space() == FunctionSpace::rQk)
{
ir = &RefinedIntRules.Get(el.GetGeomType(), order);
}
else
{
ir = &IntRules.Get(el.GetGeomType(), order);
}
}
const IntegrationRule *ir = IntRule ? IntRule : &GetRule(el, el, Trans);
elmat = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
@@ -797,13 +741,8 @@ void MassIntegrator::AssembleElementMatrix2(
shape.SetSize(tr_nd);
te_shape.SetSize(te_nd);
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int order = trial_fe.GetOrder() + test_fe.GetOrder() + Trans.OrderW();
ir = &IntRules.Get(trial_fe.GetGeomType(), order);
}
const IntegrationRule *ir = IntRule ? IntRule :
&GetRule(trial_fe, test_fe, Trans);
elmat = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
@@ -824,6 +763,20 @@ void MassIntegrator::AssembleElementMatrix2(
}
}
const IntegrationRule &MassIntegrator::GetRule(const FiniteElement &trial_fe,
const FiniteElement &test_fe,
ElementTransformation &Trans)
{
// int order = trial_fe.GetOrder() + test_fe.GetOrder();
const int order = trial_fe.GetOrder() + test_fe.GetOrder() + Trans.OrderW();
if (trial_fe.Space() == FunctionSpace::rQk)
{
return RefinedIntRules.Get(trial_fe.GetGeomType(), order);
}
return IntRules.Get(trial_fe.GetGeomType(), order);
}
void BoundaryMassIntegrator::AssembleFaceMatrix(
const FiniteElement &el1, const FiniteElement &el2,
@@ -895,7 +848,7 @@ void ConvectionIntegrator::AssembleElementMatrix(
ir = &IntRules.Get(el.GetGeomType(), order);
}
Q.Eval(Q_ir, Trans, *ir);
Q->Eval(Q_ir, Trans, *ir);
elmat = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
@@ -936,7 +889,7 @@ void GroupConvectionIntegrator::AssembleElementMatrix(
ir = &IntRules.Get(el.GetGeomType(), order);
}
Q.Eval(Q_nodal, Trans, el.GetNodes()); // sets the size of Q_nodal
Q->Eval(Q_nodal, Trans, el.GetNodes()); // sets the size of Q_nodal
elmat = 0.0;
for (int i = 0; i < ir->GetNPoints(); i++)
@@ -1417,7 +1370,7 @@ void DerivativeIntegrator::AssembleElementMatrix2 (
dshapedxi(l) = dshapedxt(l,xi);
}
shape *= Q.Eval(Trans,ip) * det * ip.weight;
shape *= Q->Eval(Trans,ip) * det * ip.weight;
AddMultVWt (shape, dshapedxi, elmat);
}
}
@@ -3263,7 +3216,7 @@ ScalarProductInterpolator::AssembleElementMatrix2(const FiniteElement &dom_fe,
ElementTransformation &Trans,
DenseMatrix &elmat)
{
internal::ShapeCoefficient dom_shape_coeff(Q, dom_fe);
internal::ShapeCoefficient dom_shape_coeff(*Q, dom_fe);
elmat.SetSize(ran_fe.GetDof(),dom_fe.GetDof());
@@ -3298,7 +3251,7 @@ ScalarVectorProductInterpolator::AssembleElementMatrix2(
}
};
VShapeCoefficient dom_shape_coeff(Q, dom_fe, Trans.GetSpaceDim());
VShapeCoefficient dom_shape_coeff(*Q, dom_fe, Trans.GetSpaceDim());
elmat.SetSize(ran_fe.GetDof(),dom_fe.GetDof());
@@ -3336,7 +3289,7 @@ VectorScalarProductInterpolator::AssembleElementMatrix2(
}
};
VecShapeCoefficient dom_shape_coeff(VQ, dom_fe);
VecShapeCoefficient dom_shape_coeff(*VQ, dom_fe);
elmat.SetSize(ran_fe.GetDof(),dom_fe.GetDof());
@@ -3383,11 +3336,11 @@ VectorCrossProductInterpolator::AssembleElementMatrix2(
}
};
VCrossVShapeCoefficient dom_shape_coeff(VQ, dom_fe);
VCrossVShapeCoefficient dom_shape_coeff(*VQ, dom_fe);
if (ran_fe.GetRangeType() == FiniteElement::SCALAR)
{
elmat.SetSize(ran_fe.GetDof()*VQ.GetVDim(),dom_fe.GetDof());
elmat.SetSize(ran_fe.GetDof()*VQ->GetVDim(),dom_fe.GetDof());
}
else
{
@@ -3436,7 +3389,7 @@ VectorInnerProductInterpolator::AssembleElementMatrix2(
ElementTransformation &Trans,
DenseMatrix &elmat)
{
internal::VDotVShapeCoefficient dom_shape_coeff(VQ, dom_fe);
internal::VDotVShapeCoefficient dom_shape_coeff(*VQ, dom_fe);
elmat.SetSize(ran_fe.GetDof(),dom_fe.GetDof());
+128 -63
View File
@@ -15,7 +15,6 @@
#include "../config/config.hpp"
#include "nonlininteg.hpp"
#include "fespace.hpp"
#include "bilininteg_ext.hpp"
namespace mfem
{
@@ -23,19 +22,45 @@ namespace mfem
/// Abstract base class BilinearFormIntegrator
class BilinearFormIntegrator : public NonlinearFormIntegrator
{
public:
BilinearFormIntegrator(const IntegrationRule *ir = NULL) :
NonlinearFormIntegrator(ir) { }
protected:
BilinearFormIntegrator(const IntegrationRule *ir = NULL)
: NonlinearFormIntegrator(ir) { }
public:
// TODO: add support for other assembly levels (in addition to PA) and their
// actions.
// TODO: for mixed meshes the quadrature rules to be used by methods like
// AssemblePA() can be given as a QuadratureSpace, e.g. using a new method:
// SetQuadratureSpace().
// TODO: the methods for the various assembly levels make sense even in the
// base class NonlinearFormIntegrator, except that not all assembly levels
// make sense for the action of the nonlinear operator (but they all make
// sense for its Jacobian).
/// Method defining partial assembly.
virtual void Assemble(const FiniteElementSpace&);
/** The result of the partial assembly is stored internally so that it can be
used later in the methods AddMultPA() and AddMultTransposePA(). */
virtual void AssemblePA(const FiniteElementSpace &fes);
/// Method for partially assembled action.
virtual void MultAssembled(Vector&, Vector&);
/** Perform the action of integrator on the input @a x and add the result to
the output @a y. Both @a x and @a y are E-vectors, i.e. they represent
the element-wise discontinuous version of the FE space.
This method can be called only after the method AssemblePA() has been
called. */
virtual void AddMultPA(const Vector &x, Vector &y) const;
/// Method for partially assembled transposed action.
virtual void MultAssembledTranspose(Vector&, Vector&);
/** Perform the transpose action of integrator on the input @a x and add the
result to the output @a y. Both @a x and @a y are E-vectors, i.e. they
represent the element-wise discontinuous version of the FE space.
This method can be called only after the method AssemblePA() has been
called. */
virtual void AddMultTransposePA(const Vector &x, Vector &y) const;
/// Given a particular Finite Element computes the element matrix elmat.
virtual void AssembleElementMatrix(const FiniteElement &el,
@@ -284,10 +309,10 @@ protected:
Vector & shape)
{ trial_fe.CalcPhysShape(Trans, shape); }
private:
Coefficient *Q;
private:
#ifndef MFEM_THREAD_SAFE
Vector test_shape;
Vector trial_shape;
@@ -358,13 +383,13 @@ protected:
DenseMatrix & shape)
{ trial_fe.CalcVShape(Trans, shape); }
private:
Coefficient *Q;
VectorCoefficient *VQ;
VectorCoefficient *DQ;
MatrixCoefficient *MQ;
private:
#ifndef MFEM_THREAD_SAFE
Vector V;
Vector D;
@@ -439,12 +464,12 @@ protected:
Vector & shape)
{ scalar_fe.CalcPhysShape(Trans, shape); }
private:
VectorCoefficient *VQ;
bool transpose;
bool cross_2d; // In 2D use a cross product rather than a dot product
private:
#ifndef MFEM_THREAD_SAFE
Vector V;
DenseMatrix vshape;
@@ -1637,27 +1662,34 @@ protected:
can be a scalar or a matrix coefficient. */
class DiffusionIntegrator: public BilinearFormIntegrator
{
protected:
Coefficient *Q;
MatrixCoefficient *MQ;
private:
Vector vec, pointflux, shape;
#ifndef MFEM_THREAD_SAFE
DenseMatrix dshape, dshapedxt, invdfdx, mq;
DenseMatrix te_dshape, te_dshapedxt;
#endif
Coefficient *Q;
MatrixCoefficient *MQ;
// PA extension
DofToQuad *maps;
GeometryExtension *geom;
const DofToQuad *maps; ///< Not owned
const GeometricFactors *geom; ///< Not owned
int dim, ne, dofs1D, quad1D;
Vector pa_data;
public:
/// Construct a diffusion integrator with coefficient Q = 1
DiffusionIntegrator() { Q = NULL; MQ = NULL; maps = NULL; geom = NULL; }
/// Construct a diffusion integrator with a scalar coefficient q
DiffusionIntegrator (Coefficient &q) : Q(&q) { MQ = NULL; maps = NULL; geom = NULL; }
DiffusionIntegrator(Coefficient &q)
: Q(&q) { MQ = NULL; maps = NULL; geom = NULL; }
/// Construct a diffusion integrator with a matrix coefficient q
DiffusionIntegrator (MatrixCoefficient &q) : MQ(&q) { Q = NULL; maps = NULL; geom = NULL; }
DiffusionIntegrator(MatrixCoefficient &q)
: MQ(&q) { Q = NULL; maps = NULL; geom = NULL; }
/** Given a particular Finite Element
computes the element stiffness matrix elmat. */
@@ -1685,11 +1717,12 @@ public:
ElementTransformation &Trans,
Vector &flux, Vector *d_energy = NULL);
/// PA extension
virtual void Assemble(const FiniteElementSpace&);
virtual void MultAssembled(Vector&, Vector&);
virtual void AssemblePA(const FiniteElementSpace&);
virtual ~DiffusionIntegrator();
virtual void AddMultPA(const Vector&, Vector&) const;
static const IntegrationRule &GetRule(const FiniteElement &trial_fe,
const FiniteElement &test_fe);
};
/** Class for local mass matrix assembling a(u,v) := (Q u, v) */
@@ -1701,13 +1734,15 @@ protected:
#endif
Coefficient *Q;
// PA extension
Vector vec;
DofToQuad *maps;
GeometryExtension *geom;
Vector pa_data;
const DofToQuad *maps; ///< Not owned
const GeometricFactors *geom; ///< Not owned
int dim, ne, nq, dofs1D, quad1D;
public:
MassIntegrator(const IntegrationRule *ir = NULL)
: BilinearFormIntegrator(ir) { Q = NULL; maps = NULL; geom = NULL; }
/// Construct a mass integrator with coefficient q
MassIntegrator(Coefficient &q, const IntegrationRule *ir = NULL)
: BilinearFormIntegrator(ir), Q(&q) { maps = NULL; geom = NULL; }
@@ -1721,11 +1756,14 @@ public:
const FiniteElement &test_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
/// PA extension
virtual void Assemble(const FiniteElementSpace&);
virtual void MultAssembled(Vector&, Vector&);
virtual ~MassIntegrator();
virtual void AssemblePA(const FiniteElementSpace&);
virtual void AddMultPA(const Vector&, Vector&) const;
static const IntegrationRule &GetRule(const FiniteElement &trial_fe,
const FiniteElement &test_fe,
ElementTransformation &Trans);
};
class BoundaryMassIntegrator : public MassIntegrator
@@ -1744,17 +1782,19 @@ public:
/// alpha (q . grad u, v)
class ConvectionIntegrator : public BilinearFormIntegrator
{
protected:
VectorCoefficient *Q;
double alpha;
private:
#ifndef MFEM_THREAD_SAFE
DenseMatrix dshape, adjJ, Q_ir;
Vector shape, vec2, BdFidxT;
#endif
VectorCoefficient &Q;
double alpha;
public:
ConvectionIntegrator(VectorCoefficient &q, double a = 1.0)
: Q(q) { alpha = a; }
: Q(&q) { alpha = a; }
virtual void AssembleElementMatrix(const FiniteElement &,
ElementTransformation &,
DenseMatrix &);
@@ -1763,15 +1803,17 @@ public:
/// alpha (q . grad u, v) using the "group" FE discretization
class GroupConvectionIntegrator : public BilinearFormIntegrator
{
protected:
VectorCoefficient *Q;
double alpha;
private:
DenseMatrix dshape, adjJ, Q_nodal, grad;
Vector shape;
VectorCoefficient &Q;
double alpha;
public:
GroupConvectionIntegrator(VectorCoefficient &q, double a = 1.0)
: Q(q) { alpha = a; }
: Q(&q) { alpha = a; }
virtual void AssembleElementMatrix(const FiniteElement &,
ElementTransformation &,
DenseMatrix &);
@@ -1787,16 +1829,17 @@ private:
Vector shape, te_shape, vec;
DenseMatrix partelmat;
DenseMatrix mcoeff;
int Q_order;
protected:
Coefficient *Q;
VectorCoefficient *VQ;
MatrixCoefficient *MQ;
int Q_order;
public:
/// Construct an integrator with coefficient 1.0
VectorMassIntegrator()
: vdim(-1), Q(NULL), VQ(NULL), MQ(NULL), Q_order(0) { }
: vdim(-1), Q_order(0), Q(NULL), VQ(NULL), MQ(NULL) { }
/** Construct an integrator with scalar coefficient q.
If possible, save memory by using a scalar integrator since
the resulting matrix is block diagonal with the same diagonal
@@ -1835,11 +1878,14 @@ public:
does NOT depend on the ElementTransformation Trans. */
class VectorFEDivergenceIntegrator : public BilinearFormIntegrator
{
private:
protected:
Coefficient *Q;
private:
#ifndef MFEM_THREAD_SAFE
Vector divshape, shape;
#endif
public:
VectorFEDivergenceIntegrator() { Q = NULL; }
VectorFEDivergenceIntegrator(Coefficient &q) { Q = &q; }
@@ -1857,14 +1903,17 @@ public:
This is equivalent to a weak divergence of the Nedelec basis functions. */
class VectorFEWeakDivergenceIntegrator: public BilinearFormIntegrator
{
private:
protected:
Coefficient *Q;
private:
#ifndef MFEM_THREAD_SAFE
DenseMatrix dshape;
DenseMatrix dshapedxt;
DenseMatrix vshape;
DenseMatrix invdfdx;
#endif
public:
VectorFEWeakDivergenceIntegrator() { Q = NULL; }
VectorFEWeakDivergenceIntegrator(Coefficient &q) { Q = &q; }
@@ -1881,13 +1930,16 @@ public:
test spaces are switched, assembles the form (u, curl v). */
class VectorFECurlIntegrator: public BilinearFormIntegrator
{
private:
protected:
Coefficient *Q;
private:
#ifndef MFEM_THREAD_SAFE
DenseMatrix curlshapeTrial;
DenseMatrix vshapeTest;
DenseMatrix curlshapeTrial_dFT;
#endif
public:
VectorFECurlIntegrator() { Q = NULL; }
VectorFECurlIntegrator(Coefficient &q) { Q = &q; }
@@ -1900,17 +1952,19 @@ public:
DenseMatrix &elmat);
};
/// Class for integrating (Q D_i(u), v); u and v are scalars
class DerivativeIntegrator : public BilinearFormIntegrator
{
protected:
Coefficient* Q;
private:
Coefficient & Q;
int xi;
DenseMatrix dshape, dshapedxt, invdfdx;
Vector shape, dshapedxi;
public:
DerivativeIntegrator(Coefficient &q, int i) : Q(q), xi(i) { }
DerivativeIntegrator(Coefficient &q, int i) : Q(&q), xi(i) { }
virtual void AssembleElementMatrix(const FiniteElement &el,
ElementTransformation &Trans,
DenseMatrix &elmat)
@@ -1930,6 +1984,8 @@ private:
DenseMatrix curlshape, curlshape_dFt, M;
DenseMatrix vshape, projcurl;
#endif
protected:
Coefficient *Q;
MatrixCoefficient *MQ;
@@ -1963,6 +2019,8 @@ private:
#ifndef MFEM_THREAD_SAFE
DenseMatrix dshape_hat, dshape, curlshape, Jadj, grad_hat, grad;
#endif
protected:
Coefficient *Q;
public:
@@ -1984,9 +2042,6 @@ public:
class VectorFEMassIntegrator: public BilinearFormIntegrator
{
private:
Coefficient *Q;
VectorCoefficient *VQ;
MatrixCoefficient *MQ;
void Init(Coefficient *q, VectorCoefficient *vq, MatrixCoefficient *mq)
{ Q = q; VQ = vq; MQ = mq; }
@@ -1998,6 +2053,11 @@ private:
DenseMatrix trial_vshape;
#endif
protected:
Coefficient *Q;
VectorCoefficient *VQ;
MatrixCoefficient *MQ;
public:
VectorFEMassIntegrator() { Init(NULL, NULL, NULL); }
VectorFEMassIntegrator(Coefficient *_q) { Init(_q, NULL, NULL); }
@@ -2020,9 +2080,10 @@ public:
scalar FE space; p is also in a (different) scalar FE space. */
class VectorDivergenceIntegrator : public BilinearFormIntegrator
{
private:
protected:
Coefficient *Q;
private:
Vector shape;
Vector divshape;
DenseMatrix dshape;
@@ -2043,9 +2104,10 @@ public:
/// (Q div u, div v) for RT elements
class DivDivIntegrator: public BilinearFormIntegrator
{
private:
protected:
Coefficient *Q;
private:
#ifndef MFEM_THREAD_SAFE
Vector divshape;
#endif
@@ -2067,9 +2129,10 @@ public:
diffusion matrix in each diagonal block. */
class VectorDiffusionIntegrator : public BilinearFormIntegrator
{
private:
protected:
Coefficient *Q;
private:
DenseMatrix Jinv;
DenseMatrix dshape;
DenseMatrix gshape;
@@ -2094,10 +2157,11 @@ public:
using multiple copies of a scalar FE space. */
class ElasticityIntegrator : public BilinearFormIntegrator
{
private:
protected:
double q_lambda, q_mu;
Coefficient *lambda, *mu;
private:
#ifndef MFEM_THREAD_SAFE
Vector shape;
DenseMatrix dshape, gshape, pelmat;
@@ -2154,11 +2218,12 @@ public:
points. */
class DGTraceIntegrator : public BilinearFormIntegrator
{
private:
protected:
Coefficient *rho;
VectorCoefficient *u;
double alpha, beta;
private:
Vector shape1, shape2;
public:
@@ -2445,7 +2510,7 @@ public:
class ScalarProductInterpolator : public DiscreteInterpolator
{
public:
ScalarProductInterpolator(Coefficient & sc) : Q(sc) { }
ScalarProductInterpolator(Coefficient & sc) : Q(&sc) { }
virtual void AssembleElementMatrix2(const FiniteElement &dom_fe,
const FiniteElement &ran_fe,
@@ -2453,7 +2518,7 @@ public:
DenseMatrix &elmat);
protected:
Coefficient &Q;
Coefficient *Q;
};
/** Interpolator of a scalar coefficient multiplied by a vector field onto
@@ -2463,14 +2528,14 @@ class ScalarVectorProductInterpolator : public DiscreteInterpolator
{
public:
ScalarVectorProductInterpolator(Coefficient & sc)
: Q(sc) { }
: Q(&sc) { }
virtual void AssembleElementMatrix2(const FiniteElement &dom_fe,
const FiniteElement &ran_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
protected:
Coefficient &Q;
Coefficient *Q;
};
/** Interpolator of a vector coefficient multiplied by a scalar field onto
@@ -2480,14 +2545,14 @@ class VectorScalarProductInterpolator : public DiscreteInterpolator
{
public:
VectorScalarProductInterpolator(VectorCoefficient & vc)
: VQ(vc) { }
: VQ(&vc) { }
virtual void AssembleElementMatrix2(const FiniteElement &dom_fe,
const FiniteElement &ran_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
protected:
VectorCoefficient &VQ;
VectorCoefficient *VQ;
};
/** Interpolator of the cross product between a vector coefficient and an
@@ -2497,14 +2562,14 @@ class VectorCrossProductInterpolator : public DiscreteInterpolator
{
public:
VectorCrossProductInterpolator(VectorCoefficient & vc)
: VQ(vc) { }
: VQ(&vc) { }
virtual void AssembleElementMatrix2(const FiniteElement &nd_fe,
const FiniteElement &rt_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
protected:
VectorCoefficient &VQ;
VectorCoefficient *VQ;
};
/** Interpolator of the inner product between a vector coefficient and an
@@ -2513,14 +2578,14 @@ protected:
class VectorInnerProductInterpolator : public DiscreteInterpolator
{
public:
VectorInnerProductInterpolator(VectorCoefficient & vc) : VQ(vc) { }
VectorInnerProductInterpolator(VectorCoefficient & vc) : VQ(&vc) { }
virtual void AssembleElementMatrix2(const FiniteElement &rt_fe,
const FiniteElement &l2_fe,
ElementTransformation &Trans,
DenseMatrix &elmat);
protected:
VectorCoefficient &VQ;
VectorCoefficient *VQ;
};
}
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
-80
View File
@@ -1,80 +0,0 @@
// Copyright (c) 2010, Lawrence Livermore National Security, LLC. Produced at
// the Lawrence Livermore National Laboratory. LLNL-CODE-443211. All Rights
// reserved. See file COPYRIGHT for details.
//
// This file is part of the MFEM library. For more information and source code
// availability see http://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the GNU Lesser General Public License (as published by the Free
// Software Foundation) version 2.1 dated February 1999.
#ifndef MFEM_BILININTEG_EXT
#define MFEM_BILININTEG_EXT
#include "fespace.hpp"
namespace mfem
{
/// GeometryExtension
class GeometryExtension
{
public:
Array<int> eMap;
Array<double> nodes;
Array<double> X, J, invJ, detJ;
static GeometryExtension* Get(const FiniteElementSpace&,
const IntegrationRule&);
static GeometryExtension* Get(const FiniteElementSpace&,
const IntegrationRule&,
const Vector&);
static void ReorderByVDim(const GridFunction*);
static void ReorderByNodes(const GridFunction*);
};
/// DofToQuad
class DofToQuad
{
private:
std::string hash;
public:
~DofToQuad();
void operator=(DofToQuad&);
void operator=(DofToQuad const&);
public:
Array<double> W, B, G, Bt, Gt;
public:
static DofToQuad* Get(const FiniteElementSpace&,
const IntegrationRule&,
const bool = false);
static DofToQuad* Get(const FiniteElementSpace&,
const FiniteElementSpace&,
const IntegrationRule&,
const bool = false);
static DofToQuad* Get(const FiniteElement&,
const FiniteElement&,
const IntegrationRule&,
const bool = false);
static DofToQuad* GetTensorMaps(const FiniteElement&,
const FiniteElement&,
const IntegrationRule&,
const bool = false);
static DofToQuad* GetD2QTensorMaps(const FiniteElement&,
const IntegrationRule&,
const bool = false);
static DofToQuad* GetSimplexMaps(const FiniteElement&,
const IntegrationRule&,
const bool = false);
static DofToQuad* GetSimplexMaps(const FiniteElement&,
const FiniteElement&,
const IntegrationRule&,
const bool = false);
static DofToQuad* GetD2QSimplexMaps(const FiniteElement&,
const IntegrationRule&,
const bool = false);
};
}
#endif
+789
View File
@@ -0,0 +1,789 @@
// Copyright (c) 2010, Lawrence Livermore National Security, LLC. Produced at
// the Lawrence Livermore National Laboratory. LLNL-CODE-443211. All Rights
// reserved. See file COPYRIGHT for details.
//
// This file is part of the MFEM library. For more information and source code
// availability see http://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the GNU Lesser General Public License (as published by the Free
// Software Foundation) version 2.1 dated February 1999.
#include "../general/forall.hpp"
#include "bilininteg.hpp"
#include "gridfunc.hpp"
using namespace std;
namespace mfem
{
// PA Mass Integrator
// PA Mass Assemble kernel
void MassIntegrator::AssemblePA(const FiniteElementSpace &fes)
{
// Assuming the same element type
Mesh *mesh = fes.GetMesh();
if (mesh->GetNE() == 0) { return; }
const FiniteElement &el = *fes.GetFE(0);
ElementTransformation *T = mesh->GetElementTransformation(0);
const IntegrationRule *ir = IntRule ? IntRule : &GetRule(el, el, *T);
dim = mesh->Dimension();
ne = fes.GetMesh()->GetNE();
nq = ir->GetNPoints();
geom = mesh->GetGeometricFactors(*ir, GeometricFactors::COORDINATES |
GeometricFactors::JACOBIANS);
maps = &el.GetDofToQuad(*ir, DofToQuad::TENSOR);
dofs1D = maps->ndof;
quad1D = maps->nqpt;
pa_data.SetSize(ne*nq, Device::GetMemoryType());
ConstantCoefficient *const_coeff = dynamic_cast<ConstantCoefficient*>(Q);
// TODO: other types of coefficients ...
if (dim==1) { MFEM_ABORT("Not supported yet... stay tuned!"); }
if (dim==2)
{
double constant = 0.0;
if (const_coeff)
{
constant = const_coeff->constant;
}
else
{
MFEM_ABORT("Coefficient type not supported");
}
const int NE = ne;
const int NQ = nq;
auto w = ir->GetWeights().Read();
auto J = Reshape(geom->J.Read(), NQ,2,2,NE);
auto v = Reshape(pa_data.Write(), NQ, NE);
MFEM_FORALL(e, NE,
{
for (int q = 0; q < NQ; ++q)
{
const double J11 = J(q,0,0,e);
const double J12 = J(q,1,0,e);
const double J21 = J(q,0,1,e);
const double J22 = J(q,1,1,e);
const double detJ = (J11*J22)-(J21*J12);
v(q,e) = w[q] * constant * detJ;
}
});
}
if (dim==3)
{
double constant = 0.0;
if (const_coeff)
{
constant = const_coeff->constant;
}
else
{
MFEM_ABORT("Coefficient type not supported");
}
const int NE = ne;
const int NQ = nq;
auto W = ir->GetWeights().Read();
auto J = Reshape(geom->J.Read(), NQ,3,3,NE);
auto v = Reshape(pa_data.Write(), NQ,NE);
MFEM_FORALL(e, NE,
{
for (int q = 0; q < NQ; ++q)
{
const double J11 = J(q,0,0,e), J12 = J(q,0,1,e), J13 = J(q,0,2,e);
const double J21 = J(q,1,0,e), J22 = J(q,1,1,e), J23 = J(q,1,2,e);
const double J31 = J(q,2,0,e), J32 = J(q,2,1,e), J33 = J(q,2,2,e);
const double detJ = J11 * (J22 * J33 - J32 * J23) -
/* */ J21 * (J12 * J33 - J32 * J13) +
/* */ J31 * (J12 * J23 - J22 * J13);
v(q,e) = W[q] * constant * detJ;
}
});
}
}
#ifdef MFEM_USE_OCCA
// OCCA PA Mass Apply 2D kernel
static void OccaPAMassApply2D(const int D1D,
const int Q1D,
const int NE,
const Array<double> &B,
const Array<double> &Bt,
const Vector &op,
const Vector &x,
Vector &y)
{
occa::properties props;
props["defines/D1D"] = D1D;
props["defines/Q1D"] = Q1D;
const occa::memory o_B = OccaMemoryRead(B.GetMemory(), B.Size());
const occa::memory o_Bt = OccaMemoryRead(Bt.GetMemory(), Bt.Size());
const occa::memory o_op = OccaMemoryRead(op.GetMemory(), op.Size());
const occa::memory o_x = OccaMemoryRead(x.GetMemory(), x.Size());
occa::memory o_y = OccaMemoryReadWrite(y.GetMemory(), y.Size());
const occa_id_t id = std::make_pair(D1D,Q1D);
if (!Device::Allows(Backend::OCCA_CUDA))
{
static occa_kernel_t OccaMassApply2D_cpu;
if (OccaMassApply2D_cpu.find(id) == OccaMassApply2D_cpu.end())
{
const occa::kernel MassApply2D_CPU =
mfem::OccaDev().buildKernel("occa://mfem/fem/occa.okl",
"MassApply2D_CPU", props);
OccaMassApply2D_cpu.emplace(id, MassApply2D_CPU);
}
OccaMassApply2D_cpu.at(id)(NE, o_B, o_Bt, o_op, o_x, o_y);
}
else
{
static occa_kernel_t OccaMassApply2D_gpu;
if (OccaMassApply2D_gpu.find(id) == OccaMassApply2D_gpu.end())
{
const occa::kernel MassApply2D_GPU =
mfem::OccaDev().buildKernel("occa://mfem/fem/occa.okl",
"MassApply2D_GPU", props);
OccaMassApply2D_gpu.emplace(id, MassApply2D_GPU);
}
OccaMassApply2D_gpu.at(id)(NE, o_B, o_Bt, o_op, o_x, o_y);
}
}
// OCCA PA Mass Apply 3D kernel
static void OccaPAMassApply3D(const int D1D,
const int Q1D,
const int NE,
const Array<double> &B,
const Array<double> &Bt,
const Vector &op,
const Vector &x,
Vector &y)
{
occa::properties props;
props["defines/D1D"] = D1D;
props["defines/Q1D"] = Q1D;
const occa::memory o_B = OccaMemoryRead(B.GetMemory(), B.Size());
const occa::memory o_Bt = OccaMemoryRead(Bt.GetMemory(), Bt.Size());
const occa::memory o_op = OccaMemoryRead(op.GetMemory(), op.Size());
const occa::memory o_x = OccaMemoryRead(x.GetMemory(), x.Size());
occa::memory o_y = OccaMemoryReadWrite(y.GetMemory(), y.Size());
const occa_id_t id = std::make_pair(D1D,Q1D);
if (!Device::Allows(Backend::OCCA_CUDA))
{
static occa_kernel_t OccaMassApply3D_cpu;
if (OccaMassApply3D_cpu.find(id) == OccaMassApply3D_cpu.end())
{
const occa::kernel MassApply3D_CPU =
mfem::OccaDev().buildKernel("occa://mfem/fem/occa.okl",
"MassApply3D_CPU", props);
OccaMassApply3D_cpu.emplace(id, MassApply3D_CPU);
}
OccaMassApply3D_cpu.at(id)(NE, o_B, o_Bt, o_op, o_x, o_y);
}
else
{
static occa_kernel_t OccaMassApply3D_gpu;
if (OccaMassApply3D_gpu.find(id) == OccaMassApply3D_gpu.end())
{
const occa::kernel MassApply3D_GPU =
mfem::OccaDev().buildKernel("occa://mfem/fem/occa.okl",
"MassApply3D_GPU", props);
OccaMassApply3D_gpu.emplace(id, MassApply3D_GPU);
}
OccaMassApply3D_gpu.at(id)(NE, o_B, o_Bt, o_op, o_x, o_y);
}
}
#endif // MFEM_USE_OCCA
template<const int T_D1D = 0,
const int T_Q1D = 0>
static void PAMassApply2D(const int NE,
const Array<double> &B_,
const Array<double> &Bt_,
const Vector &op_,
const Vector &x_,
Vector &y_,
const int d1d = 0,
const int q1d = 0)
{
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
MFEM_VERIFY(D1D <= MAX_D1D, "");
MFEM_VERIFY(Q1D <= MAX_Q1D, "");
auto B = Reshape(B_.Read(), Q1D, D1D);
auto Bt = Reshape(Bt_.Read(), D1D, Q1D);
auto op = Reshape(op_.Read(), Q1D, Q1D, NE);
auto x = Reshape(x_.Read(), D1D, D1D, NE);
auto y = Reshape(y_.ReadWrite(), D1D, D1D, NE);
MFEM_FORALL(e, NE,
{
const int D1D = T_D1D ? T_D1D : d1d; // nvcc workaround
const int Q1D = T_Q1D ? T_Q1D : q1d;
// the following variables are evaluated at compile time
constexpr int max_D1D = T_D1D ? T_D1D : MAX_D1D;
constexpr int max_Q1D = T_Q1D ? T_Q1D : MAX_Q1D;
double sol_xy[max_Q1D][max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] = 0.0;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
double sol_x[max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
sol_x[qy] = 0.0;
}
for (int dx = 0; dx < D1D; ++dx)
{
const double s = x(dx,dy,e);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] += B(qx,dx)* s;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
const double d2q = B(qy,dy);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] += d2q * sol_x[qx];
}
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] *= op(qx,qy,e);
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
double sol_x[max_D1D];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] = 0.0;
}
for (int qx = 0; qx < Q1D; ++qx)
{
const double s = sol_xy[qy][qx];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] += Bt(dx,qx) * s;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
const double q2d = Bt(dy,qy);
for (int dx = 0; dx < D1D; ++dx)
{
y(dx,dy,e) += q2d * sol_x[dx];
}
}
}
});
}
template<const int T_D1D = 0,
const int T_Q1D = 0,
const int T_NBZ = 0>
static void SmemPAMassApply2D(const int NE,
const Array<double> &b_,
const Array<double> &bt_,
const Vector &op_,
const Vector &x_,
Vector &y_,
const int d1d = 0,
const int q1d = 0)
{
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int NBZ = T_NBZ ? T_NBZ : 1;
constexpr int MQ1 = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
MFEM_VERIFY(D1D <= MD1, "");
MFEM_VERIFY(Q1D <= MQ1, "");
auto b = Reshape(b_.Read(), Q1D, D1D);
auto op = Reshape(op_.Read(), Q1D, Q1D, NE);
auto x = Reshape(x_.Read(), D1D, D1D, NE);
auto y = Reshape(y_.ReadWrite(), D1D, D1D, NE);
MFEM_FORALL_2D(e, NE, Q1D, Q1D, NBZ,
{
const int tidz = MFEM_THREAD_ID(z);
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int NBZ = T_NBZ ? T_NBZ : 1;
constexpr int MQ1 = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
constexpr int MDQ = (MQ1 > MD1) ? MQ1 : MD1;
MFEM_SHARED double BBt[MQ1*MD1];
double (*B)[MD1] = (double (*)[MD1]) BBt;
double (*Bt)[MQ1] = (double (*)[MQ1]) BBt;
MFEM_SHARED double sm0[NBZ][MDQ*MDQ];
MFEM_SHARED double sm1[NBZ][MDQ*MDQ];
double (*X)[MD1] = (double (*)[MD1]) (sm0 + tidz);
double (*DQ)[MQ1] = (double (*)[MQ1]) (sm1 + tidz);
double (*QQ)[MQ1] = (double (*)[MQ1]) (sm0 + tidz);
double (*QD)[MD1] = (double (*)[MD1]) (sm1 + tidz);
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
X[dy][dx] = x(dx,dy,e);
}
}
if (tidz == 0)
{
MFEM_FOREACH_THREAD(d,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
B[q][d] = b(q,d);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double dq = 0.0;
for (int dx = 0; dx < D1D; ++dx)
{
dq += X[dy][dx] * B[qx][dx];
}
DQ[dy][qx] = dq;
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double qq = 0.0;
for (int dy = 0; dy < D1D; ++dy)
{
qq += DQ[dy][qx] * B[qy][dy];
}
QQ[qy][qx] = qq * op(qx, qy, e);
}
}
MFEM_SYNC_THREAD;
if (tidz == 0)
{
MFEM_FOREACH_THREAD(d,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
Bt[d][q] = b(q,d);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double dq = 0.0;
for (int qx = 0; qx < Q1D; ++qx)
{
dq += QQ[qy][qx] * Bt[dx][qx];
}
QD[qy][dx] = dq;
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double dd = 0.0;
for (int qy = 0; qy < Q1D; ++qy)
{
dd += (QD[qy][dx] * Bt[dy][qy]);
}
y(dx, dy, e) += dd;
}
}
});
}
template<const int T_D1D = 0,
const int T_Q1D = 0>
static void PAMassApply3D(const int NE,
const Array<double> &B_,
const Array<double> &Bt_,
const Vector &op_,
const Vector &x_,
Vector &y_,
const int d1d = 0,
const int q1d = 0)
{
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
MFEM_VERIFY(D1D <= MAX_D1D, "");
MFEM_VERIFY(Q1D <= MAX_Q1D, "");
auto B = Reshape(B_.Read(), Q1D, D1D);
auto Bt = Reshape(Bt_.Read(), D1D, Q1D);
auto op = Reshape(op_.Read(), Q1D, Q1D, Q1D, NE);
auto x = Reshape(x_.Read(), D1D, D1D, D1D, NE);
auto y = Reshape(y_.ReadWrite(), D1D, D1D, D1D, NE);
MFEM_FORALL(e, NE,
{
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int max_D1D = T_D1D ? T_D1D : MAX_D1D;
constexpr int max_Q1D = T_Q1D ? T_Q1D : MAX_Q1D;
double sol_xyz[max_Q1D][max_Q1D][max_Q1D];
for (int qz = 0; qz < Q1D; ++qz)
{
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] = 0.0;
}
}
}
for (int dz = 0; dz < D1D; ++dz)
{
double sol_xy[max_Q1D][max_Q1D];
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] = 0.0;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
double sol_x[max_Q1D];
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] = 0;
}
for (int dx = 0; dx < D1D; ++dx)
{
const double s = x(dx,dy,dz,e);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_x[qx] += B(qx,dx) * s;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
const double wy = B(qy,dy);
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xy[qy][qx] += wy * sol_x[qx];
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
const double wz = B(qz,dz);
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] += wz * sol_xy[qy][qx];
}
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
for (int qy = 0; qy < Q1D; ++qy)
{
for (int qx = 0; qx < Q1D; ++qx)
{
sol_xyz[qz][qy][qx] *= op(qx,qy,qz,e);
}
}
}
for (int qz = 0; qz < Q1D; ++qz)
{
double sol_xy[max_D1D][max_D1D];
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
sol_xy[dy][dx] = 0;
}
}
for (int qy = 0; qy < Q1D; ++qy)
{
double sol_x[max_D1D];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] = 0;
}
for (int qx = 0; qx < Q1D; ++qx)
{
const double s = sol_xyz[qz][qy][qx];
for (int dx = 0; dx < D1D; ++dx)
{
sol_x[dx] += Bt(dx,qx) * s;
}
}
for (int dy = 0; dy < D1D; ++dy)
{
const double wy = Bt(dy,qy);
for (int dx = 0; dx < D1D; ++dx)
{
sol_xy[dy][dx] += wy * sol_x[dx];
}
}
}
for (int dz = 0; dz < D1D; ++dz)
{
const double wz = Bt(dz,qz);
for (int dy = 0; dy < D1D; ++dy)
{
for (int dx = 0; dx < D1D; ++dx)
{
y(dx,dy,dz,e) += wz * sol_xy[dy][dx];
}
}
}
}
});
}
template<const int T_D1D = 0,
const int T_Q1D = 0>
static void SmemPAMassApply3D(const int NE,
const Array<double> &b_,
const Array<double> &bt_,
const Vector &op_,
const Vector &x_,
Vector &y_,
const int d1d = 0,
const int q1d = 0)
{
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int M1Q = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int M1D = T_D1D ? T_D1D : MAX_D1D;
MFEM_VERIFY(D1D <= M1D, "");
MFEM_VERIFY(Q1D <= M1Q, "");
auto b = Reshape(b_.Read(), Q1D, D1D);
auto op = Reshape(op_.Read(), Q1D, Q1D, Q1D, NE);
auto x = Reshape(x_.Read(), D1D, D1D, D1D, NE);
auto y = Reshape(y_.ReadWrite(), D1D, D1D, D1D, NE);
MFEM_FORALL_3D(e, NE, Q1D, Q1D, Q1D,
{
const int tidz = MFEM_THREAD_ID(z);
const int D1D = T_D1D ? T_D1D : d1d;
const int Q1D = T_Q1D ? T_Q1D : q1d;
constexpr int MQ1 = T_Q1D ? T_Q1D : MAX_Q1D;
constexpr int MD1 = T_D1D ? T_D1D : MAX_D1D;
constexpr int MDQ = (MQ1 > MD1) ? MQ1 : MD1;
MFEM_SHARED double sDQ[MQ1*MD1];
double (*B)[MD1] = (double (*)[MD1]) sDQ;
double (*Bt)[MQ1] = (double (*)[MQ1]) sDQ;
MFEM_SHARED double sm0[MDQ*MDQ*MDQ];
MFEM_SHARED double sm1[MDQ*MDQ*MDQ];
double (*X)[MD1][MD1] = (double (*)[MD1][MD1]) sm0;
double (*DDQ)[MD1][MQ1] = (double (*)[MD1][MQ1]) sm1;
double (*DQQ)[MQ1][MQ1] = (double (*)[MQ1][MQ1]) sm0;
double (*QQQ)[MQ1][MQ1] = (double (*)[MQ1][MQ1]) sm1;
double (*QQD)[MQ1][MD1] = (double (*)[MQ1][MD1]) sm0;
double (*QDD)[MD1][MD1] = (double (*)[MD1][MD1]) sm1;
MFEM_FOREACH_THREAD(dz,z,D1D)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
X[dz][dy][dx] = x(dx,dy,dz,e);
}
}
}
if (tidz == 0)
{
MFEM_FOREACH_THREAD(d,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
B[q][d] = b(q,d);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dz,z,D1D)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u = 0.0;
for (int dx = 0; dx < D1D; ++dx)
{
u += X[dz][dy][dx] * B[qx][dx];
}
DDQ[dz][dy][qx] = u;
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dz,z,D1D)
{
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u = 0.0;
for (int dy = 0; dy < D1D; ++dy)
{
u += DDQ[dz][dy][qx] * B[qy][dy];
}
DQQ[dz][qy][qx] = u;
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qz,z,Q1D)
{
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(qx,x,Q1D)
{
double u = 0.0;
for (int dz = 0; dz < D1D; ++dz)
{
u += DQQ[dz][qy][qx] * B[qz][dz];
}
QQQ[qz][qy][qx] = u * op(qx,qy,qz,e);
}
}
}
MFEM_SYNC_THREAD;
if (tidz == 0)
{
MFEM_FOREACH_THREAD(d,y,D1D)
{
MFEM_FOREACH_THREAD(q,x,Q1D)
{
Bt[d][q] = b(q,d);
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qz,z,Q1D)
{
MFEM_FOREACH_THREAD(qy,y,Q1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u = 0.0;
for (int qx = 0; qx < Q1D; ++qx)
{
u += QQQ[qz][qy][qx] * Bt[dx][qx];
}
QQD[qz][qy][dx] = u;
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(qz,z,Q1D)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u = 0.0;
for (int qy = 0; qy < Q1D; ++qy)
{
u += QQD[qz][qy][dx] * Bt[dy][qy];
}
QDD[qz][dy][dx] = u;
}
}
}
MFEM_SYNC_THREAD;
MFEM_FOREACH_THREAD(dz,z,D1D)
{
MFEM_FOREACH_THREAD(dy,y,D1D)
{
MFEM_FOREACH_THREAD(dx,x,D1D)
{
double u = 0.0;
for (int qz = 0; qz < Q1D; ++qz)
{
u += QDD[qz][dy][dx] * Bt[dz][qz];
}
y(dx,dy,dz,e) += u;
}
}
}
});
}
static void PAMassApply(const int dim,
const int D1D,
const int Q1D,
const int NE,
const Array<double> &B,
const Array<double> &Bt,
const Vector &op,
const Vector &x,
Vector &y)
{
#ifdef MFEM_USE_OCCA
if (DeviceCanUseOcca())
{
if (dim == 2)
{
OccaPAMassApply2D(D1D, Q1D, NE, B, Bt, op, x, y);
return;
}
if (dim == 3)
{
OccaPAMassApply3D(D1D, Q1D, NE, B, Bt, op, x, y);
return;
}
MFEM_ABORT("OCCA PA Mass Apply unknown kernel!");
}
#endif // MFEM_USE_OCCA
if (dim == 2)
{
switch ((D1D << 4 ) | Q1D)
{
case 0x22: return SmemPAMassApply2D<2,2,16>(NE, B, Bt, op, x, y);
case 0x33: return SmemPAMassApply2D<3,3,16>(NE, B, Bt, op, x, y);
case 0x44: return SmemPAMassApply2D<4,4,8>(NE, B, Bt, op, x, y);
case 0x55: return SmemPAMassApply2D<5,5,8>(NE, B, Bt, op, x, y);
case 0x66: return SmemPAMassApply2D<6,6,4>(NE, B, Bt, op, x, y);
case 0x77: return SmemPAMassApply2D<7,7,4>(NE, B, Bt, op, x, y);
case 0x88: return SmemPAMassApply2D<8,8,2>(NE, B, Bt, op, x, y);
case 0x99: return SmemPAMassApply2D<9,9,2>(NE, B, Bt, op, x, y);
default: return PAMassApply2D(NE, B, Bt, op, x, y, D1D, Q1D);
}
}
else if (dim == 3)
{
switch ((D1D << 4 ) | Q1D)
{
case 0x23: return SmemPAMassApply3D<2,3>(NE, B, Bt, op, x, y);
case 0x34: return SmemPAMassApply3D<3,4>(NE, B, Bt, op, x, y);
case 0x45: return SmemPAMassApply3D<4,5>(NE, B, Bt, op, x, y);
case 0x56: return SmemPAMassApply3D<5,6>(NE, B, Bt, op, x, y);
case 0x67: return SmemPAMassApply3D<6,7>(NE, B, Bt, op, x, y);
case 0x78: return SmemPAMassApply3D<7,8>(NE, B, Bt, op, x, y);
case 0x89: return SmemPAMassApply3D<8,9>(NE, B, Bt, op, x, y);
default: return PAMassApply3D(NE, B, Bt, op, x, y, D1D, Q1D);
}
}
MFEM_ABORT("Unknown kernel.");
}
void MassIntegrator::AddMultPA(const Vector &x, Vector &y) const
{
PAMassApply(dim, dofs1D, quad1D, ne, maps->B, maps->Bt, pa_data, x, y);
}
} // namespace mfem
+20 -12
View File
@@ -28,11 +28,6 @@ double PWConstCoefficient::Eval(ElementTransformation & T,
return (constants(att-1));
}
DeviceFunctionCoefficientPtr FunctionCoefficient::GetDeviceFunction()
{
return DeviceFunction;
}
double FunctionCoefficient::Eval(ElementTransformation & T,
const IntegrationPoint & ip)
{
@@ -45,10 +40,6 @@ double FunctionCoefficient::Eval(ElementTransformation & T,
{
return ((*Function)(transip));
}
else if (DeviceFunction)
{
return ((*DeviceFunction)(Vector3(x)));
}
else
{
return (*TDFunction)(transip, GetTime());
@@ -134,19 +125,27 @@ void VectorFunctionCoefficient::Eval(Vector &V, ElementTransformation &T,
}
VectorArrayCoefficient::VectorArrayCoefficient (int dim)
: VectorCoefficient(dim), Coeff(dim)
: VectorCoefficient(dim), Coeff(dim), ownCoeff(dim)
{
for (int i = 0; i < dim; i++)
{
Coeff[i] = NULL;
ownCoeff[i] = true;
}
}
void VectorArrayCoefficient::Set(int i, Coefficient *c, bool own)
{
if (ownCoeff[i]) { delete Coeff[i]; }
Coeff[i] = c;
ownCoeff[i] = own;
}
VectorArrayCoefficient::~VectorArrayCoefficient()
{
for (int i = 0; i < vdim; i++)
{
delete Coeff[i];
if (ownCoeff[i]) { delete Coeff[i]; }
}
}
@@ -318,17 +317,26 @@ MatrixArrayCoefficient::MatrixArrayCoefficient (int dim)
: MatrixCoefficient (dim)
{
Coeff.SetSize(height*width);
ownCoeff.SetSize(height*width);
for (int i = 0; i < (height*width); i++)
{
Coeff[i] = NULL;
ownCoeff[i] = true;
}
}
void MatrixArrayCoefficient::Set(int i, int j, Coefficient * c, bool own)
{
if (ownCoeff[i*width+j]) { delete Coeff[i*width+j]; }
Coeff[i*width+j] = c;
ownCoeff[i*width+j] = own;
}
MatrixArrayCoefficient::~MatrixArrayCoefficient ()
{
for (int i=0; i < height*width; i++)
{
delete Coeff[i];
if (ownCoeff[i]) { delete Coeff[i]; }
}
}
+8 -22
View File
@@ -112,7 +112,6 @@ public:
const IntegrationPoint &ip);
};
typedef double (*DeviceFunctionCoefficientPtr)(const Vector3&);
/// class for C-function coefficient
class FunctionCoefficient : public Coefficient
@@ -120,7 +119,6 @@ class FunctionCoefficient : public Coefficient
protected:
double (*Function)(const Vector &);
double (*TDFunction)(const Vector &, double);
double (*DeviceFunction)(const Vector3&);
public:
/// Define a time-independent coefficient from a C-function
@@ -128,7 +126,6 @@ public:
{
Function = f;
TDFunction = NULL;
DeviceFunction = NULL;
}
/// Define a time-dependent coefficient from a C-function
@@ -136,16 +133,6 @@ public:
{
Function = NULL;
TDFunction = tdf;
DeviceFunction = NULL;
}
/// Define a time-independent coefficient from a C-function using
/// Vector3 instead of a Vector.
FunctionCoefficient(double (*df)(const Vector3 &))
{
Function = NULL;
TDFunction = NULL;
DeviceFunction = df;
}
/// (DEPRECATED) Define a time-independent coefficient from a C-function
@@ -155,7 +142,6 @@ public:
{
Function = reinterpret_cast<double(*)(const Vector&)>(f);
TDFunction = NULL;
DeviceFunction = NULL;
}
/// (DEPRECATED) Define a time-dependent coefficient from a C-function
@@ -165,17 +151,11 @@ public:
{
Function = NULL;
TDFunction = reinterpret_cast<double(*)(const Vector&,double)>(tdf);
DeviceFunction = NULL;
}
/// Evaluate coefficient
virtual double Eval(ElementTransformation &T,
const IntegrationPoint &ip);
/// Return the coefficient's C-function that uses Vector3.
/// Warning: for now, the returned function can only be used on the
/// host inside a MFEM_FORALL.
DeviceFunctionCoefficientPtr GetDeviceFunction();
};
class GridFunction;
@@ -389,6 +369,7 @@ class VectorArrayCoefficient : public VectorCoefficient
{
private:
Array<Coefficient*> Coeff;
Array<bool> ownCoeff;
public:
/// Construct vector of dim coefficients.
@@ -400,7 +381,7 @@ public:
Coefficient **GetCoeffs() { return Coeff; }
/// Sets coefficient in the vector.
void Set(int i, Coefficient *c) { delete Coeff[i]; Coeff[i] = c; }
void Set(int i, Coefficient *c, bool own=true);
/// Evaluates i'th component of the vector.
double Eval(int i, ElementTransformation &T, const IntegrationPoint &ip)
@@ -520,9 +501,13 @@ public:
void SetDeltaCoefficient(const DeltaCoefficient& _d) { d = _d; }
/// Return the associated scalar DeltaCoefficient.
DeltaCoefficient& GetDeltaCoefficient() { return d; }
void SetScale(double s) { d.SetScale(s); }
void SetDirection(const Vector& _d);
void SetDeltaCenter(const Vector& center) { d.SetDeltaCenter(center); }
void GetDeltaCenter(Vector& center) { d.GetDeltaCenter(center); }
/** @brief Return the specified direction vector multiplied by the value
returned by DeltaCoefficient::EvalDelta() of the associated scalar
DeltaCoefficient. */
@@ -648,6 +633,7 @@ class MatrixArrayCoefficient : public MatrixCoefficient
{
private:
Array<Coefficient *> Coeff;
Array<bool> ownCoeff;
public:
@@ -655,7 +641,7 @@ public:
Coefficient* GetCoeff (int i, int j) { return Coeff[i*width+j]; }
void Set(int i, int j, Coefficient * c) { delete Coeff[i*width+j]; Coeff[i*width+j] = c; }
void Set(int i, int j, Coefficient * c, bool own=true);
double Eval(int i, int j, ElementTransformation &T, const IntegrationPoint &ip)
{ return Coeff[i*width+j] ? Coeff[i*width+j] -> Eval(T, ip, GetTime()) : 0.0; }
+17 -4
View File
@@ -108,6 +108,7 @@ DataCollection::DataCollection(const std::string& collection_name, Mesh *mesh_)
precision = precision_default;
pad_digits_cycle = pad_digits_rank = pad_digits_default;
format = SERIAL_FORMAT; // use serial mesh format
compression = false;
error = NO_ERROR;
}
@@ -161,6 +162,14 @@ void DataCollection::SetFormat(int fmt)
format = fmt;
}
void DataCollection::SetCompression(bool comp)
{
compression = comp;
#ifdef MFEM_USE_GZSTREAM
MFEM_ASSERT(!compression, "GZStream not enabled in MFEM build.");
#endif
}
void DataCollection::SetPrefixPath(const std::string& prefix)
{
if (!prefix.empty())
@@ -219,7 +228,8 @@ void DataCollection::SaveMesh()
}
std::string mesh_name = GetMeshFileName();
std::ofstream mesh_file(mesh_name.c_str());
const char *mode = (compression) ? "zwb6" : "w";
ofgzstream mesh_file(mesh_name.c_str(), mode);
mesh_file.precision(precision);
#ifdef MFEM_USE_MPI
const ParMesh *pmesh = dynamic_cast<const ParMesh*>(mesh);
@@ -267,7 +277,9 @@ const
void DataCollection::SaveOneField(const FieldMapIterator &it)
{
std::ofstream field_file(GetFieldFileName(it->first).c_str());
const char *mode = (compression) ? "zwb6" : "w";
ofgzstream field_file(GetFieldFileName(it->first).c_str(), mode);
field_file.precision(precision);
(it->second)->Save(field_file);
if (!field_file)
@@ -279,7 +291,8 @@ void DataCollection::SaveOneField(const FieldMapIterator &it)
void DataCollection::SaveOneQField(const QFieldMapIterator &it)
{
std::ofstream q_field_file(GetFieldFileName(it->first).c_str());
const char *mode = (compression) ? "zwb6" : "w";
ofgzstream q_field_file(GetFieldFileName(it->first).c_str(), mode);
q_field_file.precision(precision);
(it->second)->Save(q_field_file);
if (!q_field_file)
@@ -576,7 +589,7 @@ void VisItDataCollection::LoadFields()
it != field_info_map.end(); ++it)
{
std::string fname = path_left + it->first + path_right;
std::ifstream file(fname.c_str());
ifgzstream file(fname.c_str());
// TODO: in parallel, check for errors on all processors
if (!file)
{
+4
View File
@@ -205,6 +205,7 @@ protected:
/// Output mesh format: see the #Format enumeration
int format;
bool compression;
/// Should the collection delete its mesh and fields
bool own_data;
@@ -346,6 +347,9 @@ public:
validation. */
virtual void SetFormat(int fmt);
/// Set the flag for use of gz compressed files
void SetCompression(bool comp);
/// Set the path where the DataCollection will be saved.
void SetPrefixPath(const std::string &prefix);
+105
View File
@@ -203,6 +203,22 @@ void FiniteElement::CalcPhysDShape(ElementTransformation &Trans,
Mult(vshape, Trans.InverseJacobian(), dshape);
}
const DofToQuad &FiniteElement::GetDofToQuad(const IntegrationRule &,
DofToQuad::Mode) const
{
mfem_error("FiniteElement::GetDofToQuad(...) is not implemented for "
"this element!");
return *dof2quad_array[0]; // suppress a warning
}
FiniteElement::~FiniteElement()
{
for (int i = 0; i < dof2quad_array.Size(); i++)
{
delete dof2quad_array[i];
}
}
void ScalarFiniteElement::NodalLocalInterpolation (
ElementTransformation &Trans, DenseMatrix &I,
@@ -278,6 +294,95 @@ void ScalarFiniteElement::ScalarLocalInterpolation(
}
}
const DofToQuad &ScalarFiniteElement::GetDofToQuad(const IntegrationRule &ir,
DofToQuad::Mode mode) const
{
MFEM_VERIFY(mode == DofToQuad::FULL, "invalid mode requested");
for (int i = 0; i < dof2quad_array.Size(); i++)
{
const DofToQuad &d2q = *dof2quad_array[i];
if (d2q.IntRule == &ir && d2q.mode == mode) { return d2q; }
}
DofToQuad *d2q = new DofToQuad;
const int nqpt = ir.GetNPoints();
d2q->FE = this;
d2q->IntRule = &ir;
d2q->mode = mode;
d2q->ndof = Dof;
d2q->nqpt = nqpt;
d2q->B.SetSize(nqpt*Dof);
d2q->Bt.SetSize(Dof*nqpt);
d2q->G.SetSize(nqpt*Dim*Dof);
d2q->Gt.SetSize(Dof*nqpt*Dim);
#ifdef MFEM_THREAD_SAFE
Vector c_shape(Dof);
DenseMatrix vshape(Dof, Dim);
#endif
for (int i = 0; i < nqpt; i++)
{
const IntegrationPoint &ip = ir.IntPoint(i);
CalcShape(ip, c_shape);
for (int j = 0; j < Dof; j++)
{
d2q->B[i+nqpt*j] = d2q->Bt[j+Dof*i] = c_shape(j);
}
CalcDShape(ip, vshape);
for (int d = 0; d < Dim; d++)
{
for (int j = 0; j < Dof; j++)
{
d2q->G[i+nqpt*(d+Dim*j)] = d2q->Gt[j+Dof*(i+nqpt*d)] = vshape(j,d);
}
}
}
dof2quad_array.Append(d2q);
return *d2q;
}
// protected method
const DofToQuad &ScalarFiniteElement::GetTensorDofToQuad(
const TensorBasisElement &tb,
const IntegrationRule &ir, DofToQuad::Mode mode) const
{
MFEM_VERIFY(mode == DofToQuad::TENSOR, "invalid mode requested");
for (int i = 0; i < dof2quad_array.Size(); i++)
{
const DofToQuad &d2q = *dof2quad_array[i];
if (d2q.IntRule == &ir && d2q.mode == mode) { return d2q; }
}
DofToQuad *d2q = new DofToQuad;
const Poly_1D::Basis &basis_1d = tb.GetBasis1D();
const int ndof = Order + 1;
const int nqpt = (int)floor(pow(ir.GetNPoints(), 1.0/Dim) + 0.5);
d2q->FE = this;
d2q->IntRule = &ir;
d2q->mode = mode;
d2q->ndof = ndof;
d2q->nqpt = nqpt;
d2q->B.SetSize(nqpt*ndof);
d2q->Bt.SetSize(ndof*nqpt);
d2q->G.SetSize(nqpt*ndof);
d2q->Gt.SetSize(ndof*nqpt);
Vector val(ndof), grad(ndof);
for (int i = 0; i < nqpt; i++)
{
// The first 'nqpt' points in 'ir' have the same x-coordinates as those
// of the 1D rule.
basis_1d.Eval(ir.IntPoint(i).x, val, grad);
for (int j = 0; j < ndof; j++)
{
d2q->B[i+nqpt*j] = d2q->Bt[j+ndof*i] = val(j);
d2q->G[i+nqpt*j] = d2q->Gt[j+ndof*i] = grad(j);
}
}
dof2quad_array.Append(d2q);
return *d2q;
}
void NodalFiniteElement::ProjectCurl_2D(
const FiniteElement &fe, ElementTransformation &Trans,
+124 -2
View File
@@ -116,7 +116,92 @@ public:
}
};
// Base and derived classes for finite elements
/** @brief Structure representing the matrices/tensors needed to evaluate (in
reference space) the values, gradients, divergences, or curls of a
FiniteElement at a the quadrature points of a given IntegrationRule. */
/** Object of this type are typically created and owned by the respective
FiniteElement object. */
class DofToQuad
{
public:
/// The FiniteElement that created and owns this object.
/** This pointer is not owned. */
const class FiniteElement *FE;
/** @brief IntegrationRule that defines the quadrature points at which the
basis functions of the #FE are evaluated. */
/** This pointer is not owned. */
const IntegrationRule *IntRule;
/// Type of data stored in the arrays #B, #Bt, #G, and #Gt.
enum Mode
{
/** @brief Full multidimensional representation which does not use tensor
product structure. The ordering of the degrees of freedom is as
defined by #FE */
FULL,
/** @brief Tensor product representation using 1D matrices/tensors with
dimensions using 1D number of quadrature points and degrees of
freedom. */
/** When representing a vector-valued FiniteElement, two DofToQuad objects
are used to describe the "closed" and "open" 1D basis functions
(TODO). */
TENSOR
};
/// Describes the contents of the #B, #Bt, #G, and #Gt arrays, see #Mode.
Mode mode;
/** @brief Number of degrees of freedom = number of basis functions. When
#mode is TENSOR, this is the 1D number. */
int ndof;
/** @brief Number of quadrature points. When #mode is TENSOR, this is the 1D
number. */
int nqpt;
/// Basis functions evaluated at quadrature points.
/** The storage layout is column-major with dimensions:
- #nqpt x #ndof, for scalar elements, or
- #nqpt x dim x #ndof, for vector elements, (TODO)
where
- dim = dimension of the finite element reference space when #mode is
FULL, and dim = 1 when #mode is TENSOR. */
Array<double> B;
/// Transpose of #B.
/** The storage layout is column-major with dimensions:
- #ndof x #nqpt, for scalar elements, or
- #ndof x #nqpt x dim, for vector elements (TODO). */
Array<double> Bt;
/** @brief Gradients/divergences/curls of basis functions evaluated at
quadrature points. */
/** The storage layout is column-major with dimensions:
- #nqpt x dim x #ndof, for scalar elements, or
- #nqpt x #ndof, for H(div) vector elements (TODO), or
- #nqpt x cdim x #ndof, for H(curl) vector elements (TODO),
where
- dim = dimension of the finite element reference space when #mode is
FULL, and 1 when #mode is TENSOR,
- cdim = 1/1/3 in 1D/2D/3D, respectively, when #mode is FULL, and cdim =
1 when #mode is TENSOR. */
Array<double> G;
/// Transpose of #G.
/** The storage layout is column-major with dimensions:
- #ndof x #nqpt x dim, for scalar elements, or
- #ndof x #nqpt, for H(div) vector elements (TODO), or
- #ndof x #nqpt x cdim, for H(curl) vector elements (TODO). */
Array<double> Gt;
};
/// Describes the space on each element
class FunctionSpace
@@ -136,6 +221,10 @@ class VectorCoefficient;
class MatrixCoefficient;
class KnotVector;
// Base and derived classes for finite elements
/// Abstract class for Finite Elements
class FiniteElement
{
@@ -152,6 +241,10 @@ protected:
#ifndef MFEM_THREAD_SAFE
mutable DenseMatrix vshape; // Dof x Dim
#endif
/// Container for all DofToQuad objects created by the FiniteElement.
/** Multiple DofToQuad objects may be needed when different quadrature rules
or different DofToQuad::Mode are used. */
mutable Array<DofToQuad*> dof2quad_array;
public:
/// Enumeration for RangeType and DerivRangeType
@@ -417,7 +510,13 @@ public:
ElementTransformation &Trans,
DenseMatrix &div) const;
virtual ~FiniteElement () { }
/** Return a DofToQuad structure corresponding to the given IntegrationRule
using the given DofToQuad::Mode. */
/** See the documentation for DofToQuad for more details. */
virtual const DofToQuad &GetDofToQuad(const IntegrationRule &ir,
DofToQuad::Mode mode) const;
virtual ~FiniteElement();
static bool IsClosedType(int b_type)
{
@@ -464,6 +563,10 @@ protected:
return static_cast<const ScalarFiniteElement &>(fe);
}
const DofToQuad &GetTensorDofToQuad(const class TensorBasisElement &tb,
const IntegrationRule &ir,
DofToQuad::Mode mode) const;
public:
ScalarFiniteElement(int D, Geometry::Type G, int Do, int O,
int F = FunctionSpace::Pk)
@@ -494,6 +597,9 @@ public:
void ScalarLocalInterpolation(ElementTransformation &Trans,
DenseMatrix &I,
const ScalarFiniteElement &fine_fe) const;
virtual const DofToQuad &GetDofToQuad(const IntegrationRule &ir,
DofToQuad::Mode mode) const;
};
class NodalFiniteElement : public ScalarFiniteElement
@@ -1750,6 +1856,14 @@ class NodalTensorFiniteElement : public NodalFiniteElement,
public:
NodalTensorFiniteElement(const int dims, const int p, const int btype,
const DofMapType dmtype);
const DofToQuad &GetDofToQuad(const IntegrationRule &ir,
DofToQuad::Mode mode) const
{
return (mode == DofToQuad::FULL) ?
ScalarFiniteElement::GetDofToQuad(ir, mode) :
ScalarFiniteElement::GetTensorDofToQuad(*this, ir, mode);
}
};
class PositiveTensorFiniteElement : public PositiveFiniteElement,
@@ -1758,6 +1872,14 @@ class PositiveTensorFiniteElement : public PositiveFiniteElement,
public:
PositiveTensorFiniteElement(const int dims, const int p,
const DofMapType dmtype);
const DofToQuad &GetDofToQuad(const IntegrationRule &ir,
DofToQuad::Mode mode) const
{
return (mode == DofToQuad::FULL) ?
ScalarFiniteElement::GetDofToQuad(ir, mode) :
ScalarFiniteElement::GetTensorDofToQuad(*this, ir, mode);
}
};
class H1_SegmentElement : public NodalTensorFiniteElement
+557 -6
View File
@@ -12,6 +12,7 @@
// Implementation of FiniteElementSpace
#include "../general/text.hpp"
#include "../general/forall.hpp"
#include "../mesh/mesh_headers.hpp"
#include "fem.hpp"
@@ -385,6 +386,7 @@ void FiniteElementSpace::MarkerToList(const Array<int> &marker,
Array<int> &list)
{
int num_marked = 0;
marker.HostRead(); // make sure we can read the array on host
for (int i = 0; i < marker.Size(); i++)
{
if (marker[i]) { num_marked++; }
@@ -652,9 +654,9 @@ void FiniteElementSpace::BuildConformingInterpolation() const
// create the conforming restriction matrix cR
int *cR_J;
{
int *cR_I = mfem::New<int>(n_true_dofs+1);
double *cR_A = mfem::New<double>(n_true_dofs);
cR_J = mfem::New<int>(n_true_dofs);
int *cR_I = new int[n_true_dofs+1];
double *cR_A = new double[n_true_dofs];
cR_J = new int[n_true_dofs];
for (int i = 0; i < n_true_dofs; i++)
{
cR_I[i] = i;
@@ -732,6 +734,8 @@ void FiniteElementSpace::BuildConformingInterpolation() const
MakeVDimMatrix(*cP);
MakeVDimMatrix(*cR);
}
if (Device::IsEnabled()) { cP->BuildTranspose(); }
}
void FiniteElementSpace::MakeVDimMatrix(SparseMatrix &mat) const
@@ -782,6 +786,57 @@ int FiniteElementSpace::GetNConformingDofs() const
return P ? (P->Width() / vdim) : ndofs;
}
const Operator *FiniteElementSpace::GetElementRestriction(
ElementDofOrdering e_ordering) const
{
// Check if we have a discontinuous space using the FE collection:
const L2_FECollection *dg_space = dynamic_cast<const L2_FECollection*>(fec);
if (dg_space) { return NULL; }
// TODO: support other DG collections.
if (e_ordering == ElementDofOrdering::LEXICOGRAPHIC)
{
if (L2E_lex.Ptr() == NULL)
{
L2E_lex.Reset(new ElementRestriction(*this, e_ordering));
}
return L2E_lex.Ptr();
}
// e_ordering == ElementDofOrdering::NATIVE
if (L2E_nat.Ptr() == NULL)
{
L2E_nat.Reset(new ElementRestriction(*this, e_ordering));
}
return L2E_nat.Ptr();
}
const QuadratureInterpolator *FiniteElementSpace::GetQuadratureInterpolator(
const IntegrationRule &ir) const
{
for (int i = 0; i < E2Q_array.Size(); i++)
{
const QuadratureInterpolator *qi = E2Q_array[i];
if (qi->IntRule == &ir) { return qi; }
}
QuadratureInterpolator *qi = new QuadratureInterpolator(*this, ir);
E2Q_array.Append(qi);
return qi;
}
const QuadratureInterpolator *FiniteElementSpace::GetQuadratureInterpolator(
const QuadratureSpace &qs) const
{
for (int i = 0; i < E2Q_array.Size(); i++)
{
const QuadratureInterpolator *qi = E2Q_array[i];
if (qi->qspace == &qs) { return qi; }
}
QuadratureInterpolator *qi = new QuadratureInterpolator(*this, qs);
E2Q_array.Append(qi);
return qi;
}
SparseMatrix *FiniteElementSpace::RefinementMatrix_main(
const int coarse_ndofs, const Table &coarse_elem_dof,
const DenseTensor localP[]) const
@@ -1485,6 +1540,10 @@ void FiniteElementSpace::GetElementDofs (int i, Array<int> &dofs) const
const FiniteElement *FiniteElementSpace::GetFE(int i) const
{
if (i < 0 || !mesh->GetNE()) { return NULL; }
MFEM_VERIFY(i < mesh->GetNE(),
"Invalid element id " << i << ", maximum allowed " << mesh->GetNE()-1);
const FiniteElement *FE =
fec->FiniteElementForGeometry(mesh->GetElementBaseGeometry(i));
@@ -1789,6 +1848,13 @@ void FiniteElementSpace::Destroy()
delete cR;
delete cP;
Th.Clear();
L2E_nat.Clear();
L2E_lex.Clear();
for (int i = 0; i < E2Q_array.Size(); i++)
{
delete E2Q_array[i];
}
E2Q_array.SetSize(0);
dof_elem_array.DeleteAll();
dof_ldof_array.DeleteAll();
@@ -2350,8 +2416,7 @@ const Operator &InterpolationGridTransfer::BackwardOperator()
return *B.Ptr();
}
// Construct B
// If not set, define a suitable mass_integ
// Construct B, if not set, define a suitable mass_integ
if (!mass_integ && ran_fes.GetNE() > 0)
{
const FiniteElement *f_fe_0 = ran_fes.GetFE(0);
@@ -2514,7 +2579,7 @@ void L2ProjectionGridTransfer::L2Projection::Mult(
fes_ho.GetElementVDofs(iho, vdofs);
x.GetSubVector(vdofs, xel_mat.GetData());
mfem::Mult(R(iho), xel_mat, yel_mat);
// Place result correctly into low-order vector
// Place result correctly into the low-order vector
for (int iref=0; iref<nref; ++iref)
{
int ilor = ho2lor.GetRow(iho)[iref];
@@ -2572,4 +2637,490 @@ const Operator &L2ProjectionGridTransfer::BackwardOperator()
return *B;
}
ElementRestriction::ElementRestriction(const FiniteElementSpace &f,
ElementDofOrdering e_ordering)
: fes(f),
ne(fes.GetNE()),
vdim(fes.GetVDim()),
byvdim(fes.GetOrdering() == Ordering::byVDIM),
ndofs(fes.GetNDofs()),
dof(ne > 0 ? fes.GetFE(0)->GetDof() : 0),
nedofs(ne*dof),
offsets(ndofs+1),
indices(ne*dof)
{
// Assuming all finite elements are the same.
height = vdim*ne*dof;
width = fes.GetVSize();
const bool dof_reorder = (e_ordering == ElementDofOrdering::LEXICOGRAPHIC);
const int *dof_map = NULL;
if (dof_reorder && ne > 0)
{
for (int e = 0; e < ne; ++e)
{
const FiniteElement *fe = fes.GetFE(e);
const TensorBasisElement* el =
dynamic_cast<const TensorBasisElement*>(fe);
if (el) { continue; }
mfem_error("Finite element not suitable for lexicographic ordering");
}
const FiniteElement *fe = fes.GetFE(0);
const TensorBasisElement* el =
dynamic_cast<const TensorBasisElement*>(fe);
const Array<int> &fe_dof_map = el->GetDofMap();
MFEM_VERIFY(fe_dof_map.Size() > 0, "invalid dof map");
dof_map = fe_dof_map.GetData();
}
const Table& e2dTable = fes.GetElementToDofTable();
const int* elementMap = e2dTable.GetJ();
// We will be keeping a count of how many local nodes point to its global dof
for (int i = 0; i <= ndofs; ++i)
{
offsets[i] = 0;
}
for (int e = 0; e < ne; ++e)
{
for (int d = 0; d < dof; ++d)
{
const int gid = elementMap[dof*e + d];
++offsets[gid + 1];
}
}
// Aggregate to find offsets for each global dof
for (int i = 1; i <= ndofs; ++i)
{
offsets[i] += offsets[i - 1];
}
// For each global dof, fill in all local nodes that point to it
for (int e = 0; e < ne; ++e)
{
for (int d = 0; d < dof; ++d)
{
const int did = (!dof_reorder)?d:dof_map[d];
const int gid = elementMap[dof*e + did];
const int lid = dof*e + d;
indices[offsets[gid]++] = lid;
}
}
// We shifted the offsets vector by 1 by using it as a counter.
// Now we shift it back.
for (int i = ndofs; i > 0; --i)
{
offsets[i] = offsets[i - 1];
}
offsets[0] = 0;
}
void ElementRestriction::Mult(const Vector& x, Vector& y) const
{
// Assumes all elements have the same number of dofs
const int nd = dof;
const int vd = vdim;
const bool t = byvdim;
auto d_offsets = offsets.Read();
auto d_indices = indices.Read();
auto d_x = Reshape(x.Read(), t?vd:ndofs, t?ndofs:vd);
auto d_y = Reshape(y.Write(), nd, vd, ne);
MFEM_FORALL(i, ndofs,
{
const int offset = d_offsets[i];
const int nextOffset = d_offsets[i+1];
for (int c = 0; c < vd; ++c)
{
const double dofValue = d_x(t?c:i,t?i:c);
for (int j = offset; j < nextOffset; ++j)
{
const int idx_j = d_indices[j];
d_y(idx_j % nd, c, idx_j / nd) = dofValue;
}
}
});
}
void ElementRestriction::MultTranspose(const Vector& x, Vector& y) const
{
// Assumes all elements have the same number of dofs
const int nd = dof;
const int vd = vdim;
const bool t = byvdim;
auto d_offsets = offsets.Read();
auto d_indices = indices.Read();
auto d_x = Reshape(x.Read(), nd, vd, ne);
auto d_y = Reshape(y.Write(), t?vd:ndofs, t?ndofs:vd);
MFEM_FORALL(i, ndofs,
{
const int offset = d_offsets[i];
const int nextOffset = d_offsets[i + 1];
for (int c = 0; c < vd; ++c)
{
double dofValue = 0;
for (int j = offset; j < nextOffset; ++j)
{
const int idx_j = d_indices[j];
dofValue += d_x(idx_j % nd, c, idx_j / nd);
}
d_y(t?c:i,t?i:c) = dofValue;
}
});
}
QuadratureInterpolator::QuadratureInterpolator(const FiniteElementSpace &fes,
const IntegrationRule &ir)
{
fespace = &fes;
qspace = NULL;
IntRule = &ir;
use_tensor_products = true; // not implemented yet (not used)
if (fespace->GetNE() == 0) { return; }
const FiniteElement *fe = fespace->GetFE(0);
MFEM_VERIFY(dynamic_cast<const ScalarFiniteElement*>(fe) != NULL,
"Only scalar finite elements are supported");
}
QuadratureInterpolator::QuadratureInterpolator(const FiniteElementSpace &fes,
const QuadratureSpace &qs)
{
fespace = &fes;
qspace = &qs;
IntRule = NULL;
use_tensor_products = true; // not implemented yet (not used)
if (fespace->GetNE() == 0) { return; }
const FiniteElement *fe = fespace->GetFE(0);
MFEM_VERIFY(dynamic_cast<const ScalarFiniteElement*>(fe) != NULL,
"Only scalar finite elements are supported");
}
template<const int T_VDIM, const int T_ND, const int T_NQ>
void QuadratureInterpolator::Eval2D(
const int NE,
const int vdim,
const DofToQuad &maps,
const Vector &e_vec,
Vector &q_val,
Vector &q_der,
Vector &q_det,
const int eval_flags)
{
const int nd = maps.ndof;
const int nq = maps.nqpt;
const int ND = T_ND ? T_ND : nd;
const int NQ = T_NQ ? T_NQ : nq;
const int VDIM = T_VDIM ? T_VDIM : vdim;
MFEM_VERIFY(ND <= MAX_ND2D, "");
MFEM_VERIFY(NQ <= MAX_NQ2D, "");
MFEM_VERIFY(VDIM == 2 || !(eval_flags & DETERMINANTS), "");
auto B = Reshape(maps.B.Read(), NQ, ND);
auto G = Reshape(maps.G.Read(), NQ, 2, ND);
auto E = Reshape(e_vec.Read(), ND, VDIM, NE);
auto val = Reshape(q_val.Write(), NQ, VDIM, NE);
auto der = Reshape(q_der.Write(), NQ, VDIM, 2, NE);
auto det = Reshape(q_det.Write(), NQ, NE);
MFEM_FORALL(e, NE,
{
const int ND = T_ND ? T_ND : nd;
const int NQ = T_NQ ? T_NQ : nq;
const int VDIM = T_VDIM ? T_VDIM : vdim;
constexpr int max_ND = T_ND ? T_ND : MAX_ND2D;
constexpr int max_VDIM = T_VDIM ? T_VDIM : MAX_VDIM2D;
double s_E[max_VDIM*max_ND];
for (int d = 0; d < ND; d++)
{
for (int c = 0; c < VDIM; c++)
{
s_E[c+d*VDIM] = E(d,c,e);
}
}
for (int q = 0; q < NQ; ++q)
{
if (eval_flags & VALUES)
{
double ed[max_VDIM];
for (int c = 0; c < VDIM; c++) { ed[c] = 0.0; }
for (int d = 0; d < ND; ++d)
{
const double b = B(q,d);
for (int c = 0; c < VDIM; c++) { ed[c] += b*s_E[c+d*VDIM]; }
}
for (int c = 0; c < VDIM; c++) { val(q,c,e) = ed[c]; }
}
if ((eval_flags & DERIVATIVES) || (eval_flags & DETERMINANTS))
{
// use MAX_VDIM2D to avoid "subscript out of range" warnings
double D[MAX_VDIM2D*2];
for (int i = 0; i < 2*VDIM; i++) { D[i] = 0.0; }
for (int d = 0; d < ND; ++d)
{
const double wx = G(q,0,d);
const double wy = G(q,1,d);
for (int c = 0; c < VDIM; c++)
{
double s_e = s_E[c+d*VDIM];
D[c+VDIM*0] += s_e * wx;
D[c+VDIM*1] += s_e * wy;
}
}
if (eval_flags & DERIVATIVES)
{
for (int c = 0; c < VDIM; c++)
{
der(q,c,0,e) = D[c+VDIM*0];
der(q,c,1,e) = D[c+VDIM*1];
}
}
if (VDIM == 2 && (eval_flags & DETERMINANTS))
{
// The check (VDIM == 2) should eliminate this block when VDIM is
// known at compile time and (VDIM != 2).
det(q,e) = D[0]*D[3] - D[1]*D[2];
}
}
}
});
}
template<const int T_VDIM, const int T_ND, const int T_NQ>
void QuadratureInterpolator::Eval3D(
const int NE,
const int vdim,
const DofToQuad &maps,
const Vector &e_vec,
Vector &q_val,
Vector &q_der,
Vector &q_det,
const int eval_flags)
{
const int nd = maps.ndof;
const int nq = maps.nqpt;
const int ND = T_ND ? T_ND : nd;
const int NQ = T_NQ ? T_NQ : nq;
const int VDIM = T_VDIM ? T_VDIM : vdim;
MFEM_VERIFY(ND <= MAX_ND3D, "");
MFEM_VERIFY(NQ <= MAX_NQ3D, "");
MFEM_VERIFY(VDIM == 3 || !(eval_flags & DETERMINANTS), "");
auto B = Reshape(maps.B.Read(), NQ, ND);
auto G = Reshape(maps.G.Read(), NQ, 3, ND);
auto E = Reshape(e_vec.Read(), ND, VDIM, NE);
auto val = Reshape(q_val.Write(), NQ, VDIM, NE);
auto der = Reshape(q_der.Write(), NQ, VDIM, 3, NE);
auto det = Reshape(q_det.Write(), NQ, NE);
MFEM_FORALL(e, NE,
{
const int ND = T_ND ? T_ND : nd;
const int NQ = T_NQ ? T_NQ : nq;
const int VDIM = T_VDIM ? T_VDIM : vdim;
constexpr int max_ND = T_ND ? T_ND : MAX_ND3D;
constexpr int max_VDIM = T_VDIM ? T_VDIM : MAX_VDIM3D;
double s_E[max_VDIM*max_ND];
for (int d = 0; d < ND; d++)
{
for (int c = 0; c < VDIM; c++)
{
s_E[c+d*VDIM] = E(d,c,e);
}
}
for (int q = 0; q < NQ; ++q)
{
if (eval_flags & VALUES)
{
double ed[max_VDIM];
for (int c = 0; c < VDIM; c++) { ed[c] = 0.0; }
for (int d = 0; d < ND; ++d)
{
const double b = B(q,d);
for (int c = 0; c < VDIM; c++) { ed[c] += b*s_E[c+d*VDIM]; }
}
for (int c = 0; c < VDIM; c++) { val(q,c,e) = ed[c]; }
}
if ((eval_flags & DERIVATIVES) || (eval_flags & DETERMINANTS))
{
// use MAX_VDIM3D to avoid "subscript out of range" warnings
double D[MAX_VDIM3D*3];
for (int i = 0; i < 3*VDIM; i++) { D[i] = 0.0; }
for (int d = 0; d < ND; ++d)
{
const double wx = G(q,0,d);
const double wy = G(q,1,d);
const double wz = G(q,2,d);
for (int c = 0; c < VDIM; c++)
{
double s_e = s_E[c+d*VDIM];
D[c+VDIM*0] += s_e * wx;
D[c+VDIM*1] += s_e * wy;
D[c+VDIM*2] += s_e * wz;
}
}
if (eval_flags & DERIVATIVES)
{
for (int c = 0; c < VDIM; c++)
{
der(q,c,0,e) = D[c+VDIM*0];
der(q,c,1,e) = D[c+VDIM*1];
der(q,c,2,e) = D[c+VDIM*2];
}
}
if (VDIM == 3 && (eval_flags & DETERMINANTS))
{
// The check (VDIM == 3) should eliminate this block when VDIM is
// known at compile time and (VDIM != 3).
det(q,e) = D[0] * (D[4] * D[8] - D[5] * D[7]) +
D[3] * (D[2] * D[7] - D[1] * D[8]) +
D[6] * (D[1] * D[5] - D[2] * D[4]);
}
}
}
});
}
void QuadratureInterpolator::Mult(
const Vector &e_vec, unsigned eval_flags,
Vector &q_val, Vector &q_der, Vector &q_det) const
{
const int ne = fespace->GetNE();
if (ne == 0) { return; }
const int vdim = fespace->GetVDim();
const int dim = fespace->GetMesh()->Dimension();
const FiniteElement *fe = fespace->GetFE(0);
const IntegrationRule *ir =
IntRule ? IntRule : &qspace->GetElementIntRule(0);
const DofToQuad &maps = fe->GetDofToQuad(*ir, DofToQuad::FULL);
const int nd = maps.ndof;
const int nq = maps.nqpt;
void (*eval_func)(
const int NE,
const int vdim,
const DofToQuad &maps,
const Vector &e_vec,
Vector &q_val,
Vector &q_der,
Vector &q_det,
const int eval_flags) = NULL;
if (vdim == 1)
{
if (dim == 2)
{
switch (100*nd + nq)
{
// Q0
case 101: eval_func = &Eval2D<1,1,1>; break;
case 104: eval_func = &Eval2D<1,1,4>; break;
// Q1
case 404: eval_func = &Eval2D<1,4,4>; break;
case 409: eval_func = &Eval2D<1,4,9>; break;
// Q2
case 909: eval_func = &Eval2D<1,9,9>; break;
case 916: eval_func = &Eval2D<1,9,16>; break;
// Q3
case 1616: eval_func = &Eval2D<1,16,16>; break;
case 1625: eval_func = &Eval2D<1,16,25>; break;
case 1636: eval_func = &Eval2D<1,16,36>; break;
// Q4
case 2525: eval_func = &Eval2D<1,25,25>; break;
case 2536: eval_func = &Eval2D<1,25,36>; break;
case 2549: eval_func = &Eval2D<1,25,49>; break;
case 2564: eval_func = &Eval2D<1,25,64>; break;
}
if (nq >= 100 || !eval_func)
{
eval_func = &Eval2D<1>;
}
}
else if (dim == 3)
{
switch (1000*nd + nq)
{
// Q0
case 1001: eval_func = &Eval3D<1,1,1>; break;
case 1008: eval_func = &Eval3D<1,1,8>; break;
// Q1
case 8008: eval_func = &Eval3D<1,8,8>; break;
case 8027: eval_func = &Eval3D<1,8,27>; break;
// Q2
case 27027: eval_func = &Eval3D<1,27,27>; break;
case 27064: eval_func = &Eval3D<1,27,64>; break;
// Q3
case 64064: eval_func = &Eval3D<1,64,64>; break;
case 64125: eval_func = &Eval3D<1,64,125>; break;
case 64216: eval_func = &Eval3D<1,64,216>; break;
// Q4
case 125125: eval_func = &Eval3D<1,125,125>; break;
case 125216: eval_func = &Eval3D<1,125,216>; break;
}
if (nq >= 1000 || !eval_func)
{
eval_func = &Eval3D<1>;
}
}
}
else if (vdim == dim)
{
if (dim == 2)
{
switch (100*nd + nq)
{
// Q1
case 404: eval_func = &Eval2D<2,4,4>; break;
case 409: eval_func = &Eval2D<2,4,9>; break;
// Q2
case 909: eval_func = &Eval2D<2,9,9>; break;
case 916: eval_func = &Eval2D<2,9,16>; break;
// Q3
case 1616: eval_func = &Eval2D<2,16,16>; break;
case 1625: eval_func = &Eval2D<2,16,25>; break;
case 1636: eval_func = &Eval2D<2,16,36>; break;
// Q4
case 2525: eval_func = &Eval2D<2,25,25>; break;
case 2536: eval_func = &Eval2D<2,25,36>; break;
case 2549: eval_func = &Eval2D<2,25,49>; break;
case 2564: eval_func = &Eval2D<2,25,64>; break;
}
if (nq >= 100 || !eval_func)
{
eval_func = &Eval2D<2>;
}
}
else if (dim == 3)
{
switch (1000*nd + nq)
{
// Q1
case 8008: eval_func = &Eval3D<3,8,8>; break;
case 8027: eval_func = &Eval3D<3,8,27>; break;
// Q2
case 27027: eval_func = &Eval3D<3,27,27>; break;
case 27064: eval_func = &Eval3D<3,27,64>; break;
// Q3
case 64064: eval_func = &Eval3D<3,64,64>; break;
case 64125: eval_func = &Eval3D<3,64,125>; break;
case 64216: eval_func = &Eval3D<3,64,216>; break;
// Q4
case 125125: eval_func = &Eval3D<3,125,125>; break;
case 125216: eval_func = &Eval3D<3,125,216>; break;
}
if (nq >= 1000 || !eval_func)
{
eval_func = &Eval3D<3>;
}
}
}
if (eval_func)
{
eval_func(ne, vdim, maps, e_vec, q_val, q_der, q_det, eval_flags);
}
else
{
MFEM_ABORT("case not supported yet");
}
}
void QuadratureInterpolator::MultTranspose(
unsigned eval_flags, const Vector &q_val, const Vector &q_der,
Vector &e_vec) const
{
MFEM_ABORT("this method is not implemented yet");
}
} // namespace mfem
+184
View File
@@ -59,9 +59,25 @@ Ordering::Map<Ordering::byVDIM>(int ndofs, int vdim, int dof, int vd)
}
/// Constants describing the possible orderings of the DOFs in one element.
enum class ElementDofOrdering
{
/// Native ordering as defined by the FiniteElement.
/** This ordering can be used by tensor-product elements when the
interpolation from the DOFs to quadrature points does not use the
tensor-product structure. */
NATIVE,
/// Lexicographic ordering for tensor-product FiniteElements.
/** This ordering can be used only with tensor-product elements. */
LEXICOGRAPHIC
};
// Forward declarations
class NURBSExtension;
class BilinearFormIntegrator;
class QuadratureSpace;
class QuadratureInterpolator;
/** @brief Class FiniteElementSpace - responsible for providing FEM view of the
@@ -110,6 +126,11 @@ protected:
/// Transformation to apply to GridFunctions after space Update().
OperatorHandle Th;
/// The element restriction operators, see GetElementRestriction().
mutable OperatorHandle L2E_nat, L2E_lex;
mutable Array<QuadratureInterpolator*> E2Q_array;
long sequence; // should match Mesh::GetSequence
void UpdateNURBS();
@@ -257,14 +278,60 @@ public:
bool Conforming() const { return mesh->Conforming(); }
bool Nonconforming() const { return mesh->Nonconforming(); }
/// The returned SparseMatrix is owned by the FiniteElementSpace.
const SparseMatrix *GetConformingProlongation() const;
/// The returned SparseMatrix is owned by the FiniteElementSpace.
const SparseMatrix *GetConformingRestriction() const;
/// The returned Operator is owned by the FiniteElementSpace.
virtual const Operator *GetProlongationMatrix() const
{ return GetConformingProlongation(); }
/// The returned SparseMatrix is owned by the FiniteElementSpace.
virtual const SparseMatrix *GetRestrictionMatrix() const
{ return GetConformingRestriction(); }
/// Return an Operator that converts L-vectors to E-vectors.
/** An L-vector is a vector of size GetVSize() which is the same size as a
GridFunction. An E-vector represents the element-wise discontinuous
version of the FE space.
The layout of the E-vector is: ND x VDIM x NE, where ND is the number of
degrees of freedom, VDIM is the vector dimension of the FE space, and NE
is the number of the mesh elements.
The parameter @a e_ordering describes how the local DOFs in each element
should be ordered, see ElementDofOrdering.
For discontinuous spaces, where the element-restriction is the identity,
this method will return NULL.
The returned Operator is owned by the FiniteElementSpace. */
const Operator *GetElementRestriction(ElementDofOrdering e_ordering) const;
/** @brief Return a QuadratureInterpolator that interpolates E-vectors to
quadrature point values and/or derivatives (Q-vectors). */
/** An E-vector represents the element-wise discontinuous version of the FE
space and can be obtained, for example, from a GridFunction using the
Operator returned by GetElementRestriction().
All elements will use the same IntegrationRule, @a ir as the target
quadrature points. */
const QuadratureInterpolator *GetQuadratureInterpolator(
const IntegrationRule &ir) const;
/** @brief Return a QuadratureInterpolator that interpolates E-vectors to
quadrature point values and/or derivatives (Q-vectors). */
/** An E-vector represents the element-wise discontinuous version of the FE
space and can be obtained, for example, from a GridFunction using the
Operator returned by GetElementRestriction().
The target quadrature points in the elements are described by the given
QuadratureSpace, @a qs. */
const QuadratureInterpolator *GetQuadratureInterpolator(
const QuadratureSpace &qs) const;
/// Returns vector dimension.
inline int GetVDim() const { return vdim; }
@@ -806,6 +873,123 @@ public:
virtual const Operator &BackwardOperator();
};
/// Operator that converts FiniteElementSpace L-vectors to E-vectors.
/** Objects of this type are typically created and owned by FiniteElementSpace
objects, see FiniteElementSpace::GetElementRestriction(). */
class ElementRestriction : public Operator
{
protected:
const FiniteElementSpace &fes;
const int ne;
const int vdim;
const bool byvdim;
const int ndofs;
const int dof;
const int nedofs;
Array<int> offsets;
Array<int> indices;
public:
ElementRestriction(const FiniteElementSpace&, ElementDofOrdering);
void Mult(const Vector &x, Vector &y) const;
void MultTranspose(const Vector &x, Vector &y) const;
};
/** @brief A class that performs interpolation from an E-vector to quadrature
point values and/or derivatives (Q-vectors). */
/** An E-vector represents the element-wise discontinuous version of the FE
space and can be obtained, for example, from a GridFunction using the
Operator returned by FiniteElementSpace::GetElementRestriction().
The target quadrature points in the elements can be described either by an
IntegrationRule (all mesh elements must be of the same type in this case) or
by a QuadratureSpace. */
class QuadratureInterpolator
{
protected:
friend class FiniteElementSpace; // Needs access to qspace and IntRule
const FiniteElementSpace *fespace; ///< Not owned
const QuadratureSpace *qspace; ///< Not owned
const IntegrationRule *IntRule; ///< Not owned
mutable bool use_tensor_products;
static const int MAX_NQ2D = 100;
static const int MAX_ND2D = 100;
static const int MAX_VDIM2D = 2;
static const int MAX_NQ3D = 1000;
static const int MAX_ND3D = 1000;
static const int MAX_VDIM3D = 3;
public:
enum EvalFlags
{
VALUES = 1 << 0, ///< Evaluate the values at quadrature points
DERIVATIVES = 1 << 1, ///< Evaluate the derivatives at quadrature points
/** @brief Assuming the derivative at quadrature points form a matrix,
this flag can be used to compute and store their determinants. This
flag can only be used in Mult(). */
DETERMINANTS = 1 << 2
};
QuadratureInterpolator(const FiniteElementSpace &fes,
const IntegrationRule &ir);
QuadratureInterpolator(const FiniteElementSpace &fes,
const QuadratureSpace &qs);
/** @brief Disable the use of tensor product evaluations, for tensor-product
elements, e.g. quads and hexes. */
/** Currently, tensor product evaluations are not implemented and this method
has no effect. */
void DisableTensorProducts(bool disable = true) const
{ use_tensor_products = !disable; }
/// Interpolate the E-vector @a e_vec to quadrature points.
/** The @a eval_flags are a bitwise mask of constants from the EvalFlags
enumeration. When the VALUES flag is set, the values at quadrature points
are computed and stored in the Vector @a q_val. Similarly, when the flag
DERIVATIVES is set, the derivatives are computed and stored in @a q_der.
When the DETERMINANTS flags is set, it is assumed that the derivatives
form a matrix at each quadrature point (i.e. the associated
FiniteElementSpace is a vector space) and their determinants are computed
and stored in @a q_det. */
void Mult(const Vector &e_vec, unsigned eval_flags,
Vector &q_val, Vector &q_der, Vector &q_det) const;
/// Perform the transpose operation of Mult(). (TODO)
void MultTranspose(unsigned eval_flags, const Vector &q_val,
const Vector &q_der, Vector &e_vec) const;
// Compute kernels follow (cannot be private or protected with nvcc)
/// Template compute kernel for 2D.
template<const int T_VDIM = 0, const int T_ND = 0, const int T_NQ = 0>
static void Eval2D(const int NE,
const int vdim,
const DofToQuad &maps,
const Vector &e_vec,
Vector &q_val,
Vector &q_der,
Vector &q_det,
const int eval_flags);
/// Template compute kernel for 3D.
template<const int T_VDIM = 0, const int T_ND = 0, const int T_NQ = 0>
static void Eval3D(const int NE,
const int vdim,
const DofToQuad &maps,
const Vector &e_vec,
Vector &q_val,
Vector &q_der,
Vector &q_det,
const int eval_flags);
};
}
#endif
+19 -7
View File
@@ -30,6 +30,9 @@ using namespace std;
GridFunction::GridFunction(Mesh *m, std::istream &input)
: Vector()
{
// Grid functions are stored on the device
UseDevice(true);
fes = new FiniteElementSpace;
fec = fes->Load(m, input);
@@ -60,6 +63,8 @@ GridFunction::GridFunction(Mesh *m, std::istream &input)
GridFunction::GridFunction(Mesh *m, GridFunction *gf_array[], int num_pieces)
{
UseDevice(true);
// all GridFunctions must have the same FE collection, vdim, ordering
int vdim, ordering;
@@ -163,6 +168,7 @@ void GridFunction::Update()
Vector old_data;
old_data.Swap(*this);
SetSize(T->Height());
UseDevice(true);
T->Mult(old_data, *this);
}
else
@@ -192,7 +198,9 @@ void GridFunction::MakeRef(FiniteElementSpace *f, Vector &v, int v_offset)
MFEM_ASSERT(v.Size() >= v_offset + f->GetVSize(), "");
if (f != fes) { Destroy(); }
fes = f;
NewDataAndSize((double *)v + v_offset, fes->GetVSize());
v.UseDevice(true);
NewMemoryAndSize(Memory<double>(v.GetMemory(), v_offset, fes->GetVSize()),
fes->GetVSize(), true);
sequence = fes->GetSequence();
}
@@ -215,13 +223,16 @@ void GridFunction::MakeTRef(FiniteElementSpace *f, Vector &tv, int tv_offset)
if (!f->GetProlongationMatrix())
{
MakeRef(f, tv, tv_offset);
t_vec.NewDataAndSize(data, size);
t_vec.NewMemoryAndSize(data, size, false);
}
else
{
MFEM_ASSERT(tv.Size() >= tv_offset + f->GetTrueVSize(), "");
SetSpace(f); // works in parallel
t_vec.NewDataAndSize(&tv(tv_offset), f->GetTrueVSize());
tv.UseDevice(true);
const int tv_size = f->GetTrueVSize();
t_vec.NewMemoryAndSize(Memory<double>(tv.GetMemory(), tv_offset, tv_size),
tv_size, true);
}
}
@@ -302,7 +313,7 @@ int GridFunction::VectorDim() const
{
fe = fes->GetFE(0);
}
if (fe->GetRangeType() == FiniteElement::SCALAR)
if (!fe || fe->GetRangeType() == FiniteElement::SCALAR)
{
return fes->GetVDim();
}
@@ -315,7 +326,7 @@ void GridFunction::GetTrueDofs(Vector &tv) const
if (!R)
{
// R is identity -> make tv a reference to *this
tv.NewDataAndSize(data, size);
tv.NewDataAndSize(const_cast<double*>((const double*)data), size);
}
else
{
@@ -1367,7 +1378,7 @@ void GridFunction::AccumulateAndCountBdrValues(
if (vdofs.Size() == 0) { continue; }
transf = mesh->GetEdgeTransformation(edge);
transf->Attribute = -1; // FIXME: set the boundary attribute
transf->Attribute = -1; // TODO: set the boundary attribute
fe = fes->GetEdgeElement(edge);
if (!vcoeff)
{
@@ -1471,7 +1482,7 @@ void GridFunction::AccumulateAndCountBdrTangentValues(
if (dofs.Size() == 0) { continue; }
T = mesh->GetEdgeTransformation(edge);
T->Attribute = -1; // FIXME: set the boundary attribute
T->Attribute = -1; // TODO: set the boundary attribute
fe = fes->GetEdgeElement(edge);
lvec.SetSize(fe->GetDof());
fe->Project(vcoeff, *T, lvec);
@@ -1776,6 +1787,7 @@ void GridFunction::ProjectBdrCoefficient(VectorCoefficient &vcoeff,
void GridFunction::ProjectBdrCoefficient(Coefficient *coeff[], Array<int> &attr)
{
Array<int> values_counter;
this->HostReadWrite();
AccumulateAndCountBdrValues(coeff, NULL, attr, values_counter);
ComputeMeans(ARITHMETIC, values_counter);
#ifdef MFEM_DEBUG
+12 -7
View File
@@ -68,15 +68,16 @@ protected:
public:
GridFunction() { fes = NULL; fec = NULL; sequence = 0; }
GridFunction() { fes = NULL; fec = NULL; sequence = 0; UseDevice(true); }
/// Copy constructor. The internal true-dof vector #t_vec is not copied.
GridFunction(const GridFunction &orig)
: Vector(orig), fes(orig.fes), fec(NULL), sequence(orig.sequence) { }
: Vector(orig), fes(orig.fes), fec(NULL), sequence(orig.sequence)
{ UseDevice(true); }
/// Construct a GridFunction associated with the FiniteElementSpace @a *f.
GridFunction(FiniteElementSpace *f) : Vector(f->GetVSize())
{ fes = f; fec = NULL; sequence = f->GetSequence(); }
{ fes = f; fec = NULL; sequence = f->GetSequence(); UseDevice(true); }
/// Construct a GridFunction using previously allocated array @a data.
/** The GridFunction does not assume ownership of @a data which is assumed to
@@ -84,8 +85,9 @@ public:
for externally allocated array, the pointer @a data can be NULL. The data
array can be replaced later using the method SetData().
*/
GridFunction(FiniteElementSpace *f, double *data) : Vector(data, f->GetVSize())
{ fes = f; fec = NULL; sequence = f->GetSequence(); }
GridFunction(FiniteElementSpace *f, double *data)
: Vector(data, f->GetVSize())
{ fes = f; fec = NULL; sequence = f->GetSequence(); UseDevice(true); }
/// Construct a GridFunction on the given Mesh, using the data from @a input.
/** The content of @a input should be in the format created by the method
@@ -124,6 +126,7 @@ public:
/// @brief Extract the true-dofs from the GridFunction. If all dofs are true,
/// then `tv` will be set to point to the data of `*this`.
/** @warning This method breaks const-ness when all dofs are true. */
void GetTrueDofs(Vector &tv) const;
/// Shortcut for calling GetTrueDofs() with GetTrueVector() as argument.
@@ -702,7 +705,7 @@ inline void QuadratureFunction::GetElementValues(int idx, Vector &values) const
const int s_offset = qspace->element_offsets[idx];
const int sl_size = qspace->element_offsets[idx+1] - s_offset;
values.SetSize(vdim*sl_size);
double *q = data + vdim*s_offset;
const double *q = data + vdim*s_offset;
for (int i = 0; i<values.Size(); i++)
{
values(i) = *(q++);
@@ -722,12 +725,14 @@ inline void QuadratureFunction::GetElementValues(int idx,
const int s_offset = qspace->element_offsets[idx];
const int sl_size = qspace->element_offsets[idx+1] - s_offset;
values.SetSize(vdim, sl_size);
double *q = data + vdim*s_offset;
const double *q = data + vdim*s_offset;
for (int j = 0; j<sl_size; j++)
{
for (int i = 0; i<vdim; i++)
{
values(i,j) = *(q++);
}
}
}
} // namespace mfem
+13
View File
@@ -78,6 +78,19 @@ IntegrationRule::IntegrationRule(IntegrationRule &irx, IntegrationRule &iry,
}
}
const Array<double> &IntegrationRule::GetWeights() const
{
if (weights.Size() != GetNPoints())
{
weights.SetSize(GetNPoints());
for (int i = 0; i < GetNPoints(); i++)
{
weights[i] = IntPoint(i).weight;
}
}
return weights;
}
void IntegrationRule::GrundmannMollerSimplexRule(int s, int n)
{
// for pow on older compilers
+8
View File
@@ -87,6 +87,9 @@ class IntegrationRule : public Array<IntegrationPoint>
private:
friend class IntegrationRules;
int Order;
/** @brief The quadrature weights gathered as a contiguous array. Created
by request with the method GetWeights(). */
mutable Array<double> weights;
/// Define n-simplex rule (triangle/tetrahedron for n=2/3) of order (2s+1)
void GrundmannMollerSimplexRule(int s, int n = 3);
@@ -239,6 +242,11 @@ public:
/// Returns a const reference to the i-th integration point
const IntegrationPoint &IntPoint(int i) const { return (*this)[i]; }
/// Return the quadrature weights in a contiguous array.
/** If a contiguous array is not required, the weights can be accessed with
a call like this: `IntPoint(i).weight`. */
const Array<double> &GetWeights() const;
/// Destroys an IntegrationRule object
~IntegrationRule() { }
};
+7
View File
@@ -19,6 +19,9 @@ namespace mfem
LinearForm::LinearForm(FiniteElementSpace *f, LinearForm *lf)
: Vector(f->GetVSize())
{
// Linear forms are stored on the device
UseDevice(true);
fes = f;
extern_lfs = 1;
@@ -83,6 +86,10 @@ void LinearForm::Assemble()
Vector::operator=(0.0);
// The above operation is executed on device because of UseDevice().
// The first use of AddElementVector() below will move it back to host
// because both 'vdofs' and 'elemvect' are on host.
if (dlfi.Size())
{
for (i = 0; i < fes -> GetNE(); i++)
+2 -2
View File
@@ -64,7 +64,7 @@ public:
/// Creates linear form associated with FE space @a *f.
/** The pointer @a f is not owned by the newly constructed object. */
LinearForm(FiniteElementSpace *f) : Vector(f->GetVSize())
{ fes = f; extern_lfs = 0; }
{ fes = f; extern_lfs = 0; UseDevice(true); }
/** @brief Create a LinearForm on the FiniteElementSpace @a f, using the
same integrators as the LinearForm @a lf.
@@ -79,7 +79,7 @@ public:
/** The associated FiniteElementSpace can be set later using one of the
methods: Update(FiniteElementSpace *) or
Update(FiniteElementSpace *, Vector &, int). */
LinearForm() { fes = NULL; extern_lfs = 0; }
LinearForm() { fes = NULL; extern_lfs = 0; UseDevice(true); }
/// Copy assignment. Only the data of the base class Vector is copied.
/** It is assumed that this object and @a rhs use FiniteElementSpace%s that
+3 -3
View File
@@ -181,7 +181,7 @@ void VectorDomainLFIntegrator::AssembleRHSElementVect(
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int intorder = el.GetOrder() + 1;
int intorder = 2*el.GetOrder();
ir = &IntRules.Get(el.GetGeomType(), intorder);
}
@@ -240,7 +240,7 @@ void VectorBoundaryLFIntegrator::AssembleRHSElementVect(
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int intorder = el.GetOrder() + 1;
int intorder = 2*el.GetOrder();
ir = &IntRules.Get(el.GetGeomType(), intorder);
}
@@ -275,7 +275,7 @@ void VectorBoundaryLFIntegrator::AssembleRHSElementVect(
const IntegrationRule *ir = IntRule;
if (ir == NULL)
{
int intorder = el.GetOrder() + 1;
int intorder = 2*el.GetOrder();
ir = &IntRules.Get(Tr.FaceGeom, intorder);
}
+36 -36
View File
@@ -35,11 +35,11 @@ typedef double* QLocal2D_t @dim(Q1D, Q1D, NE);
typedef double* DLocal3D_t @dim(D1D, D1D, D1D, NE);
typedef double* QLocal3D_t @dim(Q1D, Q1D, Q1D, NE);
typedef double* Jacobian2D_t @dim(2, 2, Q2D, NE);
typedef double* Jacobian3D_t @dim(3, 3, Q3D, NE);
typedef double* Jacobian2D_t @dim(Q2D, 2, 2, NE);
typedef double* Jacobian3D_t @dim(Q3D, 3, 3, NE);
typedef double* SymmOperator2D_t @dim(3, Q2D, NE);
typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
typedef double* SymmOperator2D_t @dim(Q2D, 3, NE);
typedef double* SymmOperator3D_t @dim(Q3D, 6, NE);
@kernel void DiffusionSetup2D(const int NE,
@restrict const double *W,
@@ -48,12 +48,12 @@ typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
@restrict SymmOperator2D_t op) {
for (int e = 0; e < NE; ++e; @outer) {
for (int q = 0; q < Q2D; ++q; @inner) {
const double J11 = J(0, 0, q, e), J12 = J(1, 0, q, e);
const double J21 = J(0, 1, q, e), J22 = J(1, 1, q, e);
const double J11 = J(q, 0, 0, e), J12 = J(q, 1, 0, e);
const double J21 = J(q, 0, 1, e), J22 = J(q, 1, 1, e);
const double c_detJ = W[q] * COEFF / ((J11 * J22) - (J21 * J12));
op(0, q, e) = c_detJ * (J21*J21 + J22*J22); // (1,1)
op(1, q, e) = -c_detJ * (J21*J11 + J22*J12); // (1,2), (2,1)
op(2, q, e) = c_detJ * (J11*J11 + J12*J12); // (2,2)
op(q, 0, e) = c_detJ * (J21*J21 + J22*J22); // (1,1)
op(q, 1, e) = -c_detJ * (J21*J11 + J22*J12); // (1,2), (2,1)
op(q, 2, e) = c_detJ * (J11*J11 + J12*J12); // (2,2)
}
}
}
@@ -65,9 +65,9 @@ typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
@restrict SymmOperator3D_t op) {
for (int e = 0; e < NE; ++e; @outer) {
for (int q = 0; q < Q3D; ++q; @inner) {
const double J11 = J(0, 0, q, e), J12 = J(1, 0, q, e), J13 = J(2, 0, q, e);
const double J21 = J(0, 1, q, e), J22 = J(1, 1, q, e), J23 = J(2, 1, q, e);
const double J31 = J(0, 2, q, e), J32 = J(1, 2, q, e), J33 = J(2, 2, q, e);
const double J11 = J(q, 0, 0, e), J12 = J(q, 1, 0, e), J13 = J(q, 2, 0, e);
const double J21 = J(q, 0, 1, e), J22 = J(q, 1, 1, e), J23 = J(q, 2, 1, e);
const double J31 = J(q, 0, 2, e), J32 = J(q, 1, 2, e), J33 = J(q, 2, 2, e);
const double detJ = ((J11 * J22 * J33) + (J12 * J23 * J31) + (J13 * J21 * J32) -
(J13 * J22 * J31) - (J12 * J21 * J33) - (J11 * J23 * J32));
@@ -88,12 +88,12 @@ typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
const double A33 = (J11 * J22) - (J12 * J21);
// adj(J)^Tadj(J)
op(0, q, e) = c_detJ * (A11*A11 + A21*A21 + A31*A31); // (1,1)
op(1, q, e) = c_detJ * (A11*A12 + A21*A22 + A31*A32); // (1,2), (2,1)
op(2, q, e) = c_detJ * (A11*A13 + A21*A23 + A31*A33); // (1,3), (3,1)
op(3, q, e) = c_detJ * (A12*A12 + A22*A22 + A32*A32); // (2,2)
op(4, q, e) = c_detJ * (A12*A13 + A22*A23 + A32*A33); // (2,3), (3,2)
op(5, q, e) = c_detJ * (A13*A13 + A23*A23 + A33*A33); // (3,3)
op(q, 0, e) = c_detJ * (A11*A11 + A21*A21 + A31*A31); // (1,1)
op(q, 1, e) = c_detJ * (A11*A12 + A21*A22 + A31*A32); // (1,2), (2,1)
op(q, 2, e) = c_detJ * (A11*A13 + A21*A23 + A31*A33); // (1,3), (3,1)
op(q, 3, e) = c_detJ * (A12*A12 + A22*A22 + A32*A32); // (2,2)
op(q, 4, e) = c_detJ * (A12*A13 + A22*A23 + A32*A33); // (2,3), (3,2)
op(q, 5, e) = c_detJ * (A13*A13 + A23*A23 + A33*A33); // (3,3)
}
}
}
@@ -146,9 +146,9 @@ typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
for (int qy = 0; qy < Q1D; ++qy) {
for (int qx = 0; qx < Q1D; ++qx) {
const int q = QUAD_2D_ID(qx, qy);
const double O11 = op(0, q, e);
const double O12 = op(1, q, e);
const double O22 = op(2, q, e);
const double O11 = op(q, 0, e);
const double O12 = op(q, 1, e);
const double O22 = op(q, 2, e);
const double gradX = grad[qy][qx][0];
const double gradY = grad[qy][qx][1];
@@ -255,9 +255,9 @@ typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
}
const int q = QUAD_2D_ID(qx, qy);
const double O11 = op(0, q, e);
const double O12 = op(1, q, e);
const double O22 = op(2, q, e);
const double O11 = op(q, 0, e);
const double O12 = op(q, 1, e);
const double O22 = op(q, 2, e);
s_grad(0, qx, qy) = (O11 * gradX) + (O12 * gradY);
s_grad(1, qx, qy) = (O12 * gradX) + (O22 * gradY);
@@ -382,12 +382,12 @@ typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
for (int qy = 0; qy < Q1D; ++qy) {
for (int qx = 0; qx < Q1D; ++qx) {
const int q = QUAD_3D_ID(qx, qy, qz);
const double O11 = op(0, q, e);
const double O12 = op(1, q, e);
const double O13 = op(2, q, e);
const double O22 = op(3, q, e);
const double O23 = op(4, q, e);
const double O33 = op(5, q, e);
const double O11 = op(q, 0, e);
const double O12 = op(q, 1, e);
const double O13 = op(q, 2, e);
const double O22 = op(q, 3, e);
const double O23 = op(q, 4, e);
const double O33 = op(q, 5, e);
const double gradX = grad[qz][qy][qx][0];
const double gradY = grad[qz][qy][qx][1];
@@ -557,12 +557,12 @@ typedef double* SymmOperator3D_t @dim(6, Q3D, NE);
}
const int q = QUAD_3D_ID(qx, qy, qz);
const double O11 = op(0, q, e);
const double O12 = op(1, q, e);
const double O13 = op(2, q, e);
const double O22 = op(3, q, e);
const double O23 = op(4, q, e);
const double O33 = op(5, q, e);
const double O11 = op(q, 0, e);
const double O12 = op(q, 1, e);
const double O13 = op(q, 2, e);
const double O22 = op(q, 3, e);
const double O23 = op(q, 4, e);
const double O33 = op(q, 5, e);
const double qDxyz = (O11 * Dxyz) + (O12 * xDyz) + (O13 * xyDz);
const double qxDyz = (O12 * Dxyz) + (O22 * xDyz) + (O23 * xyDz);
+8 -1
View File
@@ -203,7 +203,14 @@ void ParBilinearForm::AssembleSharedFaces(int skip_zeros)
vdofs1.Copy(vdofs_all);
for (int j = 0; j < vdofs2.Size(); j++)
{
vdofs2[j] += height;
if (vdofs2[j] >= 0)
{
vdofs2[j] += height;
}
else
{
vdofs2[j] -= height;
}
}
vdofs_all.Append(vdofs2);
for (int k = 0; k < fbfi.Size(); k++)
+306 -36
View File
@@ -14,6 +14,7 @@
#ifdef MFEM_USE_MPI
#include "pfespace.hpp"
#include "../general/forall.hpp"
#include "../general/sort_pairs.hpp"
#include "../mesh/mesh_headers.hpp"
#include "../general/binaryio.hpp"
@@ -613,15 +614,15 @@ void ParFiniteElementSpace::Build_Dof_TrueDof_Matrix() const // matrix P
int ldof = GetVSize();
int ltdof = TrueVSize();
HYPRE_Int *i_diag = mfem::New<HYPRE_Int>(ldof+1);
HYPRE_Int *j_diag = mfem::New<HYPRE_Int>(ltdof);
HYPRE_Int *i_diag = new HYPRE_Int[ldof+1];
HYPRE_Int *j_diag = new HYPRE_Int[ltdof];
int diag_counter;
HYPRE_Int *i_offd = mfem::New<HYPRE_Int>(ldof+1);
HYPRE_Int *j_offd = mfem::New<HYPRE_Int>(ldof-ltdof);
HYPRE_Int *i_offd = new HYPRE_Int[ldof+1];
HYPRE_Int *j_offd = new HYPRE_Int[ldof-ltdof];
int offd_counter;
HYPRE_Int *cmap = mfem::New<HYPRE_Int>(ldof-ltdof);
HYPRE_Int *cmap = new HYPRE_Int[ldof-ltdof];
HYPRE_Int *col_starts = GetTrueDofOffsets();
HYPRE_Int *row_starts = GetDofOffsets();
@@ -747,12 +748,14 @@ void ParFiniteElementSpace::GetEssentialTrueDofs(const Array<int>
// Verify that in boolean arithmetic: P^T ess_dofs = R ess_dofs.
Array<int> true_ess_dofs2(true_ess_dofs.Size());
HypreParMatrix *Pt = Dof_TrueDof_Matrix()->Transpose();
Pt->BooleanMult(1, ess_dofs, 0, true_ess_dofs2);
const int *ess_dofs_data = ess_dofs.HostRead();
Pt->BooleanMult(1, ess_dofs_data, 0, true_ess_dofs2);
delete Pt;
int counter = 0;
const int *ted = true_ess_dofs.HostRead();
for (int i = 0; i < true_ess_dofs.Size(); i++)
{
if (bool(true_ess_dofs[i]) != bool(true_ess_dofs2[i])) { counter++; }
if (bool(ted[i]) != bool(true_ess_dofs2[i])) { counter++; }
}
MFEM_VERIFY(counter == 0, "internal MFEM error: counter = " << counter);
#endif
@@ -854,7 +857,20 @@ const Operator *ParFiniteElementSpace::GetProlongationMatrix() const
{
if (Conforming())
{
if (!Pconf) { Pconf = new ConformingProlongationOperator(*this); }
if (!Pconf)
{
if (!Device::Allows(Backend::DEVICE_MASK))
{
Pconf = new ConformingProlongationOperator(*this);
}
else
{
if (NRanks > 1)
{
Pconf = new DeviceConformingProlongationOperator(*this);
}
}
}
return Pconf;
}
else
@@ -902,11 +918,15 @@ void ParFiniteElementSpace::ExchangeFaceNbrData()
{
GetElementVDofs(my_elems[i], ldofs);
for (int j = 0; j < ldofs.Size(); j++)
if (ldof_marker[ldofs[j]] != fn)
{
int ldof = (ldofs[j] >= 0 ? ldofs[j] : -1-ldofs[j]);
if (ldof_marker[ldof] != fn)
{
ldof_marker[ldofs[j]] = fn;
ldof_marker[ldof] = fn;
send_face_nbr_ldof.AddAColumnInRow(fn);
}
}
send_nbr_elem_dof.AddColumnsInRow(send_el_off[fn] + i, ldofs.Size());
}
@@ -960,9 +980,11 @@ void ParFiniteElementSpace::ExchangeFaceNbrData()
GetElementVDofs(my_elems[i], ldofs);
for (int j = 0; j < ldofs.Size(); j++)
{
if (ldof_marker[ldofs[j]] != fn)
int ldof = (ldofs[j] >= 0 ? ldofs[j] : -1-ldofs[j]);
if (ldof_marker[ldof] != fn)
{
ldof_marker[ldofs[j]] = fn;
ldof_marker[ldof] = fn;
send_face_nbr_ldof.AddConnection(fn, ldofs[j]);
}
}
@@ -983,12 +1005,14 @@ void ParFiniteElementSpace::ExchangeFaceNbrData()
for (int i = 0; i < num_ldofs; i++)
{
ldof_marker[ldofs[i]] = i;
int ldof = (ldofs[i] >= 0 ? ldofs[i] : -1-ldofs[i]);
ldof_marker[ldof] = i;
}
for ( ; j < j_end; j++)
{
send_J[j] = ldof_marker[send_J[j]];
int ldof = (send_J[j] >= 0 ? send_J[j] : -1-send_J[j]);
send_J[j] = (send_J[j] >= 0 ? ldof_marker[ldof] : -1-ldof_marker[ldof]);
}
}
@@ -1023,7 +1047,14 @@ void ParFiniteElementSpace::ExchangeFaceNbrData()
for ( ; j < j_end; j++)
{
recv_J[j] += shift;
if (recv_J[j] >= 0)
{
recv_J[j] += shift;
}
else
{
recv_J[j] -= shift;
}
}
}
@@ -1072,8 +1103,15 @@ void ParFiniteElementSpace::ExchangeFaceNbrData()
for (int fn = 0, j = 0; fn < num_face_nbrs; fn++)
{
for (int j_end = face_nbr_ldof.GetI()[fn+1]; j < j_end; j++)
face_nbr_glob_dof_map[j] =
dof_face_nbr_offsets[fn] + face_nbr_ldof.GetJ()[j];
{
int ldof = face_nbr_ldof.GetJ()[j];
if (ldof < 0)
{
ldof = -1-ldof;
}
face_nbr_glob_dof_map[j] = dof_face_nbr_offsets[fn] + ldof;
}
}
MPI_Waitall(num_face_nbrs, send_requests, statuses);
@@ -2249,7 +2287,7 @@ HypreParMatrix* ParFiniteElementSpace
}
// create offd column mapping
HYPRE_Int *cmap = mfem::New<HYPRE_Int>(col_map.size());
HYPRE_Int *cmap = new HYPRE_Int[col_map.size()];
int offd_col = 0;
for (std::map<HYPRE_Int, int>::iterator
it = col_map.begin(); it != col_map.end(); ++it)
@@ -2258,14 +2296,14 @@ HypreParMatrix* ParFiniteElementSpace
it->second = offd_col++;
}
HYPRE_Int *I_diag = mfem::New<HYPRE_Int>(vdim*local_rows + 1);
HYPRE_Int *I_offd = mfem::New<HYPRE_Int>(vdim*local_rows + 1);
HYPRE_Int *I_diag = new HYPRE_Int[vdim*local_rows + 1];
HYPRE_Int *I_offd = new HYPRE_Int[vdim*local_rows + 1];
HYPRE_Int *J_diag = mfem::New<HYPRE_Int>(nnz_diag);
HYPRE_Int *J_offd = mfem::New<HYPRE_Int>(nnz_offd);
HYPRE_Int *J_diag = new HYPRE_Int[nnz_diag];
HYPRE_Int *J_offd = new HYPRE_Int[nnz_offd];
double *A_diag = mfem::New<double>(nnz_diag);
double *A_offd = mfem::New<double>(nnz_offd);
double *A_diag = new double[nnz_diag];
double *A_offd = new double[nnz_offd];
int vdim1 = bynodes ? vdim : 1;
int vdim2 = bynodes ? 1 : vdim;
@@ -2316,7 +2354,7 @@ HypreParMatrix* ParFiniteElementSpace
static HYPRE_Int* make_i_array(int nrows)
{
HYPRE_Int *I = mfem::New<HYPRE_Int>(nrows+1);
HYPRE_Int *I = new HYPRE_Int[nrows+1];
for (int i = 0; i <= nrows; i++) { I[i] = -1; }
return I;
}
@@ -2328,7 +2366,7 @@ static HYPRE_Int* make_j_array(HYPRE_Int* I, int nrows)
{
if (I[i] >= 0) { nnz++; }
}
HYPRE_Int *J = mfem::New<HYPRE_Int>(nnz);
HYPRE_Int *J = new HYPRE_Int[nnz];
I[nrows] = -1;
for (int i = 0, k = 0; i <= nrows; i++)
@@ -2427,7 +2465,7 @@ ParFiniteElementSpace::RebalanceMatrix(int old_ndofs,
}
SortPairs<HYPRE_Int, int>(cmap_offd, offd_cols);
HYPRE_Int* cmap = mfem::New<HYPRE_Int>(offd_cols);
HYPRE_Int* cmap = new HYPRE_Int[offd_cols];
for (int i = 0; i < offd_cols; i++)
{
cmap[i] = cmap_offd[i].one;
@@ -2623,7 +2661,7 @@ ParFiniteElementSpace::ParallelDerefinementMatrix(int old_ndofs,
offd->SetWidth(col_map.size());
// create offd column mapping for use by hypre
HYPRE_Int *cmap = mfem::New<HYPRE_Int>(offd->Width());
HYPRE_Int *cmap = new HYPRE_Int[offd->Width()];
for (std::map<HYPRE_Int, int>::iterator
it = col_map.begin(); it != col_map.end(); ++it)
{
@@ -2863,9 +2901,8 @@ void ConformingProlongationOperator::Mult(const Vector &x, Vector &y) const
MFEM_ASSERT(x.Size() == Width(), "");
MFEM_ASSERT(y.Size() == Height(), "");
const double *xdata = x.GetData();
double *ydata = y.GetData();
x.Pull();
const double *xdata = x.HostRead();
double *ydata = y.HostWrite();
const int m = external_ldofs.Size();
const int in_layout = 2; // 2 - input is ltdofs array
@@ -2882,7 +2919,6 @@ void ConformingProlongationOperator::Mult(const Vector &x, Vector &y) const
const int out_layout = 0; // 0 - output is ldofs array
gc.BcastEnd(ydata, out_layout);
y.Push();
}
void ConformingProlongationOperator::MultTranspose(
@@ -2891,9 +2927,8 @@ void ConformingProlongationOperator::MultTranspose(
MFEM_ASSERT(x.Size() == Height(), "");
MFEM_ASSERT(y.Size() == Width(), "");
const double *xdata = x.GetData();
double *ydata = y.GetData();
x.Pull();
const double *xdata = x.HostRead();
double *ydata = y.HostWrite();
const int m = external_ldofs.Size();
gc.ReduceBegin(xdata);
@@ -2909,7 +2944,242 @@ void ConformingProlongationOperator::MultTranspose(
const int out_layout = 2; // 2 - output is an array on all ltdofs
gc.ReduceEnd<double>(ydata, out_layout, GroupCommunicator::Sum);
y.Push();
}
DeviceConformingProlongationOperator::DeviceConformingProlongationOperator(
const ParFiniteElementSpace &pfes) :
ConformingProlongationOperator(pfes),
mpi_gpu_aware(Device::GetGPUAwareMPI())
{
MFEM_ASSERT(pfes.Conforming(), "internal error");
const SparseMatrix *R = pfes.GetRestrictionMatrix();
MFEM_ASSERT(R->Finalized(), "");
const int tdofs = R->Height();
MFEM_ASSERT(tdofs == pfes.GetTrueVSize(), "");
MFEM_ASSERT(tdofs == R->GetI()[tdofs], "");
ltdof_ldof = Array<int>(const_cast<int*>(R->GetJ()), tdofs);
ltdof_ldof.UseDevice();
{
Table nbr_ltdof;
gc.GetNeighborLTDofTable(nbr_ltdof);
const int nb_connections = nbr_ltdof.Size_of_connections();
shr_ltdof.SetSize(nb_connections);
shr_ltdof.CopyFrom(nbr_ltdof.GetJ());
shr_buf.SetSize(nb_connections);
shr_buf.UseDevice(true);
shr_buf_offsets = nbr_ltdof.GetI();
{
Array<int> shr_ltdof(nbr_ltdof.GetJ(), nb_connections);
Array<int> unique_ltdof(shr_ltdof);
unique_ltdof.Sort();
unique_ltdof.Unique();
// Note: the next loop modifies the J array of nbr_ltdof
for (int i = 0; i < shr_ltdof.Size(); i++)
{
shr_ltdof[i] = unique_ltdof.FindSorted(shr_ltdof[i]);
MFEM_ASSERT(shr_ltdof[i] != -1, "internal error");
}
Table unique_shr;
Transpose(shr_ltdof, unique_shr, unique_ltdof.Size());
unq_ltdof = Array<int>(unique_ltdof, unique_ltdof.Size());
unq_shr_i = Array<int>(unique_shr.GetI(), unique_shr.Size()+1);
unq_shr_j = Array<int>(unique_shr.GetJ(), unique_shr.Size_of_connections());
}
delete [] nbr_ltdof.GetJ();
nbr_ltdof.LoseData();
}
{
Table nbr_ldof;
gc.GetNeighborLDofTable(nbr_ldof);
const int nb_connections = nbr_ldof.Size_of_connections();
ext_ldof.SetSize(nb_connections);
ext_ldof.CopyFrom(nbr_ldof.GetJ());
ext_buf.SetSize(nb_connections);
ext_buf.UseDevice(true);
ext_buf_offsets = nbr_ldof.GetI();
delete [] nbr_ldof.GetJ();
nbr_ldof.LoseData();
}
const GroupTopology &gtopo = gc.GetGroupTopology();
int req_counter = 0;
for (int nbr = 1; nbr < gtopo.GetNumNeighbors(); nbr++)
{
const int send_offset = shr_buf_offsets[nbr];
const int send_size = shr_buf_offsets[nbr+1] - send_offset;
if (send_size > 0) { req_counter++; }
const int recv_offset = ext_buf_offsets[nbr];
const int recv_size = ext_buf_offsets[nbr+1] - recv_offset;
if (recv_size > 0) { req_counter++; }
}
requests = new MPI_Request[req_counter];
}
static void ExtractSubVector(const int N,
const Array<int> &indices,
const Vector &in, Vector &out)
{
auto y = out.Write();
const auto x = in.Read();
const auto I = indices.Read();
MFEM_FORALL(i, N, y[i] = x[I[i]];); // indices can be repeated
}
void DeviceConformingProlongationOperator::BcastBeginCopy(
const Vector &x) const
{
// shr_buf[i] = src[shr_ltdof[i]]
if (shr_ltdof.Size() == 0) { return; }
ExtractSubVector(shr_ltdof.Size(), shr_ltdof, x, shr_buf);
// If the above kernel is executed asynchronously, we should wait for it to
// complete
if (mpi_gpu_aware) { Device::Synchronize(); }
}
static void SetSubVector(const int N,
const Array<int> &indices,
const Vector &in, Vector &out)
{
auto y = out.Write();
const auto x = in.Read();
const auto I = indices.Read();
MFEM_FORALL(i, N, y[I[i]] = x[i];);
}
void DeviceConformingProlongationOperator::BcastLocalCopy(
const Vector &x, Vector &y) const
{
// dst[ltdof_ldof[i]] = src[i]
if (ltdof_ldof.Size() == 0) { return; }
SetSubVector(ltdof_ldof.Size(), ltdof_ldof, x, y);
}
void DeviceConformingProlongationOperator::BcastEndCopy(
Vector &y) const
{
// dst[ext_ldof[i]] = ext_buf[i]
if (ext_ldof.Size() == 0) { return; }
SetSubVector(ext_ldof.Size(), ext_ldof, ext_buf, y);
}
void DeviceConformingProlongationOperator::Mult(const Vector &x,
Vector &y) const
{
const GroupTopology &gtopo = gc.GetGroupTopology();
BcastBeginCopy(x); // copy to 'shr_buf'
int req_counter = 0;
for (int nbr = 1; nbr < gtopo.GetNumNeighbors(); nbr++)
{
const int send_offset = shr_buf_offsets[nbr];
const int send_size = shr_buf_offsets[nbr+1] - send_offset;
if (send_size > 0)
{
auto send_buf = mpi_gpu_aware ? shr_buf.Read() : shr_buf.HostRead();
MPI_Isend(send_buf + send_offset, send_size, MPI_DOUBLE,
gtopo.GetNeighborRank(nbr), 41822,
gtopo.GetComm(), &requests[req_counter++]);
}
const int recv_offset = ext_buf_offsets[nbr];
const int recv_size = ext_buf_offsets[nbr+1] - recv_offset;
if (recv_size > 0)
{
auto recv_buf = mpi_gpu_aware ? ext_buf.Write() : ext_buf.HostWrite();
MPI_Irecv(recv_buf + recv_offset, recv_size, MPI_DOUBLE,
gtopo.GetNeighborRank(nbr), 41822,
gtopo.GetComm(), &requests[req_counter++]);
}
}
BcastLocalCopy(x, y);
MPI_Waitall(req_counter, requests, MPI_STATUSES_IGNORE);
BcastEndCopy(y); // copy from 'ext_buf'
}
DeviceConformingProlongationOperator::~DeviceConformingProlongationOperator()
{
delete [] requests;
delete [] ext_buf_offsets;
delete [] shr_buf_offsets;
}
void DeviceConformingProlongationOperator::ReduceBeginCopy(
const Vector &x) const
{
// ext_buf[i] = src[ext_ldof[i]]
if (ext_ldof.Size() == 0) { return; }
ExtractSubVector(ext_ldof.Size(), ext_ldof, x, ext_buf);
// If the above kernel is executed asynchronously, we should wait for it to
// complete
if (mpi_gpu_aware) { Device::Synchronize(); }
}
void DeviceConformingProlongationOperator::ReduceLocalCopy(
const Vector &x, Vector &y) const
{
// dst[i] = src[ltdof_ldof[i]]
if (ltdof_ldof.Size() == 0) { return; }
ExtractSubVector(ltdof_ldof.Size(), ltdof_ldof, x, y);
}
static void AddSubVector(const int num_unique_dst_indices,
const Array<int> &unique_dst_indices,
const Array<int> &unique_to_src_offsets,
const Array<int> &unique_to_src_indices,
const Vector &src,
Vector &dst)
{
auto y = dst.Write();
const auto x = src.Read();
const auto DST_I = unique_dst_indices.Read();
const auto SRC_O = unique_to_src_offsets.Read();
const auto SRC_I = unique_to_src_indices.Read();
MFEM_FORALL(i, num_unique_dst_indices,
{
const int dst_idx = DST_I[i];
double sum = y[dst_idx];
const int end = SRC_O[i+1];
for (int j = SRC_O[i]; j != end; ++j) { sum += x[SRC_I[j]]; }
y[dst_idx] = sum;
});
}
void DeviceConformingProlongationOperator::ReduceEndAssemble(Vector &y) const
{
// dst[shr_ltdof[i]] += shr_buf[i]
const int unq_ltdof_size = unq_ltdof.Size();
if (unq_ltdof_size == 0) { return; }
AddSubVector(unq_ltdof_size, unq_ltdof, unq_shr_i, unq_shr_j, shr_buf, y);
}
void DeviceConformingProlongationOperator::MultTranspose(const Vector &x,
Vector &y) const
{
const GroupTopology &gtopo = gc.GetGroupTopology();
ReduceBeginCopy(x); // copy to 'ext_buf'
int req_counter = 0;
for (int nbr = 1; nbr < gtopo.GetNumNeighbors(); nbr++)
{
const int send_offset = ext_buf_offsets[nbr];
const int send_size = ext_buf_offsets[nbr+1] - send_offset;
if (send_size > 0)
{
auto send_buf = mpi_gpu_aware ? ext_buf.Read() : ext_buf.HostRead();
MPI_Isend(send_buf + send_offset, send_size, MPI_DOUBLE,
gtopo.GetNeighborRank(nbr), 41823,
gtopo.GetComm(), &requests[req_counter++]);
}
const int recv_offset = shr_buf_offsets[nbr];
const int recv_size = shr_buf_offsets[nbr+1] - recv_offset;
if (recv_size > 0)
{
auto recv_buf = mpi_gpu_aware ? shr_buf.Write() : shr_buf.HostWrite();
MPI_Irecv(recv_buf + recv_offset, recv_size, MPI_DOUBLE,
gtopo.GetNeighborRank(nbr), 41823,
gtopo.GetComm(), &requests[req_counter++]);
}
}
ReduceLocalCopy(x, y);
MPI_Waitall(req_counter, requests, MPI_STATUSES_IGNORE);
ReduceEndAssemble(y); // assemble from 'shr_buf'
}
} // namespace mfem
+46
View File
@@ -387,6 +387,52 @@ public:
virtual void MultTranspose(const Vector &x, Vector &y) const;
};
/// Auxiliary device class used by ParFiniteElementSpace.
class DeviceConformingProlongationOperator: public
ConformingProlongationOperator
{
protected:
bool mpi_gpu_aware;
Array<int> shr_ltdof, ext_ldof;
mutable Vector shr_buf, ext_buf;
int *shr_buf_offsets, *ext_buf_offsets;
Array<int> ltdof_ldof, unq_ltdof;
Array<int> unq_shr_i, unq_shr_j;
MPI_Request *requests;
// Kernel: copy ltdofs from 'src' to 'shr_buf' - prepare for send.
// shr_buf[i] = src[shr_ltdof[i]]
void BcastBeginCopy(const Vector &src) const;
// Kernel: copy ltdofs from 'src' to ldofs in 'dst'.
// dst[ltdof_ldof[i]] = src[i]
void BcastLocalCopy(const Vector &src, Vector &dst) const;
// Kernel: copy ext. dofs from 'ext_buf' to 'dst' - after recv.
// dst[ext_ldof[i]] = ext_buf[i]
void BcastEndCopy(Vector &dst) const;
// Kernel: copy ext. dofs from 'src' to 'ext_buf' - prepare for send.
// ext_buf[i] = src[ext_ldof[i]]
void ReduceBeginCopy(const Vector &src) const;
// Kernel: copy owned ldofs from 'src' to ltdofs in 'dst'.
// dst[i] = src[ltdof_ldof[i]]
void ReduceLocalCopy(const Vector &src, Vector &dst) const;
// Kernel: assemble dofs from 'shr_buf' into to 'dst' - after recv.
// dst[shr_ltdof[i]] += shr_buf[i]
void ReduceEndAssemble(Vector &dst) const;
public:
DeviceConformingProlongationOperator(const ParFiniteElementSpace &pfes);
virtual ~DeviceConformingProlongationOperator();
virtual void Mult(const Vector &x, Vector &y) const;
virtual void MultTranspose(const Vector &x, Vector &y) const;
};
}
#endif // MFEM_USE_MPI
+15 -14
View File
@@ -367,10 +367,10 @@ void ParGridFunction::ProjectDiscCoefficient(Coefficient &coeff, AvgType type)
GroupCommunicator &gcomm = pfes->GroupComm();
gcomm.Reduce<int>(zones_per_vdof, GroupCommunicator::Sum);
gcomm.Bcast(zones_per_vdof);
// Accumulate for all tdofs.
HypreParVector *tv = this->ParallelAssemble();
this->Distribute(tv);
delete tv;
// Accumulate for all vdofs.
gcomm.Reduce<double>(data, GroupCommunicator::Sum);
gcomm.Bcast<double>(data);
ComputeMeans(type, zones_per_vdof);
}
@@ -389,10 +389,10 @@ void ParGridFunction::ProjectDiscCoefficient(VectorCoefficient &vcoeff,
GroupCommunicator &gcomm = pfes->GroupComm();
gcomm.Reduce<int>(zones_per_vdof, GroupCommunicator::Sum);
gcomm.Bcast(zones_per_vdof);
// Accumulate for all tdofs.
HypreParVector *tv = this->ParallelAssemble();
this->Distribute(tv);
delete tv;
// Accumulate for all vdofs.
gcomm.Reduce<double>(data, GroupCommunicator::Sum);
gcomm.Bcast<double>(data);
ComputeMeans(type, zones_per_vdof);
}
@@ -425,8 +425,8 @@ void ParGridFunction::ProjectBdrCoefficient(
}
else
{
// FIXME: same as the conforming case after 'cut-mesh-groups-dev-*' is
// merged?
// TODO: is this the same as the conforming case (after the merge of
// cut-mesh-groups-dev)?
ComputeMeans(ARITHMETIC, values_counter);
}
#ifdef MFEM_DEBUG
@@ -469,8 +469,8 @@ void ParGridFunction::ProjectBdrCoefficientTangent(VectorCoefficient &vcoeff,
}
else
{
// FIXME: same as the conforming case after 'cut-mesh-groups-dev-*' is
// merged?
// TODO: is this the same as the conforming case (after the merge of
// cut-mesh-groups-dev)?
ComputeMeans(ARITHMETIC, values_counter);
}
#ifdef MFEM_DEBUG
@@ -487,16 +487,17 @@ void ParGridFunction::ProjectBdrCoefficientTangent(VectorCoefficient &vcoeff,
void ParGridFunction::Save(std::ostream &out) const
{
double *data_ = const_cast<double*>(HostRead());
for (int i = 0; i < size; i++)
{
if (pfes->GetDofSign(i) < 0) { data[i] = -data[i]; }
if (pfes->GetDofSign(i) < 0) { data_[i] = -data_[i]; }
}
GridFunction::Save(out);
for (int i = 0; i < size; i++)
{
if (pfes->GetDofSign(i) < 0) { data[i] = -data[i]; }
if (pfes->GetDofSign(i) < 0) { data_[i] = -data_[i]; }
}
}
+6 -6
View File
@@ -956,13 +956,13 @@ double TMOP_Integrator::GetElementEnergy(const FiniteElement &el,
Tpr->Attribute = T.Attribute;
Tpr->GetPointMat().Transpose(PMatI); // PointMat = PMatI^T
}
// FIXME: computing the coefficients 'coeff1' and 'coeff0' in physical
// coordinates means that, generally, the gradient and Hessian of the
// TMOP_Integrator will depend on the derivatives of the coefficients.
// TODO: computing the coefficients 'coeff1' and 'coeff0' in physical
// coordinates means that, generally, the gradient and Hessian of the
// TMOP_Integrator will depend on the derivatives of the coefficients.
//
// In some cases the coefficients are independent of any movement of
// the physical coordinates (i.e. changes in 'elfun'), e.g. when the
// coefficient is a ConstantCoefficient or a GridFunctionCoefficient.
// In some cases the coefficients are independent of any movement of
// the physical coordinates (i.e. changes in 'elfun'), e.g. when the
// coefficient is a ConstantCoefficient or a GridFunctionCoefficient.
for (int i = 0; i < ir->GetNPoints(); i++)
{
+5 -43
View File
@@ -19,54 +19,12 @@
namespace mfem
{
BaseArray::BaseArray(int asize, int ainc, int elementsize)
{
if (asize > 0)
{
data = mfem::New<char>(asize * elementsize);
size = allocsize = asize;
}
else
{
data = 0;
size = allocsize = 0;
}
inc = ainc;
}
BaseArray::~BaseArray()
{
if (allocsize > 0)
{
mfem::Delete((char*)data);
}
}
void BaseArray::GrowSize(int minsize, int elementsize)
{
void *p;
int nsize = (inc > 0) ? abs(allocsize) + inc : 2 * abs(allocsize);
if (nsize < minsize) { nsize = minsize; }
p = mfem::New<char>(nsize * elementsize);
if (size > 0)
{
mfem::Memcpy(p, data, size * elementsize);
}
if (allocsize > 0)
{
mfem::Delete((char*)data);
}
data = p;
allocsize = nsize;
}
template <class T>
void Array<T>::Print(std::ostream &out, int width) const
{
for (int i = 0; i < size; i++)
{
out << ((T*)data)[i];
out << data[i];
if ( !((i+1) % width) || i+1 == size )
{
out << '\n';
@@ -113,10 +71,12 @@ T Array<T>::Max() const
T max = operator[](0);
for (int i = 1; i < size; i++)
{
if (max < operator[](i))
{
max = operator[](i);
}
}
return max;
}
@@ -128,10 +88,12 @@ T Array<T>::Min() const
T min = operator[](0);
for (int i = 1; i < size; i++)
{
if (operator[](i) < min)
{
min = operator[](i);
}
}
return min;
}
+188 -108
View File
@@ -14,6 +14,7 @@
#include "../config/config.hpp"
#include "mem_manager.hpp"
#include "device.hpp"
#include "error.hpp"
#include "globals.hpp"
@@ -25,31 +26,6 @@
namespace mfem
{
/// Base class for array container.
class BaseArray
{
protected:
/// Pointer to data
void *data;
/// Size of the array
int size;
/// Size of the allocated memory
int allocsize;
/** Increment of allocated memory on overflow,
inc = 0 doubles the array */
int inc;
BaseArray() { }
/// Creates array of asize elements of size elementsize
BaseArray(int asize, int ainc, int elmentsize);
/// Free the allocated memory
~BaseArray();
/** Increases the allocsize of the array to be at least minsize.
The current content of the array is copied to the newly allocated
space. minsize must be > abs(allocsize). */
void GrowSize(int minsize, int elementsize);
};
template <class T>
class Array;
@@ -65,70 +41,81 @@ void Swap(Array<T> &, Array<T> &);
The elements can be accessed by the [] operator, the range is 0 to size-1.
*/
template <class T>
class Array : public BaseArray
class Array
{
protected:
/// Pointer to data
Memory<T> data;
/// Size of the array
int size;
inline void GrowSize(int minsize);
public:
friend void Swap<T>(Array<T> &, Array<T> &);
/// Creates an empty array
inline Array() : size(0) { data.Reset(); }
/// Creates array of asize elements
explicit inline Array(int asize = 0, int ainc = 0)
: BaseArray(asize, ainc, sizeof (T)) { }
explicit inline Array(int asize)
: size(asize) { asize > 0 ? data.New(asize) : data.Reset(); }
/** Creates array using an existing c-array of asize elements;
allocsize is set to -asize to indicate that the data will not
be deleted. */
inline Array(T *_data, int asize, int ainc = 0)
{ data = _data; size = asize; allocsize = -asize; inc = ainc; }
inline Array(T *_data, int asize)
{ data.Wrap(_data, asize, false); size = asize; }
/// Copy constructor: deep copy
Array(const Array<T> &src)
: BaseArray(src.size, 0, sizeof(T))
{ mfem::Memcpy(data, src.data, size*sizeof(T)); }
/** This method supports source arrays using any MemoryType. */
inline Array(const Array &src);
/// Copy constructor (deep copy) from an Array of convertable type
template <typename CT>
Array(const Array<CT> &src)
: BaseArray(src.Size(), 0, sizeof(T))
{ for (int i = 0; i < size; i++) { (*this)[i] = T(src[i]); } }
inline Array(const Array<CT> &src);
/// Destructor
inline ~Array() { }
inline ~Array() { data.Delete(); }
/// Assignment operator: deep copy
Array<T> &operator=(const Array<T> &src) { src.Copy(*this); return *this; }
/// Assignment operator (deep copy) from an Array of convertable type
template <typename CT>
Array<T> &operator=(const Array<CT> &src)
{
SetSize(src.Size());
for (int i = 0; i < size; i++) { (*this)[i] = T(src[i]); }
return *this;
}
inline Array &operator=(const Array<CT> &src);
/// Return the data as 'T *'
inline operator T *() { return (T *)data; }
inline operator T *() { return data; }
/// Return the data as 'const T *'
inline operator const T *() const { return (const T *)data; }
inline operator const T *() const { return data; }
/// Returns the data
inline T *GetData() { return (T *)data; }
inline T *GetData() { return data; }
/// Returns the data
inline const T *GetData() const { return (T *)data; }
inline const T *GetData() const { return data; }
/// Return a reference to the Memory object used by the Array.
Memory<T> &GetMemory() { return data; }
/// Return a reference to the Memory object used by the Array, const version.
const Memory<T> &GetMemory() const { return data; }
/// Return the device flag of the Memory object used by the Array
bool UseDevice() const { return data.UseDevice(); }
/// Return true if the data will be deleted by the array
inline bool OwnsData() const { return (allocsize > 0); }
inline bool OwnsData() const { return data.OwnsHostPtr(); }
/// Changes the ownership of the data
inline void StealData(T **p)
{ *p = (T*)data; data = 0; size = allocsize = 0; }
inline void StealData(T **p) { *p = data; data.Reset(); size = 0; }
/// NULL-ifies the data
inline void LoseData() { data = 0; size = allocsize = 0; }
inline void LoseData() { data.Reset(); size = 0; }
/// Make the Array own the data
void MakeDataOwner() { allocsize = abs(allocsize); }
void MakeDataOwner() const { data.SetHostPtrOwner(true); }
/// Logical size of the array
inline int Size() const { return size; }
@@ -139,13 +126,18 @@ public:
/// Same as SetSize(int) plus initialize new entries with 'initval'
inline void SetSize(int nsize, const T &initval);
/** @brief Resize the array to size @a nsize using MemoryType @a mt. Note
that unlike the other versions of SetSize(), the current content of the
array is not preserved. */
inline void SetSize(int nsize, MemoryType mt);
/** Maximum number of entries the array can store without allocating more
memory. */
inline int Capacity() const { return abs(allocsize); }
inline int Capacity() const { return data.Capacity(); }
/// Ensures that the allocated size is at least the given size.
inline void Reserve(int capacity)
{ if (capacity > abs(allocsize)) { GrowSize(capacity, sizeof(T)); } }
{ if (capacity > Capacity()) { GrowSize(capacity); } }
/// Access element
inline T & operator[](int i);
@@ -188,11 +180,7 @@ public:
inline void DeleteAll();
/// Create a copy of the current array
inline void Copy(Array &copy) const
{
copy.SetSize(Size());
mfem::Memcpy(copy.GetData(), data, Size()*sizeof(T));
}
inline void Copy(Array &copy) const;
/// Make this Array a reference to a pointer
inline void MakeRef(T *, int);
@@ -200,7 +188,7 @@ public:
/// Make this Array a reference to 'master'
inline void MakeRef(const Array &master);
inline void GetSubArray(int offset, int sa_size, Array<T> &sa);
inline void GetSubArray(int offset, int sa_size, Array<T> &sa) const;
/// Prints array to stream with width elements per row
void Print(std::ostream &out = mfem::out, int width = 4) const;
@@ -235,18 +223,18 @@ public:
T Min() const;
/// Sorts the array. This requires operator< to be defined for T.
void Sort() { std::sort((T*) data, (T*) data + size); }
void Sort() { std::sort((T*)data, data + size); }
/// Sorts the array using the supplied comparison function object.
template<class Compare>
void Sort(Compare cmp) { std::sort((T*) data, (T*) data + size, cmp); }
void Sort(Compare cmp) { std::sort((T*)data, data + size, cmp); }
/** Removes duplicities from a sorted array. This requires operator== to be
defined for T. */
void Unique()
{
T* end = std::unique((T*) data, (T*) data + size);
SetSize(end - (T*) data);
T* end = std::unique((T*)data, data + size);
SetSize(end - data);
}
/// return true if the array is sorted.
@@ -266,11 +254,41 @@ public:
template <typename U>
inline void CopyTo(U *dest) { std::copy(begin(), end(), dest); }
template <typename U>
inline void CopyFrom(const U *src)
{ std::memcpy(begin(), src, MemoryUsage()); }
// STL-like begin/end
inline T* begin() const { return (T*) data; }
inline T* end() const { return (T*) data + size; }
inline T* begin() { return data; }
inline T* end() { return data + size; }
inline const T* begin() const { return data; }
inline const T* end() const { return data + size; }
long MemoryUsage() const { return Capacity() * sizeof(T); }
/// Shortcut for mfem::Read(a.GetMemory(), a.Size(), on_dev).
const T *Read(bool on_dev = true) const
{ return mfem::Read(data, size, on_dev); }
/// Shortcut for mfem::Read(a.GetMemory(), a.Size(), false).
const T *HostRead() const
{ return mfem::Read(data, size, false); }
/// Shortcut for mfem::Write(a.GetMemory(), a.Size(), on_dev).
T *Write(bool on_dev = true)
{ return mfem::Write(data, size, on_dev); }
/// Shortcut for mfem::Write(a.GetMemory(), a.Size(), false).
T *HostWrite()
{ return mfem::Write(data, size, false); }
/// Shortcut for mfem::ReadWrite(a.GetMemory(), a.Size(), on_dev).
T *ReadWrite(bool on_dev = true)
{ return mfem::ReadWrite(data, size, on_dev); }
/// Shortcut for mfem::ReadWrite(a.GetMemory(), a.Size(), false).
T *HostReadWrite()
{ return mfem::ReadWrite(data, size, false); }
};
template <class T>
@@ -278,7 +296,9 @@ inline bool operator==(const Array<T> &LHS, const Array<T> &RHS)
{
if ( LHS.Size() != RHS.Size() ) { return false; }
for (int i=0; i<LHS.Size(); i++)
{
if ( LHS[i] != RHS[i] ) { return false; }
}
return true;
}
@@ -565,17 +585,51 @@ inline void Swap(Array<T> &a, Array<T> &b)
{
Swap(a.data, b.data);
Swap(a.size, b.size);
Swap(a.allocsize, b.allocsize);
Swap(a.inc, b.inc);
}
template <class T>
inline Array<T>::Array(const Array &src)
: size(src.Size())
{
size > 0 ? data.New(size, src.data.GetMemoryType()) : data.Reset();
data.CopyFrom(src.data, size);
data.UseDevice(src.data.UseDevice());
}
template <typename T> template <typename CT>
inline Array<T>::Array(const Array<CT> &src)
: size(src.Size())
{
size > 0 ? data.New(size) : data.Reset();
for (int i = 0; i < size; i++) { (*this)[i] = T(src[i]); }
}
template <class T>
inline void Array<T>::GrowSize(int minsize)
{
const int nsize = std::max(minsize, 2 * data.Capacity());
Memory<T> p(nsize, data.GetMemoryType());
p.CopyFrom(data, size);
p.UseDevice(data.UseDevice());
data.Delete();
data = p;
}
template <typename T> template <typename CT>
inline Array<T> &Array<T>::operator=(const Array<CT> &src)
{
SetSize(src.Size());
for (int i = 0; i < size; i++) { (*this)[i] = T(src[i]); }
return *this;
}
template <class T>
inline void Array<T>::SetSize(int nsize)
{
MFEM_ASSERT( nsize>=0, "Size must be non-negative. It is " << nsize );
if (nsize > abs(allocsize))
if (nsize > Capacity())
{
GrowSize(nsize, sizeof(T));
GrowSize(nsize);
}
size = nsize;
}
@@ -586,24 +640,51 @@ inline void Array<T>::SetSize(int nsize, const T &initval)
MFEM_ASSERT( nsize>=0, "Size must be non-negative. It is " << nsize );
if (nsize > size)
{
if (nsize > abs(allocsize))
if (nsize > Capacity())
{
GrowSize(nsize, sizeof(T));
GrowSize(nsize);
}
for (int i = size; i < nsize; i++)
{
((T*)data)[i] = initval;
data[i] = initval;
}
}
size = nsize;
}
template <class T>
inline void Array<T>::SetSize(int nsize, MemoryType mt)
{
MFEM_ASSERT(nsize >= 0, "invalid new size: " << nsize);
if (mt == data.GetMemoryType())
{
if (nsize <= Capacity())
{
size = nsize;
return;
}
}
const bool use_dev = data.UseDevice();
data.Delete();
if (nsize > 0)
{
data.New(nsize, mt);
size = nsize;
}
else
{
data.Reset();
size = 0;
}
data.UseDevice(use_dev);
}
template <class T>
inline T &Array<T>::operator[](int i)
{
MFEM_ASSERT( i>=0 && i<size,
"Access element " << i << " of array, size = " << size );
return ((T*)data)[i];
return data[i];
}
template <class T>
@@ -611,14 +692,14 @@ inline const T &Array<T>::operator[](int i) const
{
MFEM_ASSERT( i>=0 && i<size,
"Access element " << i << " of array, size = " << size );
return ((T*)data)[i];
return data[i];
}
template <class T>
inline int Array<T>::Append(const T &el)
{
SetSize(size+1);
((T*)data)[size-1] = el;
data[size-1] = el;
return size;
}
@@ -630,7 +711,7 @@ inline int Array<T>::Append(const T *els, int nels)
SetSize(size + nels);
for (int i = 0; i < nels; i++)
{
((T*)data)[old_size+i] = els[i];
data[old_size+i] = els[i];
}
return size;
}
@@ -641,9 +722,9 @@ inline int Array<T>::Prepend(const T &el)
SetSize(size+1);
for (int i = size-1; i > 0; i--)
{
((T*)data)[i] = ((T*)data)[i-1];
data[i] = data[i-1];
}
((T*)data)[0] = el;
data[0] = el;
return size;
}
@@ -651,21 +732,21 @@ template <class T>
inline T &Array<T>::Last()
{
MFEM_ASSERT(size > 0, "Array size is zero: " << size);
return ((T*)data)[size-1];
return data[size-1];
}
template <class T>
inline const T &Array<T>::Last() const
{
MFEM_ASSERT(size > 0, "Array size is zero: " << size);
return ((T*)data)[size-1];
return data[size-1];
}
template <class T>
inline int Array<T>::Union(const T &el)
{
int i = 0;
while ((i < size) && (((T*)data)[i] != el)) { i++; }
while ((i < size) && (data[i] != el)) { i++; }
if (i == size)
{
Append(el);
@@ -678,7 +759,7 @@ inline int Array<T>::Find(const T &el) const
{
for (int i = 0; i < size; i++)
{
if (((T*)data)[i] == el) { return i; }
if (data[i] == el) { return i; }
}
return -1;
}
@@ -686,7 +767,7 @@ inline int Array<T>::Find(const T &el) const
template <class T>
inline int Array<T>::FindSorted(const T &el) const
{
const T *begin = (const T*) data, *end = begin + size;
const T *begin = data, *end = begin + size;
const T* first = std::lower_bound(begin, end, el);
if (first == end || !(*first == el)) { return -1; }
return first - begin;
@@ -697,11 +778,11 @@ inline void Array<T>::DeleteFirst(const T &el)
{
for (int i = 0; i < size; i++)
{
if (((T*)data)[i] == el)
if (data[i] == el)
{
for (i++; i < size; i++)
{
((T*)data)[i-1] = ((T*)data)[i];
data[i-1] = data[i];
}
size--;
return;
@@ -712,41 +793,40 @@ inline void Array<T>::DeleteFirst(const T &el)
template <class T>
inline void Array<T>::DeleteAll()
{
if (allocsize > 0)
{
mfem::Delete((char*)data);
}
data = NULL;
size = allocsize = 0;
const bool use_dev = data.UseDevice();
data.Delete();
data.Reset();
size = 0;
data.UseDevice(use_dev);
}
template <typename T>
inline void Array<T>::Copy(Array &copy) const
{
copy.SetSize(Size(), data.GetMemoryType());
data.CopyTo(copy.data, Size());
copy.data.UseDevice(data.UseDevice());
}
template <class T>
inline void Array<T>::MakeRef(T *p, int s)
{
if (allocsize > 0)
{
mfem::Delete((char*)data);
}
data = p;
data.Delete();
data.Wrap(p, s, false);
size = s;
allocsize = -s;
}
template <class T>
inline void Array<T>::MakeRef(const Array &master)
{
if (allocsize > 0)
{
mfem::Delete((char*)data);
}
data = master.data;
data.Delete();
data = master.data; // note: copies the device flag
size = master.size;
allocsize = -abs(master.allocsize);
inc = master.inc;
data.ClearOwnerFlags();
}
template <class T>
inline void Array<T>::GetSubArray(int offset, int sa_size, Array<T> &sa)
inline void Array<T>::GetSubArray(int offset, int sa_size, Array<T> &sa) const
{
sa.SetSize(sa_size);
for (int i = 0; i < sa_size; i++)
@@ -760,14 +840,14 @@ inline void Array<T>::operator=(const T &a)
{
for (int i = 0; i < size; i++)
{
((T*)data)[i] = a;
data[i] = a;
}
}
template <class T>
inline void Array<T>::Assign(const T *p)
{
memcpy(data, p, Size()*sizeof(T));
data.CopyFromHost(p, Size());
}
+72
View File
@@ -513,6 +513,78 @@ void GroupCommunicator::SetLTDofTable(const Array<int> &ldof_ltdof)
group_ltdof.ShiftUpI();
}
void GroupCommunicator::GetNeighborLTDofTable(Table &nbr_ltdof) const
{
nbr_ltdof.MakeI(nbr_send_groups.Size());
for (int nbr = 1; nbr < nbr_send_groups.Size(); nbr++)
{
const int num_send_groups = nbr_send_groups.RowSize(nbr);
if (num_send_groups > 0)
{
const int *grp_list = nbr_send_groups.GetRow(nbr);
for (int i = 0; i < num_send_groups; i++)
{
const int group = grp_list[i];
const int nltdofs = group_ltdof.RowSize(group);
nbr_ltdof.AddColumnsInRow(nbr, nltdofs);
}
}
}
nbr_ltdof.MakeJ();
for (int nbr = 1; nbr < nbr_send_groups.Size(); nbr++)
{
const int num_send_groups = nbr_send_groups.RowSize(nbr);
if (num_send_groups > 0)
{
const int *grp_list = nbr_send_groups.GetRow(nbr);
for (int i = 0; i < num_send_groups; i++)
{
const int group = grp_list[i];
const int nltdofs = group_ltdof.RowSize(group);
const int *ltdofs = group_ltdof.GetRow(group);
nbr_ltdof.AddConnections(nbr, ltdofs, nltdofs);
}
}
}
nbr_ltdof.ShiftUpI();
}
void GroupCommunicator::GetNeighborLDofTable(Table &nbr_ldof) const
{
nbr_ldof.MakeI(nbr_recv_groups.Size());
for (int nbr = 1; nbr < nbr_recv_groups.Size(); nbr++)
{
const int num_recv_groups = nbr_recv_groups.RowSize(nbr);
if (num_recv_groups > 0)
{
const int *grp_list = nbr_recv_groups.GetRow(nbr);
for (int i = 0; i < num_recv_groups; i++)
{
const int group = grp_list[i];
const int nldofs = group_ldof.RowSize(group);
nbr_ldof.AddColumnsInRow(nbr, nldofs);
}
}
}
nbr_ldof.MakeJ();
for (int nbr = 1; nbr < nbr_recv_groups.Size(); nbr++)
{
const int num_recv_groups = nbr_recv_groups.RowSize(nbr);
if (num_recv_groups > 0)
{
const int *grp_list = nbr_recv_groups.GetRow(nbr);
for (int i = 0; i < num_recv_groups; i++)
{
const int group = grp_list[i];
const int nldofs = group_ldof.RowSize(group);
const int *ldofs = group_ldof.GetRow(group);
nbr_ldof.AddConnections(nbr, ldofs, nldofs);
}
}
}
nbr_ldof.ShiftUpI();
}
template <class T>
T *GroupCommunicator::CopyGroupToBuffer(const T *ldata, T *buf, int group,
int layout) const
+6
View File
@@ -179,6 +179,12 @@ public:
/// Get a const reference to the associated GroupTopology object
const GroupTopology &GetGroupTopology() const { return gtopo; }
/// Dofs to be sent to communication neighbors
void GetNeighborLTDofTable(Table &nbr_ltdof) const;
/// Dofs to be received from communication neighbors
void GetNeighborLDofTable(Table &nbr_ldof) const;
/** @brief Data structure on which we define reduce operations.
The data is associated with (and the operation is performed on) one group
+74 -42
View File
@@ -10,17 +10,38 @@
// Software Foundation) version 2.1 dated February 1999.
#include "cuda.hpp"
#include "globals.hpp"
namespace mfem
{
// Internal debug option, useful for tracking CUDA allocations, deallocations
// and transfers.
// #define MFEM_TRACK_CUDA_MEM
#ifdef MFEM_USE_CUDA
void mfem_cuda_error(cudaError_t err, const char *expr, const char *func,
const char *file, int line)
{
mfem::err << "\n\nCUDA error: (" << expr << ") failed with error:\n --> "
<< cudaGetErrorString(err)
<< "\n ... in function: " << func
<< "\n ... in file: " << file << ':' << line << '\n';
mfem_error();
}
#endif
void* CuMemAlloc(void** dptr, size_t bytes)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS != ::cuMemAlloc((CUdeviceptr*)dptr, bytes))
{
mfem_error("Error in CuMemAlloc");
}
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "CuMemAlloc(): allocating " << bytes << " bytes ... "
<< std::flush;
#endif
MFEM_GPU_CHECK(cudaMalloc(dptr, bytes));
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "done: " << *dptr << std::endl;
#endif
#endif
return *dptr;
}
@@ -28,10 +49,14 @@ void* CuMemAlloc(void** dptr, size_t bytes)
void* CuMemFree(void *dptr)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS != ::cuMemFree((CUdeviceptr)dptr))
{
mfem_error("Error in CuMemFree");
}
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "CuMemFree(): deallocating memory @ " << dptr << " ... "
<< std::flush;
#endif
MFEM_GPU_CHECK(cudaFree(dptr));
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dptr;
}
@@ -39,72 +64,79 @@ void* CuMemFree(void *dptr)
void* CuMemcpyHtoD(void* dst, const void* src, size_t bytes)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS != ::cuMemcpyHtoD((CUdeviceptr)dst, src, bytes))
{
mfem_error("Error in CuMemcpyHtoD");
}
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "CuMemcpyHtoD(): copying " << bytes << " bytes from "
<< src << " to " << dst << " ... " << std::flush;
#endif
MFEM_GPU_CHECK(cudaMemcpy(dst, src, bytes, cudaMemcpyHostToDevice));
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dst;
}
void* CuMemcpyHtoDAsync(void* dst, const void* src, size_t bytes, void *s)
void* CuMemcpyHtoDAsync(void* dst, const void* src, size_t bytes)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS !=
::cuMemcpyHtoDAsync((CUdeviceptr)dst, src, bytes, (CUstream)s))
{
mfem_error("Error in CuMemcpyHtoDAsync");
}
MFEM_GPU_CHECK(cudaMemcpyAsync(dst, src, bytes, cudaMemcpyHostToDevice));
#endif
return dst;
}
void* CuMemcpyDtoD(void* dst, void* src, size_t bytes)
void* CuMemcpyDtoD(void *dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS !=
::cuMemcpyDtoD((CUdeviceptr)dst, (CUdeviceptr)src, bytes))
{
mfem_error("Error in CuMemcpyDtoD");
}
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "CuMemcpyDtoD(): copying " << bytes << " bytes from "
<< src << " to " << dst << " ... " << std::flush;
#endif
MFEM_GPU_CHECK(cudaMemcpy(dst, src, bytes, cudaMemcpyDeviceToDevice));
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dst;
}
void* CuMemcpyDtoDAsync(void* dst, void* src, size_t bytes, void *s)
void* CuMemcpyDtoDAsync(void* dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS !=
::cuMemcpyDtoDAsync((CUdeviceptr)dst, (CUdeviceptr)src,
bytes, (CUstream)s))
{
mfem_error("Error in CuMemcpyDtoDAsync");
}
MFEM_GPU_CHECK(cudaMemcpyAsync(dst, src, bytes, cudaMemcpyDeviceToDevice));
#endif
return dst;
}
void* CuMemcpyDtoH(void *dst, void *src, size_t bytes)
void* CuMemcpyDtoH(void *dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS != ::cuMemcpyDtoH(dst, (CUdeviceptr)src, bytes))
{
mfem_error("Error in CuMemcpyDtoH");
}
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "CuMemcpyDtoH(): copying " << bytes << " bytes from "
<< src << " to " << dst << " ... " << std::flush;
#endif
MFEM_GPU_CHECK(cudaMemcpy(dst, src, bytes, cudaMemcpyDeviceToHost));
#ifdef MFEM_TRACK_CUDA_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dst;
}
void* CuMemcpyDtoHAsync(void* dst, void* src, size_t bytes, void *s)
void* CuMemcpyDtoHAsync(void *dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_CUDA
if (CUDA_SUCCESS !=
::cuMemcpyDtoHAsync(dst, (CUdeviceptr)src, bytes, (CUstream)s))
{
mfem_error("Error in CuMemcpyDtoHAsync");
}
MFEM_GPU_CHECK(cudaMemcpyAsync(dst, src, bytes, cudaMemcpyDeviceToHost));
#endif
return dst;
}
int CuGetDeviceCount()
{
int num_gpus = -1;
#ifdef MFEM_USE_CUDA
MFEM_GPU_CHECK(cudaGetDeviceCount(&num_gpus));
#endif
return num_gpus;
}
} // namespace mfem
+37 -75
View File
@@ -24,91 +24,51 @@
#define MFEM_CUDA_BLOCKS 256
#ifdef MFEM_USE_CUDA
#define MFEM_ATTR_DEVICE __device__
#define MFEM_ATTR_HOST_DEVICE __host__ __device__
// Define the CUDA debug macros:
// - MFEM_CUDA_CHECK_DRV(x) where 'x' returns/is type 'CUresult'
// - MFEM_CUDA_CHECK_RT(x) where 'x' returns/is type 'cudaError_t'
#ifdef MFEM_DEBUG
#define MFEM_CUDA_CHECK_DRV(x) \
do \
{ \
CUresult err = (x); \
if (err != CUDA_SUCCESS) \
{ \
const char *error_string; \
cuGetErrorString(err, &error_string); \
_MFEM_MESSAGE("CUDA error: (" << #x \
<< ") failed with error:\n --> " \
<< error_string, 0); \
} \
} \
while (0)
#define MFEM_CUDA_CHECK_RT(x) \
#define MFEM_DEVICE __device__
#define MFEM_HOST_DEVICE __host__ __device__
// Define a CUDA error check macro, MFEM_GPU_CHECK(x), where x returns/is of
// type 'cudaError_t'. This macro evaluates 'x' and raises an error if the
// result is not cudaSuccess.
#define MFEM_GPU_CHECK(x) \
do \
{ \
cudaError_t err = (x); \
if (err != cudaSuccess) \
{ \
_MFEM_MESSAGE("CUDA error: (" << #x \
<< ") failed with error:\n --> " \
<< cudaGetErrorString(err), 0); \
mfem_cuda_error(err, #x, _MFEM_FUNC_NAME, __FILE__, __LINE__); \
} \
} \
while (0)
#define MFEM_DEVICE_SYNC MFEM_GPU_CHECK(cudaDeviceSynchronize())
#else
#define MFEM_CUDA_CHECK_DRV(x) x
#define MFEM_CUDA_CHECK_RT(x) x
#endif
#else // MFEM_USE_CUDA
#define MFEM_ATTR_DEVICE
#define MFEM_ATTR_HOST_DEVICE
typedef int CUdevice;
typedef int CUcontext;
typedef void* CUstream;
#define MFEM_DEVICE
#define MFEM_HOST_DEVICE
#define MFEM_DEVICE_SYNC
#endif // MFEM_USE_CUDA
// Define the MFEM inner threading macros
#if defined(MFEM_USE_CUDA) && defined(__CUDA_ARCH__)
#define MFEM_SHARED __shared__
#define MFEM_SYNC_THREAD __syncthreads()
#define MFEM_THREAD_ID(k) threadIdx.k
#define MFEM_THREAD_SIZE(k) blockDim.k
#define MFEM_FOREACH_THREAD(i,k,N) for(int i=threadIdx.k; i<N; i+=blockDim.k)
#else
#define MFEM_SHARED
#define MFEM_SYNC_THREAD
#define MFEM_THREAD_ID(k) 0
#define MFEM_THREAD_SIZE(k) 1
#define MFEM_FOREACH_THREAD(i,k,N) for(int i=0; i<N; i++)
#endif
namespace mfem
{
// Define 'atomicAdd' function.
#ifdef __CUDA_ARCH__
#if __CUDA_ARCH__ < 600
static __device__ inline double atomicAdd(double* address, double val)
{
unsigned long long int* address_as_ull = (unsigned long long int*)address;
unsigned long long int old = *address_as_ull, assumed;
do
{
assumed = old;
old =
atomicCAS(address_as_ull, assumed,
__double_as_longlong(val +
__longlong_as_double(assumed)));
// Note: uses integer comparison to avoid hang in case of NaN
// (since NaN != NaN)
}
while (assumed != old);
return __longlong_as_double(old);
}
#endif // __CUDA_ARCH__ < 600
template<typename T> MFEM_ATTR_DEVICE
inline T AtomicAdd(T volatile *address, T val)
{
return atomicAdd((T *)address, val);
}
#else // __CUDA_ARCH__
template<typename T> inline T AtomicAdd(T volatile *address, T val)
{
#ifdef MFEM_USE_OPENMP
#pragma omp atomic
#ifdef MFEM_USE_CUDA
// Function used by the macro MFEM_GPU_CHECK.
void mfem_cuda_error(cudaError_t err, const char *expr, const char *func,
const char *file, int line);
#endif
*address += val;
return *address;
}
#endif // __CUDA_ARCH__
/// Allocates device memory
void* CuMemAlloc(void **d_ptr, size_t bytes);
@@ -120,20 +80,22 @@ void* CuMemFree(void *d_ptr);
void* CuMemcpyHtoD(void *d_dst, const void *h_src, size_t bytes);
/// Copies memory from Host to Device
void* CuMemcpyHtoDAsync(void *d_dst, const void *h_src,
size_t bytes, void *stream);
void* CuMemcpyHtoDAsync(void *d_dst, const void *h_src, size_t bytes);
/// Copies memory from Device to Device
void* CuMemcpyDtoD(void *d_dst, void *d_src, size_t bytes);
void* CuMemcpyDtoD(void *d_dst, const void *d_src, size_t bytes);
/// Copies memory from Device to Device
void* CuMemcpyDtoDAsync(void *d_dst, void *d_src, size_t bytes, void *stream);
void* CuMemcpyDtoDAsync(void *d_dst, const void *d_src, size_t bytes);
/// Copies memory from Device to Host
void* CuMemcpyDtoH(void *h_dst, void *d_src, size_t bytes);
void* CuMemcpyDtoH(void *h_dst, const void *d_src, size_t bytes);
/// Copies memory from Device to Host
void* CuMemcpyDtoHAsync(void *h_dst, void *d_src, size_t bytes, void *stream);
void* CuMemcpyDtoHAsync(void *h_dst, const void *d_src, size_t bytes);
/// Get the number of CUDA devices
int CuGetDeviceCount();
} // namespace mfem
+74 -36
View File
@@ -24,15 +24,16 @@ namespace mfem
namespace internal
{
CUstream *cuStream = NULL;
static CUdevice cuDevice;
static CUcontext cuContext;
OccaDevice occaDevice;
#ifdef MFEM_USE_OCCA
// Default occa::device used by MFEM.
occa::device occaDevice;
#endif
// Backends listed by priority, high to low:
static const Backend::Id backend_list[Backend::NUM_BACKENDS] =
{
Backend::OCCA_CUDA, Backend::RAJA_CUDA, Backend::CUDA,
Backend::HIP,
Backend::OCCA_OMP, Backend::RAJA_OMP, Backend::OMP,
Backend::OCCA_CPU, Backend::RAJA_CPU, Backend::CPU
};
@@ -40,12 +41,22 @@ static const Backend::Id backend_list[Backend::NUM_BACKENDS] =
// Backend names listed by priority, high to low:
static const char *backend_name[Backend::NUM_BACKENDS] =
{
"occa-cuda", "raja-cuda", "cuda", "occa-omp", "raja-omp", "omp",
"occa-cuda", "raja-cuda", "cuda", "hip", "occa-omp", "raja-omp", "omp",
"occa-cpu", "raja-cpu", "cpu"
};
} // namespace mfem::internal
// Initialize the unique global Device variable.
Device Device::device_singleton;
Device::~Device()
{
if (destroy_mm) { mm.Destroy(); }
}
void Device::Configure(const std::string &device, const int dev)
{
std::map<std::string, Backend::Id> bmap;
@@ -67,18 +78,22 @@ void Device::Configure(const std::string &device, const int dev)
}
// OCCA_CUDA needs CUDA or RAJA_CUDA:
Get().allowed_backends = Get().backends;
if (Allows(Backend::OCCA_CUDA) && !Allows(Backend::RAJA_CUDA))
{
Get().MarkBackend(Backend::CUDA);
}
// Activate all backends for Setup().
Get().allowed_backends = Get().backends;
// Perform setup.
Get().Setup(dev);
// Enable only the default host CPU backend.
Get().allowed_backends = Backend::CPU;
// Enable the device
Enable();
// Copy all data members from the global 'singleton_device' into '*this'.
std::memcpy(this, &Get(), sizeof(Device));
// Only '*this' will call the MemoryManager::Destroy() method.
destroy_mm = true;
}
void Device::Print(std::ostream &out)
@@ -87,7 +102,7 @@ void Device::Print(std::ostream &out)
bool add_comma = false;
for (int i = 0; i < Backend::NUM_BACKENDS; i++)
{
if (Get().backends & internal::backend_list[i])
if (backends & internal::backend_list[i])
{
if (add_comma) { out << ','; }
add_comma = true;
@@ -97,17 +112,35 @@ void Device::Print(std::ostream &out)
out << '\n';
}
void Device::UpdateMemoryTypeAndClass()
{
if (Device::Allows(Backend::DEVICE_MASK))
{
mem_type = MemoryType::CUDA;
mem_class = MemoryClass::CUDA;
}
else
{
mem_type = MemoryType::HOST;
mem_class = MemoryClass::HOST;
}
}
void Device::Enable()
{
if (Get().backends & ~Backend::CPU)
{
Get().mode = Device::ACCELERATED;
Get().UpdateMemoryTypeAndClass();
}
}
#ifdef MFEM_USE_CUDA
static void DeviceSetup(const int dev, int &ngpu)
{
cudaGetDeviceCount(&ngpu);
MFEM_VERIFY(ngpu>0, "No CUDA device found!");
cuInit(0);
cuDeviceGet(&internal::cuDevice, dev);
cuCtxCreate(&internal::cuContext, CU_CTX_SCHED_AUTO, internal::cuDevice);
internal::cuStream = new CUstream;
MFEM_VERIFY(internal::cuStream, "CUDA stream could not be created!");
cuStreamCreate(internal::cuStream, CU_STREAM_DEFAULT);
ngpu = CuGetDeviceCount();
MFEM_VERIFY(ngpu > 0, "No CUDA device found!");
MFEM_GPU_CHECK(cudaSetDevice(dev));
}
#endif
@@ -118,6 +151,18 @@ static void CudaDeviceSetup(const int dev, int &ngpu)
#endif
}
static void HipDeviceSetup(const int dev, int &ngpu)
{
#ifdef MFEM_USE_HIP
int deviceId;
MFEM_GPU_CHECK(hipGetDevice(&deviceId));
hipDeviceProp_t props;
MFEM_GPU_CHECK(hipGetDeviceProperties(&props, deviceId));
MFEM_VERIFY(dev==deviceId,"");
ngpu = 1;
#endif
}
static void RajaDeviceSetup(const int dev, int &ngpu)
{
#ifdef MFEM_USE_CUDA
@@ -125,7 +170,7 @@ static void RajaDeviceSetup(const int dev, int &ngpu)
#endif
}
static void OccaDeviceSetup(CUdevice cu_dev, CUcontext cu_ctx)
static void OccaDeviceSetup(const int dev)
{
#ifdef MFEM_USE_OCCA
const int cpu = Device::Allows(Backend::OCCA_CPU);
@@ -138,7 +183,8 @@ static void OccaDeviceSetup(CUdevice cu_dev, CUcontext cu_ctx)
if (cuda)
{
#if OCCA_CUDA_ENABLED
internal::occaDevice = occa::cuda::wrapDevice(cu_dev, cu_ctx);
std::string mode("mode: 'CUDA', device_id : ");
internal::occaDevice.setup(mode.append(1,'0'+dev));
#else
MFEM_ABORT("the OCCA CUDA backend requires OCCA built with CUDA!");
#endif
@@ -183,11 +229,14 @@ void Device::Setup(const int device)
ngpu = 0;
dev = device;
#ifndef MFEM_USE_CUDA
MFEM_VERIFY(!Allows(Backend::CUDA_MASK),
"the CUDA backends require MFEM built with MFEM_USE_CUDA=YES");
#endif
#ifndef MFEM_USE_HIP
MFEM_VERIFY(!Allows(Backend::HIP_MASK),
"the HIP backends require MFEM built with MFEM_USE_HIP=YES");
#endif
#ifndef MFEM_USE_RAJA
MFEM_VERIFY(!Allows(Backend::RAJA_MASK),
"the RAJA backends require MFEM built with MFEM_USE_RAJA=YES");
@@ -197,22 +246,11 @@ void Device::Setup(const int device)
"the OpenMP and RAJA OpenMP backends require MFEM built with"
" MFEM_USE_OPENMP=YES");
#endif
// The check for MFEM_USE_OCCA is in the function OccaDeviceSetup().
// We initialize CUDA and/or RAJA_CUDA first so OccaDeviceSetup() can reuse
// the same initialized cuDevice and cuContext objects when OCCA_CUDA is
// enabled.
if (Allows(Backend::CUDA)) { CudaDeviceSetup(dev, ngpu); }
if (Allows(Backend::HIP)) { HipDeviceSetup(dev, ngpu); }
if (Allows(Backend::RAJA_CUDA)) { RajaDeviceSetup(dev, ngpu); }
if (Allows(Backend::OCCA_MASK))
{
OccaDeviceSetup(internal::cuDevice, internal::cuContext);
}
}
Device::~Device()
{
delete internal::cuStream;
// The check for MFEM_USE_OCCA is in the function OccaDeviceSetup().
if (Allows(Backend::OCCA_MASK)) { OccaDeviceSetup(dev); }
}
} // mfem
+182 -62
View File
@@ -12,7 +12,9 @@
#ifndef MFEM_DEVICE_HPP
#define MFEM_DEVICE_HPP
#include "cuda.hpp"
#include "globals.hpp"
#include "mem_manager.hpp"
namespace mfem
{
@@ -34,23 +36,25 @@ struct Backend
OMP = 1 << 1,
/// [device] CUDA backend. Enabled when MFEM_USE_CUDA = YES.
CUDA = 1 << 2,
/// [device] HIP backend. Enabled when MFEM_USE_HIP = YES.
HIP = 1 << 3,
/** @brief [host] RAJA CPU backend: sequential execution on each MPI rank.
Enabled when MFEM_USE_RAJA = YES. */
RAJA_CPU = 1 << 3,
RAJA_CPU = 1 << 4,
/** @brief [host] RAJA OpenMP backend. Enabled when MFEM_USE_RAJA = YES
and MFEM_USE_OPENMP = YES. */
RAJA_OMP = 1 << 4,
RAJA_OMP = 1 << 5,
/** @brief [device] RAJA CUDA backend. Enabled when MFEM_USE_RAJA = YES
and MFEM_USE_CUDA = YES. */
RAJA_CUDA = 1 << 5,
RAJA_CUDA = 1 << 6,
/** @brief [host] OCCA CPU backend: sequential execution on each MPI rank.
Enabled when MFEM_USE_OCCA = YES. */
OCCA_CPU = 1 << 6,
OCCA_CPU = 1 << 7,
/// [host] OCCA OpenMP backend. Enabled when MFEM_USE_OCCA = YES.
OCCA_OMP = 1 << 7,
OCCA_OMP = 1 << 8,
/** @brief [device] OCCA CUDA backend. Enabled when MFEM_USE_OCCA = YES
and MFEM_USE_CUDA = YES. */
OCCA_CUDA = 1 << 8
OCCA_CUDA = 1 << 9
};
/** @brief Additional useful constants. For example, the *_MASK constants can
@@ -58,26 +62,34 @@ struct Backend
enum
{
/// Number of backends: from (1 << 0) to (1 << (NUM_BACKENDS-1)).
NUM_BACKENDS = 9,
NUM_BACKENDS = 10,
/// Biwise-OR of all CPU backends
CPU_MASK = CPU | RAJA_CPU | OCCA_CPU,
/// Biwise-OR of all CUDA backends
CUDA_MASK = CUDA | RAJA_CUDA | OCCA_CUDA,
/// Biwise-OR of all RAJA backends
RAJA_MASK = RAJA_CPU | RAJA_OMP | RAJA_CUDA,
/// Biwise-OR of all OCCA backends
OCCA_MASK = OCCA_CPU | OCCA_OMP | OCCA_CUDA,
/// Biwise-OR of all HIP backends
HIP_MASK = HIP,
/// Biwise-OR of all OpenMP backends
OMP_MASK = OMP | RAJA_OMP | OCCA_OMP,
/// Biwise-OR of all device backends
DEVICE_MASK = CUDA_MASK
DEVICE_MASK = CUDA_MASK | HIP_MASK,
/// Biwise-OR of all RAJA backends
RAJA_MASK = RAJA_CPU | RAJA_OMP | RAJA_CUDA,
/// Biwise-OR of all OCCA backends
OCCA_MASK = OCCA_CPU | OCCA_OMP | OCCA_CUDA
};
};
/** @brief The MFEM Device class abstracts hardware devices, such as GPUs, as
well as programming models, such as CUDA, OCCA, RAJA and OpenMP. */
/** @brief The MFEM Device class abstracts hardware devices such as GPUs, as
well as programming models such as CUDA, OCCA, RAJA and OpenMP. */
/** This class represents a "virtual device" with the following properties:
- There a single object of this class which is controlled by its static
methods.
- At most one object of this class can be constructed and that object is
controlled by its static methods.
- If no Device object is constructed, the static methods will use a default
global object which is never configured and always uses Backend::CPU.
- Once configured, the object cannot be re-configured during the program
lifetime.
- MFEM classes use this object to determine where (host or device) to
@@ -85,37 +97,81 @@ struct Backend
- Multiple backends can be configured at the same time; currently, a fixed
priority order is used to select a specific backend from the list of
configured backends. See the Backend class and the Configure() method in
this class for details.
- The device can be disabled to restrict the backend selection to only the
default host CPU backend, see the methods Enable() and Disable(). */
this class for details. */
class Device
{
private:
enum MODES {SEQUENTIAL, ACCELERATED};
static Device device_singleton;
MODES mode;
int dev = 0; ///< Device ID of the configured device.
int ngpu = -1; ///< Number of detected devices; -1: not initialized.
unsigned long backends; ///< Bitwise-OR of all configured backends.
/** Bitwise-OR mask of all allowed backends. All backends are active when the
Device is enabled. When the Device is disabled, only the host CPU backend
is allowed. */
unsigned long allowed_backends;
/// Set to true during configuration, except in 'device_singleton'.
bool destroy_mm;
bool mpi_gpu_aware;
MemoryType mem_type; ///< Current Device MemoryType
MemoryClass mem_class; ///< Current Device MemoryClass
Device()
: mode(Device::SEQUENTIAL),
backends(Backend::CPU),
allowed_backends(backends) { }
Device(Device const&);
void operator=(Device const&);
static Device& Get() { static Device singleton; return singleton; }
static Device& Get() { return device_singleton; }
/// Setup switcher based on configuration settings
void Setup(const int dev = 0);
void MarkBackend(Backend::Id b) { backends |= b; }
void UpdateMemoryTypeAndClass();
/// Enable the use of the configured device in the code that follows.
/** After this call MFEM classes will use the backend kernels whenever
possible, transferring data automatically to the device, if necessary.
If the only configured backend is the default host CPU one, the device
will remain disabled.
If the device is actually enabled, this method will also update the
current MemoryType and MemoryClass. */
static void Enable();
public:
/** @brief Default constructor. Unless Configure() is called later, the
default Backend::CPU will be used. */
/** @note At most one Device object can be constructed during the lifetime of
a program.
@note This object should be destroyed after all other MFEM objects that
use the Device are destroyed. */
Device()
: mode(Device::SEQUENTIAL),
backends(Backend::CPU),
destroy_mm(false),
mpi_gpu_aware(false),
mem_type(MemoryType::HOST),
mem_class(MemoryClass::HOST)
{ }
/** @brief Construct a Device and configure it based on the @a device string.
See Configure() for more details. */
/** @note At most one Device object can be constructed during the lifetime of
a program.
@note This object should be destroyed after all other MFEM objects that
use the Device are destroyed. */
Device(const std::string &device, const int dev = 0)
: mode(Device::SEQUENTIAL),
backends(Backend::CPU),
destroy_mm(false),
mpi_gpu_aware(false),
mem_type(MemoryType::HOST),
mem_class(MemoryClass::HOST)
{ Configure(device, dev); }
/// Destructor.
~Device();
/// Configure the Device backends.
/** The string parameter @a device must be a comma-separated list of backend
string names (see below). The @a dev argument specifies the ID of the
@@ -126,17 +182,16 @@ public:
string name of 'RAJA_CPU' is 'raja-cpu'.
* The 'cpu' backend is always enabled with lowest priority.
* The current backend priority from highest to lowest is: 'occa-cuda',
'raja-cuda', 'cuda', 'occa-omp', 'raja-omp', 'omp', 'occa-cpu',
'raja-cuda', 'cuda', 'hip', 'occa-omp', 'raja-omp', 'omp', 'occa-cpu',
'raja-cpu', 'cpu'.
* Multiple backends can be configured at the same time.
* Only one 'occa-*' backend can be configured at a time.
* The backend 'occa-cuda' enables the 'cuda' backend unless 'raja-cuda'
is already enabled.
* After this call, the Device will be disabled. */
static void Configure(const std::string &device, const int dev = 0);
is already enabled. */
void Configure(const std::string &device, const int dev = 0);
/// Print the configuration of the MFEM virtual device object.
static void Print(std::ostream &out = mfem::out);
void Print(std::ostream &out = mfem::out);
/// Return true if Configure() has been called previously.
static inline bool IsConfigured() { return Get().ngpu >= 0; }
@@ -144,47 +199,112 @@ public:
/// Return true if an actual device (e.g. GPU) has been configured.
static inline bool IsAvailable() { return Get().ngpu > 0; }
/// Enable the use of the configured device in the code that follows.
/** After this call MFEM classes will use the backend kernels whenever
possible, transferring data automatically to the device, if necessary.
If the only configured backend is the default host CPU one, the device
will remain disabled. */
static inline void Enable()
{
if (Get().backends & ~Backend::CPU)
{
Get().mode = Device::ACCELERATED;
Get().allowed_backends = Get().backends;
}
}
/// Disable the use of the configured device in the code that follows.
/** After this call MFEM classes will only use default CPU kernels,
transferring data automatically from the device, if necessary. */
static inline void Disable()
{
Get().mode = Device::SEQUENTIAL;
Get().allowed_backends = Backend::CPU;
}
/// Return true if the Device is enabled.
/// Return true if any backend other than Backend::CPU is enabled.
static inline bool IsEnabled() { return Get().mode == ACCELERATED; }
/// The opposite of IsEnabled().
static inline bool IsDisabled() { return !IsEnabled(); }
/** @brief Return true if any of the backends in the backend mask, @a b_mask,
are allowed. The allowed backends are all configured backends minus the
device backends when the Device is disabled. */
are allowed. */
/** This method can be used with any of the Backend::Id constants, the
Backend::*_MASK, or combinations of those. */
static inline bool Allows(unsigned long b_mask)
{ return Get().allowed_backends & b_mask; }
{ return Get().backends & b_mask; }
~Device();
/** @brief Get the current Device MemoryType. This is the MemoryType used by
most MFEM classes when allocating memory to be used with device kernels.
*/
static inline MemoryType GetMemoryType() { return Get().mem_type; }
/** @brief Get the current Device MemoryClass. This is the MemoryClass used
by most MFEM device kernels to access Memory objects. */
static inline MemoryClass GetMemoryClass() { return Get().mem_class; }
static void SetGPUAwareMPI(const bool force = true)
{ Get().mpi_gpu_aware = force; }
static bool GetGPUAwareMPI() { return Get().mpi_gpu_aware; }
static void Synchronize() { MFEM_DEVICE_SYNC; }
};
// Inline Memory access functions using the mfem::Device MemoryClass or
// MemoryClass::HOST.
/** @brief Get a pointer for read access to @a mem with the mfem::Device
MemoryClass, if @a on_dev = true, or MemoryClass::HOST, otherwise. */
/** Also, if @a on_dev = true, the device flag of @a mem will be set. */
template <typename T>
inline const T *Read(const Memory<T> &mem, int size, bool on_dev = true)
{
if (!on_dev)
{
return mem.Read(MemoryClass::HOST, size);
}
else
{
mem.UseDevice(true);
return mem.Read(Device::GetMemoryClass(), size);
}
}
/** @brief Shortcut to Read(const Memory<T> &mem, int size, false) */
template <typename T>
inline const T *HostRead(const Memory<T> &mem, int size)
{
return mfem::Read(mem, size, false);
}
/** @brief Get a pointer for write access to @a mem with the mfem::Device
MemoryClass, if @a on_dev = true, or MemoryClass::HOST, otherwise. */
/** Also, if @a on_dev = true, the device flag of @a mem will be set. */
template <typename T>
inline T *Write(Memory<T> &mem, int size, bool on_dev = true)
{
if (!on_dev)
{
return mem.Write(MemoryClass::HOST, size);
}
else
{
mem.UseDevice(true);
return mem.Write(Device::GetMemoryClass(), size);
}
}
/** @brief Shortcut to Write(const Memory<T> &mem, int size, false) */
template <typename T>
inline const T *HostWrite(const Memory<T> &mem, int size)
{
return mfem::Write(mem, size, false);
}
/** @brief Get a pointer for read+write access to @a mem with the mfem::Device
MemoryClass, if @a on_dev = true, or MemoryClass::HOST, otherwise. */
/** Also, if @a on_dev = true, the device flag of @a mem will be set. */
template <typename T>
inline T *ReadWrite(Memory<T> &mem, int size, bool on_dev = true)
{
if (!on_dev)
{
return mem.ReadWrite(MemoryClass::HOST, size);
}
else
{
mem.UseDevice(true);
return mem.ReadWrite(Device::GetMemoryClass(), size);
}
}
/** @brief Shortcut to ReadWrite(const Memory<T> &mem, int size, false) */
template <typename T>
inline const T *HostReadWrite(const Memory<T> &mem, int size)
{
return mfem::ReadWrite(mem, size, false);
}
} // mfem
#endif // MFEM_DEVICE_HPP
+288 -37
View File
@@ -15,6 +15,7 @@
#include "../config/config.hpp"
#include "error.hpp"
#include "cuda.hpp"
#include "hip.hpp"
#include "occa.hpp"
#include "device.hpp"
#include "mem_manager.hpp"
@@ -22,19 +23,48 @@
#ifdef MFEM_USE_RAJA
#include "RAJA/RAJA.hpp"
#if defined(RAJA_ENABLE_CUDA) && !defined(MFEM_USE_CUDA)
#error When RAJA is built with CUDA, MFEM_USE_CUDA=YES is required
#endif
#endif
namespace mfem
{
// Maximum size of dofs and quads in 1D.
const int MAX_D1D = 16;
const int MAX_Q1D = 16;
// Implementation of MFEM's "parallel for" (forall) device/host kernel
// interfaces supporting RAJA, CUDA, OpenMP, and sequential backends.
// The MFEM_FORALL wrapper
#define MFEM_FORALL(i,N,...) \
ForallWrap(N, \
[=] MFEM_ATTR_DEVICE (int i) {__VA_ARGS__}, \
[&] (int i) {__VA_ARGS__})
#define MFEM_FORALL(i,N,...) \
ForallWrap<1>(true,N, \
[=] MFEM_DEVICE (int i) {__VA_ARGS__}, \
[&] (int i) {__VA_ARGS__})
// MFEM_FORALL with a 2D CUDA block
#define MFEM_FORALL_2D(i,N,X,Y,BZ,...) \
ForallWrap<2>(true,N, \
[=] MFEM_DEVICE (int i) {__VA_ARGS__}, \
[&] (int i) {__VA_ARGS__}, \
X,Y,BZ)
// MFEM_FORALL with a 3D CUDA block
#define MFEM_FORALL_3D(i,N,X,Y,Z,...) \
ForallWrap<3>(true,N, \
[=] MFEM_DEVICE (int i) {__VA_ARGS__}, \
[&] (int i) {__VA_ARGS__}, \
X,Y,Z)
// MFEM_FORALL that uses the basic CPU backend when use_dev is false. See for
// example the functions in vector.cpp, where we don't want to use the mfem
// device for operations on small vectors.
#define MFEM_FORALL_SWITCH(use_dev,i,N,...) \
ForallWrap<1>(use_dev,N, \
[=] MFEM_DEVICE (int i) {__VA_ARGS__}, \
[&] (int i) {__VA_ARGS__})
/// OpenMP backend
@@ -54,28 +84,100 @@ void OmpWrap(const int N, HBODY &&h_body)
/// RAJA Cuda backend
template <int BLOCKS, typename DBODY>
void RajaCudaWrap(const int N, DBODY &&d_body)
{
#if defined(MFEM_USE_RAJA) && defined(RAJA_ENABLE_CUDA)
using RAJA::statement::Segs;
template <const int BLOCKS = MFEM_CUDA_BLOCKS, typename DBODY>
void RajaCudaWrap1D(const int N, DBODY &&d_body)
{
RAJA::forall<RAJA::cuda_exec<BLOCKS>>(RAJA::RangeSegment(0,N),d_body);
#else
MFEM_ABORT("RAJA::Cuda requested but RAJA::Cuda is not enabled!");
#endif
}
template <typename DBODY>
void RajaCudaWrap2D(const int N, DBODY &&d_body,
const int X, const int Y, const int BZ)
{
MFEM_VERIFY(N>0, "");
MFEM_VERIFY(BZ>0, "");
const int G = (N+BZ-1)/BZ;
RAJA::kernel<RAJA::KernelPolicy<
RAJA::statement::CudaKernel<
RAJA::statement::For<0, RAJA::cuda_block_x_loop,
RAJA::statement::For<1, RAJA::cuda_thread_x_direct,
RAJA::statement::For<2, RAJA::cuda_thread_y_direct,
RAJA::statement::For<3, RAJA::cuda_thread_z_direct,
RAJA::statement::Lambda<0, Segs<0>>>>>>>>>
(RAJA::make_tuple(RAJA::RangeSegment(0,G), RAJA::RangeSegment(0,X),
RAJA::RangeSegment(0,Y), RAJA::RangeSegment(0,BZ)),
[=] RAJA_DEVICE (const int n)
{
const int k = n*BZ + threadIdx.z;
if (k >= N) { return; }
d_body(k);
MFEM_SYNC_THREAD;
});
MFEM_GPU_CHECK(cudaGetLastError());
}
template <typename DBODY>
void RajaCudaWrap3D(const int N, DBODY &&d_body,
const int X, const int Y, const int Z)
{
MFEM_VERIFY(N>0, "");
RAJA::kernel<RAJA::KernelPolicy<
RAJA::statement::CudaKernel<
RAJA::statement::For<0, RAJA::cuda_block_x_loop,
RAJA::statement::For<1, RAJA::cuda_thread_x_direct,
RAJA::statement::For<2, RAJA::cuda_thread_y_direct,
RAJA::statement::For<3, RAJA::cuda_thread_z_direct,
RAJA::statement::Lambda<0, Segs<0>>>>>>>>>
(RAJA::make_tuple(RAJA::RangeSegment(0,N), RAJA::RangeSegment(0,X),
RAJA::RangeSegment(0,Y), RAJA::RangeSegment(0,Z)),
[=] RAJA_DEVICE (const int k) { d_body(k); MFEM_SYNC_THREAD; });
MFEM_GPU_CHECK(cudaGetLastError());
}
#endif
/// RAJA OpenMP backend
template <typename HBODY>
void RajaOmpWrap(const int N, HBODY &&h_body)
{
#if defined(MFEM_USE_RAJA) && defined(RAJA_ENABLE_OPENMP)
using RAJA::statement::Segs;
template <typename HBODY>
void RajaOmpWrap1D(const int N, HBODY &&h_body)
{
RAJA::forall<RAJA::omp_parallel_for_exec>(RAJA::RangeSegment(0,N), h_body);
#else
MFEM_ABORT("RAJA::OpenMP requested but RAJA::OpenMP is not enabled!");
#endif
}
template <typename HBODY>
void RajaOmpWrap2D(const int N, HBODY &&h_body,
const int X, const int Y, const int BZ)
{
RAJA::kernel<RAJA::KernelPolicy<
RAJA::statement::For<0, RAJA::omp_parallel_for_exec,
RAJA::statement::Lambda<0, Segs<0>>>>>
(RAJA::make_tuple(RAJA::RangeSegment(0,N), RAJA::RangeSegment(0,X),
RAJA::RangeSegment(0,Y), RAJA::RangeSegment(0,BZ)),
[=] (int k) { h_body(k); });
}
template <typename HBODY>
void RajaOmpWrap3D(const int N, HBODY &&h_body,
const int X, const int Y, const int Z)
{
RAJA::kernel<RAJA::KernelPolicy<
RAJA::statement::For<0, RAJA::omp_parallel_for_exec,
RAJA::statement::Lambda<0, Segs<0>>>>>
(RAJA::make_tuple(RAJA::RangeSegment(0,N), RAJA::RangeSegment(0,X),
RAJA::RangeSegment(0,Y), RAJA::RangeSegment(0,Z)),
[=] (int k) { h_body(k); });
}
#endif
/// RAJA sequential loop backend
template <typename HBODY>
@@ -93,47 +195,196 @@ void RajaSeqWrap(const int N, HBODY &&h_body)
#ifdef MFEM_USE_CUDA
template <typename BODY> __global__ static
void CuKernel(const int N, BODY body)
void CuKernel1D(const int N, BODY body)
{
const int k = blockDim.x*blockIdx.x + threadIdx.x;
if (k >= N) { return; }
body(k);
}
template <int BLOCKS, typename DBODY>
void CuWrap(const int N, DBODY &&d_body)
template <typename BODY> __global__ static
void CuKernel2D(const int N, BODY body, const int BZ)
{
if (N==0) { return; }
const int GRID = (N+BLOCKS-1)/BLOCKS;
CuKernel<<<GRID,BLOCKS>>>(N,d_body);
const cudaError_t last = cudaGetLastError();
MFEM_VERIFY(last == cudaSuccess, cudaGetErrorString(last));
const int k = blockIdx.x*BZ + threadIdx.z;
if (k >= N) { return; }
body(k);
}
#else // MFEM_USE_CUDA
template <typename BODY> __global__ static
void CuKernel3D(const int N, BODY body)
{
const int k = blockIdx.x;
if (k >= N) { return; }
body(k);
}
template <int BLOCKS, typename DBODY>
void CuWrap(const int N, DBODY &&d_body) {}
template <const int BLCK = MFEM_CUDA_BLOCKS, typename DBODY>
void CuWrap1D(const int N, DBODY &&d_body)
{
if (N==0) { return; }
const int GRID = (N+BLCK-1)/BLCK;
CuKernel1D<<<GRID,BLCK>>>(N, d_body);
MFEM_GPU_CHECK(cudaGetLastError());
}
#endif
template <typename DBODY>
void CuWrap2D(const int N, DBODY &&d_body,
const int X, const int Y, const int BZ)
{
if (N==0) { return; }
MFEM_VERIFY(BZ>0, "");
const int GRID = (N+BZ-1)/BZ;
const dim3 BLCK(X,Y,BZ);
CuKernel2D<<<GRID,BLCK>>>(N,d_body,BZ);
MFEM_GPU_CHECK(cudaGetLastError());
}
template <typename DBODY>
void CuWrap3D(const int N, DBODY &&d_body,
const int X, const int Y, const int Z)
{
if (N==0) { return; }
const int GRID = N;
const dim3 BLCK(X,Y,Z);
CuKernel3D<<<GRID,BLCK>>>(N,d_body);
MFEM_GPU_CHECK(cudaGetLastError());
}
#endif // MFEM_USE_CUDA
/// HIP backend
#ifdef MFEM_USE_HIP
template <typename BODY> __global__ static
void HipKernel1D(const int N, BODY body)
{
const int k = hipBlockDim_x*hipBlockIdx_x + hipThreadIdx_x;
if (k >= N) { return; }
body(k);
}
template <typename BODY> __global__ static
void HipKernel2D(const int N, BODY body, const int BZ)
{
const int k = hipBlockIdx_x*BZ + hipThreadIdx_z;
if (k >= N) { return; }
body(k);
}
template <typename BODY> __global__ static
void HipKernel3D(const int N, BODY body)
{
const int k = hipBlockIdx_x;
if (k >= N) { return; }
body(k);
}
template <const int BLCK = MFEM_HIP_BLOCKS, typename DBODY>
void HipWrap1D(const int N, DBODY &&d_body)
{
if (N==0) { return; }
const int GRID = (N+BLCK-1)/BLCK;
hipLaunchKernelGGL(HipKernel1D,GRID,BLCK,0,0,N,d_body);
MFEM_GPU_CHECK(hipGetLastError());
}
template <typename DBODY>
void HipWrap2D(const int N, DBODY &&d_body,
const int X, const int Y, const int BZ)
{
if (N==0) { return; }
const int GRID = (N+BZ-1)/BZ;
const dim3 BLCK(X,Y,BZ);
hipLaunchKernelGGL(HipKernel2D,GRID,BLCK,0,0,N,d_body,BZ);
MFEM_GPU_CHECK(hipGetLastError());
}
template <typename DBODY>
void HipWrap3D(const int N, DBODY &&d_body,
const int X, const int Y, const int Z)
{
if (N==0) { return; }
const int GRID = N;
const dim3 BLCK(X,Y,Z);
hipLaunchKernelGGL(HipKernel3D,GRID,BLCK,0,0,N,d_body);
MFEM_GPU_CHECK(hipGetLastError());
}
#endif // MFEM_USE_HIP
/// The forall kernel body wrapper
template <typename DBODY, typename HBODY>
void ForallWrap(const int N, DBODY &&d_body, HBODY &&h_body)
template <const int DIM, typename DBODY, typename HBODY>
inline void ForallWrap(const bool use_dev, const int N,
DBODY &&d_body, HBODY &&h_body,
const int X=0, const int Y=0, const int Z=0)
{
if (Device::Allows(Backend::RAJA_CUDA))
{ return RajaCudaWrap<MFEM_CUDA_BLOCKS>(N, d_body); }
if (!use_dev) { goto backend_cpu; }
if (Device::Allows(Backend::CUDA))
{ return CuWrap<MFEM_CUDA_BLOCKS>(N, d_body); }
#if defined(MFEM_USE_RAJA) && defined(RAJA_ENABLE_CUDA)
// Handle all allowed CUDA backends except Backend::CUDA
if (DIM == 1 && Device::Allows(Backend::CUDA_MASK & ~Backend::CUDA))
{ return RajaCudaWrap1D(N, d_body); }
if (Device::Allows(Backend::RAJA_OMP)) { return RajaOmpWrap(N, h_body); }
if (DIM == 2 && Device::Allows(Backend::CUDA_MASK & ~Backend::CUDA))
{ return RajaCudaWrap2D(N, d_body, X, Y, Z); }
if (Device::Allows(Backend::OMP)) { return OmpWrap(N, h_body); }
if (DIM == 3 && Device::Allows(Backend::CUDA_MASK & ~Backend::CUDA))
{ return RajaCudaWrap3D(N, d_body, X, Y, Z); }
#endif
if (Device::Allows(Backend::RAJA_CPU)) { return RajaSeqWrap(N, h_body); }
#ifdef MFEM_USE_CUDA
// Handle all allowed CUDA backends
if (DIM == 1 && Device::Allows(Backend::CUDA_MASK))
{ return CuWrap1D(N, d_body); }
if (DIM == 2 && Device::Allows(Backend::CUDA_MASK))
{ return CuWrap2D(N, d_body, X, Y, Z); }
if (DIM == 3 && Device::Allows(Backend::CUDA_MASK))
{ return CuWrap3D(N, d_body, X, Y, Z); }
#endif
#ifdef MFEM_USE_HIP
// Handle all allowed HIP backends
if (DIM == 1 && Device::Allows(Backend::HIP_MASK))
{ return HipWrap1D(N, d_body); }
if (DIM == 2 && Device::Allows(Backend::HIP_MASK))
{ return HipWrap2D(N, d_body, X, Y, Z); }
if (DIM == 3 && Device::Allows(Backend::HIP_MASK))
{ return HipWrap3D(N, d_body, X, Y, Z); }
#endif
#if defined(MFEM_USE_RAJA) && defined(RAJA_ENABLE_OPENMP)
// Handle all allowed OpenMP backends except Backend::OMP
if (DIM == 1 && Device::Allows(Backend::OMP_MASK & ~Backend::OMP))
{ return RajaOmpWrap1D(N, h_body); }
if (DIM == 2 && Device::Allows(Backend::OMP_MASK & ~Backend::OMP))
{ return RajaOmpWrap2D(N, h_body, X, Y, Z); }
if (DIM == 3 && Device::Allows(Backend::OMP_MASK & ~Backend::OMP))
{ return RajaOmpWrap3D(N, h_body, X, Y, Z); }
#endif
#ifdef MFEM_USE_OPENMP
// Handle all allowed OpenMP backends
if (Device::Allows(Backend::OMP_MASK)) { return OmpWrap(N, h_body); }
#endif
#ifdef MFEM_USE_RAJA
// Handle all allowed CPU backends except Backend::CPU
if (Device::Allows(Backend::CPU_MASK & ~Backend::CPU))
{ return RajaSeqWrap(N, h_body); }
#endif
backend_cpu:
// Handle Backend::CPU. This is also a fallback for any allowed backends not
// handled above, e.g. OCCA_CPU with configuration 'occa-cpu,cpu', or
// OCCA_OMP with configuration 'occa-omp,cpu'.
for (int k = 0; k < N; k++) { h_body(k); }
}
+6
View File
@@ -31,6 +31,12 @@ std::string MakeParFilename(const std::string &prefix, const int myid,
return fname.str();
}
#ifdef MFEM_COUNT_FLOPS
namespace internal
{
long long flop_count;
}
#endif
#ifdef MFEM_USE_MPI
+133
View File
@@ -0,0 +1,133 @@
// Copyright (c) 2010, Lawrence Livermore National Security, LLC. Produced at
// the Lawrence Livermore National Laboratory. LLNL-CODE-443211. All Rights
// reserved. See file COPYRIGHT for details.
//
// This file is part of the MFEM library. For more information and source code
// availability see http://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the GNU Lesser General Public License (as published by the Free
// Software Foundation) version 2.1 dated February 1999.
#include "hip.hpp"
#include "globals.hpp"
namespace mfem
{
// Internal debug option, useful for tracking HIP allocations, deallocations
// and transfers.
// #define MFEM_TRACK_HIP_MEM
#ifdef MFEM_USE_HIP
void mfem_hip_error(hipError_t err, const char *expr, const char *func,
const char *file, int line)
{
mfem::err << "\n\nHIP error: (" << expr << ") failed with error:\n --> "
<< hipGetErrorString(err)
<< "\n ... in function: " << func
<< "\n ... in file: " << file << ':' << line << '\n';
mfem_error();
}
#endif
void* HipMemAlloc(void** dptr, size_t bytes)
{
#ifdef MFEM_USE_HIP
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "HipMemAlloc(): allocating " << bytes << " bytes ... "
<< std::flush;
#endif
MFEM_GPU_CHECK(hipMalloc(dptr, bytes));
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "done: " << *dptr << std::endl;
#endif
#endif
return *dptr;
}
void* HipMemFree(void *dptr)
{
#ifdef MFEM_USE_HIP
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "HipMemFree(): deallocating memory @ " << dptr << " ... "
<< std::flush;
#endif
MFEM_GPU_CHECK(hipFree(dptr));
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dptr;
}
void* HipMemcpyHtoD(void* dst, const void* src, size_t bytes)
{
#ifdef MFEM_USE_HIP
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "HipMemcpyHtoD(): copying " << bytes << " bytes from "
<< src << " to " << dst << " ... " << std::flush;
#endif
MFEM_GPU_CHECK(hipMemcpy(dst, src, bytes, hipMemcpyHostToDevice));
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dst;
}
void* HipMemcpyHtoDAsync(void* dst, const void* src, size_t bytes)
{
#ifdef MFEM_USE_HIP
MFEM_GPU_CHECK(hipMemcpyAsync(dst, src, bytes, hipMemcpyHostToDevice));
#endif
return dst;
}
void* HipMemcpyDtoD(void *dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_HIP
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "HipMemcpyDtoD(): copying " << bytes << " bytes from "
<< src << " to " << dst << " ... " << std::flush;
#endif
MFEM_GPU_CHECK(hipMemcpy(dst, src, bytes, hipMemcpyDeviceToDevice));
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dst;
}
void* HipMemcpyDtoDAsync(void* dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_HIP
MFEM_GPU_CHECK(hipMemcpyAsync(dst, src, bytes, hipMemcpyDeviceToDevice));
#endif
return dst;
}
void* HipMemcpyDtoH(void *dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_HIP
#ifdef MFEM_TRACK_HPI_MEM
mfem::out << "HipMemcpyDtoH(): copying " << bytes << " bytes from "
<< src << " to " << dst << " ... " << std::flush;
#endif
MFEM_GPU_CHECK(hipMemcpy(dst, src, bytes, hipMemcpyDeviceToHost));
#ifdef MFEM_TRACK_HIP_MEM
mfem::out << "done." << std::endl;
#endif
#endif
return dst;
}
void* HipMemcpyDtoHAsync(void *dst, const void *src, size_t bytes)
{
#ifdef MFEM_USE_HIP
MFEM_GPU_CHECK(hipMemcpyAsync(dst, src, bytes, hipMemcpyDeviceToHost));
#endif
return dst;
}
} // namespace mfem
+85
View File
@@ -0,0 +1,85 @@
// Copyright (c) 2010, Lawrence Livermore National Security, LLC. Produced at
// the Lawrence Livermore National Laboratory. LLNL-CODE-443211. All Rights
// reserved. See file COPYRIGHT for details.
//
// This file is part of the MFEM library. For more information and source code
// availability see http://mfem.org.
//
// MFEM is free software; you can redistribute it and/or modify it under the
// terms of the GNU Lesser General Public License (as published by the Free
// Software Foundation) version 2.1 dated February 1999.
#ifndef MFEM_HIP_HPP
#define MFEM_HIP_HPP
#include "../config/config.hpp"
#include "error.hpp"
#ifdef MFEM_USE_HIP
#include <hip/hip_runtime.h>
#endif
// HIP block size used by MFEM.
#define MFEM_HIP_BLOCKS 256
#ifdef MFEM_USE_HIP
// Define a HIP error check macro, MFEM_GPU_CHECK(x), where x returns/is of
// type 'hipError_t'. This macro evaluates 'x' and raises an error if the
// result is not hipSuccess.
#define MFEM_GPU_CHECK(x) \
do \
{ \
hipError_t err = (x); \
if (err != hipSuccess) \
{ \
mfem_hip_error(err, #x, _MFEM_FUNC_NAME, __FILE__, __LINE__); \
} \
} \
while (0)
#endif // MFEM_USE_HIP
// Define the MFEM inner threading macros
#if defined(MFEM_USE_HIP) && defined(__ROCM_ARCH__)
#define MFEM_SHARED __shared__
#define MFEM_SYNC_THREAD __syncthreads()
#define MFEM_THREAD_ID(k) hipThreadIdx_ ##k
#define MFEM_THREAD_SIZE(k) hipBlockDim_ ##k
#define MFEM_FOREACH_THREAD(i,k,N) for(int i=hipThreadIdx_ ##k; i<N; i+=hipBlockDim_ ##k)
#endif
namespace mfem
{
#ifdef MFEM_USE_HIP
// Function used by the macro MFEM_GPU_CHECK.
void mfem_hip_error(hipError_t err, const char *expr, const char *func,
const char *file, int line);
#endif
/// Allocates device memory
void* HipMemAlloc(void **d_ptr, size_t bytes);
/// Frees device memory
void* HipMemFree(void *d_ptr);
/// Copies memory from Host to Device
void* HipMemcpyHtoD(void *d_dst, const void *h_src, size_t bytes);
/// Copies memory from Host to Device
void* HipMemcpyHtoDAsync(void *d_dst, const void *h_src, size_t bytes);
/// Copies memory from Device to Device
void* HipMemcpyDtoD(void *d_dst, const void *d_src, size_t bytes);
/// Copies memory from Device to Device
void* HipMemcpyDtoDAsync(void *d_dst, const void *d_src, size_t bytes);
/// Copies memory from Device to Host
void* HipMemcpyDtoH(void *h_dst, const void *d_src, size_t bytes);
/// Copies memory from Device to Host
void* HipMemcpyDtoHAsync(void *h_dst, const void *d_src, size_t bytes);
} // namespace mfem
#endif // MFEM_HIP_HPP
+563 -209
View File
@@ -15,10 +15,48 @@
#include <list>
#include <unordered_map>
#include <algorithm> // std::max
namespace mfem
{
#ifdef MFEM_USE_HIP
#define MFEM_GPU(...) Hip ## __VA_ARGS__
#else
#define MFEM_GPU(...) Cu ## __VA_ARGS__
#endif
MemoryType GetMemoryType(MemoryClass mc)
{
switch (mc)
{
case MemoryClass::HOST: return MemoryType::HOST;
case MemoryClass::HOST_32: return MemoryType::HOST_32;
case MemoryClass::HOST_64: return MemoryType::HOST_64;
case MemoryClass::CUDA: return MemoryType::CUDA;
case MemoryClass::CUDA_UVM: return MemoryType::CUDA_UVM;
}
return MemoryType::HOST;
}
MemoryClass operator*(MemoryClass mc1, MemoryClass mc2)
{
// | HOST HOST_32 HOST_64 CUDA CUDA_UVM
// ---------+--------------------------------------------------
// HOST | HOST HOST_32 HOST_64 CUDA CUDA_UVM
// HOST_32 | HOST_32 HOST_32 HOST_64 CUDA CUDA_UVM
// HOST_64 | HOST_64 HOST_64 HOST_64 CUDA CUDA_UVM
// CUDA | CUDA CUDA CUDA CUDA CUDA_UVM
// CUDA_UVM | CUDA_UVM CUDA_UVM CUDA_UVM CUDA_UVM CUDA_UVM
// Using the enumeration ordering:
// HOST < HOST_32 < HOST_64 < CUDA < CUDA_UVM,
// the above table is simply: a*b = max(a,b).
return std::max(mc1, mc2);
}
namespace internal
{
@@ -28,17 +66,15 @@ struct Alias;
/// Memory class that holds:
/// - a boolean telling which memory space is being used
/// - the size in bytes of this memory region,
/// - the host and the device pointer,
/// - a list of all aliases seen using this region (used only to free them).
/// - the host and the device pointer.
struct Memory
{
bool host;
const std::size_t bytes;
void *const h_ptr;
void *d_ptr;
std::list<const void*> aliases;
Memory(void* const h, const std::size_t size):
host(true), bytes(size), h_ptr(h), d_ptr(nullptr), aliases() {}
host(true), bytes(size), h_ptr(h), d_ptr(nullptr) {}
};
/// Alias class that holds the base memory region and the offset
@@ -46,10 +82,13 @@ struct Alias
{
Memory *const mem;
const long offset;
unsigned long counter;
};
typedef std::unordered_map<const void*, Memory> MemoryMap;
typedef std::unordered_map<const void*, const Alias*> AliasMap;
// TODO: use 'Alias' or 'const Alias' as the mapped type in the AliasMap instead
// of 'Alias*'
typedef std::unordered_map<const void*, Alias*> AliasMap;
struct Ledger
{
@@ -64,270 +103,213 @@ static internal::Ledger *maps;
MemoryManager::MemoryManager()
{
exists = true;
enabled = true;
maps = new internal::Ledger();
}
MemoryManager::~MemoryManager()
{
if (exists) { Destroy(); }
}
void MemoryManager::Destroy()
{
MFEM_VERIFY(exists, "MemoryManager has been destroyed already!");
for (auto& n : maps->memories)
{
internal::Memory &mem = n.second;
if (mem.d_ptr) { MFEM_GPU(MemFree)(mem.d_ptr); }
}
for (auto& n : maps->aliases)
{
delete n.second;
}
delete maps;
exists = false;
}
void* MemoryManager::Insert(void *ptr, const std::size_t bytes)
{
if (!UsingMM()) { return ptr; }
const bool known = IsKnown(ptr);
if (known)
if (ptr == NULL)
{
MFEM_VERIFY(bytes == 0, "Trying to add NULL with size " << bytes);
return NULL;
}
auto res = maps->memories.emplace(ptr, internal::Memory(ptr, bytes));
if (res.second == false)
{
mfem_error("Trying to add an already present address!");
}
maps->memories.emplace(ptr, internal::Memory(ptr, bytes));
return ptr;
}
void *MemoryManager::Erase(void *ptr)
void MemoryManager::InsertDevice(void *ptr, void *h_ptr, size_t bytes)
{
MFEM_VERIFY(ptr != NULL, "cannot register NULL device pointer");
MFEM_VERIFY(h_ptr != NULL, "internal error");
auto res = maps->memories.emplace(h_ptr, internal::Memory(h_ptr, bytes));
if (res.second == false)
{
mfem_error("Trying to add an already present address!");
}
res.first->second.d_ptr = ptr;
}
void *MemoryManager::Erase(void *ptr, bool free_dev_ptr)
{
if (!UsingMM()) { return ptr; }
if (!ptr) { return ptr; }
const bool known = IsKnown(ptr);
if (!known)
auto mem_map_iter = maps->memories.find(ptr);
if (mem_map_iter == maps->memories.end())
{
mfem_error("Trying to erase an unknown pointer!");
}
internal::Memory &mem = maps->memories.at(ptr);
if (mem.d_ptr) { CuMemFree(mem.d_ptr); }
for (const void *alias : mem.aliases)
{
maps->aliases.erase(maps->aliases.find(alias));
}
mem.aliases.clear();
maps->memories.erase(maps->memories.find(ptr));
internal::Memory &mem = mem_map_iter->second;
if (mem.d_ptr && free_dev_ptr) { MFEM_GPU(MemFree)(mem.d_ptr); }
maps->memories.erase(mem_map_iter);
return ptr;
}
void MemoryManager::SetHostDevicePtr(void *h_ptr, void *d_ptr, const bool host)
{
internal::Memory &base = maps->memories.at(h_ptr);
base.d_ptr = d_ptr;
base.host = host;
}
bool MemoryManager::IsKnown(const void *ptr)
{
return maps->memories.find(ptr) != maps->memories.end();
}
bool MemoryManager::IsOnHost(const void *ptr)
{
return maps->memories.at(ptr).host;
}
std::size_t MemoryManager::Bytes(const void *ptr)
{
return maps->memories.at(ptr).bytes;
}
void *MemoryManager::GetDevicePtr(const void *ptr)
void *MemoryManager::GetDevicePtr(const void *ptr, size_t bytes, bool copy_data)
{
if (!ptr)
{
MFEM_VERIFY(bytes == 0, "Trying to access NULL with size " << bytes);
return NULL;
}
internal::Memory &base = maps->memories.at(ptr);
const size_t bytes = base.bytes;
if (!base.d_ptr)
{
CuMemAlloc(&base.d_ptr, bytes);
CuMemcpyHtoD(base.d_ptr, ptr, bytes);
MFEM_GPU(MemAlloc)(&base.d_ptr, base.bytes);
}
if (copy_data)
{
MFEM_ASSERT(bytes <= base.bytes, "invalid copy size");
MFEM_GPU(MemcpyHtoD)(base.d_ptr, ptr, bytes);
base.host = false;
}
return base.d_ptr;
}
// Looks if ptr is an alias of one memory
static const void* AliasBaseMemory(const internal::Ledger *maps,
const void *ptr)
void MemoryManager::InsertAlias(const void *base_ptr, void *alias_ptr,
bool base_is_alias)
{
for (internal::MemoryMap::const_iterator mem = maps->memories.begin();
mem != maps->memories.end(); mem++)
long offset = static_cast<const char*>(alias_ptr) -
static_cast<const char*>(base_ptr);
if (!base_ptr)
{
const void *b_ptr = mem->first;
if (b_ptr > ptr) { continue; }
const void *end = static_cast<const char*>(b_ptr) + mem->second.bytes;
if (ptr < end) { return b_ptr; }
MFEM_VERIFY(offset == 0,
"Trying to add alias to NULL at offset " << offset);
return;
}
if (base_is_alias)
{
const internal::Alias *alias = maps->aliases.at(base_ptr);
base_ptr = alias->mem->h_ptr;
offset += alias->offset;
}
internal::Memory &mem = maps->memories.at(base_ptr);
auto res = maps->aliases.emplace(alias_ptr, nullptr);
if (res.second == false) // alias_ptr was already in the map
{
if (res.first->second->mem != &mem || res.first->second->offset != offset)
{
mfem_error("alias already exists with different base/offset!");
}
else
{
res.first->second->counter++;
}
}
else
{
res.first->second = new internal::Alias{&mem, offset, 1};
}
return nullptr;
}
bool MemoryManager::IsAlias(const void *ptr)
void MemoryManager::EraseAlias(void *alias_ptr)
{
const internal::AliasMap::const_iterator found = maps->aliases.find(ptr);
if (found != maps->aliases.end()) { return true; }
MFEM_ASSERT(!IsKnown(ptr), "Ptr is an already known address!");
const void *base = AliasBaseMemory(maps, ptr);
if (!base) { return false; }
internal::Memory &mem = maps->memories.at(base);
const long offset = static_cast<const char*>(ptr) -
static_cast<const char*> (base);
const internal::Alias *alias = new internal::Alias{&mem, offset};
maps->aliases.emplace(ptr, alias);
mem.aliases.push_back(ptr);
return true;
if (!alias_ptr) { return; }
auto alias_map_iter = maps->aliases.find(alias_ptr);
if (alias_map_iter == maps->aliases.end())
{
mfem_error("alias not found");
}
internal::Alias *alias = alias_map_iter->second;
if (--alias->counter) { return; }
// erase the alias from the alias map:
maps->aliases.erase(alias_map_iter);
delete alias;
}
static inline bool MmDeviceIniFilter(void)
void *MemoryManager::GetAliasDevicePtr(const void *alias_ptr, size_t bytes,
bool copy_data)
{
if (!mm.UsingMM()) { return true; }
if (!mm.IsEnabled()) { return true; }
if (!Device::IsAvailable()) { return true; }
if (!Device::IsConfigured()) { return true; }
return false;
if (!alias_ptr)
{
MFEM_VERIFY(bytes == 0, "Trying to access NULL with size " << bytes);
return NULL;
}
auto &alias_map = maps->aliases;
auto alias_map_iter = alias_map.find(alias_ptr);
if (alias_map_iter == alias_map.end())
{
mfem_error("alias not found");
}
const internal::Alias *alias = alias_map_iter->second;
internal::Memory &base = *alias->mem;
MFEM_ASSERT((char*)base.h_ptr + alias->offset == alias_ptr,
"internal error");
if (!base.d_ptr)
{
MFEM_GPU(MemAlloc)(&base.d_ptr, base.bytes);
}
if (copy_data)
{
MFEM_GPU(MemcpyHtoD)((char*)base.d_ptr + alias->offset, alias_ptr, bytes);
base.host = false;
}
return (char*)base.d_ptr + alias->offset;
}
// Turn a known address into the right host or device address. Alloc, Push, or
// Pull it if necessary.
static void *PtrKnown(internal::Ledger *maps, void *ptr)
static void PullKnown(internal::Ledger *maps,
const void *ptr, const std::size_t bytes, bool copy_data)
{
internal::Memory &base = maps->memories.at(ptr);
const bool ptr_on_host = base.host;
const std::size_t bytes = base.bytes;
const bool run_on_device = Device::Allows(Backend::DEVICE_MASK);
if (ptr_on_host && !run_on_device) { return ptr; }
if (bytes==0) { mfem_error("PtrKnown bytes==0"); }
if (!base.d_ptr) { CuMemAlloc(&base.d_ptr, bytes); }
if (!base.d_ptr) { mfem_error("PtrKnown !base->d_ptr"); }
if (!ptr_on_host && run_on_device) { return base.d_ptr; }
if (!ptr) { mfem_error("PtrKnown !ptr"); }
if (!ptr_on_host && !run_on_device) // Pull
MFEM_ASSERT(base.h_ptr == ptr, "internal error");
// There are cases where it is OK if base.d_ptr is not allocated yet:
// for example, when requesting read-write access on host to memory created
// as device memory.
if (copy_data && base.d_ptr)
{
CuMemcpyDtoH(ptr, base.d_ptr, bytes);
MFEM_GPU(MemcpyDtoH)(base.h_ptr, base.d_ptr, bytes);
base.host = true;
return ptr;
}
// Push
if (!(ptr_on_host && run_on_device)) { mfem_error("PtrKnown !(host && gpu)"); }
CuMemcpyHtoD(base.d_ptr, ptr, bytes);
base.host = false;
return base.d_ptr;
}
// Turn an alias into the right host or device address. Alloc, Push, or Pull it
// if necessary.
static void *PtrAlias(internal::Ledger *maps, void *ptr)
{
const bool gpu = Device::Allows(Backend::DEVICE_MASK);
const internal::Alias *alias = maps->aliases.at(ptr);
const internal::Memory *base = alias->mem;
const bool host = base->host;
const bool device = !base->host;
const std::size_t bytes = base->bytes;
if (host && !gpu) { return ptr; }
if (bytes==0) { mfem_error("PtrAlias bytes==0"); }
if (!base->d_ptr) { CuMemAlloc(&(alias->mem->d_ptr), bytes); }
if (!base->d_ptr) { mfem_error("PtrAlias !base->d_ptr"); }
void *a_ptr = static_cast<char*>(base->d_ptr) + alias->offset;
if (device && gpu) { return a_ptr; }
if (!base->h_ptr) { mfem_error("PtrAlias !base->h_ptr"); }
if (device && !gpu) // Pull
{
CuMemcpyDtoH(base->h_ptr, base->d_ptr, bytes);
alias->mem->host = true;
return ptr;
}
// Push
if (!(host && gpu)) { mfem_error("PtrAlias !(host && gpu)"); }
CuMemcpyHtoD(base->d_ptr, base->h_ptr, bytes);
alias->mem->host = false;
return a_ptr;
}
void *MemoryManager::Ptr(void *ptr)
{
if (ptr==NULL) { return NULL; };
if (MmDeviceIniFilter()) { return ptr; }
if (IsKnown(ptr)) { return PtrKnown(maps, ptr); }
if (IsAlias(ptr)) { return PtrAlias(maps, ptr); }
if (Device::Allows(Backend::DEVICE_MASK))
{
mfem_error("Trying to use unknown pointer on the DEVICE!");
}
return ptr;
}
const void *MemoryManager::Ptr(const void *ptr)
{
return static_cast<const void*>(Ptr(const_cast<void*>(ptr)));
}
static void PushKnown(internal::Ledger *maps,
const void *ptr, const std::size_t bytes)
{
internal::Memory &base = maps->memories.at(ptr);
if (!base.d_ptr) { CuMemAlloc(&base.d_ptr, base.bytes); }
CuMemcpyHtoD(base.d_ptr, ptr, bytes == 0 ? base.bytes : bytes);
}
static void PushAlias(const internal::Ledger *maps,
const void *ptr, const std::size_t bytes)
{
const internal::Alias *alias = maps->aliases.at(ptr);
void *dst = static_cast<char*>(alias->mem->d_ptr) + alias->offset;
CuMemcpyHtoD(dst, ptr, bytes);
}
void MemoryManager::Push(const void *ptr, const std::size_t bytes)
{
if (MmDeviceIniFilter()) { return; }
if (IsKnown(ptr)) { return PushKnown(maps, ptr, bytes); }
if (IsAlias(ptr)) { return PushAlias(maps, ptr, bytes); }
if (Device::Allows(Backend::DEVICE_MASK))
{ mfem_error("Unknown pointer to push to!"); }
}
static void PullKnown(const internal::Ledger *maps,
const void *ptr, const std::size_t bytes)
{
const internal::Memory &base = maps->memories.at(ptr);
const bool host = base.host;
if (host) { return; }
CuMemcpyDtoH(base.h_ptr, base.d_ptr, bytes == 0 ? base.bytes : bytes);
}
static void PullAlias(const internal::Ledger *maps,
const void *ptr, const std::size_t bytes)
const void *ptr, const std::size_t bytes, bool copy_data)
{
const internal::Alias *alias = maps->aliases.at(ptr);
const bool host = alias->mem->host;
if (host) { return; }
if (!ptr) { mfem_error("PullAlias !ptr"); }
if (!alias->mem->d_ptr) { mfem_error("PullAlias !alias->mem->d_ptr"); }
CuMemcpyDtoH(const_cast<void*>(ptr),
static_cast<char*>(alias->mem->d_ptr) + alias->offset,
bytes);
}
void MemoryManager::Pull(const void *ptr, const std::size_t bytes)
{
if (MmDeviceIniFilter()) { return; }
if (IsKnown(ptr)) { return PullKnown(maps, ptr, bytes); }
if (IsAlias(ptr)) { return PullAlias(maps, ptr, bytes); }
if (Device::Allows(Backend::DEVICE_MASK))
{ mfem_error("Unknown pointer to pull from!"); }
}
namespace internal { extern CUstream *cuStream; }
void* MemoryManager::Memcpy(void *dst, const void *src,
const std::size_t bytes, const bool async)
{
void *d_dst = Ptr(dst);
void *d_src = const_cast<void*>(Ptr(src));
if (bytes == 0) { return dst; }
const bool run_on_host = !Device::Allows(Backend::DEVICE_MASK);
if (run_on_host) { return std::memcpy(dst, src, bytes); }
if (!async) { return CuMemcpyDtoD(d_dst, d_src, bytes); }
return CuMemcpyDtoDAsync(d_dst, d_src, bytes, internal::cuStream);
MFEM_ASSERT((char*)alias->mem->h_ptr + alias->offset == ptr,
"internal error");
// There are cases where it is OK if alias->mem->d_ptr is not allocated yet:
// for example, when requesting read-write access on host to memory created
// as device memory.
if (copy_data && alias->mem->d_ptr)
{
MFEM_GPU(MemcpyDtoH)(const_cast<void*>(ptr),
static_cast<char*>(alias->mem->d_ptr) + alias->offset,
bytes);
}
}
void MemoryManager::RegisterCheck(void *ptr)
{
if (ptr != NULL && UsingMM())
if (ptr != NULL)
{
if (!IsKnown(ptr))
{
@@ -347,17 +329,389 @@ void MemoryManager::PrintPtrs(void)
<< "h_ptr " << mem.h_ptr << ", "
<< "d_ptr " << mem.d_ptr;
}
mfem::out << std::endl;
}
void MemoryManager::GetAll(void)
// Static private MemoryManager methods used by class Memory
void *MemoryManager::New_(void *h_ptr, std::size_t size, MemoryType mt,
unsigned &flags)
{
for (const auto& n : maps->memories)
// TODO: save the types of the pointers ...
flags = Mem::REGISTERED | Mem::OWNS_INTERNAL;
switch (mt)
{
const void *ptr = n.first;
Ptr(ptr);
case MemoryType::HOST: return nullptr; // case is handled outside
case MemoryType::HOST_32:
case MemoryType::HOST_64:
mfem_error("New_(): aligned host types are not implemented yet");
return nullptr;
case MemoryType::CUDA:
mm.Insert(h_ptr, size);
flags = flags | Mem::OWNS_HOST | Mem::OWNS_DEVICE | Mem::VALID_DEVICE;
return h_ptr;
case MemoryType::CUDA_UVM:
mfem_error("New_(): CUDA UVM allocation is not implemented yet");
return nullptr;
}
return nullptr;
}
void *MemoryManager::Register_(void *ptr, void *h_ptr, std::size_t capacity,
MemoryType mt, bool own, bool alias,
unsigned &flags)
{
// TODO: save the type of the registered pointer ...
MFEM_VERIFY(alias == false, "cannot register an alias!");
flags = flags | (Mem::REGISTERED | Mem::OWNS_INTERNAL);
if (IsHostMemory(mt))
{
mm.Insert(ptr, capacity);
flags = (own ? flags | Mem::OWNS_HOST : flags & ~Mem::OWNS_HOST) |
Mem::OWNS_DEVICE | Mem::VALID_HOST;
return ptr;
}
MFEM_VERIFY(mt == MemoryType::CUDA, "Only CUDA pointers are supported");
mm.InsertDevice(ptr, h_ptr, capacity);
flags = (own ? flags | Mem::OWNS_DEVICE : flags & ~Mem::OWNS_DEVICE) |
Mem::OWNS_HOST | Mem::VALID_DEVICE;
return h_ptr;
}
void MemoryManager::Alias_(void *base_h_ptr, std::size_t offset,
std::size_t size, unsigned base_flags,
unsigned &flags)
{
// TODO: store the 'size' in the MemoryManager?
mm.InsertAlias(base_h_ptr, (char*)base_h_ptr + offset,
base_flags & Mem::ALIAS);
flags = (base_flags | Mem::ALIAS | Mem::OWNS_INTERNAL) &
~(Mem::OWNS_HOST | Mem::OWNS_DEVICE);
}
MemoryType MemoryManager::Delete_(void *h_ptr, unsigned flags)
{
// TODO: this logic needs to be updated when support for HOST_32 and HOST_64
// memory types is added.
MFEM_ASSERT(!(flags & Mem::OWNS_DEVICE) || (flags & Mem::OWNS_INTERNAL),
"invalid Memory state");
if (mm.exists && (flags & Mem::OWNS_INTERNAL))
{
if (flags & Mem::ALIAS)
{
mm.EraseAlias(h_ptr);
}
else
{
mm.Erase(h_ptr, flags & Mem::OWNS_DEVICE);
}
}
return MemoryType::HOST;
}
void *MemoryManager::ReadWrite_(void *h_ptr, MemoryClass mc,
std::size_t size, unsigned &flags)
{
switch (mc)
{
case MemoryClass::HOST:
if (!(flags & Mem::VALID_HOST))
{
if (flags & Mem::ALIAS) { PullAlias(maps, h_ptr, size, true); }
else { PullKnown(maps, h_ptr, size, true); }
}
flags = (flags | Mem::VALID_HOST) & ~Mem::VALID_DEVICE;
return h_ptr;
case MemoryClass::HOST_32:
// TODO: check that the host pointer is MemoryType::HOST_32 or
// MemoryType::HOST_64
return h_ptr;
case MemoryClass::HOST_64:
// TODO: check that the host pointer is MemoryType::HOST_64
return h_ptr;
case MemoryClass::CUDA:
{
// TODO: check that the device pointer is MemoryType::CUDA or
// MemoryType::CUDA_UVM
const bool need_copy = !(flags & Mem::VALID_DEVICE);
flags = (flags | Mem::VALID_DEVICE) & ~Mem::VALID_HOST;
// TODO: add support for UVM
if (flags & Mem::ALIAS)
{
return mm.GetAliasDevicePtr(h_ptr, size, need_copy);
}
return mm.GetDevicePtr(h_ptr, size, need_copy);
}
case MemoryClass::CUDA_UVM:
// TODO: check that the host+device pointers are MemoryType::CUDA_UVM
// Do we need to update the validity flags?
return h_ptr; // the host and device pointers are the same
}
return nullptr;
}
const void *MemoryManager::Read_(void *h_ptr, MemoryClass mc,
std::size_t size, unsigned &flags)
{
switch (mc)
{
case MemoryClass::HOST:
if (!(flags & Mem::VALID_HOST))
{
if (flags & Mem::ALIAS) { PullAlias(maps, h_ptr, size, true); }
else { PullKnown(maps, h_ptr, size, true); }
}
flags = flags | Mem::VALID_HOST;
return h_ptr;
case MemoryClass::HOST_32:
// TODO: check that the host pointer is MemoryType::HOST_32 or
// MemoryType::HOST_64
return h_ptr;
case MemoryClass::HOST_64:
// TODO: check that the host pointer is MemoryType::HOST_64
return h_ptr;
case MemoryClass::CUDA:
{
// TODO: check that the device pointer is MemoryType::CUDA or
// MemoryType::CUDA_UVM
const bool need_copy = !(flags & Mem::VALID_DEVICE);
flags = flags | Mem::VALID_DEVICE;
// TODO: add support for UVM
if (flags & Mem::ALIAS)
{
return mm.GetAliasDevicePtr(h_ptr, size, need_copy);
}
return mm.GetDevicePtr(h_ptr, size, need_copy);
}
case MemoryClass::CUDA_UVM:
// TODO: check that the host+device pointers are MemoryType::CUDA_UVM
// Do we need to update the validity flags?
return h_ptr; // the host and device pointers are the same
}
return nullptr;
}
void *MemoryManager::Write_(void *h_ptr, MemoryClass mc, std::size_t size,
unsigned &flags)
{
switch (mc)
{
case MemoryClass::HOST:
flags = (flags | Mem::VALID_HOST) & ~Mem::VALID_DEVICE;
return h_ptr;
case MemoryClass::HOST_32:
// TODO: check that the host pointer is MemoryType::HOST_32 or
// MemoryType::HOST_64
flags = (flags | Mem::VALID_HOST) & ~Mem::VALID_DEVICE;
return h_ptr;
case MemoryClass::HOST_64:
// TODO: check that the host pointer is MemoryType::HOST_64
flags = (flags | Mem::VALID_HOST) & ~Mem::VALID_DEVICE;
return h_ptr;
case MemoryClass::CUDA:
// TODO: check that the device pointer is MemoryType::CUDA or
// MemoryType::CUDA_UVM
flags = (flags | Mem::VALID_DEVICE) & ~Mem::VALID_HOST;
// TODO: add support for UVM
if (flags & Mem::ALIAS)
{
return mm.GetAliasDevicePtr(h_ptr, size, false);
}
return mm.GetDevicePtr(h_ptr, size, false);
case MemoryClass::CUDA_UVM:
// TODO: check that the host+device pointers are MemoryType::CUDA_UVM
// Do we need to update the validity flags?
return h_ptr; // the host and device pointers are the same
}
return nullptr;
}
void MemoryManager::SyncAlias_(const void *base_h_ptr, void *alias_h_ptr,
size_t alias_size, unsigned base_flags,
unsigned &alias_flags)
{
// This is called only when (base_flags & Mem::REGISTERED) is true.
// Note that (alias_flags & REGISTERED) may not be true.
MFEM_ASSERT(alias_flags & Mem::ALIAS, "not an alias");
if ((base_flags & Mem::VALID_HOST) && !(alias_flags & Mem::VALID_HOST))
{
PullAlias(maps, alias_h_ptr, alias_size, true);
}
if ((base_flags & Mem::VALID_DEVICE) && !(alias_flags & Mem::VALID_DEVICE))
{
if (!(alias_flags & Mem::REGISTERED))
{
mm.InsertAlias(base_h_ptr, alias_h_ptr, base_flags & Mem::ALIAS);
alias_flags = (alias_flags | Mem::REGISTERED | Mem::OWNS_INTERNAL) &
~(Mem::OWNS_HOST | Mem::OWNS_DEVICE);
}
mm.GetAliasDevicePtr(alias_h_ptr, alias_size, true);
}
alias_flags = (alias_flags & ~(Mem::VALID_HOST | Mem::VALID_DEVICE)) |
(base_flags & (Mem::VALID_HOST | Mem::VALID_DEVICE));
}
MemoryType MemoryManager::GetMemoryType_(void *h_ptr, unsigned flags)
{
// TODO: support other memory types
if (flags & Mem::VALID_DEVICE) { return MemoryType::CUDA; }
return MemoryType::HOST;
}
void MemoryManager::Copy_(void *dest_h_ptr, const void *src_h_ptr,
std::size_t size, unsigned src_flags,
unsigned &dest_flags)
{
// Type of copy to use based on the src and dest validity flags:
// | src
// | h | d | hd
// -----------+-----+-----+------
// h | h2h d2h h2h
// dest d | h2d d2d d2d
// hd | h2h d2d d2d
const bool src_on_host =
(src_flags & Mem::VALID_HOST) &&
(!(src_flags & Mem::VALID_DEVICE) ||
((dest_flags & Mem::VALID_HOST) && !(dest_flags & Mem::VALID_DEVICE)));
const bool dest_on_host =
(dest_flags & Mem::VALID_HOST) &&
(!(dest_flags & Mem::VALID_DEVICE) ||
((src_flags & Mem::VALID_HOST) && !(src_flags & Mem::VALID_DEVICE)));
const void *src_d_ptr = src_on_host ? NULL :
((src_flags & Mem::ALIAS) ?
mm.GetAliasDevicePtr(src_h_ptr, size, false) :
mm.GetDevicePtr(src_h_ptr, size, false));
if (dest_on_host)
{
if (src_on_host)
{
if (dest_h_ptr != src_h_ptr && size != 0)
{
MFEM_ASSERT((char*)dest_h_ptr + size <= src_h_ptr ||
(char*)src_h_ptr + size <= dest_h_ptr,
"data overlaps!");
std::memcpy(dest_h_ptr, src_h_ptr, size);
}
}
else
{
MFEM_GPU(MemcpyDtoH)(dest_h_ptr, src_d_ptr, size);
}
}
else
{
void *dest_d_ptr = (dest_flags & Mem::ALIAS) ?
mm.GetAliasDevicePtr(dest_h_ptr, size, false) :
mm.GetDevicePtr(dest_h_ptr, size, false);
if (src_on_host)
{
MFEM_GPU(MemcpyHtoD)(dest_d_ptr, src_h_ptr, size);
}
else
{
MFEM_GPU(MemcpyDtoD)(dest_d_ptr, src_d_ptr, size);
}
}
dest_flags = dest_flags &
~(dest_on_host ? Mem::VALID_DEVICE : Mem::VALID_HOST);
}
void MemoryManager::CopyToHost_(void *dest_h_ptr, const void *src_h_ptr,
std::size_t size, unsigned src_flags)
{
const bool src_on_host = src_flags & Mem::VALID_HOST;
if (src_on_host)
{
if (dest_h_ptr != src_h_ptr && size != 0)
{
MFEM_ASSERT((char*)dest_h_ptr + size <= src_h_ptr ||
(char*)src_h_ptr + size <= dest_h_ptr,
"data overlaps!");
std::memcpy(dest_h_ptr, src_h_ptr, size);
}
}
else
{
const void *src_d_ptr = (src_flags & Mem::ALIAS) ?
mm.GetAliasDevicePtr(src_h_ptr, size, false) :
mm.GetDevicePtr(src_h_ptr, size, false);
MFEM_GPU(MemcpyDtoH)(dest_h_ptr, src_d_ptr, size);
}
}
void MemoryManager::CopyFromHost_(void *dest_h_ptr, const void *src_h_ptr,
std::size_t size, unsigned &dest_flags)
{
const bool dest_on_host = dest_flags & Mem::VALID_HOST;
if (dest_on_host)
{
if (dest_h_ptr != src_h_ptr && size != 0)
{
MFEM_ASSERT((char*)dest_h_ptr + size <= src_h_ptr ||
(char*)src_h_ptr + size <= dest_h_ptr,
"data overlaps!");
std::memcpy(dest_h_ptr, src_h_ptr, size);
}
}
else
{
void *dest_d_ptr = (dest_flags & Mem::ALIAS) ?
mm.GetAliasDevicePtr(dest_h_ptr, size, false) :
mm.GetDevicePtr(dest_h_ptr, size, false);
MFEM_GPU(MemcpyHtoD)(dest_d_ptr, src_h_ptr, size);
}
dest_flags = dest_flags &
~(dest_on_host ? Mem::VALID_DEVICE : Mem::VALID_HOST);
}
void MemoryPrintFlags(unsigned flags)
{
typedef Memory<int> Mem;
mfem::out
<< " registered = " << bool(flags & Mem::REGISTERED)
<< "\n owns host = " << bool(flags & Mem::OWNS_HOST)
<< "\n owns device = " << bool(flags & Mem::OWNS_DEVICE)
<< "\n owns internal = " << bool(flags & Mem::OWNS_INTERNAL)
<< "\n valid host = " << bool(flags & Mem::VALID_HOST)
<< "\n valid device = " << bool(flags & Mem::VALID_DEVICE)
<< "\n alias = " << bool(flags & Mem::ALIAS)
<< "\n device flag = " << bool(flags & Mem::USE_DEVICE)
<< std::endl;
}
MemoryManager mm;
bool MemoryManager::exists = false;
+688 -135
View File
@@ -13,6 +13,9 @@
#define MFEM_MEM_MANAGER_HPP
#include "globals.hpp"
#include "error.hpp"
#include <cstring> // std::memcpy
#include <type_traits> // std::is_const
namespace mfem
{
@@ -20,167 +23,717 @@ namespace mfem
// Implementation of MFEM's lightweight device/host memory manager designed to
// work seamlessly with the OCCA, RAJA, and other kernels supported by MFEM.
/// Memory types supported by MFEM.
enum class MemoryType
{
HOST, ///< Host memory; using new[] and delete[]
HOST_32, ///< Host memory aligned at 32 bytes (not supported yet)
HOST_64, ///< Host memory aligned at 64 bytes (not supported yet)
CUDA, ///< cudaMalloc, cudaFree
CUDA_UVM ///< cudaMallocManaged, cudaFree (not supported yet)
};
/// Memory classes identify subsets of memory types.
/** This type is used by kernels that can work with multiple MemoryType%s. For
example, kernels that can use CUDA or CUDA_UVM memory types should use
MemoryClass::CUDA for their inputs. */
enum class MemoryClass
{
HOST, ///< Memory types: { HOST, HOST_32, HOST_64, CUDA_UVM }
HOST_32, ///< Memory types: { HOST_32, HOST_64 }
HOST_64, ///< Memory types: { HOST_64 }
CUDA, ///< Memory types: { CUDA, CUDA_UVM }
CUDA_UVM ///< Memory types: { CUDA_UVM }
};
/// Return true if the given memory type is in MemoryClass::HOST.
inline bool IsHostMemory(MemoryType mt) { return mt <= MemoryType::HOST_64; }
/// Return a suitable MemoryType for a given MemoryClass.
MemoryType GetMemoryType(MemoryClass mc);
/// Return a suitable MemoryClass from a pair of MemoryClass%es.
/** Note: this operation is commutative, i.e. a*b = b*a, associative, i.e.
(a*b)*c = a*(b*c), and has an identity element: MemoryClass::HOST.
Currently, the operation is defined as a*b := max(a,b) where the max
operation is based on the enumeration ordering:
HOST < HOST_32 < HOST_64 < CUDA < CUDA_UVM. */
MemoryClass operator*(MemoryClass mc1, MemoryClass mc2);
/// Class used by MFEM to store pointers to host and/or device memory.
/** The template class parameter, T, must be a plain-old-data (POD) type.
In many respects this class behaves like a pointer:
* When destroyed, a Memory object does NOT automatically delete any
allocated memory.
* Only the method Delete() will deallocate a Memory object.
* Other methods that modify the object (e.g. New(), Wrap(), etc) will simply
overwrite the old contents.
* One difference with a pointer is that a const Memory object does not allow
modification of the content (unlike e.g. a const pointer).
A Memory object stores up to two different pointers: one host pointer (with
MemoryType from MemoryClass::HOST) and one device pointer (currently one of
MemoryType::CUDA or MemoryTyep::CUDA_UVM).
A Memory object can hold (wrap) an externally allocated pointer with any
given MemoryType.
Access to the content of the Memory object can be requested with any given
MemoryClass through the methods ReadWrite(), Read(), and Write().
Requesting such access may result in additional (internally handled)
memory allocation and/or memory copy.
* When ReadWrite() is called, the returned pointer becomes the only
valid pointer.
* When Read() is called, the returned pointer becomes valid, however
the other pointer (host or device) may remain valid as well.
* When Write() is called, the returned pointer becomes the only valid
pointer, however, unlike ReadWrite(), no memory copy will be performed.
The host memory (pointer from MemoryClass::HOST) can be accessed through the
inline methods: `operator[]()`, `operator*()`, the implicit conversion
functions `operator T*()`, `operator const T*()`, and the explicit
conversion template functions `operator U*()`, `operator const U*()` (with
any suitable type U). In certain cases, using these methods may have
undefined behavior, e.g. if the host pointer is not currently valid. */
template <typename T>
class Memory
{
protected:
friend class MemoryManager;
friend void MemoryPrintFlags(unsigned flags);
enum FlagMask
{
REGISTERED = 1, ///< #h_ptr is registered with the MemoryManager
OWNS_HOST = 2, ///< The host pointer will be deleted by Delete()
OWNS_DEVICE = 4, ///< The device pointer will be deleted by Delete()
OWNS_INTERNAL = 8, ///< Ownership flag for internal Memory data
VALID_HOST = 16, ///< Host pointer is valid
VALID_DEVICE = 32, ///< Device pointer is valid
ALIAS = 64,
/// Internal device flag, see e.g. Vector::UseDevice()
USE_DEVICE = 128
};
/// Pointer to host memory. Not owned.
/** When the pointer is not registered with the MemoryManager, this pointer
has type MemoryType::HOST. When the pointer is registered, it can be any
type from MemoryClass::HOST. */
T *h_ptr;
int capacity;
mutable unsigned flags;
// 'flags' is mutable so that it can be modified in Set{Host,Device}PtrOwner,
// Copy{From,To}, {ReadWrite,Read,Write}.
public:
/// Default constructor: no initialization.
Memory() { }
/// Copy constructor: default.
Memory(const Memory &orig) = default;
/// Move constructor: default.
Memory(Memory &&orig) = default;
/// Copy-assignment operator: default.
Memory &operator=(const Memory &orig) = default;
/// Move-assignment operator: default.
Memory &operator=(Memory &&orig) = default;
/// Allocate host memory for @a size entries.
explicit Memory(int size) { New(size); }
/** @brief Allocate memory for @a size entries with the given MemoryType
@a mt. */
/** The newly allocated memory is not initialized, however the given
MemoryType is still set as valid. */
Memory(int size, MemoryType mt) { New(size, mt); }
/** @brief Wrap an externally allocated host pointer, @a ptr with type
MemoryType::HOST. */
/** The parameter @a own determines whether @a ptr will be deleted (using
operator delete[]) when the method Delete() is called. */
explicit Memory(T *ptr, int size, bool own) { Wrap(ptr, size, own); }
/// Wrap an externally allocated pointer, @a ptr, of the given MemoryType.
/** The new memory object will have the given MemoryType set as valid.
The given @a ptr must be allocated appropriately for the given
MemoryType.
The parameter @a own determines whether @a ptr will be deleted when the
method Delete() is called. */
Memory(T *ptr, int size, MemoryType mt, bool own)
{ Wrap(ptr, size, mt, own); }
/** @brief Alias constructor. Create a Memory object that points inside the
Memory object @a base. */
/** The new Memory object uses the same MemoryType(s) as @a base. */
Memory(const Memory &base, int offset, int size)
{ MakeAlias(base, offset, size); }
/// Destructor: default.
/** @note The destructor will NOT delete the current memory. */
~Memory() = default;
/** @brief Return true if the host pointer is owned. Ownership indicates
whether the pointer will be deleted by the method Delete(). */
bool OwnsHostPtr() const { return flags & OWNS_HOST; }
/** @brief Set/clear the ownership flag for the host pointer. Ownership
indicates whether the pointer will be deleted by the method Delete(). */
void SetHostPtrOwner(bool own) const
{ flags = own ? (flags | OWNS_HOST) : (flags & ~OWNS_HOST); }
/** @brief Return true if the device pointer is owned. Ownership indicates
whether the pointer will be deleted by the method Delete(). */
bool OwnsDevicePtr() const { return flags & OWNS_DEVICE; }
/** @brief Set/clear the ownership flag for the device pointer. Ownership
indicates whether the pointer will be deleted by the method Delete(). */
void SetDevicePtrOwner(bool own) const
{ flags = own ? (flags | OWNS_DEVICE) : (flags & ~OWNS_DEVICE); }
/** @brief Clear the ownership flags for the host and device pointers, as
well as any internal data allocated by the Memory object. */
void ClearOwnerFlags() const
{ flags = flags & ~(OWNS_HOST | OWNS_DEVICE | OWNS_INTERNAL); }
/// Read the internal device flag.
bool UseDevice() const { return flags & USE_DEVICE; }
/// Set the internal device flag.
void UseDevice(bool use_dev) const
{ flags = use_dev ? (flags | USE_DEVICE) : (flags & ~USE_DEVICE); }
/// Return the size of the allocated memory.
int Capacity() const { return capacity; }
/// Reset the memory to be empty, ensuring that Delete() will be a no-op.
/** This is the Memory class equivalent to setting a pointer to NULL, see
Empty().
@note The current memory is NOT deleted by this method. */
void Reset() { h_ptr = NULL; capacity = 0; flags = 0; }
/// Return true if the Memory object is empty, see Reset().
/** Default-constructed objects are uninitialized, so they are not guaranteed
to be empty. */
bool Empty() const { return h_ptr == NULL; }
/// Allocate host memory for @a size entries with type MemoryType::HOST.
/** @note The current memory is NOT deleted by this method. */
void New(int size)
{ h_ptr = new T[size]; capacity = size; flags = OWNS_HOST | VALID_HOST; }
/// Allocate memory for @a size entries with the given MemoryType.
/** The newly allocated memory is not initialized, however the given
MemoryType is still set as valid.
@note The current memory is NOT deleted by this method. */
inline void New(int size, MemoryType mt);
/** @brief Wrap an externally allocated host pointer, @a ptr with type
MemoryType::HOST. */
/** The parameter @a own determines whether @a ptr will be deleted (using
operator delete[]) when the method Delete() is called.
@note The current memory is NOT deleted by this method. */
inline void Wrap(T *ptr, int size, bool own)
{ h_ptr = ptr; capacity = size; flags = (own ? OWNS_HOST : 0) | VALID_HOST; }
/// Wrap an externally allocated pointer, @a ptr, of the given MemoryType.
/** The new memory object will have the given MemoryType set as valid.
The given @a ptr must be allocated appropriately for the given
MemoryType.
The parameter @a own determines whether @a ptr will be deleted when the
method Delete() is called.
@note The current memory is NOT deleted by this method. */
inline void Wrap(T *ptr, int size, MemoryType mt, bool own);
/// Create a memory object that points inside the memory object @a base.
/** The new Memory object uses the same MemoryType(s) as @a base.
@note The current memory is NOT deleted by this method. */
inline void MakeAlias(const Memory &base, int offset, int size);
/// Delete the owned pointers. The Memory is not reset by this method.
inline void Delete();
/// Array subscript operator for host memory.
inline T &operator[](int idx);
/// Array subscript operator for host memory, const version.
inline const T &operator[](int idx) const;
/// Direct access to the host memory as T* (implicit conversion).
/** When the type T is const-qualified, this method can be used only if the
host pointer is currently valid (the device pointer may be valid or
invalid).
When the type T is not const-qualified, this method can be used only if
the host pointer is the only valid pointer.
When the Memory is empty, this method can be used and it returns NULL. */
inline operator T*();
/// Direct access to the host memory as const T* (implicit conversion).
/** This method can be used only if the host pointer is currently valid (the
device pointer may be valid or invalid).
When the Memory is empty, this method can be used and it returns NULL. */
inline operator const T*() const;
/// Direct access to the host memory via explicit typecast.
/** A pointer to type T must be reinterpret_cast-able to a pointer to type U.
In particular, this method cannot be used to cast away const-ness from
the base type T.
When the type U is const-qualified, this method can be used only if the
host pointer is currently valid (the device pointer may be valid or
invalid).
When the type U is not const-qualified, this method can be used only if
the host pointer is the only valid pointer.
When the Memory is empty, this method can be used and it returns NULL. */
template <typename U>
inline explicit operator U*();
/// Direct access to the host memory via explicit typecast, const version.
/** A pointer to type T must be reinterpret_cast-able to a pointer to type
const U.
This method can be used only if the host pointer is currently valid (the
device pointer may be valid or invalid).
When the Memory is empty, this method can be used and it returns NULL. */
template <typename U>
inline explicit operator const U*() const;
/// Get read-write access to the memory with the given MemoryClass.
/** If only read or only write access is needed, then the methods
Read() or Write() should be used instead of this method.
The parameter @a size must not exceed the Capacity(). */
inline T *ReadWrite(MemoryClass mc, int size);
/// Get read-only access to the memory with the given MemoryClass.
/** The parameter @a size must not exceed the Capacity(). */
inline const T *Read(MemoryClass mc, int size) const;
/// Get write-only access to the memory with the given MemoryClass.
/** The parameter @a size must not exceed the Capacity().
The contents of the returned pointer is undefined, unless it was
validated by a previous call to Read() or ReadWrite() with
the same MemoryClass. */
inline T *Write(MemoryClass mc, int size);
/// Copy the host/device pointer validity flags from @a other to @a *this.
/** This method synchronizes the pointer validity flags of two Memory objects
that use the same host/device pointers, or when @a *this is an alias
(sub-Memory) of @a other. Typically, this method should be called after
@a other is manipulated in a way that changes its pointer validity flags
(e.g. it was moved from device to host memory). */
inline void Sync(const Memory &other) const;
/** @brief Update the alias Memory @a *this to match the memory location (all
valid locations) of its base Memory, @a base. */
/** This method is useful when alias Memory is moved and manipulated in a
different memory space. Such operations render the pointer validity flags
of the base incorrect. Calling this method will ensure that @a base is
up-to-date. Note that this is achieved by moving/copying @a *this (if
necessary), and not @a base. */
inline void SyncAlias(const Memory &base, int alias_size) const;
/** @brief Return a MemoryType that is currently valid. If both the host and
the device pointers are currently valid, then the device memory type is
returned. */
inline MemoryType GetMemoryType() const;
/// Copy @a size entries from @a src to @a *this.
/** The given @a size should not exceed the Capacity() of the source @a src
and the destination, @a *this. */
inline void CopyFrom(const Memory &src, int size);
/// Copy @a size entries from the host pointer @a src to @a *this.
/** The given @a size should not exceed the Capacity() of @a *this. */
inline void CopyFromHost(const T *src, int size);
/// Copy @a size entries from @a *this to @a dest.
/** The given @a size should not exceed the Capacity() of @a *this and the
destination, @a dest. */
inline void CopyTo(Memory &dest, int size) const
{ dest.CopyFrom(*this, size); }
/// Copy @a size entries from @a *this to the host pointer @a dest.
/** The given @a size should not exceed the Capacity() of @a *this. */
inline void CopyToHost(T *dest, int size) const;
};
/// The memory manager class
class MemoryManager
{
private:
/// Allow to enable/disable the Ptr, Pull and Push functionalities
/// New and Delete will still continue to register the pointers
bool enabled;
template <typename T> friend class Memory;
// Used by the private static methods called by class Memory:
typedef Memory<int> Mem;
/// Allow to detect if a global memory manager instance exists
static bool exists;
// Methods used by class Memory
// Allocate and register a new pointer. Return the host pointer.
// h_ptr must be already allocated using new T[] if mt is a pure device
// memory type, e.g. CUDA (mt will not be HOST).
static void *New_(void *h_ptr, std::size_t size, MemoryType mt,
unsigned &flags);
// Register an external pointer of the given MemoryType. Return the host
// pointer.
static void *Register_(void *ptr, void *h_ptr, std::size_t capacity,
MemoryType mt, bool own, bool alias, unsigned &flags);
// Register an alias. Return the host pointer. Note: base_h_ptr may be an
// alias.
static void Alias_(void *base_h_ptr, std::size_t offset, std::size_t size,
unsigned base_flags, unsigned &flags);
// Un-register and free memory identified by its host pointer. Returns the
// memory type of the host pointer.
static MemoryType Delete_(void *h_ptr, unsigned flags);
// Return a pointer to the memory identified by the host pointer h_ptr for
// access with the given MemoryClass.
static void *ReadWrite_(void *h_ptr, MemoryClass mc, std::size_t size,
unsigned &flags);
static const void *Read_(void *h_ptr, MemoryClass mc, std::size_t size,
unsigned &flags);
static void *Write_(void *h_ptr, MemoryClass mc, std::size_t size,
unsigned &flags);
static void SyncAlias_(const void *base_h_ptr, void *alias_h_ptr,
size_t alias_size, unsigned base_flags,
unsigned &alias_flags);
// Return the type the of the currently valid memory. If more than one types
// are valid, return a device type.
static MemoryType GetMemoryType_(void *h_ptr, unsigned flags);
// Copy entries from valid memory type to valid memory type. Both dest_h_ptr
// and src_h_ptr are registered host pointers.
static void Copy_(void *dest_h_ptr, const void *src_h_ptr, std::size_t size,
unsigned src_flags, unsigned &dest_flags);
// Copy entries from valid memory type to host memory, where dest_h_ptr is
// not a registered host pointer and src_h_ptr is a registered host pointer.
static void CopyToHost_(void *dest_h_ptr, const void *src_h_ptr,
std::size_t size, unsigned src_flags);
// Copy entries from host memory to valid memory type, where dest_h_ptr is a
// registered host pointer and src_h_ptr is not a registered host pointer.
static void CopyFromHost_(void *dest_h_ptr, const void *src_h_ptr,
std::size_t size, unsigned &dest_flags);
/// Adds an address in the map
void *Insert(void *ptr, const std::size_t bytes);
void InsertDevice(void *ptr, void *h_ptr, size_t bytes);
/// Remove the address from the map, as well as all its aliases
void *Erase(void *ptr, bool free_dev_ptr = true);
/// Return the corresponding device pointer of ptr, allocating and moving the
/// data if needed (used in OccaPtr)
void *GetDevicePtr(const void *ptr, size_t bytes, bool copy_data);
void InsertAlias(const void *base_ptr, void *alias_ptr, bool base_is_alias);
void EraseAlias(void *alias_ptr);
void *GetAliasDevicePtr(const void *alias_ptr, size_t bytes, bool copy_data);
/// Return true if the pointer has been registered
bool IsKnown(const void *ptr);
public:
MemoryManager();
~MemoryManager();
/// Adds an address in the map
void *Insert(void *ptr, const std::size_t bytes);
/// Remove the address from the map, as well as all its aliases
void *Erase(void *ptr);
/// Return true if the memory manager is used: pointers seen by mfem::New and
/// mfem::Delete will be inserted in the ledger and erased from it
static inline bool UsingMM()
{
#ifdef MFEM_USE_MM
return true;
#else
return false;
#endif
}
/// Disable the memory manager: Ptr, Push and Pull will be no-op
void Disable() { enabled = false; }
/// Enable the memory manager: Ptr, Push and Pull wont be no-op
void Enable() { enabled = true; }
/// Return true if the memory manager is used and enabled
bool IsEnabled() { return UsingMM() && enabled; }
/// The opposite of IsEnabled().
bool IsDisabled() { return !IsEnabled(); }
void Destroy();
/// Return true if a global memory manager instance exists
static bool Exists() { return exists; }
/** @brief Translates ptr to host or device address, depending on what
backends are currently allowed by the Device class and on the ptr
state. */
void *Ptr(void *ptr);
const void *Ptr(const void *ptr);
/// Data will be pushed/pulled before the copy happens on the H or the D
void* Memcpy(void *dst, const void *src,
std::size_t bytes, const bool async = false);
/// Return the bytes of the memory region which base address is ptr
std::size_t Bytes(const void *ptr);
/// Return true if the registered pointer is on the host side
bool IsOnHost(const void *ptr);
/// Return true if the pointer has been registered
bool IsKnown(const void *ptr);
/// Return true if the pointer is an alias inside a registered memory region
bool IsAlias(const void *ptr);
/// Push the data to the device
void Push(const void *ptr, const std::size_t bytes =0);
/// Pull the data from the device
void Pull(const void *ptr, const std::size_t bytes =0);
/// Return the corresponding device pointer of ptr, allocating and moving the
/// data if needed (used in OccaPtr)
void *GetDevicePtr(const void *ptr);
/// Registers external host pointer in the memory manager which will manage
/// the corresponding device pointer, but not the provided host pointer.
template<class T>
void RegisterHostPtr(T *ptr_host, const std::size_t size)
{
Insert(ptr_host, size*sizeof(T));
#ifdef MFEM_DEBUG
RegisterCheck(ptr_host);
#endif
}
/// Registers external host and device pointers in the memory manager.
template<class T>
void RegisterHostAndDevicePtr(T *ptr_host, T *ptr_device,
const std::size_t size, const bool host)
{
RegisterHostPtr(ptr_host, size);
SetHostDevicePtr(ptr_host, ptr_device, host);
}
/// Set the host h_ptr, device d_ptr and mode host of the memory region just
/// been registered with h_ptr (see RegisterHostAndDevicePtr)
void SetHostDevicePtr(void *h_ptr, void *d_ptr, const bool host);
/// Unregisters the host pointer from the memory manager. To be used with
/// memory not allocated by the memory manager.
template<class T>
void UnregisterHostPtr(T *ptr) { Erase(ptr); }
/// Check if pointer has been registered in the memory manager
void RegisterCheck(void *ptr);
/// Prints all pointers known by the memory manager
void PrintPtrs(void);
/// Copies all memory to the current memory space
void GetAll(void);
};
// Inline methods
template <typename T>
inline void Memory<T>::New(int size, MemoryType mt)
{
if (mt == MemoryType::HOST)
{
New(size);
}
else
{
// Allocate the host pointer with new T[] if 'mt' is a pure device memory
// type, e.g. CUDA.
T *tmp = (mt == MemoryType::CUDA) ? new T[size] : NULL;
h_ptr = (T*)MemoryManager::New_(tmp, size*sizeof(T), mt, flags);
capacity = size;
}
}
template <typename T>
inline void Memory<T>::Wrap(T *ptr, int size, MemoryType mt, bool own)
{
if (mt == MemoryType::HOST)
{
Wrap(ptr, size, own);
}
else
{
// Allocate the host pointer with new T[] if 'mt' is a pure device memory
// type, e.g. CUDA.
T *tmp = (mt == MemoryType::CUDA) ? new T[size] : NULL;
h_ptr = (T*)MemoryManager::Register_(ptr, tmp, size*sizeof(T), mt, own,
false, flags);
capacity = size;
}
}
template <typename T>
inline void Memory<T>::MakeAlias(const Memory &base, int offset, int size)
{
h_ptr = base.h_ptr + offset;
capacity = size;
if (!(base.flags & REGISTERED))
{
flags = (base.flags | ALIAS) & ~(OWNS_HOST | OWNS_DEVICE);
}
else
{
MemoryManager::Alias_(base.h_ptr, offset*sizeof(T), size*sizeof(T),
base.flags, flags);
}
}
template <typename T>
inline void Memory<T>::Delete()
{
if (!(flags & REGISTERED) ||
MemoryManager::Delete_((void*)h_ptr, flags) == MemoryType::HOST)
{
if (flags & OWNS_HOST) { delete [] h_ptr; }
}
}
template <typename T>
inline T &Memory<T>::operator[](int idx)
{
MFEM_ASSERT((flags & VALID_HOST) && !(flags & VALID_DEVICE),
"invalid host pointer access");
return h_ptr[idx];
}
template <typename T>
inline const T &Memory<T>::operator[](int idx) const
{
MFEM_ASSERT((flags & VALID_HOST), "invalid host pointer access");
return h_ptr[idx];
}
template <typename T>
inline Memory<T>::operator T*()
{
MFEM_ASSERT(Empty() ||
((flags & VALID_HOST) &&
(std::is_const<T>::value || !(flags & VALID_DEVICE))),
"invalid host pointer access");
return h_ptr;
}
template <typename T>
inline Memory<T>::operator const T*() const
{
MFEM_ASSERT(Empty() || (flags & VALID_HOST), "invalid host pointer access");
return h_ptr;
}
template <typename T> template <typename U>
inline Memory<T>::operator U*()
{
MFEM_ASSERT(Empty() ||
((flags & VALID_HOST) &&
(std::is_const<U>::value || !(flags & VALID_DEVICE))),
"invalid host pointer access");
return reinterpret_cast<U*>(h_ptr);
}
template <typename T> template <typename U>
inline Memory<T>::operator const U*() const
{
MFEM_ASSERT(Empty() || (flags & VALID_HOST), "invalid host pointer access");
return reinterpret_cast<U*>(h_ptr);
}
template <typename T>
inline T *Memory<T>::ReadWrite(MemoryClass mc, int size)
{
if (!(flags & REGISTERED))
{
if (mc == MemoryClass::HOST) { return h_ptr; }
MemoryManager::Register_(h_ptr, NULL, capacity*sizeof(T),
MemoryType::HOST, flags & OWNS_HOST,
flags & ALIAS, flags);
}
return (T*)MemoryManager::ReadWrite_(h_ptr, mc, size*sizeof(T), flags);
}
template <typename T>
inline const T *Memory<T>::Read(MemoryClass mc, int size) const
{
if (!(flags & REGISTERED))
{
if (mc == MemoryClass::HOST) { return h_ptr; }
MemoryManager::Register_((void*)h_ptr, NULL, capacity*sizeof(T),
MemoryType::HOST, flags & OWNS_HOST,
flags & ALIAS, flags);
}
return (const T *)MemoryManager::Read_(
(void*)h_ptr, mc, size*sizeof(T), flags);
}
template <typename T>
inline T *Memory<T>::Write(MemoryClass mc, int size)
{
if (!(flags & REGISTERED))
{
if (mc == MemoryClass::HOST) { return h_ptr; }
MemoryManager::Register_(h_ptr, NULL, capacity*sizeof(T),
MemoryType::HOST, flags & OWNS_HOST,
flags & ALIAS, flags);
}
return (T*)MemoryManager::Write_(h_ptr, mc, size*sizeof(T), flags);
}
template <typename T>
inline void Memory<T>::Sync(const Memory &other) const
{
if (!(flags & REGISTERED) && (other.flags & REGISTERED))
{
MFEM_ASSERT(h_ptr == other.h_ptr &&
(flags & ALIAS) == (other.flags & ALIAS),
"invalid input");
flags = (flags | REGISTERED) & ~(OWNS_DEVICE | OWNS_INTERNAL);
}
flags = (flags & ~(VALID_HOST | VALID_DEVICE)) |
(other.flags & (VALID_HOST | VALID_DEVICE));
}
template <typename T>
inline void Memory<T>::SyncAlias(const Memory &base, int alias_size) const
{
// Assuming that if *this is registered then base is also registered.
MFEM_ASSERT(!(flags & REGISTERED) || (base.flags & REGISTERED),
"invalid base state");
if (!(base.flags & REGISTERED)) { return; }
MemoryManager::SyncAlias_(base.h_ptr, h_ptr, alias_size*sizeof(T),
base.flags, flags);
}
template <typename T>
inline MemoryType Memory<T>::GetMemoryType() const
{
if (!(flags & REGISTERED)) { return MemoryType::HOST; }
return MemoryManager::GetMemoryType_(h_ptr, flags);
}
template <typename T>
inline void Memory<T>::CopyFrom(const Memory &src, int size)
{
if (!(flags & REGISTERED) && !(src.flags & REGISTERED))
{
if (h_ptr != src.h_ptr && size != 0)
{
MFEM_ASSERT(h_ptr + size <= src || src + size <= h_ptr,
"data overlaps!");
std::memcpy(h_ptr, src, size*sizeof(T));
}
// *this is not registered, so (flags & VALID_HOST) must be true
}
else
{
MemoryManager::Copy_(h_ptr, src.h_ptr, size*sizeof(T), src.flags, flags);
}
}
template <typename T>
inline void Memory<T>::CopyFromHost(const T *src, int size)
{
if (!(flags & REGISTERED))
{
if (h_ptr != src && size != 0)
{
MFEM_ASSERT(h_ptr + size <= src || src + size <= h_ptr,
"data overlaps!");
std::memcpy(h_ptr, src, size*sizeof(T));
}
// *this is not registered, so (flags & VALID_HOST) must be true
}
else
{
MemoryManager::CopyFromHost_(h_ptr, src, size*sizeof(T), flags);
}
}
template <typename T>
inline void Memory<T>::CopyToHost(T *dest, int size) const
{
if (!(flags & REGISTERED))
{
if (h_ptr != dest && size != 0)
{
MFEM_ASSERT(h_ptr + size <= dest || dest + size <= h_ptr,
"data overlaps!");
std::memcpy(dest, h_ptr, size*sizeof(T));
}
}
else
{
MemoryManager::CopyToHost_(dest, h_ptr, size*sizeof(T), flags);
}
}
/** @brief Print the state of a Memory object based on its internal flags.
Useful in a debugger. */
extern void MemoryPrintFlags(unsigned flags);
/// The (single) global memory manager object
extern MemoryManager mm;
/// Main memory allocation template function. Allocates n*size bytes and returns
/// a pointer to the allocated memory.
template<class T>
inline T *New(const std::size_t n)
{
T *ptr = new T[n];
if (!MemoryManager::Exists()) { return ptr; }
return static_cast<T*>(mm.Insert(ptr, n*sizeof(T)));
}
/// Frees the memory space pointed to by ptr, which must have been returned by a
/// previous call to mfem::New.
template<class T>
inline void Delete(T *ptr)
{
static_assert(!std::is_void<T>::value, "Cannot Delete a void pointer. "
"Explicitly provide the correct type as a template parameter.");
if (!ptr) { return; }
delete [] ptr;
if (!MemoryManager::Exists()) { return; }
mm.Erase(ptr);
}
/// Return a host or device address corresponding to current memory space
template <class T>
inline T *Ptr(T *a) { return static_cast<T*>(mm.Ptr(a)); }
/// Data will be pushed/pulled before the copy happens on the host or the device
inline void* Memcpy(void *dst, const void *src,
std::size_t bytes, const bool async = false)
{ return mm.Memcpy(dst, src, bytes, async); }
/// Push the data to the device
inline void Push(const void *ptr, const std::size_t bytes = 0)
{ return mm.Push(ptr, bytes); }
/// Pull the data from the device
inline void Pull(const void *ptr, const std::size_t bytes = 0)
{ return mm.Pull(ptr, bytes); }
} // namespace mfem
#endif // MFEM_MEM_MANAGER_HPP
+16 -33
View File
@@ -9,53 +9,36 @@
// terms of the GNU Lesser General Public License (as published by the Free
// Software Foundation) version 2.1 dated February 1999.
#include "forall.hpp"
#include "occa.hpp"
#ifdef MFEM_USE_OCCA
#include "device.hpp"
#if defined(MFEM_USE_CUDA) && OCCA_CUDA_ENABLED
#include <occa/modes/cuda/utils.hpp>
#endif
namespace mfem
{
// This variable is defined in device.cpp:
namespace internal { extern OccaDevice occaDevice; }
namespace internal { extern occa::device occaDevice; }
static OccaMemory OccaWrapMemory(const OccaDevice dev, const void *d_adrs,
const size_t bytes)
occa::device &OccaDev() { return internal::occaDevice; }
occa::memory OccaMemoryWrap(void *ptr, std::size_t bytes)
{
// This function is called when an OCCA kernel is going to be used.
#ifdef MFEM_USE_OCCA
void *adrs = const_cast<void*>(d_adrs);
#if defined(MFEM_USE_CUDA) && OCCA_CUDA_ENABLED
// If OCCA_CUDA is allowed, it will be used since it has the highest priority
if (Device::Allows(Backend::OCCA_CUDA))
{
return occa::cuda::wrapMemory(dev, adrs, bytes);
return occa::cuda::wrapMemory(internal::occaDevice, ptr, bytes);
}
#endif // MFEM_USE_CUDA && OCCA_CUDA_ENABLED
// otherwise, fallback to occa::cpu address space
return occa::cpu::wrapMemory(dev, adrs, bytes);
#else // MFEM_USE_OCCA
return (void*)NULL;
#endif
return occa::cpu::wrapMemory(internal::occaDevice, ptr, bytes);
}
OccaMemory OccaPtr(const void *ptr)
{
// This function is called when 'ptr' needs to be passed to an OCCA kernel.
OccaDevice dev = internal::occaDevice;
if (!mm.UsingMM()) { return OccaWrapMemory(dev, ptr, 0); }
const bool known = mm.IsKnown(ptr);
if (!known) { mfem_error("OccaPtr: Unknown address!"); }
const bool ptr_on_host = mm.IsOnHost(ptr);
const size_t bytes = mm.Bytes(ptr);
const bool run_on_host = !Device::Allows(Backend::DEVICE_MASK);
// If the priority of a host OCCA backend is higher than all device OCCA
// backends, then we will need to run-on-host even if the Device allows a
// device backend.
if (ptr_on_host && run_on_host) { return OccaWrapMemory(dev, ptr, bytes); }
if (run_on_host) { mfem_error("OccaPtr: !ptr_on_host && run_on_host"); }
void *d_ptr = mm.GetDevicePtr(ptr);
return OccaWrapMemory(dev, d_ptr, bytes);
}
OccaDevice OccaDev() { return internal::occaDevice; }
} // namespace mfem
#endif // MFEM_USE_OCCA

Some files were not shown because too many files have changed in this diff Show More