zhangyiss/mlx - mlx - Gitea for Geophysics

mirror of https://github.com/ml-explore/mlx.git synced 2025-10-22 11:14:32 +08:00

Author	SHA1	Message	Date
Cheng	b8f76f717a	Print exceptions in eval_cpu/eval_gpu and abort (#1754 )	2025-01-08 06:31:09 -08:00
Awni Hannun	d1766f2c70	Add boolean mask support in vector SDPA (#1757 )	2025-01-07 20:24:53 -08:00
Awni Hannun	516ded618b	Dynamic slicing (#1741 ) * dynamic slice and slice update * python bindings + tests + fix set item * fix compile issue * comment * fix jit	2025-01-07 14:02:16 -08:00
Awni Hannun	058d6ce683	mpi send use input as output (#1750 ) * mpi send use input as output * move earlier	2025-01-06 06:08:43 -08:00
Awni Hannun	259025100e	Fix nd ternary on GPU (#1746 )	2025-01-03 11:52:17 -08:00
Awni Hannun	6fa0501387	Fix concatenate/slice_update vjp + reduce binary size (#1735 ) * fix concatenate vjp + reduce binary size * also cast in slice update	2025-01-02 16:36:33 -08:00
Valentin Roussellet	88f993da38	Explicit parentheses around some logical operators (#1732 ) * fix some warnings * format	2024-12-24 07:02:20 -08:00
Awni Hannun	ebfe64b92d	shapeless slice update and broadcast when possible (#1727 )	2024-12-23 11:25:15 -08:00
Awni Hannun	0308e9af71	Allow offset to be an mx.array for `mx.fast.rope` (#1724 ) * allow offset for rope * comment	2024-12-19 15:51:44 -08:00
Awni Hannun	e03f0372b1	More shape type (#1705 ) * more shape type * fix	2024-12-19 08:08:20 -08:00
Awni Hannun	7480059306	track resource limit and throw if exceeded (#1718 )	2024-12-18 18:45:58 -08:00
Awni Hannun	9111999af3	Fix small sort with metal validation (#1695 )	2024-12-12 09:21:45 -08:00
Awni Hannun	6bd28d246e	Allow no copy negative strides in as_strided and slice (#1688 ) * allow no copy negative strides in as_strided and slice * fix jit * fix jit	2024-12-12 08:59:45 -08:00
Awni Hannun	4e1e9520e1	Flatten and unflatten (#1692 ) * flatten and unflatten * fix grad * fix shape infer * use squeeze + unsqueeze in get_item	2024-12-11 21:51:37 -08:00
Awni Hannun	f76a49e555	`ExpandDims` primitive (#1687 ) * add squeeze primitive * simplify squeeze, use in gather * fix * fix * fix * fix * fix no cpu * use squeeze in matmul and friends * expand dims primitive * comment	2024-12-10 16:39:07 -08:00
Awni Hannun	40c62c1321	Use int64 stride everywhere (#1671 ) * use int64 stride everywhere * fix ext * fix ext * more shape + cleanup * one more * few more	2024-12-09 11:09:02 -08:00
Alex Barron	95c4a2e3af	add back conditionaltype (#1655 )	2024-12-06 11:12:01 -08:00
Jagrit Digani	9d40e521d7	Stop matrix copies with new attention kernel (#1639 )	2024-12-02 14:12:38 -08:00
Jesper Stemann Andersen	e4eeb4e910	Added missing unordered_map includes (#1635 ) * Added missing includes in mlx/io.h and mlx/backend/metal/metal.h * Added additional missing unordered_map includes that fixes build on FreeBSD	2024-12-02 07:03:03 -08:00
Ikko Eltociear Ashimine	9bc2183a31	docs: update device.cpp (#1632 ) unecessary -> unnecessary	2024-11-27 20:58:26 -08:00
Awni Hannun	d4b222b6d3	Fix some leaks and races (#1629 ) * fix leak and fix potential race * more leak fixes * fix one more	2024-11-27 20:01:20 -08:00
Awni Hannun	211411faf2	fix large ops (#1620 )	2024-11-24 09:17:10 -08:00
Alex Barron	6f7986d592	Cleaner `qmv`/`qvm` (#1616 )	2024-11-22 11:14:08 -08:00
Jagrit Digani	02bec0bb6d	Matrix Attention kernel (#1610 ) * Rough INIT * [WIP]: Loading and Matmuls added * [WIP]: Reductions and min working aligned kernel at headdim = 64 * [WIP] Added headdim 80 for testing * [WIP] Update dispatch params for testing * [WIP] Add support for unaligned seq lengths - still looks messy * Update sdpa_benchmarks * Update sdpa_benchmarks * Update sdpa_benchmarks * Enable gqa support * Update benchmark and switch off 128 headdim * Update headdim 128 tuning * Remove older fast attention code. Write out O strided * Disable hd=128 until further optimizations * Enable bf16 * Fix data size bug * Enable attn build outside of jit	2024-11-22 10:34:05 -08:00
Alex Barron	c79f6a4a8c	3 and 6 bit quantization (#1613 ) * Support 3 and 6 bit quantization	2024-11-22 10:22:13 -08:00
Awni Hannun	0c5eea226b	Reduce specializations (#1607 ) * start of reduce specializations * fix all reduce * fix many dims * fix * non-jit tests clear * cleanup instantiations * cpu merges * change dim specializations * optimize * fix jit * fix jit * use higher precision for integer sum+prod * fixes	2024-11-21 19:53:00 -08:00
Awni Hannun	dcca0d7477	contiguous op / prim (#1612 )	2024-11-21 19:51:49 -08:00
Awni Hannun	61d787726a	Fix view scalar bug segfault (#1603 ) * fix view scalar bug * fix view scalar bug * one more fix	2024-11-19 10:54:05 -08:00
Awni Hannun	2419edd5b2	Faster indexing math in a few kernels (#1589 ) * wip: faster compiled kernels * faster general unary with uint specialization * index type in compiled, unary, binary, ternary, copy * fix jit * jit fix * specialize gather + scatter * nit in docs	2024-11-18 19:52:00 -08:00
Awni Hannun	9d7fa6b8e6	Use osx deployment target to pick Metal version (#1595 ) * choose metal based on deployment target rather than system version * nit * unused compile def	2024-11-18 19:16:49 -08:00
Angelos Katharopoulos	073076ac7d	2-Pass Sdpa Inference Kernel (#1597 )	2024-11-18 17:31:53 -08:00
Awni Hannun	9bd03dd9b4	More buffer donation with no-ops (#1591 ) * more donation * fix test * fix build	2024-11-18 08:35:41 -08:00
Awni Hannun	6931f84412	fix dispatch threads for a few kernels (#1594 )	2024-11-18 08:35:25 -08:00
Awni Hannun	610af352d4	Dispatch bf16 at run time when using the JIT (#1584 ) * Dispatch bf16 at run time when using the JIT * fix extension * fix extension build * fix extension build * Update utils.h	2024-11-15 16:54:36 -08:00
Awni Hannun	b35f1e3c9c	fix donation in sdpa (#1587 )	2024-11-13 17:21:13 -08:00
Alex Barron	a4c47b0276	OOB QMV fix (#1579 ) * fix oob access in qmv * skip more * fix small case	2024-11-08 17:59:45 -08:00
Alex Barron	111fefd5e9	Fix OOB access in qmv (#1577 ) * fix oob access in qmv * skip more	2024-11-08 15:41:30 -08:00
Awni Hannun	c1fe1ef081	Bfs width limit (#1568 ) * width limit * fix * large limit * put env vars in env namespace	2024-11-08 15:00:46 -08:00
Awni Hannun	9f0d5c12fc	Fully wrap the command encoder (#1572 ) * fully wrap the command encoder * use consistent style + fix extensions	2024-11-08 11:50:21 -08:00
Awni Hannun	9a3842a2d9	fix (#1566 )	2024-11-06 17:10:33 -08:00
Alex Barron	26be608470	Add split_k `qvm` for long context (#1564 ) * Add splitk qvm * configurable splitk * tuning * remove extra instantiation * remove refactor * separate test * cpu tolerance	2024-11-05 11:25:19 -08:00
Angelos Katharopoulos	248431eb3c	Reductions update (#1351 )	2024-11-04 22:25:16 -08:00
Awni Hannun	f1951d6cce	Use fewer barriers (#1561 ) * use fewer barriers * comment	2024-11-04 10:26:49 -08:00
Angelos Katharopoulos	62f297b51d	Sdpa fix (#1558 )	2024-11-02 21:25:46 -07:00
Awni Hannun	4f72c66911	improvements to scatter / gather (#1541 )	2024-10-30 19:30:54 -07:00
Jagrit Digani	960e3f0f05	Gemm update (#1518 )	2024-10-30 19:30:28 -07:00
Awni Hannun	884af42da2	Fix thread group for large arrays (#1543 ) * fix thread group for large arrays * comment * one more	2024-10-30 16:25:12 -07:00
Carlo Cabrera	1a992e31e8	Skip using Residency sets in VMs (#1537 ) * Skip using Residency sets in VMs Attempting to use residency sets in a VM throws[^1] libc++abi: terminating due to uncaught exception of type std::runtime_error: [metal::Device] Unable to construct residency set. Not quite sure if this is the best fix, but it does make the error go away. Note that it was previously possible to run simple programs that used mlx in a VM prior to `0eb56d5be0`. See related discussion at Homebrew/homebrew-core#195627. [^1]: https://github.com/Homebrew/homebrew-core/actions/runs/11525831492/job/32105148462#step:3:56 Co-authored-by: Awni Hannun <awni.hannun@gmail.com> * change residency check --------- Co-authored-by: Awni Hannun <awni.hannun@gmail.com> Co-authored-by: Awni Hannun <awni@apple.com>	2024-10-29 19:37:23 -07:00
Awni Hannun	015c247393	change wino dispatch conditoin (#1534 )	2024-10-28 11:13:44 -07:00
Awni Hannun	d3cd26820e	Faster bits and bernoulli (#1535 ) * faster bits and bernoulli * fix bernoulli	2024-10-28 11:11:00 -07:00

... 2 3 4 5 6 ...

440 Commits